Technical SEO
Soft 404 Errors: What They Are and How to Fix Them in Search Console
A soft 404 is a URL that answers with a 200 (success) status code while its content tells the visitor the page does not exist, or shows almost nothing at all. Google excludes these pages from Search and lists them under “Soft 404” in the Search Console Page indexing report.
This is the third entry in my Search Console status series, after Discovered – currently not indexed and Crawled – currently not indexed. It covers what Google means by a soft 404, where they come from, the single-page application case in detail (with a small test you can run yourself), how to fix each cause, and how to validate the fix.
Key takeaways
- A soft 404 is a mismatch between the status code and the content: the server says “success”, the page says “missing” or shows nothing useful.
- Google decides this from the content, not the status code, so a real page can be flagged too if it renders blank or shows an error to Googlebot.
- The classic source on JavaScript frameworks is the SPA catch-all rule that serves
index.htmlwith a200for every path, including paths that do not exist. - The fix depends on the page: return
404or410if it is gone,301to a genuine equivalent if it moved, and fix rendering if the page is real. - Soft 404s are not just a report row. Google says they continue to be crawled and waste your budget.
What a soft 404 is, in Google’s words
Google’s crawling documentation defines it precisely: a soft 404 is a URL that returns a page telling the user that the page does not exist and also a 200 (success) status code. The same page adds that “in some cases, it might be a page with no main content or empty page.”
That second clause matters: an empty shell or a page whose main content failed to load can land in the same bucket as a friendly error page.
Google’s status code reference explains how the label gets applied. For a 2xx response, Google considers the content for indexing, and if the content suggests an error, an empty page or an error message, Search Console will show a soft 404 error. The same page is explicit that a 2xx status code does not guarantee indexing.
Soft 404 compared with the statuses around it
| Response | What the server says | What Google does |
|---|---|---|
200 with real content | The page exists | Considers it for indexing |
200 with an error message or empty content | The page exists | Treats it as a soft 404 and excludes it |
404 or 410 | The page does not exist | Does not index it; drops it if it was indexed |
301 to an equivalent page | The page has moved | Processes the redirect target instead |
The 4xx and 301 rows come from the same reference: Google doesn’t index URLs that return a 4xx status code, and treats a 301 as a strong signal to process the target.
The practical difference between a 404 and a soft 404 is that the first is an honest answer and the second makes Google work it out. Google’s crawl budget guide puts the cost plainly: a 404 is “a strong signal not to crawl that URL again”, while soft 404 pages “will continue to be crawled, and waste your budget.”
Where soft 404s come from
Google lists four ways a server or CMS can produce one: a missing server-side include file, a broken connection to the database, an empty internal search result page, and an unloaded or otherwise missing JavaScript file.
Deleted content that still returns 200. A product is removed from the database, but the template still renders with an empty body, a “no longer available” message and a 200. The URL looks alive to Google and reads as dead.
Redirects to an irrelevant page. Sending every removed URL to the homepage or a broad category is a common habit. Google’s fix guidance reserves the 301 for a page that has moved or has “a clear replacement”, and a homepage is neither, so the old URL can end up treated as a soft 404 instead of passing anything on. I cover the redirect decision in detail in SEO redirects, and the migration version of the mistake in site migration.
Empty listing pages. Internal search results with zero matches, a tag with no posts, a filter combination that returns nothing. Each renders a valid page whose main content is “nothing here”.
Pages that render blank for Googlebot. The page is real, but a script failed, a resource was blocked in robots.txt or an API call timed out during rendering. What Google received was a shell.
Single-page applications with a catch-all fallback. This one produces soft 404s at scale, so it gets its own section.
The SPA case: client-rendered “not found” pages that return 200
On the frameworks this site writes about most (React, Vue, Angular, and any client-rendered app on a static host), the typical soft 404 is not a content problem. It is a routing configuration that makes the server incapable of saying “404”.
Why every missing URL returns 200
A client-side router needs every path to load the same index.html, so the app can boot and decide what to show. Vue Router’s own guide tells you to add a simple catch-all fallback route to your server: “If the URL doesn’t match any static assets, it should serve the same index.html page that your app lives in.”
The same guide is honest about the consequence: “Your server will no longer report 404 errors as all not-found paths now serve up your index.html file.”
Hosting platforms describe the same rule. Netlify’s documentation gives /* /index.html 200 as the rewrite for single-page apps and says it will effectively serve the index.html instead of giving a 404 no matter what URL the browser requests.
Follow a request for a URL that does not exist. The server returns the shell with a 200, the app boots, the router finds no match and renders a not-found component. A human sees a helpful error page; Google sees a successful response whose rendered content is an error message, which is the textbook soft 404.
The router’s catch-all route does not help, because it runs in the browser after the status code has already been sent. I walk through the Vue version of this, including the SSR route out, in Vue SEO.
A test you can run in two minutes
To show the mechanism, I built the smallest possible version: one index.html shell and two Node servers. The first is the naive catch-all, the same shape as the plain Node example in Vue Router’s guide. The second checks the path against the routes that exist before it answers.
// naive.mjs: every path gets the app shell with a 200
import http from 'node:http';
import { readFileSync } from 'node:fs';
const shell = readFileSync('index.html', 'utf8');
http.createServer((req, res) => {
res.writeHead(200, { 'Content-Type': 'text/html; charset=utf-8' });
res.end(shell);
}).listen(4001);
// fixed.mjs: same shell, but the server knows which paths are real
import http from 'node:http';
import { readFileSync } from 'node:fs';
const shell = readFileSync('index.html', 'utf8');
const staticRoutes = new Set(['/', '/pricing', '/blog']);
const products = new Set(['blue-widget', 'red-widget']); // in production: a DB or API lookup
function routeExists(pathname) {
if (staticRoutes.has(pathname)) return true;
const m = pathname.match(/^\/products\/([a-z0-9-]+)$/);
return Boolean(m && products.has(m[1]));
}
http.createServer((req, res) => {
const { pathname } = new URL(req.url, 'http://localhost');
const status = routeExists(pathname) ? 200 : 404;
res.writeHead(status, { 'Content-Type': 'text/html; charset=utf-8' });
// The app still boots on a 404, so the visitor sees your not-found view and navigation.
res.end(shell);
}).listen(4002);
Running both on Node 22 and requesting four paths with curl -s -o /dev/null -w '%{http_code}' gave this output:
/pricing naive: 200 fixed: 200
/products/blue-widget naive: 200 fixed: 200
/products/discontinued-widget naive: 200 fixed: 404
/this-does-not-exist naive: 200 fixed: 404
The naive server answers 200 for a product that no longer exists and for a path that never existed. The fixed server sends the same shell to every visitor, so the user experience is identical, but the status line is now true.
Run the same check against your own site. Request a URL you know is fake and read the status line, not the page:
curl -sI https://example.com/this-does-not-exist | head -1
If it prints 200, every broken inbound link, typo and deleted route on the site is a soft 404 waiting to be reported.
Fixing it, in order of preference
1. Let the server or edge decide the status. The fixed server above is the principle: something that runs before the response is sent must know whether the route exists. For static routes that can be a build-time manifest; for dynamic routes it means a lookup at the edge or in a server function. Server rendering gives you this for free, one more reason the approaches in single-page application SEO matter.
2. Scope the fallback instead of applying it to everything. If only /app/ is client-rendered, rewrite /app/* rather than /* and let the host’s normal 404 handling cover the rest. Netlify’s docs use a scoped rule of this shape, /app/* /app/index.html 200!, in their shadowing example (the ! forces the rewrite even when a file exists at that path). A static-first build avoids the problem entirely: this site’s Astro build emits a real 404.html for the host to serve on unknown paths, with no SPA rewrite in front of it.
3. If you cannot change the server, use one of Google’s two client-side fallbacks. Google’s JavaScript SEO guide accepts that in client-rendered apps using meaningful HTTP status codes can be impossible or impractical, and recommends one of two strategies:
- a JavaScript redirect to a URL for which the server responds with a
404(for example/not-found), or - adding
<meta name="robots" content="noindex">to error pages using JavaScript.
Google’s sample for the first approach, lightly trimmed:
fetch(`/api/products/${productId}`)
.then(response => response.json())
.then(product => {
if (product.exists) {
showProductDetails(product);
} else {
// this product does not exist, so this is an error page
window.location.href = '/not-found'; // redirect to a URL the server answers with 404
}
});
The redirect only works if /not-found itself returns a 404 from the server, which on a blanket /* /index.html 200 rule it will not. Exclude that one path from the rewrite, or serve it as a static file with the correct status. It also depends on rendering: Google’s redirects guide warns that if you set a JavaScript redirect, Google might never see it if rendering of the content failed. The noindex route is simpler to ship but leaves the URL returning 200, so Google keeps fetching it to find the tag.
Framework notes
Next.js App Router. Calling notFound() renders the 404 UI, and Next.js injects a <meta name="robots" content="noindex" /> tag so the page is not indexed. There is a catch the docs spell out: if the existence check runs inside a <Suspense> boundary, “the response has already begun streaming as a 200, and the status can’t change once streaming has started.” The noindex keeps it out of results, but for a real 404 status the check has to run before the response streams. More on the App Router in Next.js SEO.
Vue and Nuxt. Vue Router’s guide points Node users to the real fix: use the router on the server side to match the incoming URL and respond with 404 if no route is matched. The Nuxt trap of pointing a static host at 200.html for every path is covered in Nuxt SEO.
React and Angular. Same mechanics, same escape routes: see React SEO, Angular SEO and the wider JavaScript SEO guide.
How to find soft 404s in Search Console
Open the Page indexing report and look in the “Why pages aren’t indexed” table for the Soft 404 row. Google’s description of that row: the page request returns what we think is a soft 404 response, meaning it returns a user-friendly “not found” message but not a 404 HTTP response code.
Click the row to see the affected URLs, then group them by template or URL pattern before touching anything. A fix usually lives in a template or a routing rule, not in individual pages.
For each group, inspect one representative URL. Google recommends you run a live URL inspection test and click View tested page to see a screenshot of how Google renders the page. That screenshot answers the first question immediately: is this page really gone, or did it just fail to render for Googlebot?
Two cheaper checks catch soft 404s before Google does:
- The fake-URL test from the SPA section, run against every host and subdomain you own.
- A crawl that flags thin
200responses. Any crawler that reports word count and title lets you filter for200pages with very little content, or titles containing phrases like “not found”, “no results” or “unavailable”. Those are soft 404 candidates.
Server logs help too: soft 404s show up as Googlebot requests to URLs that should no longer exist. The method is in log file analysis for SEO.
How to fix a soft 404, by cause
Google’s fix guidance splits the problem by the state of the page. Decide which row each URL group belongs to before you change anything.
| The page is… | Fix | Why |
|---|---|---|
| Gone, with no replacement | Return 404 or 410 | Tells Google the page doesn’t exist and you don’t want it indexed |
| Moved, or has a clear equivalent | 301 to the equivalent page | Google treats a 301 as a strong signal to process the target |
| Real, but flagged anyway | Fix what Googlebot sees when it renders | The content failed to load for Google |
| Real, but intentionally empty (zero results, empty tag) | noindex, or 404 when truly empty | An empty state is not a page worth indexing |
Gone: return 404 or 410
If you removed the page and there is no replacement, Google’s guidance is to return a 404 (not found) or 410 (gone) response. Build a useful custom 404 page for people, with navigation and links to popular content, but make sure the server still sends the 404 status with it. Google’s guidance says it directly: custom 404 pages are created solely for users.
Moved: 301 to a genuine equivalent
If the page has moved or has a clear replacement, return a 301. The key word is “clear”. If nothing on the site satisfies the same intent, let the URL 404.
Real page, flagged anyway: fix the render
This is the case that confuses people, because the page works in their browser. Google’s explanation: if an otherwise good page was flagged, it’s likely it didn’t load properly for Googlebot, it was missing critical resources, or it displayed a prominent error message during rendering.
The same section lists the usual culprits: resources blocked by robots.txt, too many resources on a page, server errors, and slow or very large resources. Work through them with the URL Inspection live test:
- Open View tested page and check the screenshot and rendered HTML. If the main content is missing, the page failed to render.
- Check the page resources for anything blocked or failed, including the API calls the page depends on. A blocked JavaScript bundle on a client-rendered page produces an empty shell.
- Look for error-like copy in the main content area, such as a region block or a large “something went wrong” banner.
Short pages can trip the classifier too. Jan-Willem Bobbink asked on LinkedIn how to fix soft 404 detection on a page about soft 404s, and the replies suggested adding more substantive content so the page reads as being about the topic rather than an instance of it. If a thin page is dominated by error-like wording, give it more to say.
Intentionally empty: noindex or 404
Zero-result search pages, empty tags and empty filter combinations are real URLs with nothing on them, and Google names empty internal search results as a soft 404 source. Keep internal search out of the index with noindex, stop linking to empty filters, and return a 404 for a category that is permanently empty.
Validating the fix
Once a template is fixed, tell Search Console. Open the Soft 404 row in the Page indexing report and start validation. Two details from Google’s documentation change how you should do it.
First, validation is all-or-nothing per issue: if you missed a fix, validation will stop when Google finds a single remaining instance of that issue. Fix every URL in the group before you click.
Second, you can scope it. Google’s tip is to submit a sitemap containing only your most important pages, filter the report by that sitemap and validate against that subset, which can complete faster than a request that includes all affected URLs.
Validation is optional: the same help page notes that Google updates the instance count whenever it crawls a page with a known issue. What validation adds is an email and a progress log, useful when you need to show that a release worked.
For URLs that now return a real 404, expect them to reappear under Not found (404) rather than vanish. Google’s description of that row says 404 responses are not necessarily a problem, if the page has been removed without any replacement. That is the intended outcome, not a new problem.
Why soft 404s are worth fixing
The first cost is crawl waste. A 200 never tells Google to stop, so soft 404 URLs keep being fetched. On a small site that is noise; on a large site, or an SPA where every broken inbound link resolves to a 200, it drains crawl capacity your real pages need. The mechanics are in crawl budget.
The second is diagnostic noise. When a real page is flagged because it rendered blank, you want it to stand out, not sit among hundreds of dead URLs.
Soft 404 checks sit in the crawl and indexing stage of my 12-phase SEO and GEO audit, next to the fake-URL status test and the rendering checks it depends on. If you want the full pass in order, the technical SEO audit checklist lays it out. If your framework is creating soft 404s faster than you can clear them, that routing and rendering work is what my GEO and technical SEO service covers.
Frequently asked questions
What is a soft 404?
A soft 404 is a page that returns a 200 (success) status code while its content says the page does not exist, or has no main content at all. Google detects the mismatch from the content and excludes the page from Search. Search Console reports it under “Soft 404” in the Page indexing report.
Is a soft 404 bad for SEO?
It is not a penalty, but it has real costs. The flagged URLs are excluded from Search, and Google says soft 404 pages continue to be crawled and waste your crawl budget.
What is the difference between a 404 and a soft 404?
A 404 is a status code the server sends to say the page does not exist, and Google drops such URLs from the index and crawls them less over time. A soft 404 is Google’s label for a page that returns 200 but reads as an error or empty page. The first is an honest answer; the second leaves Google to work it out from the content.
Why does my single-page application have soft 404 errors?
Most SPAs use a server rule that serves index.html with a 200 for every path, so the client-side router can take over. That means a URL that does not exist also returns 200, and the app then renders a not-found view. The fix is to let the server or edge return a real 404 for unknown routes, or to use Google’s client-side fallbacks: a redirect to a URL that returns 404, or a JavaScript-added noindex.
Why is Google reporting a soft 404 on a page that works?
Usually because the page did not render properly for Googlebot. Google lists blocked resources, too many resources, server errors and slow or very large resources as common reasons. Run a live URL Inspection test, open View tested page and compare the screenshot with what you see in your browser.