Technical SEO

X-Robots-Tag: The HTTP Header for Noindex on PDFs, Images and Anything Without a <head>

· · 19 min read

The X-Robots-Tag is an HTTP response header that carries the same indexing rules as the robots meta tag (noindex, nofollow, nosnippet and the rest), but in the response headers instead of the HTML. It is how you apply those rules to files that have no <head> to put a meta tag in: PDFs, images, video, JSON and feeds.

This guide is written for the developer who has to ship it. It covers when the header is the right tool, the exact rules Google supports today, working configuration for Apache, Nginx, Netlify, Vercel, Next.js, Cloudflare Workers and Express, and how to confirm it with curl. The Nginx and Express examples were run against local servers, and the output shown is what they returned, including two ways the header silently goes missing.

Key takeaways

  • The header and the meta tag are equivalent. Google’s specification says “any rule that can be used in a robots meta tag can also be specified as an X-Robots-Tag.” Use the header when there is no HTML, or when one server rule is easier than editing many templates.
  • A URL blocked in robots.txt never shows Google its header. If the file is disallowed, the noindex is never read, so keep anything you want de-indexed crawlable.
  • noarchive no longer does anything in Google Search. Google now lists it with nocache and nositelinkssearchbox as rules it ignores.
  • The header is easy to lose. Nginx drops server-level add_header lines in any location that sets its own, and a later res.set() in Express overwrites an earlier one. Both are shown below with real output.
  • Check with curl, not with a browser extension. Request the exact URL, follow redirects, and read the X-Robots-Tag lines that come back.

What the X-Robots-Tag header is

A robots meta tag lives in the HTML: <meta name="robots" content="noindex">. The X-Robots-Tag moves the same instruction into the HTTP response, where it applies to whatever the response contains:

HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex

Google treats the two methods as interchangeable. Its guide to blocking indexing says they “have the same effect; choose the method that is more convenient for your site and appropriate for the content type.” What differs is where you set them and what they can reach.

Robots meta tagX-Robots-Tag header
Where it livesIn the HTML <head>In the HTTP response headers
Works on PDFs, images, video, JSONNoYes
Set byTemplates, CMS, SEO pluginWeb server, CDN, platform config or app code
ScopeOne page at a timeAny pattern: a file type, a directory, a whole host
Visible in page sourceYesNo, only in the response headers
Survives a CDN or proxyYes, it is part of the bodyOnly if nothing strips or replaces it

That last row matters. A header can be added, replaced or dropped by every layer between your application and the crawler, which is why this guide spends as much time on verifying as on setting it.

If you are still deciding whether a URL should be de-indexed at all or consolidated into another URL, that is a separate question, covered in canonical vs noindex. This page assumes the decision is made and you need to implement it.

When the meta tag cannot do the job

Reach for the header in these situations:

  • PDFs, Word documents and spreadsheets. A PDF has no <head>. Google’s spec says directly that “to block indexing of non-HTML resources, such as PDF files, video files, or image files, use the X-Robots-Tag response header instead.”
  • Images and video. Old product shots, internal diagrams or licensed images you do not want in image search. noindex on the image response removes that file, and noimageindex on an HTML page stops the images on it being indexed.
  • JSON, XML feeds and API responses. Public API endpoints and data files get discovered through links and sitemaps. A noindex header keeps them out of results without touching the payload.
  • Staging and preview environments. One rule at the host level covers every response, including the assets and files a template-level meta tag would miss.
  • Whole directories. /downloads/, /internal/, /print/ or an old campaign path, handled in one server rule instead of in every file.
  • Pages you cannot template. A third-party app mounted on a path, or a legacy system where editing the <head> means a release you cannot get.

For ordinary HTML pages you control, either method works.

The rules Google supports

The list below is taken from Google’s robots meta tag, data-nosnippet and X-Robots-Tag specification, which I read on 28 September 2026 (the page shows a last update of 24 March 2026). Google also publishes the list in a machine-readable format.

RuleWhat it does in Google Search
allNo restrictions. The default, so listing it has no effect.
noindexDo not show this page, media or resource in search results.
nofollowDo not follow the links on this page.
noneEquivalent to noindex, nofollow.
nosnippetNo text snippet or video preview. Google says this applies to web search, Google Images, Discover, AI Overviews and AI Mode, and also stops the content being used as a direct input for AI Overviews and AI Mode.
indexifembeddedAllows the content to be indexed when embedded in another page through an iframe, despite noindex. Only has an effect alongside noindex.
max-snippet: [number]Caps the text snippet at that many characters. 0 means no snippet, -1 lets Google choose.
max-image-preview: [setting]none, standard or large. large is the setting Google Discover eligibility depends on.
max-video-preview: [number]Caps a video preview at that many seconds. 0 allows a static image at most, -1 means no limit.
notranslateDo not offer a translation of this result.
noimageindexDo not index the images on this page.
unavailable_after: [date/time]Do not show this page after the given date. Google accepts widely used formats such as RFC 822, RFC 850 and ISO 8601.

The spec also states that the header name, the user agent name and the values “are not case sensitive”, so x-robots-tag: NoIndex is read the same way as X-Robots-Tag: noindex.

Rules that no longer do anything

Google keeps a short list of historical rules it now ignores. noarchive “is no longer used by Google Search to control whether a cached link is shown in search results, as the cached link feature no longer exists.” nocache “isn’t used by Google Search”, and nositelinkssearchbox went with the sitelinks search box.

Leaving noarchive in place does no harm, and other engines interpret their own rules (Bing’s use of it for Copilot is covered in getting cited by Microsoft Copilot). Just do not rely on it for anything in Google.

Combining rules and targeting one crawler

You can put several rules in one comma-separated header or send the header more than once. Google’s own example combines two headers:

X-Robots-Tag: noimageindex
X-Robots-Tag: unavailable_after: 25 Jun 2010 15:00:00 PST

To target one crawler, put its user agent token before the rules. Rules without a token apply to every crawler:

X-Robots-Tag: googlebot: nofollow
X-Robots-Tag: otherbot: noindex, nofollow

When several rules apply to Googlebot, Google combines them. The spec’s example is a page with nofollow for all robots and noindex for googlebot; Googlebot reads that as noindex, nofollow. It adds that “the search engine will use the sum of the negative rules.”

The spec notes that noindex “applies to search engine crawlers”; non-search crawlers such as AdsBot-Google may need rules addressed to them by name. AI training crawlers are controlled mostly through robots.txt tokens, listed in robots.txt for developers.

Server and platform configuration

Every example below sets noindex on a file type or directory. Swap in whichever rules you need. Where a pattern has a trap, the trap follows the code.

Apache

Google’s documentation gives this for all PDFs, in .htaccess or httpd.conf, with mod_headers enabled:

<Files ~ "\.pdf$">
  Header set X-Robots-Tag "noindex, nofollow"
</Files>

And for images:

<Files ~ "\.(png|jpe?g|gif)$">
  Header set X-Robots-Tag "noindex"
</Files>

Use set, not add. The mod_headers documentation warns that add can produce “two (or more) headers having the same name”, while set replaces any previous header with that name.

The same page explains Apache’s two internal header tables, onsuccess (the default) and always. Headers from a CGI or mod_proxy_fcgi backend (a typical PHP setup) live in the always table, so a plain Header set can leave a duplicated header. Apache’s documented pattern for that case is:

Header onsuccess unset X-Robots-Tag
Header always set X-Robots-Tag "noindex"

Nginx

Google’s equivalent for Nginx:

location ~* \.pdf$ {
  add_header X-Robots-Tag "noindex, nofollow";
}

The trap is inheritance. The Nginx headers module documentation says add_header directives “are inherited from the previous configuration level if and only if there are no add_header directives defined on the current level.” Add any header to a location block, and every add_header from the server block stops applying there.

I tested this on Nginx 1.24.0 with a staging-style config: X-Robots-Tag at server level, and an /img/ location that someone later gave a Cache-Control header.

server {
  listen 8081;
  add_header X-Robots-Tag "noindex, nofollow";

  location /img/ {
    add_header Cache-Control "public, max-age=86400";
  }
}

The output:

$ curl -sI localhost:8081/
HTTP/1.1 200 OK
Content-Type: text/html
X-Robots-Tag: noindex, nofollow

$ curl -sI localhost:8081/img/diagram.png
HTTP/1.1 200 OK
Content-Type: image/png
Cache-Control: public, max-age=86400

The image has lost its noindex, and nothing in the config looks wrong. The fix is to repeat the header in every location that has its own add_header. After that change the same request returned both Cache-Control and X-Robots-Tag: noindex, nofollow. On Nginx 1.29.3 and later, the docs describe a new add_header_inherit merge; directive that appends inherited values instead of dropping them.

Two more details from the same page. add_header only applies to responses with status 200, 201, 204, 206, 301, 302, 303, 304, 307 or 308 unless you add the always parameter. And with no add_header in a location at all, the server-level header is inherited as you would expect.

Netlify

Netlify reads headers from a _headers file in the publish directory or from [[headers]] tables in netlify.toml, according to its custom headers documentation. For a downloads directory:

# _headers
/downloads/*
  X-Robots-Tag: noindex
# netlify.toml
[[headers]]
  for = "/downloads/*"
  [headers.values]
    X-Robots-Tag = "noindex"

Two limitations from the same page matter here. Custom headers “apply only to files Netlify serves from our own backing store”, so they are not applied to proxied content or to URLs handled by a function or edge function, such as server-rendered pages. Those responses have to set the header themselves. And headers declared this way are “global for all builds and cannot be scoped for specific branches or deploy contexts.”

For staging, Netlify documents a workaround: keep per-context header files in a separate folder and copy the right one into the publish directory in that context’s build command.

[context.staging]
  command = "npm run build && cp ./custom-headers/_stagingHeaders ./dist/_headers"

Here _stagingHeaders would contain /* with X-Robots-Tag: noindex.

Vercel

In vercel.json, headers are a list of source patterns with key and value pairs, per the vercel.json reference:

{
  "headers": [
    {
      "source": "/downloads/(.*)",
      "headers": [
        { "key": "X-Robots-Tag", "value": "noindex" }
      ]
    }
  ]
}

Vercel already sends X-Robots-Tag: noindex on some deployments. Its response headers reference says the header is present on preview deployments and on outdated production deployments. Vercel’s knowledge base guide on preview indexing adds the exception: Vercel “omits X-Robots-Tag: noindex when a custom domain is assigned to a non-production branch.” A staging.example.com domain on a branch is indexable unless you add the header yourself. The guide’s vercel.json approach matches on host:

{
  "headers": [
    {
      "source": "/(.*)",
      "has": [{ "type": "host", "value": "staging.example.com" }],
      "headers": [{ "key": "X-Robots-Tag", "value": "noindex" }]
    }
  ]
}

Next.js

next.config.js takes a headers() function that returns source patterns, documented in the Next.js headers reference. The docs note that headers “are checked before the filesystem which includes pages and /public files”, so the rule covers static files in /public too.

// next.config.js
module.exports = {
  async headers() {
    const rules = [
      {
        source: '/downloads/:path*',
        headers: [{ key: 'X-Robots-Tag', value: 'noindex' }],
      },
    ];

    // Vercel sets VERCEL_ENV to "preview" on preview builds.
    if (process.env.VERCEL_ENV === 'preview') {
      rules.push({
        source: '/:path*',
        headers: [{ key: 'X-Robots-Tag', value: 'noindex' }],
      });
    }

    return rules;
  },
};

The preview check follows the pattern in Vercel’s guide above. Watch the ordering rule: Next.js says that if two entries match the same path and set the same key, “the last header key will override the first.” Put broad rules first and specific ones last, or the broad one will win. For page-level metadata, the wider framework picture is in the Next.js SEO guide.

Cloudflare Workers

A Worker can add the header to whatever the origin returns. Responses from fetch() are immutable, so Cloudflare’s alter headers example clones the response before changing it:

export default {
  async fetch(request) {
    const response = await fetch(request);
    const url = new URL(request.url);

    // Clone the response so its headers can be changed
    const newResponse = new Response(response.body, response);

    if (url.pathname.endsWith('.pdf') || url.pathname.startsWith('/downloads/')) {
      newResponse.headers.set('X-Robots-Tag', 'noindex');
    }

    return newResponse;
  },
};

This is useful when you cannot change the origin, such as a hosted CMS or a storage bucket. It also works the other way. newResponse.headers.delete('X-Robots-Tag') will strip a header the origin sends, which is exactly what you do not want a forgotten Worker doing to production.

Express

In Express, set the header in middleware for broad rules and in express.static’s setHeaders option for files. This is the server I ran with Express 5.2.1 on Node 22:

import express from 'express';

const app = express();
const isStaging = process.env.APP_ENV === 'staging';

// Staging: noindex every response
if (isStaging) {
  app.use((req, res, next) => {
    res.set('X-Robots-Tag', 'noindex, nofollow');
    next();
  });
}

// Production: noindex PDFs and images served from /public
app.use(express.static('public', {
  setHeaders(res, filePath) {
    if (/\.(pdf|png|jpe?g|gif|webp)$/i.test(filePath)) {
      res.set('X-Robots-Tag', 'noindex');
    }
  },
}));

// A JSON endpoint
app.get('/api/prices', (req, res) => {
  res.set('X-Robots-Tag', 'noindex');
  res.json({ plan: 'pro' });
});

// A Googlebot-only rule plus a rule for every crawler
app.get('/offer', (req, res) => {
  res.append('X-Robots-Tag', 'googlebot: nosnippet');
  res.append('X-Robots-Tag', 'unavailable_after: 2026-12-31');
  res.type('html').send('<p>Ends soon</p>');
});

app.listen(process.env.PORT || 3000);

On the production instance, each rule came back as intended:

$ curl -sI localhost:3001/docs/price-list.pdf
HTTP/1.1 200 OK
X-Robots-Tag: noindex
Content-Type: application/pdf

$ curl -sI localhost:3001/offer
HTTP/1.1 200 OK
X-Robots-Tag: googlebot: nosnippet
X-Robots-Tag: unavailable_after: 2026-12-31
Content-Type: text/html; charset=utf-8

On the staging instance, the PDF did not get the staging rule (output trimmed to the header lines):

$ curl -sI localhost:3002/
X-Robots-Tag: noindex, nofollow

$ curl -sI localhost:3002/docs/price-list.pdf
X-Robots-Tag: noindex

res.set() replaces a header, so the static handler’s noindex overwrote the middleware’s noindex, nofollow. Here it is harmless, because noindex survives. Reverse the two rules (a broad production max-image-preview:large followed by a narrow noindex, or the other way round) and the one you needed is the one that disappears. Switching to res.append() returned both headers, which Google combines:

X-Robots-Tag: noindex, nofollow
X-Robots-Tag: noindex

How to verify the header with curl

Page source will not show the header. Ask the server directly:

curl -sI https://example.com/downloads/price-list.pdf | grep -i x-robots-tag

-I sends a HEAD request. Most servers answer HEAD with the same headers as GET, but not all application code does, so for anything served by your app, check a real GET that throws away the body:

curl -s -o /dev/null -D - https://example.com/api/prices | grep -iE '^(HTTP/|x-robots-tag:)'

For more than a handful of URLs, a small script avoids mistakes. This one sends a GET with a Googlebot user agent string, follows redirects, and prints the status line, any Location and every X-Robots-Tag for each hop:

#!/usr/bin/env bash
# Usage: ./check-xrobots.sh urls.txt
while read -r url; do
  [ -z "$url" ] && continue
  echo "== $url"
  curl -s -o /dev/null -D - -L \
    -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
    "$url" | grep -iE '^(HTTP/|x-robots-tag:|location:)'
done < "${1:-urls.txt}"

Against the local test servers it printed:

== http://localhost:3001/docs/price-list.pdf
HTTP/1.1 200 OK
X-Robots-Tag: noindex
== http://localhost:3001/
HTTP/1.1 200 OK
== http://localhost:3002/
HTTP/1.1 200 OK
X-Robots-Tag: noindex, nofollow

Following redirects matters because the header that counts is the one on the final 200 response. A noindex on a 301 tells Google nothing about the destination. Sending a Googlebot user agent string shows you what a user-agent-sniffing CDN or firewall rule would serve. It does not prove what Google sees, because a server can also treat requests differently by IP.

For that, use Search Console. Google’s noindex debugging guidance recommends the URL Inspection tool to see what Googlebot received, and the Page Indexing report to monitor pages from which Googlebot extracted a noindex rule. It also warns that “depending on the importance of the page on the internet, it may take months for Googlebot to revisit a page”, so a correct header will not remove anything until the next crawl. Server logs show when that crawl happens; log file analysis covers how to read them.

robots.txt stops Google seeing the header

This is the mistake that makes a correct header useless. Google’s spec is explicit: robots meta tags and X-Robots-Tag headers “are discovered when a URL is crawled. If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.”

The common version: a team wants old PDFs out of Google, adds noindex for /downloads/, and also adds Disallow: /downloads/ to be thorough. Googlebot now never requests the files, never sees the header, and any PDF that was already indexed can stay there, because Google has been told not to look.

The order that works:

  1. Remove any Disallow rule covering the URLs.
  2. Serve X-Robots-Tag: noindex on them and confirm it with curl.
  3. Wait for a recrawl, or request one through URL Inspection for the URLs that matter.
  4. Only once they have dropped out, add a Disallow if you also want to save crawl activity on them. The trade-offs of that step are covered in crawl budget.

The full crawl-versus-index model, with the precedence rules for robots.txt, is in robots.txt for developers.

Staging and preview environments

Staging is where the header earns its place, because a host-level rule covers HTML, assets, files and API responses at once. It is also the source of the most expensive mistake on this page: the rule that ships to production.

Make the rule depend on the environment rather than on memory:

  • Key it to the host or an environment variable, never to a file someone has to remember to delete. The Vercel has: host rule, the Next.js VERCEL_ENV check, the Netlify per-context header file and the Express APP_ENV check above all do this.
  • Put a production check in CI. After each deploy, run the curl script against a few production URLs and fail the pipeline if any X-Robots-Tag line contains noindex where it should not.
  • Do not rely on noindex for privacy. Vercel’s guide makes the point for its own previews: the header “only asks search engines not to index the deployment. Anyone with the URL can still open it.” Anything confidential belongs behind authentication.

If staging has already been indexed, the header alone removes it only as each URL is recrawled. During a site migration, check the new production host for a leftover staging header before launch day, not after.

Common mistakes

These are the failure modes I check for in the crawling and indexing phase of a technical SEO audit:

MistakeWhat happensFix
Staging noindex left on productionPages drop out of the index as they are recrawledEnvironment-keyed rules plus a CI check on production headers
Disallow on the same URLsGoogle never reads the noindexAllow crawling until the URLs have dropped out
Nginx location with its own add_headerServer-level X-Robots-Tag silently missing on that pathRepeat the header in the location, or use add_header_inherit merge on 1.29.3+
Two rules calling set on the same headerThe later rule wins and the other is lostUse append or merge, or combine the values into one header
Header set on the redirect, not the destinationThe final URL has no ruleCheck the 200 response with curl -L
Netlify or similar headers on SSR or function routesPlatform header rules are not appliedSet the header in the function’s own response
A proxy, Worker or CDN rule rewrites headersThe origin sends the header but the crawler never gets itTest the public URL, not the origin
Meta tag and header disagreeHard to audit, and a restrictive rule you meant to remove can still applyPick one source of truth per URL type
Relying on noarchiveNo effect in Google SearchRemove it from your Google checklist

Frequently asked questions

What is the X-Robots-Tag?

It is an HTTP response header that carries robots rules such as noindex, nofollow and nosnippet. Google says any rule that works in a robots meta tag also works as an X-Robots-Tag. Because it sits in the response headers rather than in the HTML, it works for PDFs, images, video, JSON and other files that cannot hold a meta tag.

Is X-Robots-Tag better than the meta robots tag?

Neither is stronger. Google states the two methods have the same effect. Use the header for non-HTML files, for whole directories or hosts, and for staging environments. Use the meta tag when the rule belongs to a specific page and editors need to be able to see it in the source.

How do I noindex a PDF?

Serve the PDF with an X-Robots-Tag: noindex header from your web server, CDN or platform config, then confirm it with curl -sI on the file’s URL. Make sure the PDF is not blocked in robots.txt, or Google will never read the header. An indexed PDF drops out only after Google recrawls it.

Can I target Googlebot only with the X-Robots-Tag?

Yes. Put the user agent token before the rules, for example X-Robots-Tag: googlebot: noindex. Rules without a token apply to every crawler. If both are present, Google combines the restrictive rules that apply to it.

Does Google still support noarchive?

No. Google’s specification lists noarchive among historical rules it ignores, because the cached link feature it controlled no longer exists. It does no harm to leave it in place, and other search engines may still use it, but it has no effect in Google Search.

Why does Google still show a URL I set to noindex?

Usually because Google has not recrawled it since the header was added, or because robots.txt blocks the URL so the header is never seen. Check the live response with curl, remove any Disallow covering it, and request a recrawl through URL Inspection. If the header is missing on the public URL but present at the origin, something in between is removing it.

If you want the indexing rules on your site checked end to end, from server config through CDN to what Googlebot actually receives, that is part of my technical SEO and GEO work.