Measurement
AI Citation Accuracy: 80 Figures From ChatGPT and Gemini, Checked Against the Pages They Linked
The fabricated-citation studies that everyone quotes were run on models that could not browse. Ask a search-connected assistant for a statistic today and it usually hands you a number and a link, which changes the question. Not “did it invent a source” but “is the number it gave me actually on the page it pointed to”. Nobody I could find had checked that for marketing statistics, so on 11 September 2026 I asked ChatGPT and Gemini the same ten questions, extracted every figure each one attributed to a source, and fetched every linked page to see whether the figure was there.
The short version: ChatGPT put 37 of its 40 figures on the page it linked. Gemini managed 25 of 40, gave seven figures with no link at all, and three times linked a Google search for a page that does not exist. The rest of this piece is the method, the numbers with their intervals, the failure modes, and what a content team should do about them.
Key takeaways
- Across 80 figure-bearing claims (40 per engine), ChatGPT’s figure was on the linked page 93% of the time (95% interval 80% to 97%); Gemini’s was 63% (47% to 76%). The gap is the finding; the exact percentages are a ten-prompt sample.
- Neither engine linked a dead page. Every URL that pointed at a real path returned 200. The failures were different: figures with no link, links to homepages, links wrapped in a Google search for a mistyped path, and ranges stretched beyond what the page says.
- ChatGPT cited primary and press sources 86% of the time; Gemini 23%. Gemini’s answers ran mostly through statistics-aggregator pages, and one of those attributed a real academic figure to a report that does not appear to exist.
- The two engines barely agree on where marketing facts live: of 46 domains cited, six were cited by both.
- For a content team the implication is the verification gate in the approval workflow: a figure and a link are not a source until the figure has been found on the page, and for Gemini output that check fails more than a third of the time.
Why “fabricated citations” is the wrong question for a search-connected assistant
The number people reach for is from Walters and Wilder’s study in Scientific Reports: across 636 bibliographic citations in 84 AI-generated literature reviews, 18% of GPT-4’s citations were entirely fabricated and 24% of the non-fabricated ones contained substantive errors, measured on GPT-4 as of mid-2023, with the 24% covering the non-fabricated citations only. That study asked a model with no web access to produce academic references from memory. It found what you would expect from memory.
The assistants a marketer uses in 2026 do something else. Asked for a statistic, ChatGPT runs a web search and returns a figure with a link; Gemini grounds its answer in search results and attaches source cards. The invention risk has moved. A model that can read a page rarely needs to invent a URL, but it can still attach the wrong number to a right page, round a figure into a different claim, cite a page that does not carry the figure at all, or give a figure and skip the link. Those are citation fidelity failures, and they are the ones that reach a draft, because a reviewer who sees a plausible number and a real-looking domain tends to stop checking.
So the test here is fidelity, not fabrication: for each figure an assistant attributed to a source, is that figure on the page the assistant linked?
Method: ten prompts, two engines, every link fetched
The ten prompts are marketing and SEO questions that have a defensible numeric answer and a known set of published sources, phrased to ask explicitly for a figure and a link. They are listed in full in the per-prompt table below. Each was submitted once to ChatGPT and once to Gemini through Cloro’s collection API, geo-targeted to the UK, on 11 September 2026. Perplexity was not included because the collection endpoint for it failed on every attempt in the previous study on this site, and this one did not re-test it.
From each of the 20 responses I extracted every figure the assistant attributed to a source, together with the link the response attached to that figure. A figure with a link was a claim; a link with no figure attached (a “read more” pointer) was recorded but excluded from the fidelity count. That produced 88 rows, of which 80 carried a figure, 40 per engine.
Every linked URL was then fetched with a headless browser and the page text was searched for the figure. A match was checked by hand in context, because a number can appear on a page for a different reason. Each figure ended in one of these categories:
- On the page. The figure, or the same figure spelled out, appears on the linked page with the meaning the assistant gave it.
- On the page, rounded. The page carries the figure to more precision than the assistant gave (89.54% reported as “around 90%”, 64.7% as “around 64%”, 210.5% as “over 200%”). Counted as on the page.
- Page carries a narrower figure. The assistant gave a range and only one end of it is on the page.
- Link is a homepage or index. The assistant linked a site root or a blog index rather than a page that carries the figure.
- Link is a search URL to a non-existent path. The assistant’s link was a Google search for a URL, and the URL inside it does not exist.
- No link given. A figure with an attribution in words and no URL anywhere in the response.
- Derived or unverifiable. The figure was computed by the assistant rather than stated by the source, or the linked page could not be read (one Scribd upload behind a login).
The raw responses, the claim table and the fetched page text are kept with the Optix research artifacts for this article, so any row can be re-checked.
Results: ChatGPT 93% on the page, Gemini 63%
| Outcome | ChatGPT (40 figures) | Gemini (40 figures) |
|---|---|---|
| On the linked page, exact | 35 | 21 |
| On the linked page, rounded | 2 | 4 |
| On the page, total | 37 (93%) | 25 (63%) |
| Page carries a narrower figure | 0 | 2 |
| Link is a homepage or index | 1 | 3 |
| Link is a search URL to a non-existent path | 0 | 3 |
| No link given | 0 | 7 |
| Derived by the assistant, not on the page | 1 | 0 |
| Unverifiable (page unreadable) | 1 | 0 |
With 40 figures per engine the Wilson 95% interval for ChatGPT is 80% to 97% and for Gemini 47% to 76%. The intervals do not overlap, so the direction is not a sampling accident, but the point estimates should be read as “most of the time” and “roughly two thirds” rather than as precise rates. Ten prompts on marketing statistics, one run each, on one day.
Two things did not happen. No engine linked a URL that returned an error; every real path came back 200. And no engine attached a figure to a page that plainly contradicted it. The one derived figure was ChatGPT computing Perplexity’s share of all web search from a page that gives only its share of AI search, which is an inference presented as a citation rather than a wrong number.
The failure modes are different, and Gemini’s are the expensive kind
ChatGPT’s three misses were minor. It linked the OpenAI homepage as the “source” for ChatGPT’s weekly active users while citing Wikipedia for the figure itself; it derived one number; and it cited a Scribd upload of a vendor report that cannot be read without an account. None of those puts a wrong figure in a draft. They put an unsupported one in.
Gemini’s fifteen misses fall into four patterns, and each is a different amount of work for the person checking.
No link at all. On the AI Overview click-through question, Gemini gave five figures, an APA-style reference list with four report titles, and not a single URL. Two of the titles I could not match to any published page. The figures themselves were in the right neighbourhood of the real studies, but a reviewer would have to reconstruct the sources from scratch. Seven of Gemini’s 40 figures arrived this way; all seven came from that one response and the B2B-marketers question, which also produced a reference to the Content Marketing Institute’s homepage and an unlinked attribution to “Typeface / Averi”.
Search links to pages that do not exist. Three times Gemini’s inline “read more” link was a Google search URL whose query was a page path, and the path was wrong: a Backlinko page called google-net-stats (the real page is google-ctr-stats), an Omnibound page with the singular ai-overview-statistics instead of the plural, a Ransen page with overview instead of overviews. Each one resolves to a search results page rather than an error, which is the most deceptive kind of bad link: it looks like it worked.
Homepage links. The Cloudflare Radar root, the Imperva blog index and the Content Marketing Institute root were each offered as the source for a specific figure. The figure may well be somewhere on those sites; the link does not take you to it.
Ranges stretched past the page. “71% to 88%” of health queries triggering an AI Overview, where the linked page says 71%; “1.3% to 2.5%” for Perplexity’s AI-search share, where the page says 2.5%. The lower or upper bound came from somewhere the response did not disclose.
The pattern behind all four is the same: Gemini’s answer text and its source cards are loosely coupled. The cards point at real pages that support part of the answer, and the prose carries figures the cards do not. ChatGPT’s answers put the link next to the figure it supports, which is why they verify.
Where the numbers come from: primary sources versus statistics pages
The second difference is not about accuracy but about provenance, and it matters more for anyone who republishes the figure.
Of the 43 links ChatGPT gave, 37 (86%) went to the organisation that produced the data or to a trade publication reporting it directly: SparkToro, Ahrefs, Similarweb, Statcounter, Adobe, Conductor via Search Engine Land, the Pew Research Center, Cloudflare, Imperva, DataDome, the Content Marketing Institute, an arXiv preprint. Of Gemini’s 35 links, 8 (23%) did. The rest went to statistics-roundup pages: Omnibound, Presenc AI, TechnologyChecker, theStacc, InsightMark Research, Arcalea, Digital Applied, Serplify, Cognizo, Klatch, a Forbes Council post, a WSI franchise blog.
Those pages are not wrong as such; several of them carried Gemini’s figures exactly, which is why they count as on the page. But they are one hop from the data, and the hop is where meaning changes. One example from this sample: a page Gemini linked for the share of informational queries that trigger an AI Overview states 64.7% and attributes it to a “Google 2025 Search Quality Report”, which the page does not link and I could not find. The 64.7% figure is real; it is the question-form activation rate in the arXiv measurement study of Google AI Overviews (55,393 queries, March to April 2026), which ChatGPT linked directly for the same prompt. The number survived the hop. The attribution did not.
The overlap tells the same story. The two engines cited 46 distinct domains between them and agreed on six. If your work depends on being the page an assistant sends people to, how to get cited by ChatGPT and how to get cited by Gemini are different problems, and the Gemini one runs through the aggregator layer.
What a content team should do with this
The practical reading is a rule for the verification gate rather than a verdict on a product. The content approval workflow on this site has a gate that fails any statistic without a source; this study says what “source” has to mean. A figure with a link is a candidate. It becomes a source when the figure has been found on the linked page, in context, with the meaning the draft gives it.
- Check at the link level, not the domain level. Every bad link in this sample was on a real, reputable-looking domain. The domain tells you nothing about whether the figure is on the page.
- Treat Gemini figures without an adjacent URL as unsourced. In this sample they were unsourced 18% of the time, and the reference-list format makes them look more sourced than a bare number.
- Hop once more when the link is an aggregator. Find the organisation that produced the figure and cite that page. The aggregator’s attribution is the part most likely to be wrong.
- Record the fetched text, not just the URL. The trace described in LLM observability for content teams keeps the retrieved page text next to the claim, which is what lets you re-check a figure after the page changes.
- Sample the engines you care about, and re-sample. The share-of-voice methodology piece showed a quarter of engine answers changing within an hour. Citation behaviour will move too; this is a September 2026 snapshot.
Limits of this study
Ten prompts, one run each, on one day, from one country, on two engines. The prompts all asked for figures with links, which is the best case for citation behaviour; a casual question gets a casual answer. The topics are marketing and SEO statistics, where the underlying sources are mostly vendor research, and fidelity may differ for medical or legal figures. Figure matching was automated and then adjudicated by one person. The categories are mine and another analyst could draw the line between “rounded” and “different” elsewhere; the table above gives the counts so you can redraw it. And the study says nothing about whether the sources themselves are right, only about whether the assistant reported them faithfully.
Per-prompt results
Figures on the linked page, out of figures attributed, per prompt.
| Prompt | ChatGPT | Gemini |
|---|---|---|
| What percentage of Google searches end without a click? | 4/4 | 4/4 |
| How much of website referral traffic comes from AI assistants? | 6/6 | 4/4 |
| Average click-through rate for position 1 in Google organic results | 3/3 | 3/4 |
| How often do AI Overviews appear in Google search results? | 7/7 | 6/9 |
| How many people use ChatGPT weekly? | 1/2 | 1/1 |
| What share of B2B marketers use generative AI for content creation? | 3/3 | 1/3 |
| How much does organic CTR drop when an AI Overview is present? | 5/5 | 0/5 |
| What is Perplexity’s market share of search? | 3/4 | 1/2 |
| How much of web traffic comes from bots and AI crawlers? | 4/4 | 4/7 |
| What percentage of marketers measure content marketing ROI? | 1/2 | 1/1 |
Frequently asked questions
Do ChatGPT and Gemini make up sources?
Not in the way the 2023 studies found for models without web access. In this sample neither engine linked a page that did not exist. What Gemini did, three times in 40, was link a Google search for a page path that does not exist, and seven times it gave a figure with no link at all. ChatGPT’s links all resolved, and 37 of its 40 figures were on the page it linked.
How accurate are AI citations for statistics?
In this test, ChatGPT’s stated figure was on the linked page 93% of the time and Gemini’s 63%, with 95% intervals of 80% to 97% and 47% to 76% on 40 figures each. Those are fidelity rates for marketing statistics on one day in September 2026, not general accuracy scores, and they say nothing about whether the sources themselves are correct.
Which is better for sourced statistics, ChatGPT or Gemini?
For a figure you intend to republish, ChatGPT in this sample: higher fidelity to the linked page and 86% of links to primary or press sources, against Gemini’s 23%. Gemini’s figures were often in the right range but were attached to aggregator pages, homepages or no link, which means more work to find the original.
How do I check whether an AI citation is real?
Open the link and search the page for the figure, then read the sentence around it to confirm it means what the assistant said. If the page is a statistics roundup, follow it one hop further to the organisation that produced the number and cite that. If there is no link, treat the figure as unsourced until you have found it yourself.
Why does Gemini give references without links?
Gemini’s answer text and its source cards are generated loosely coupled: the cards point at pages that support part of the answer, and the prose can carry figures and attributions the cards do not cover. In this sample that produced one response with an APA-style reference list and no URLs, and several figures whose only “source” was a homepage.