A July 2026 Microsoft preprint audited 48,095 accepted papers from four of computing's most selective venues and found that hallucinated citations have entered the permanent scholarly record: references that look valid on the page, passed peer review, and yet point to no real work or to the wrong authors. At the reference level the failure rate is under 1%; at the paper level it reaches roughly one in four accepted NeurIPS 2025 papers. The gap between those two numbers is the whole story.
The paper is "Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences" by Mark Russinovich, Ram Shankar Siva Kumar and Ahmed Salem of Microsoft (arXiv:2607.00738, posted 1 July 2026). It is a preprint, not itself peer-reviewed yet, and its authors are careful to frame their counts as a conservative lower bound. We treat every figure below as single-study, self-reported and directional, and we read it for the durable lesson it carries for anyone publishing content that AI systems will judge.
That lesson is not really about academia. It is about what happens to trust when AI hallucination leaks out of a chatbot session and into the record, and it maps directly onto how AI search engines decide whom to cite.
What did the Microsoft study actually measure?
It measured citation identity, not truth. The authors built an open-source pipeline called RefChecker that treats each bibliography entry as a checkable claim: does this reference resolve to a real work by the named authors? A citation is counted as likely hallucinated only when it fails at the level of identity. One of two ways:
- Fabrication: no indexed work matches the cited title, and the listed authors never produced anything matching it. The reference points to nothing.
- Author-identity corruption: a real paper with that title exists, but it's credited to a substantially different set of authors.
Crucially, the study excludes ordinary bibliographic drift: a wrong year, a venue change, a preprint that later became a proceedings paper, minor name variants. Those are logged but never counted. As the authors put it, the numbers "understate any broader notion of citation noise and are best read as a lower bound."
The scope matters, and the paper is honest about it: this checks whether a cited work exists and is correctly attributed, not whether the citing sentence is faithfully supported, nor whether the paper's science is correct. Citation-layer verification is necessary but not sufficient for integrity.
How does the verification pipeline work?
By trying to corroborate every reference against multiple independent authorities, then escalating only the suspicious ones. This mechanism is the part worth studying closely, because it is a clean model of how machine trust is actually established.
RefChecker resolves each normalized reference against Semantic Scholar, OpenAlex, CrossRef, DBLP and the ACL Anthology, using DOIs, arXiv IDs and cited URLs where present. A reference is escalated when it can't be verified, when author overlap falls below a 60% threshold (for references with three or more authors), or when a DOI, arXiv ID or URL resolves to a different work. Escalated cases go to an LLM with web-search capability, which must find a dedicated source page for the cited work, not merely another paper that repeats the same citation. Before the metadata is re-checked. Only references that stay unsupported, or whose best match conflicts substantially on authorship, are counted.
Two design choices stand out. First, no single source is treated as ground truth: when the authorities disagree, the pipeline preserves the evidence rather than trusting one. Second, the LLM is used as an escalation-and-evidence step, never as the sole labeler; the final call rests on external database matches. In the study's configuration, the extraction model was Gemini 3.1 Flash Lite and the deep-search model was Claude Haiku 4.5. The whole scan of one venue's 3,703 papers and 221,281 references cost about $157. Roughly four cents per paper.
That "corroborate against many independent sources, distrust what you can't verify" logic is the same principle behind why AI search resists manipulation and why self-serving content gets discounted. It's corroboration as a trust filter, applied to a bibliography.
Why does the failure look tiny and huge at the same time?
Because the answer changes with the denominator. This is the central finding, and it's easy to misread in either direction.
| Venue (2025) | Hallucinated refs, all references | Papers with ≥1 hallucinated ref | Papers with ≥2 (academic refs only) |
|---|---|---|---|
| ICLR | 0.38% | 18.7% | 1.9% |
| ICML | 0.54% | 23.3% | 3.4% |
| NeurIPS | 0.68% | 26.2% | 5.1% |
| USENIX Security | 0.81% | 34.9% | 4.8% |
Source: Russinovich, Siva Kumar & Salem, arXiv:2607.00738 (2026); 2025 venue-years, self-reported.
A sub-1% reference rate sounds negligible for a single bibliography. But a conference is not a single bibliography. ICLR 2025 alone held 221,281 extracted references across 3,703 papers: so even 0.38% works out to 835 flagged references in 692 distinct papers. Raising the bar to two flagged references in the same paper (a far more conservative signal, since one stray false positive no longer tips a paper) still leaves roughly one in twenty accepted NeurIPS and USENIX Security 2025 papers carrying at least two.
"A conference is not one reference. It is hundreds of thousands. At that scale, the same rate yields an unverifiable citation in roughly a quarter of accepted NeurIPS 2025 papers.". Phantom References, arXiv:2607.00738
The authors also report a post-ChatGPT rise. The share of papers with two or more hallucinated references sits about 1.9× higher than the pre-2023 baseline: and a high-count tail: single papers with 20 (NeurIPS), 13 (ICML and USENIX Security) and 8 (ICLR) likely hallucinated references. They are deliberately cautious about causation: the timing "is consistent with a change in authoring practice" but is not offered as proof that LLM-assisted writing caused the increase.
Why didn't peer review catch them?
Because reviewers don't verify bibliographies, and the data shows it plainly. This is the finding with the sharpest implications beyond academia.
When the authors joined each paper's hallucination count to its review scores, the difference between clean papers and affected ones was essentially zero: a mean-rating gap of +0.02 at ICLR, +0.005 at ICML and +0.04 at NeurIPS, with affected papers rated, if anything, marginally higher. Slicing by acceptance tier told the same story (posters 20.4%, spotlights 23.5%, orals 20.3% affected). Most tellingly, on ICLR 2023 the accepted papers (mean reviewer rating 6.61) were affected 16.0% of the time and the rejected papers (rating 4.67, nearly two points lower) 16.9% of the time. Almost identical, despite a large gap in perceived quality.
The conclusion the authors draw is narrow and important: bibliographic integrity is orthogonal to the quality signal reviewers extract. Whatever peer review is good at, "does this citation resolve to a real work" is not one of those things. Independent audits agree. GPTZero reported finding at least one human-verified hallucinated citation in 50 of 300 scanned ICLR 2026 submissions, each of which had already received three to five expert reviews, and the GhostCite analysis of 2.2M citations found 76.7% of surveyed reviewers admit they don't check references. Convergent measurements from independent teams are what make the phenomenon credible.
Is this just an academic problem?
No. It's a preview of the trust economy every publisher now operates in. Strip away the conference setting and the study describes a general failure mode: AI tools make it trivial to produce polished, confident, well-formed references (and claims) that don't survive verification: and the humans downstream rarely check. Substitute "buyer reading an AI answer" for "peer reviewer" and the risk is the same shape.
Author feedback in the study reinforces the point: the papers with the most failures weren't fraud, they were workflow breakdowns: bibliography entries generated or reformatted by LLM-based tools and then never verified, sometimes introduced at the camera-ready stage. Nobody in the pipeline owned the check. That is precisely the trap for a marketing or content team using AI to draft at scale.
Two durable lessons for AI visibility follow directly:
- Verify AI-assisted content before you publish. A fabricated statistic, a misattributed quote, or a link that resolves to the wrong thing is the content-marketing version of a phantom citation. It's cheap to make and expensive to trust. This is the concrete version of our standing rule that mass-produced AI content can hurt you. The risk isn't the tool, it's the unverified output. The industry now has a name for the low-effort end of this: ICML's 2026 review policy explicitly treats "AI slop" as interference.
- Be the entity that survives the check. The references that passed weren't the ones with the best prose. They were the ones that resolved cleanly against multiple independent authorities. For a brand, that means being findable, consistent and corroborated across sources you don't control, so that when an engine tries to verify a claim about you, it succeeds. That's the slow, durable work behind entity strength and why AI cites one brand over another.
How should you read a single-study result like this?
Carefully, and with its own caveats foregrounded. The same discipline we ask of any AI-visibility case study. The authors are unusually candid about the limits, and honesty about them is what makes the work trustworthy:
- False positives are the primary limitation. Most raw flags, on manual inspection, were not hallucinations. Mangled author strings, truncated titles, references corrupted during PDF extraction. A single automated flag "should never be treated as conclusive without review." The higher-threshold signals (two-or-more, and the five-plus tail) are the reliable ones.
- "Does not exist" and "not yet indexed" are indistinguishable to a scanner. Genuinely unpublished or very recently accepted work can be flagged even though it's real. A structural reason to keep humans in the loop.
- It's a preprint, LLM-dependent, and single-team. Results depend on the configured models and can shift as those models change; the paper is version 1, not yet through peer review itself.
None of that undoes the headline, because the headline rests on the conservative end of the measurement and is corroborated by independent audits. But it models the right posture: specific, dated, hedged, and cross-checked. That posture is the actual product here, for researchers and for brands alike.
The through-line is simple. Machine trust is earned by corroboration and lost by unverifiable claims, whether the claim is a citation in a NeurIPS paper or a statistic on your pricing page. Watching whether AI engines actually cite you, and whether the facts they repeat about you are the true ones. Is what Buffy Intel measures. Questions: [email protected].