Back
Science

Rise in Fabricated Citations in Academic Literature: A 2023–2025 Analysis

View source

Rise of Fabricated Academic Citations Tied to AI Tools

A series of studies published in 2025 have documented an increasing number of references to non-existent academic papers in biomedical and machine-learning literature. The findings, published in The Lancet and reported by the startup GPTZero, indicate a rise in fabricated citations since 2023, with researchers citing the potential role of generative AI tools and systemic issues in peer review.

Scope and Methodology

Biomedical Literature Audit

A study published in The Lancet on May 7, 2025, conducted by researchers at Columbia University, analyzed 2.5 million papers from the PubMed Central database. The audit screened 125.6 million references, focusing on 97 million citations that contained valid DOIs or PubMed IDs.

Using large language models to cross-reference citation titles against four scholarly databases, the pipeline flagged titles that could not be verified in any database. Manual checks by three independent reviewers confirmed fabrication in approximately 70% of flagged cases.

Key data points from the Lancet study:

  • Over 4,000 fabricated citations were identified across nearly 3,000 papers
  • The rate of papers containing at least one fabricated citation increased from 1 in 2,828 in 2023 to 1 in 458 in 2025
  • In the first seven weeks of 2024, the rate rose to 1 in 277
  • There were 12 times more publications with fabricated citations in 2025 compared to 2023
  • More than one-third of fabricated citations originated from two large open-access publishers, not named in the study

Machine Learning Conference Analysis

Separate analyses by GPTZero, a Canadian startup, examined papers submitted to top AI conferences. The company identified hundreds of AI-hallucinated citations in over 53 papers accepted and presented at the NeurIPS 2025 conference in San Diego. GPTZero also reported having previously uncovered 50 fabricated citations in papers under review for another AI research conference, ICLR.

The company's tool scans citations by verifying authors, titles, publication venues, and links against the open web and academic databases.

Types of identified hallucinations included:

  • Completely fabricated citations (nonexistent authors, paper titles, journals, or URLs)
  • Blended or paraphrased elements from multiple real papers, creating plausible but incorrect citations
  • Subtle alterations to real papers, such as expanding author initials or adding/dropping coauthors
  • Plainly incorrect author names, such as "John Smith" or "Jane Doe"

Limitations and Caveats

Study co-author Maxim Topaz described the findings as "conservative underestimates" and the "lower bound of true prevalence."

Kathryn Weber-Boer, director of scientometrics at Digital Science, noted that Google Scholar is not a reliable source for verification because some fabricated references appear there but do not trace back to genuine publications.

GPTZero reports its hallucination checker is over 99% accurate. For the NeurIPS analysis, a human expert from the company's machine-learning team manually verified every flagged citation.

Potential Causes

The growth in fabricated references suggests a possible role of generative AI, though it remains unclear whether references are fabricated by AI or by humans. Lead author Maxim Topaz reported that his own investigation was prompted by a personal experience: an AI app he used to polish a scientific paper inserted a fabricated citation, which passed through several layers of peer review before being caught by an editor.

He explained that such errors can occur when an author asks AI for a citation to support a statement, either inventing research attributed to a real author or fabricating a citation entirely.

Mohammad Hosseini (Northwestern University) described the practice as a sign of a "flawed scholarly evaluation model" and superficial engagement with literature.

Misha Teplitskiy (University of Michigan) called the findings "a signal of slop" from LLM use.

Publisher and Conference Responses

Organization Response Science family of journals Uses automated tools to check references; reports no issues detected New England Journal of Medicine and JAMA Both have citation validation processes; authors are held responsible for accuracy PLOS Reports "numerous" unverifiable references in submissions; is piloting tools to detect them NeurIPS board Acknowledged evolving LLM use; instructed reviewers to flag hallucinations in 2025. Noted that even a small percentage of papers with incorrect references does not necessarily invalidate paper content. Affirmed commitment to evolving review and authorship processes ICLR Hired GPTZero to check future submissions following prior detection of 50 hallucinated citations in papers under review

Implications

Researchers warn that fabricated citations can compromise systematic reviews and clinical guidelines. Topaz stated that errors identified in the study have not been corrected or retracted.

In academic norms, fabricated citations are typically grounds for rejection, as references are crucial for anchoring research and demonstrating engagement with existing work. Approximately half of the papers with hallucinated citations identified by GPTZero were also flagged as likely having high AI generation or usage.

The large volume of submissions (21,575 for NeurIPS 2025) presents a challenge for deep scrutiny by volunteer reviewers.