Citation hallucination is one of the clearest risks in AI-assisted academic work because the error hides inside a familiar scholarly form. A generated paragraph may sound suspicious. A generated reference often looks credible: author names, journal title, year, DOI, and formatting all appear to belong in a bibliography.
The numbers now make the risk hard to ignore. In a JMIR study of systematic-review-style reference generation, tested systems produced hallucination rates of 39.6%, 28.6%, and 91.4%. In another attribution study, ChatGPT produced correct or partially correct answers 50.6% of the time, while its suggested references existed only 14% of the time.
Those figures do not mean every AI tool has the same risk. They mean researchers should stop treating citations as decoration. If a tool gives you a reference, the next question is not "Does it look academic?" The next question is "Can I verify that the source exists and supports the claim?" The practical workflow is covered in detail in how to verify AI-generated citations; this article focuses on what the reported rates actually mean.
What Is A Citation Hallucination?
A citation hallucination is a reference error generated by an AI system. The error can be simple, such as a paper that does not exist, or mixed, such as a real journal paired with a fabricated title.
Researchers should separate these failure types:
- Non-existent citation: the paper cannot be found in credible scholarly indexes.
- Corrupted metadata: the title, author, year, venue, or DOI does not match the real paper.
- Misattributed claim: the paper exists but does not support the sentence where it is cited.
- Wrong identifier: the DOI or URL resolves to a different paper.
- Post-publication risk: the paper exists but has a retraction, correction, or expression of concern that changes how it should be used.
The third type is the easiest to miss. A model can cite a real source and still attach it to the wrong claim. For literature reviews, that is often worse than a fake title because it can pass a quick bibliography scan.
This is why citation hallucination rates should not be read as a simple leaderboard. A study that only checks whether a title exists will report a different rate than a study that checks whether the cited paper supports the generated claim.
Why Reported Rates Differ So Much
Citation hallucination rates vary because studies use different tasks, prompts, domains, models, and definitions of error. A systematic review task is not the same as a general question-answering task. A mental-health literature review is not the same as a legal research question. A bibliography check is not the same as a claim-support check.
The most important differences are:
- Task: citation generation, answer attribution, research synthesis, or retrieval-based question answering.
- Grounding: whether the tool retrieves from a source corpus before answering.
- Domain: common topics usually have more visible metadata than niche or newer topics.
- Evaluation rule: whether evaluators check existence only, metadata accuracy, or claim support.
- Human review: whether the output is machine-checked, manually checked, or both.
That means a lower number in one study does not automatically make one tool safer than another. The rate must be read with the method.
For researchers, the useful distinction is practical: general-purpose chat outputs carry high citation risk when asked to invent or attach references. Retrieval-based research tools can reduce some risks because they start from visible sources, but they still need checking. A grounded answer can still cite the wrong source, summarize badly, or overstate what a paper found.
Citation Hallucination Rates In Recent Studies
The table below summarizes several widely cited findings. It is not a ranking of products. It is a map of risk across task types.
| Source | What was tested | Reported result | What it means |
|---|---|---|---|
| JMIR systematic-review reference study | Large language models asked to reproduce references for systematic-review-style tasks | Hallucination rates of 39.6%, 28.6%, and 91.4% across tested systems | General models can produce highly plausible but invalid references in evidence-synthesis contexts. |
| Zuccon, Koopman, and Shaik attribution study | ChatGPT asked to answer domain-specific questions and provide external references | Answers were correct or partially correct 50.6% of the time, but references existed only 14% of the time | A model can answer plausibly while failing to provide real supporting evidence. |
| GhostCite preprint | Model-generated citations across research domains | The benchmark tested 13 models across 40 domains and reported citation hallucination rates from 14.23% to 94.93% | Citation risk varies strongly by model and domain, even under the same benchmark. |
| GhostCite archival audit | Published papers in AI, machine learning, and security venues | The study analyzed 2.2 million citations from 56,381 papers and reported 604 papers with invalid citations, or 1.07% of papers | Fabricated or invalid citations are not only a model-output problem; some reach the scholarly record. |
| Large-scale "in the wild" preprint | References across preprints and papers in major scholarly repositories | The study audited 111 million references across 2.5 million papers and estimated 146,932 hallucinated citations in 2025 | The issue appears large enough to matter at scholarly-infrastructure scale. |
| Legal RAG evaluation | AI legal research tools on real legal tasks | Commercial legal AI tools hallucinated between 17% and 33% of the time | Retrieval and proprietary corpora can reduce risk, but do not remove it. |
The main pattern is consistent. Citation risk falls when a system retrieves and exposes sources, but it does not disappear. The safest tool is not the one with the most confident answer. It is the one that lets you inspect the source, check the metadata, and confirm the claim.
What The JMIR Study Shows
The JMIR paper is useful because it tested a research task that resembles academic evidence work. The study used 11 systematic reviews across 4 fields, created 33 prompts, and analyzed 471 references. Papers were considered hallucinated when key bibliographic information did not match, such as title, first author, or year.
The reported rates were stark: 39.6% for one tested system, 28.6% for another, and 91.4% for another. The same study also reported low precision in matching the original systematic-review references: 9.4%, 13.4%, and 0% across the tested systems.
The implication for researchers is direct. A general model may produce a bibliography-shaped answer, but that is not the same as retrieving the evidence base for a review. If the task is a formal or semi-formal review, you need a documented search and screening process, not a generated reference list.
This is especially relevant when people use ChatGPT Deep Research for literature review. Even when an AI workflow is useful for scoping or drafting, the researcher remains responsible for confirming which papers exist, which papers were screened, and which papers support the final argument.
What Attribution Studies Show
The Zuccon, Koopman, and Shaik study is important because it separates answer quality from evidence quality. In their setup, ChatGPT gave correct or partially correct answers 50.6% of the time, but the suggested references existed only 14% of the time.
That gap is the heart of citation hallucination. A model may know enough about a topic to answer in a plausible way, yet still fail to attribute the answer to real evidence. A reader sees the answer and the reference together and assumes the second supports the first.
For literature reviews, this means you should not ask AI to "add citations" to a paragraph. That prompt rewards plausible attribution, not verified evidence. A safer approach is to ask for search terms, author names, databases to search, or criteria for what kind of study would support the claim.
The actual citation should come from a source you can open. If you are collecting papers from free discovery tools, a free AI literature review stack can help separate search, screening, reference management, and citation checking.
What Large-Scale Studies Add
Small evaluation studies show what models can do in controlled tasks. Large-scale audits show whether the problem reaches the scholarly record.
GhostCite reported a benchmark across 13 state-of-the-art models and 40 research domains, with hallucination rates ranging from 14.23% to 94.93%. It also audited 2.2 million citations from 56,381 papers and confirmed 604 papers with invalid citations, equal to 1.07% of the audited papers.
Another large-scale preprint audited 111 million references across 2.5 million papers and estimated 146,932 hallucinated citations in 2025. Nature reported the same broad concern in 2026, noting that tens of thousands of publications from 2025 might include invalid AI-generated references.
These numbers should be read with caution because preprints can change and detection methods have limits. But they point in the same direction: citation hallucination is not only a chatbot-output problem. Once unchecked references enter papers, they can be copied, cited, and indexed by downstream systems.
That is why academic search beyond Google Scholar should still be paired with verification. Better search helps, but verification is the layer that keeps a paper record from becoming false support.
Are Grounded Research Tools Safer?
Grounded tools are usually safer than ungrounded chat outputs because they retrieve or expose source records. If a tool gives you a title, DOI, abstract, PDF link, or library item, you can inspect the evidence. That is a structural advantage.
But grounded does not mean hallucination-free. The legal RAG evaluation is a useful cautionary example: even domain-specific retrieval tools hallucinated between 17% and 33% of the time. The domain is different from academic literature review, but the lesson transfers: retrieval reduces some errors and introduces a better audit trail, yet every claim still needs human verification.
For research tools, the safer pattern is source-first:
- Search or retrieve papers from scholarly records.
- Save the source metadata.
- Read or inspect the relevant passage.
- Then write the claim and attach the citation.
The unsafe pattern is prose-first:
- Ask for a polished paragraph.
- Ask AI to add citations afterward.
- Keep citations because they look plausible.
- Only verify if something seems wrong.
That second pattern is where hallucinated citations thrive. The first pattern is slower, but it produces a defensible literature review.
How Researchers Should Use The Rates
Do not use hallucination-rate studies to create a universal ranking of AI tools. Use them to decide how much verification a workflow needs.
The risk level rises when:
- The AI is asked to generate a bibliography from memory.
- The topic is niche, new, or poorly indexed.
- The output includes many specific claims.
- The tool hides or weakly exposes source records.
- The paper will be submitted to a journal, thesis committee, grant funder, or formal review process.
The risk level falls when:
- The tool starts from a real paper database or your own library.
- The answer links to inspectable source records.
- The workflow keeps search, screening, extraction, and writing separate.
- Every citation is checked against title, DOI, metadata, and claim support.
- The final bibliography is rebuilt from trusted records.
This also changes how to evaluate AI research assistants. Instead of asking whether a tool gives citations, ask whether it gives verifiable citations. A source card, DOI, PDF link, or BibTeX export is not enough by itself, but it gives you something to audit.
If you are already extracting evidence from papers, connect the extraction table to citation checking. A table in extracting data from research papers should include enough source detail to let you trace each claim back to a paper section, table, or figure.
A Risk-Based Citation Verification Rule
Not every source needs the same amount of review. A background paper in a low-stakes class assignment and a central source in a systematic review deserve different treatment.
Use this rule:
| Citation use | Minimum check | Extra check |
|---|---|---|
| Background reading | Confirm title and source record exist. | Check DOI if you plan to cite it. |
| Draft citation | Confirm title, DOI, authors, year, and venue. | Read the abstract and relevant section. |
| Key evidence claim | Confirm metadata and claim support. | Check citation context and post-publication updates. |
| Systematic or rapid review | Confirm metadata, inclusion reason, and extraction fields. | Keep a review log and verify final bibliography. |
| AI-generated bibliography | Treat every reference as unverified. | Rebuild each citation from a trusted record. |
This is not about distrusting every tool. It is about assigning verification effort where the downside is highest. A mistaken background source wastes time. A fake central citation can damage the credibility of the whole review.
For publication workflows, this also connects to disclosure. If AI helped generate or inspect citations, the final paper may need an AI-use statement depending on the journal or institution. The guide on how to disclose AI use to a journal can help with that language.

Where TrueCite Fits
TrueCite, powered by WisPaper, checks BibTeX files against real academic databases to flag hallucinated references. It is useful when a researcher needs a focused citation-verification step before relying on AI-generated references.




