A fabricated reference is hard to spot by eye. The authors are often real researchers in the right field. The journal exists. The title sounds like something they would write. The volume, pages and DOI have the correct format. Only the paper is missing.
That is a predictable result of how language models work, and it is why the University of Toronto's guidance for graduate students warns that "Generative AI tools have also been shown to reference scholarly works that do not exist." The responsibility for catching them lies with you.
How big the problem is
Chatbots without search. Walters and Wilder (2023) asked ChatGPT to write short literature reviews on 42 topics and checked all 636 references. With GPT-3.5, 55% were fabricated. With GPT-4, 18% were. Among the references that were real, 43% (GPT-3.5) and 24% (GPT-4) had substantive errors such as wrong authors, dates, volumes or page numbers.
Search-enabled models and research agents. Rao, Wong and Callison-Burch (2026) examined more than 221,000 citation URLs produced by commercial models and research agents, ten of them on one benchmark and three on another. Between 3% and 13% of URLs were hallucinated, and 5% to 18% did not resolve. Deep research agents produced more citations per answer and had higher hallucination rates than search-augmented chat models.
The trend is improving, and the residual rate is still far too high to skip checking.
Four kinds of bad citation
| Type | What it looks like | How it is caught |
|---|---|---|
| Fully fabricated | Nothing matches: no such title, and the DOI fails or leads elsewhere | Steps 1 and 2 |
| Chimera | Real authors and a real journal with an invented title, or a real title attached to the wrong authors and year | Steps 2 and 3 |
| Real but garbled | The paper exists but the year, volume, pages, DOI or spelling of an author's name is wrong | Step 3 |
| Real but misused | The paper exists and the details are right, but it does not say what the AI claims | Step 5 only |
The fourth type is the hardest, because every mechanical check passes.
The five-step check
Step 1: Resolve the DOI
If the reference has a DOI, paste it after https://doi.org/ in your browser.
- It resolves to a paper with the same title and authors: good, go to step 3.
- It resolves to a different paper: the DOI is invented or borrowed. Go to step 2 to see whether the cited title exists at all.
- "DOI not found": strong sign of fabrication, though typing errors happen. Go to step 2.
No DOI is not suspicious in itself. Books, many reports, and older or smaller journals do not have them.
Step 2: Search the exact title
Put the full title in quotation marks and search:
- Crossref (search.crossref.org), which indexes DOI-registered scholarly works.
- Google Scholar, which has the widest coverage, including books, theses and reports.
- A discipline database where relevant, such as PubMed (its Single Citation Matcher accepts partial details), ERIC or IEEE Xplore.
- WorldCat or a national library catalogue for books. Check the ISBN too.
If an exact-title search finds nothing, try the first author's name plus two distinctive title words. A real paper may have been slightly misquoted. If the authors are real, look at their publication list on Google Scholar, ORCID or their university page. A paper that is absent from its supposed first author's own publication list is very unlikely to exist.
Time limit: if two or three minutes across these sources turn up nothing, treat the reference as fabricated and remove it. Do not go looking for a similar real paper to swap in unless you are going to read that paper.
Step 3: Confirm the details at the source
Open the publisher's page or the PDF and check authors and their order, year, journal title, volume, issue, pages or article number, and DOI. Correct your reference from the source, not from the chatbot's version. Better still, import the record into your reference manager by DOI so that the metadata comes from the registry.
Step 4: Check the paper's status
- Look for a retraction, correction or expression of concern notice on the publisher page.
- Search the Retraction Watch Database, which has been freely available since Crossref acquired it in 2023. Zotero checks your library against it automatically and flags retracted items.
- Glance at the journal. If you have never heard of it, run the ten-minute legitimacy check.
Step 5: Read it, and compare it with the claim
Find the passage that supports the statement you want to make. Ask:
- Does the paper actually report this finding, in this population, with this direction of effect?
- Is it the paper's own result, or something mentioned in its introduction while citing someone else? If the latter, find the original.
- Is the claim as strong as the AI made it? "Associated with" often becomes "causes", and "in this sample of 84 students" becomes "students".
A first pass of five to ten minutes using the three-pass method is usually enough to settle it. If the paper will carry weight in your argument, read it properly.
Worked examples
We do not invent a fake reference to demonstrate on. The failing examples below are real ones, documented by a United States federal court in Mata v. Avianca, Inc. (S.D.N.Y. June 22, 2023), where a lawyer filed a brief citing judicial opinions that ChatGPT had generated. The quotations come from the court's opinion and order on sanctions, which is linked under Sources. Legal citations use a reporter volume and page where a journal article uses a DOI, and the checks map across one to one. The passing example is an ordinary journal article checked in Crossref and OpenAlex.
1. A fabricated reference: "Varghese"
The brief cited this as an Eleventh Circuit decision:
What the court found when it checked:
- The locator leads somewhere else. The order states: "The Federal Reporter citation for “Varghese” is associated with J.D. v Azar, 925 F.3d 1291 (D.C. Cir. 2019)." This is the legal equivalent of a DOI that resolves to a different paper (step 1).
- The issuing body has no record of it. The Clerk of the Eleventh Circuit "confirmed that the decision is not an authentic ruling of the Court". The docket number printed on the opinion belonged to an unrelated case. This is the equivalent of browsing the journal's volume and issue and finding no such article (step 2).
- The names were partly real. Two of the three judges named on the fake opinion do sit on the Eleventh Circuit. The third is a judge of the Fifth Circuit. Real names on a reference prove nothing about the reference.
- Asking the chatbot did not help. The lawyer asked ChatGPT "Is Varghese a real case". The order records that "ChatGPT responded that it had supplied 'real' authorities that could be found through Westlaw, LexisNexis and the Federal Reporter." A model cannot verify its own output. Check against an independent record.
The order goes on to record that the lawyers themselves eventually accepted the position: six of the decisions they had cited, "'Varghese', 'Miller', 'Petersen', 'Shaboon', 'Martinez' and 'Durden'", "were generated by ChatGPT and do not exist."
2. A chimera: a real name with an invented locator
The fake "Varghese" opinion itself cited:
The court's finding: this "does not exist as cited." A Supreme Court decision with the same name is real: Zicherman v. Korean Air Lines Co., 516 U.S. 217 (1996). It concerns a different question, and the reporter citation given belongs to another case altogether (Miccosukee Tribe v. United States, 516 F.3d 1235).
This is the pattern that catches careful people. A quick search for the name finds something real, and that seems to confirm the reference. It does not. Step 3 exists for this case: open the record and compare every element, including court or journal, year, volume and page. If they do not match, you have not found the cited work.
3. Real and correctly cited, but not saying what was claimed
The same order lists decisions cited in "Varghese" that, in the court's words, "have correct names and citations but do not contain the language quoted or support the propositions for which they are offered." One of them, In re Rimstat, Ltd., 212 F.3d 1039 (7th Cir. 2000), was offered on the bankruptcy stay. The court notes that it is a decision about sanctions for attorney misconduct and does not discuss the stay at all.
Every mechanical check passes here. Only step 5, reading the source, catches it.
The scholarly version of the same error is easy to make with a genuine paper. Take the sentence "ChatGPT fabricates more than half of its citations (Walters and Wilder, 2023)." The paper is real and the reference is accurate. But the 55% figure applies to GPT-3.5 only. For GPT-4 the same paper reports 18%. Citing it for a general claim about ChatGPT misrepresents it, and only reading the abstract and results reveals that.
4. What a pass looks like
Take the Walters and Wilder reference from the Sources list below and run the routine on it.
- Resolve the DOI.
https://doi.org/10.1038/s41598-023-41032-5opens an article page at nature.com with the same title and the same two authors. - Search the exact title. A Crossref search for "Fabrication and errors in the bibliographic citations generated by ChatGPT" returns that DOI as the first result. OpenAlex has the same work as record W4386510404.
- Confirm the details. The Crossref record gives Scientific Reports, volume 13, article number 14045, published 7 September 2023. The journal uses article numbers, so there is no page range to check.
- Check the status. OpenAlex lists the work as not retracted, and the article page carries no correction or retraction notice (checked 21 September 2026).
- Read it. The abstract reports the figures quoted in this article: 636 references in 84 generated papers, 55% fabricated for GPT-3.5 and 18% for GPT-4.
Walters and Wilder also describe what garbled references usually look like. Among the real works ChatGPT cited, the most common errors were wrong volume, issue or page numbers and wrong years. Errors in titles and author names were less common and mostly minor. So when a reference checks out in steps 1 and 2, look hardest at the numbers in step 3.
Checking a whole list quickly
- Paste the reference list into Crossref Simple Text Query (apps.crossref.org/SimpleTextQuery). It returns matched DOIs. Unmatched items form your shortlist.
- Import the matched DOIs into Zotero (the "Add item by identifier" button accepts several at once). You now have registry metadata and retraction flags.
- Work through the shortlist by hand with steps 2 and 3. Expect books, chapters and reports here as well as fabrications.
- Step 5 cannot be batched. Schedule the reading.
Prevention
- Do not ask a general chatbot for references. Ask it for search terms, then search a bibliographic database yourself.
- Prefer tools that retrieve before they generate. A tool that searches an index of real papers and links every statement to a record cannot invent the record. It can still misread it, so step 5 stays.
- Never paste a reference you have not opened. This is the same rule that applies to copying citations from other people's reference lists.
- Keep a log of which sources came from which tool. It helps with the disclosure many institutions require. See what university policies allow.
- Cite the source, not the tool. How to handle this in each style is covered in how to cite sources you found through AI tools.
Sources
- Crossref. Simple Text Query. https://apps.crossref.org/SimpleTextQuery
- Crossref. (2023, September 12). News: Crossref and Retraction Watch. https://www.crossref.org/blog/news-crossref-and-retraction-watch/
- Mata v. Avianca, Inc., No. 22-cv-1461 (PKC) (S.D.N.Y. June 22, 2023) (opinion and order on sanctions), ECF No. 54. https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts.nysd.575368.54.0_3.pdf
- National Library of Medicine. PubMed Single Citation Matcher. https://pubmed.ncbi.nlm.nih.gov/citmatch/
- Rao, D., Wong, E., & Callison-Burch, C. (2026). Detecting and correcting reference hallucinations in commercial LLMs and deep research agents. arXiv:2604.03173. https://arxiv.org/abs/2604.03173
- Retraction Watch Database. http://retractiondatabase.org/
- University of Toronto, School of Graduate Studies. Guidance on the appropriate use of generative artificial intelligence for graduate academic milestones. https://www.sgs.utoronto.ca/about/guidance-on-the-use-of-generative-artificial-intelligence/
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. https://doi.org/10.1038/s41598-023-41032-5
- Zotero. Adding items to Zotero: Add item by identifier. https://www.zotero.org/support/adding_items_to_zotero
- Zotero. Retracted item notifications. https://www.zotero.org/blog/retracted-item-notifications/





