Wonders

The Blog · Musings on Academia

How to Check Whether an AI-Generated Citation Is Real

Chatbots invent references that look authentic. A five-step verification routine using DOI lookup, Crossref, Google Scholar, library catalogues and the Retraction Watch database, with worked examples taken from the fabricated citations documented in Mata v. Avianca and from a published study of ChatGPT references.

Sep 21, 2026
How to Check Whether an AI-Generated Citation Is Real

TL;DR

Language models produce references by predicting plausible text, so a citation can have real authors, a real journal and a convincing title and still not exist. Check every AI-supplied reference in five steps: resolve the DOI, search the exact title in Crossref or Google Scholar, confirm the details on the publisher page, check for retraction, and then read the paper to confirm it supports the claim. The last step catches the most common problem with newer tools, which is a real paper cited for something it does not say. If you cannot find a reference in two or three minutes of searching, treat it as fabricated.

A fabricated reference is hard to spot by eye. The authors are often real researchers in the right field. The journal exists. The title sounds like something they would write. The volume, pages and DOI have the correct format. Only the paper is missing.

That is a predictable result of how language models work, and it is why the University of Toronto's guidance for graduate students warns that "Generative AI tools have also been shown to reference scholarly works that do not exist." The responsibility for catching them lies with you.

How big the problem is

Chatbots without search. Walters and Wilder (2023) asked ChatGPT to write short literature reviews on 42 topics and checked all 636 references. With GPT-3.5, 55% were fabricated. With GPT-4, 18% were. Among the references that were real, 43% (GPT-3.5) and 24% (GPT-4) had substantive errors such as wrong authors, dates, volumes or page numbers.

Search-enabled models and research agents. Rao, Wong and Callison-Burch (2026) examined more than 221,000 citation URLs produced by commercial models and research agents, ten of them on one benchmark and three on another. Between 3% and 13% of URLs were hallucinated, and 5% to 18% did not resolve. Deep research agents produced more citations per answer and had higher hallucination rates than search-augmented chat models.

The trend is improving, and the residual rate is still far too high to skip checking.

Four kinds of bad citation

TypeWhat it looks likeHow it is caught
Fully fabricatedNothing matches: no such title, and the DOI fails or leads elsewhereSteps 1 and 2
ChimeraReal authors and a real journal with an invented title, or a real title attached to the wrong authors and yearSteps 2 and 3
Real but garbledThe paper exists but the year, volume, pages, DOI or spelling of an author's name is wrongStep 3
Real but misusedThe paper exists and the details are right, but it does not say what the AI claimsStep 5 only

The fourth type is the hardest, because every mechanical check passes.

The five-step check

Step 1: Resolve the DOI

If the reference has a DOI, paste it after https://doi.org/ in your browser.

No DOI is not suspicious in itself. Books, many reports, and older or smaller journals do not have them.

Step 2: Search the exact title

Put the full title in quotation marks and search:

If an exact-title search finds nothing, try the first author's name plus two distinctive title words. A real paper may have been slightly misquoted. If the authors are real, look at their publication list on Google Scholar, ORCID or their university page. A paper that is absent from its supposed first author's own publication list is very unlikely to exist.

Time limit: if two or three minutes across these sources turn up nothing, treat the reference as fabricated and remove it. Do not go looking for a similar real paper to swap in unless you are going to read that paper.

Step 3: Confirm the details at the source

Open the publisher's page or the PDF and check authors and their order, year, journal title, volume, issue, pages or article number, and DOI. Correct your reference from the source, not from the chatbot's version. Better still, import the record into your reference manager by DOI so that the metadata comes from the registry.

Step 4: Check the paper's status

Step 5: Read it, and compare it with the claim

Find the passage that supports the statement you want to make. Ask:

A first pass of five to ten minutes using the three-pass method is usually enough to settle it. If the paper will carry weight in your argument, read it properly.

Worked examples

We do not invent a fake reference to demonstrate on. The failing examples below are real ones, documented by a United States federal court in Mata v. Avianca, Inc. (S.D.N.Y. June 22, 2023), where a lawyer filed a brief citing judicial opinions that ChatGPT had generated. The quotations come from the court's opinion and order on sanctions, which is linked under Sources. Legal citations use a reporter volume and page where a journal article uses a DOI, and the checks map across one to one. The passing example is an ordinary journal article checked in Crossref and OpenAlex.

1. A fabricated reference: "Varghese"

The brief cited this as an Eleventh Circuit decision:

Varghese v. China Southern Airlines Co., Ltd., 925 F.3d 1339 (11th Cir. 2019)

What the court found when it checked:

  1. The locator leads somewhere else. The order states: "The Federal Reporter citation for “Varghese” is associated with J.D. v Azar, 925 F.3d 1291 (D.C. Cir. 2019)." This is the legal equivalent of a DOI that resolves to a different paper (step 1).
  2. The issuing body has no record of it. The Clerk of the Eleventh Circuit "confirmed that the decision is not an authentic ruling of the Court". The docket number printed on the opinion belonged to an unrelated case. This is the equivalent of browsing the journal's volume and issue and finding no such article (step 2).
  3. The names were partly real. Two of the three judges named on the fake opinion do sit on the Eleventh Circuit. The third is a judge of the Fifth Circuit. Real names on a reference prove nothing about the reference.
  4. Asking the chatbot did not help. The lawyer asked ChatGPT "Is Varghese a real case". The order records that "ChatGPT responded that it had supplied 'real' authorities that could be found through Westlaw, LexisNexis and the Federal Reporter." A model cannot verify its own output. Check against an independent record.

The order goes on to record that the lawyers themselves eventually accepted the position: six of the decisions they had cited, "'Varghese', 'Miller', 'Petersen', 'Shaboon', 'Martinez' and 'Durden'", "were generated by ChatGPT and do not exist."

2. A chimera: a real name with an invented locator

The fake "Varghese" opinion itself cited:

Zicherman v. Korean Air Lines Co., 516 F.3d 1237, 1254 (11th Cir. 2008)

The court's finding: this "does not exist as cited." A Supreme Court decision with the same name is real: Zicherman v. Korean Air Lines Co., 516 U.S. 217 (1996). It concerns a different question, and the reporter citation given belongs to another case altogether (Miccosukee Tribe v. United States, 516 F.3d 1235).

This is the pattern that catches careful people. A quick search for the name finds something real, and that seems to confirm the reference. It does not. Step 3 exists for this case: open the record and compare every element, including court or journal, year, volume and page. If they do not match, you have not found the cited work.

3. Real and correctly cited, but not saying what was claimed

The same order lists decisions cited in "Varghese" that, in the court's words, "have correct names and citations but do not contain the language quoted or support the propositions for which they are offered." One of them, In re Rimstat, Ltd., 212 F.3d 1039 (7th Cir. 2000), was offered on the bankruptcy stay. The court notes that it is a decision about sanctions for attorney misconduct and does not discuss the stay at all.

Every mechanical check passes here. Only step 5, reading the source, catches it.

The scholarly version of the same error is easy to make with a genuine paper. Take the sentence "ChatGPT fabricates more than half of its citations (Walters and Wilder, 2023)." The paper is real and the reference is accurate. But the 55% figure applies to GPT-3.5 only. For GPT-4 the same paper reports 18%. Citing it for a general claim about ChatGPT misrepresents it, and only reading the abstract and results reveals that.

4. What a pass looks like

Take the Walters and Wilder reference from the Sources list below and run the routine on it.

  1. Resolve the DOI. https://doi.org/10.1038/s41598-023-41032-5 opens an article page at nature.com with the same title and the same two authors.
  2. Search the exact title. A Crossref search for "Fabrication and errors in the bibliographic citations generated by ChatGPT" returns that DOI as the first result. OpenAlex has the same work as record W4386510404.
  3. Confirm the details. The Crossref record gives Scientific Reports, volume 13, article number 14045, published 7 September 2023. The journal uses article numbers, so there is no page range to check.
  4. Check the status. OpenAlex lists the work as not retracted, and the article page carries no correction or retraction notice (checked 21 September 2026).
  5. Read it. The abstract reports the figures quoted in this article: 636 references in 84 generated papers, 55% fabricated for GPT-3.5 and 18% for GPT-4.

Walters and Wilder also describe what garbled references usually look like. Among the real works ChatGPT cited, the most common errors were wrong volume, issue or page numbers and wrong years. Errors in titles and author names were less common and mostly minor. So when a reference checks out in steps 1 and 2, look hardest at the numbers in step 3.

Checking a whole list quickly

  1. Paste the reference list into Crossref Simple Text Query (apps.crossref.org/SimpleTextQuery). It returns matched DOIs. Unmatched items form your shortlist.
  2. Import the matched DOIs into Zotero (the "Add item by identifier" button accepts several at once). You now have registry metadata and retraction flags.
  3. Work through the shortlist by hand with steps 2 and 3. Expect books, chapters and reports here as well as fabrications.
  4. Step 5 cannot be batched. Schedule the reading.

Prevention

Wonders answers from real scholarly records: it searches OpenAlex and Semantic Scholar, more than 320 million works, so the references it shows are records retrieved from those indexes rather than text a model composed. Every paper links to its source, and a quote in a summary is presented as a quote only after it has been found in the paper’s text; where it is not found, it is demoted. Steps 1 to 4 are still worth running on anything you did not retrieve yourself, and step 5, reading the paper to confirm it supports your point, is always your job.

Sources

Frequently asked questions

Why do chatbots make up references?

A language model generates text that is statistically likely given its training data. It has learned what references look like, including author name patterns, journal titles and DOI formats, and can assemble those elements into a citation without retrieving any record. Walters and Wilder describe ChatGPT as a language-processing tool, not an information-processing tool. Tools that search a bibliographic database first and then answer from the results fabricate far less, but they can still attach a real paper to a claim it does not support.

Are newer AI models still hallucinating citations?

Less often, but yes. Walters and Wilder found fabricated citations fell from 55% with GPT-3.5 to 18% with GPT-4 in 2023. A 2026 study by Rao, Wong and Callison-Burch of commercial search-enabled models and deep research agents found that 3 to 13% of the citation URLs they produced were hallucinated and 5 to 18% did not resolve. At a 3% rate, a list of 40 citation links is more likely than not to contain at least one bad link, so checking remains necessary.

What is the fastest way to check a long reference list?

Paste the whole list into Crossref's Simple Text Query, which returns a DOI for each reference it can match. Anything unmatched goes on a shortlist for manual checking in Google Scholar and library catalogues, since books, reports and older articles often have no DOI. Importing by DOI into Zotero is also efficient: the metadata comes from the registry, not from the chatbot, and Zotero flags retracted items.

The paper exists. Do I still need to read it?

Yes. Existence is the lowest bar. The more common failure now is a genuine paper cited for a finding it does not contain, or for a stronger claim than the authors make. You are responsible for every citation in your work, and the only way to know what a source says is to read the relevant parts of it.

What happens if a fabricated reference ends up in my submitted work?

You are responsible for every source you cite. The University of Toronto tells graduate students that they are "ultimately responsible for the content" when they include AI output in their work, and how a fabricated reference is handled depends on your institution's academic integrity policy, so read it. In professional settings the consequences have included court sanctions: in Mata v. Avianca (2023) the court imposed a 5,000 US dollar penalty, jointly, on two lawyers and their firm after a brief cited non-existent cases generated by ChatGPT. If you find a fabricated reference after submission, tell your instructor or editor promptly.

More from the blog

Using AI in a Literature Review: What University and Publisher Policies Actually Allow

Using AI in a Literature Review: What University and Publisher Policies Actually Allow

What Cambridge, Stanford, Toronto, Sydney, the Russell Group, Elsevier and the ICMJE say about generative AI, translated into a task-by-task guide for literature reviews, with disclosure statement templates and a record-keeping routine.

Elicit Student Discount & Other AI Research Tool Discounts

✦ WONDERSThe Rabbit Hole
Research guides
readwonders.com

Elicit Student Discount & Other AI Research Tool Discounts

Elicit has academic pricing but no student code. Consensus: up to 40% off. Paperguide: 30%. Wonders: 50%. Each price checked on the tool's own site, 21 Sept 2026.

Research Tools for Graduate Students: A 2026 Toolkit

Research Tools for Graduate Students: A 2026 Toolkit

A toolkit for MA and PhD students, by job: reference managers, literature search, writing, notes and data analysis. Prices shown only where we checked them, Sept 2026.

Best Literature Review Tools in 2026 (AI & Traditional)

Best Literature Review Tools in 2026 (AI & Traditional)

Literature review tools compared by job: finding papers, managing references, reading and writing. Free tiers and prices checked on each tool's site, Sept 2026.

Best Elicit Alternatives in 2026: 6 Tools Compared

Best Elicit Alternatives in 2026: 6 Tools Compared

Elicit alternatives by use case: Wonders (guided workspace), Scite (citation checks), Consensus (quick answers), SciSpace (PDFs), Semantic Scholar (free).

National University Streamlines Literature Review with Wonders

National University Streamlines Literature Review with Wonders

How National University put Wonders in front of its doctoral cohort — speeding up topic discovery, raising citation quality, and easing supervision load, with students reporting nine hours saved each week.

Put it into practice.

Try these techniques in Wonders — an AI workspace for literature review. 14 days free. Students get 50% off.

Start free →