Reading takes up much of the time in a literature review, and, as Keshav observes, it is a skill that is rarely taught. New researchers tend to open a PDF, start at the first word of the introduction and push through to the references, taking the same care over a marginal paper as over a central one. An hour later they have read one paper and do not know whether it mattered.
S. Keshav, a computer scientist then at the University of Waterloo, described a better routine in a two-page article, "How to Read a Paper" (2007). The idea is simple: read in up to three passes, each with a specific goal, and stop as soon as the paper has given you what you need.
The method at a glance
| First pass | Second pass | Third pass | |
|---|---|---|---|
| Goal | Bird's-eye view. Decide whether to continue | Grasp the content, without every detail | Understand in depth. Find hidden assumptions and flaws |
| Time (Keshav's estimates) | 5–10 minutes | Up to 1 hour | 4–5 hours for beginners, about 1 hour when experienced |
| Read | Title, abstract, introduction, headings, conclusions. Glance at references | Everything except proofs and fine detail. Figures, tables and graphs carefully | Everything, while mentally re-creating the work |
| Output | Answers to the five Cs. A keep, maybe or discard decision | A summary you could give to someone else, with the supporting evidence. A list of references to follow up | Ability to reconstruct the paper's structure from memory. Its strengths and weaknesses. Ideas for future work |
| Use for | Every paper you consider | Papers relevant to your topic | The handful central to your study, and papers you review |
First pass: five to ten minutes
Steps, as Keshav gives them:
- Read the title, abstract and introduction carefully.
- Read the section and subsection headings and ignore everything else.
- Read the conclusions.
- Glance over the references and mentally tick off the ones you have already read.
Then answer the five Cs:
- Category. What type of paper is it? An experiment, a survey, a qualitative study, a review, a theoretical argument, a description of a system?
- Context. Which other papers is it related to? Which theory does it use?
- Correctness. Do the assumptions appear valid?
- Contributions. What are the main contributions?
- Clarity. Is it well written?
With those answers you decide whether to read further. Keshav lists three reasons to stop: the paper does not interest you, you do not know enough about the area to follow it, or the authors make invalid assumptions. He adds that the first pass is adequate for papers that are not in your research area but may someday prove relevant.
Adapting the first pass for empirical papers
Keshav wrote for computer science, where papers often describe systems. For empirical work in health, social science or education, add three quick checks:
- Design and sample. Find the sentence in the abstract or methods that states the design and the number of participants. A single-site cross-sectional survey and a multi-site randomized trial can bear very different amounts of weight.
- Main result with a number. Locate one effect size, percentage or central theme. If the abstract has only adjectives such as "significant improvement", look at the first table.
- Limitations paragraph. It is often near the end of the discussion. Authors tell you there what they could not rule out.
A practical tip: if you have found a recent review article on your topic, read that first with a full second pass. Keshav's advice for literature surveys is blunt: "If you can find such a survey, you are done." Read it and you have a map for everything else.
Second pass: up to an hour
Read the paper with greater care but ignore details such as proofs. Keshav's two instructions:
- Look carefully at the figures, diagrams and other illustrations. Are the axes labelled? Are results shown with error bars, so that conclusions are statistically meaningful? Mistakes here separate rushed work from excellent work.
- Mark relevant unread references for further reading.
Take notes as you go: key points, and comments in the margin. At the end you should be able to summarize the main thrust of the paper, with supporting evidence, to someone else. If you cannot, Keshav's options are to set the paper aside, to return to it after reading background material, or to persevere into a third pass.
For empirical papers, a useful reading order on the second pass is: abstract, figures and tables, results, methods, then introduction and discussion. The discussion is where authors interpret, and it is easier to judge their interpretation when you have already looked at the data yourself. For clinical papers, Trisha Greenhalgh's BMJ series "How to read a paper" (1997) is the classic companion.
A note template for the second pass
Fill this in immediately after reading, in your own words, in your reference manager or notes app. It takes five minutes and saves a reread later. The fields map directly onto the columns of a synthesis matrix.
Citation:
Read on (date) / pass reached (1, 2, 3):
Question or aim:
Category (type of paper / design):
Sample, setting, data:
Theory or framework used:
Main findings (with numbers and page refs):
1.
2.
Authors' stated limitations:
My concerns (things they did not mention):
How it relates to other papers I have read (agrees / contradicts / extends):
Relevance to my project (which section or theme):
Quotes worth keeping (exact words, in quotation marks, with page):
References to follow up:
Third pass: virtually re-implement the paper
The third pass is for papers you need to understand fully: the two or three your study builds on, a paper whose method you will reuse, or one you are peer reviewing. Keshav's key instruction is to virtually re-implement the paper: making the same assumptions as the authors, re-create the work, then compare your re-creation with the actual paper. The differences expose the paper's innovations and its hidden failings and assumptions.
In a non-computational field, "re-implementing" means asking at each step what you would have done. Given this research question, which design would I choose? Whom would I sample? Which measure would I use, and what does theirs miss? Which analysis follows from that design? What else could explain this result? Where the authors did something different from what you would have done, either they know something you do not or you have found a weakness. Both are worth writing down.
During this pass you identify and challenge every assumption, and note ideas for future work. At the end you should be able to reconstruct the entire structure of the paper from memory and identify its strong and weak points, including implicit assumptions, missing citations and problems with the experimental or analytical technique.
A worked first pass
The paper: Walters and Wilder (2023), "Fabrication and errors in the bibliographic citations generated by ChatGPT", Scientific Reports, 13, 14045. It is open access, so you can follow along.
| Minute | What I read | What I noted |
|---|---|---|
| 0–2 | Title and abstract | They had ChatGPT (GPT-3.5 and GPT-4) write short literature reviews on 42 topics and checked the 636 citations produced. 55% of GPT-3.5 citations and 18% of GPT-4 citations were fabricated. Among the real ones, 43% (GPT-3.5) and 24% (GPT-4) contained substantive errors |
| 2–5 | Introduction | The authors frame ChatGPT as a language-processing tool, not an information-processing tool: it mimics texts, not necessarily their content. They study one kind of hallucination, fabricated citations, and give reasons it matters for scientific integrity and for teaching students the limits of the software. A "Previous research" subsection follows |
| 5–6 | Headings | Methods (paper topics and prompts, data compilation and analysis). Results (extent of fabrication, substantive citation errors, formatting errors and hyperlinks). Discussion (why fabricated citations persist, implications). Standard empirical structure |
| 6–8 | End of the discussion | GPT-4 is a major improvement over GPT-3.5, but problems remain. Verification is still required |
| 8–9 | References | 45 entries. A quick scan for earlier studies of chatbot-generated references to follow up |
Five Cs. Category: empirical evaluation study. Context: the literature on large language model hallucination and on citation accuracy. Correctness: the design is simple and transparent, though the models tested are now old, so the percentages will have changed. Contributions: fabrication rates by model version with a clear method for classifying citations. Clarity: good.
Decision. For a thesis on AI in academic writing, this goes to a second pass. For a thesis on sleep and grades, note it and move on. The decision took nine minutes.
Using the method for a literature review
Keshav also describes how to use the passes to survey an unfamiliar field:
- Use an academic search engine and well-chosen keywords to find three to five recent papers. Do one pass on each, then read their related work sections. If they point to a recent survey, read the survey.
- Otherwise, find shared citations and repeated author names in the bibliographies. These are the key papers and researchers. Download the key papers, and check where the key researchers have published recently.
- Look through the recent proceedings or issues of those top venues for related high-quality work.
- Make two passes through the resulting set. If they all cite a key paper you did not find earlier, obtain and read it, iterating as necessary.
An illustrative funnel, using our numbers and not Keshav's: a few hundred search results screened by title and abstract, 60 to 100 first passes, 25 to 40 second passes, and three to five third passes. At Keshav's timings, 100 first passes take up to about 17 hours and 40 second passes take up to 40 hours. Giving all 100 papers a full second pass would take up to 100 hours.
Common mistakes
- Reading linearly and spending the same effort on every paper.
- Skipping the decision. The first pass ends with keep, maybe or discard. Record it, so you do not first-pass the same paper three times.
- Highlighting instead of noting. A PDF covered in yellow records what looked important, not what you understood. Write a sentence in your own words.
- Trusting the abstract for specifics. An abstract is the authors' own short summary and leaves out detail. Check the results section for any finding you plan to cite.
- Never doing a third pass. The papers your study depends on deserve it.
- Doing a third pass on everything, which is how literature reviews never get finished.
References
- Greenhalgh, T. (1997). How to read a paper: getting your bearings (deciding what the paper is about). BMJ, 315(7102), 243–246. https://doi.org/10.1136/bmj.315.7102.243
- Keshav, S. (2007). How to read a paper. ACM SIGCOMM Computer Communication Review, 37(3), 83–84. https://doi.org/10.1145/1273445.1273458. Copy on a Stanford course page: https://web.stanford.edu/class/ee384m/Handouts/HowtoReadPaper.pdf
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. https://doi.org/10.1038/s41598-023-41032-5

