Explainable AI tries to make the behaviour of machine-learning models understandable to people. Start with the surveys: Adadi and Berrada (2018) and Barredo Arrieta et al. (2020) define the vocabulary and taxonomy. Doshi-Velez and Kim (2017) ask what it would mean to evaluate interpretability rigorously. Then read the critics: Ghassemi, Oakden-Rayner and Beam (2021) on the false hope of current approaches in health care, and Rudin et al. (2022) on the principles of interpretable machine learning.
How the literature is organised
DARPA's Explainable Artificial Intelligence programme, which aims at AI systems whose models and decisions end users can understand and appropriately trust, is described by Gunning and Aha (2019); Gunning et al. (2019) is a short overview in Science Robotics. The most-cited works are surveys that organise a fragmented area. Adadi and Berrada (2018) surveyed the field in IEEE Access. Barredo Arrieta et al. (2020) provided concepts, taxonomies, opportunities and challenges under the heading of responsible AI. Linardatos, Papastefanopoulos and Kotsiantis (2020) review interpretability methods with a taxonomy and links to implementations.
Methods papers are fewer in this list because their titles seldom include the field's name. Lundberg et al. (2020) show how to go from local explanations to global understanding for tree models, and Lundberg et al. (2018) applied explanations to predicting hypoxaemia during surgery.
Foundational critique comes from Doshi-Velez and Kim (2017), a position paper that defines interpretability and asks for a rigorous science of it, and from Murdoch et al. (2019).
Main debates
The sharpest debate is whether current explanation methods deliver what is claimed for them. Ghassemi, Oakden-Rayner and Beam (2021) argue that the expectation that explainable AI will engender trust, provide transparency and mitigate bias in health care is a false hope with current methods. Rudin et al. (2022) hold that interpretability is crucial for high-stakes decisions and set out principles and ten technical challenges for interpretable machine learning. A second debate is evaluation: Doshi-Velez and Kim (2017) found little consensus on how interpretability should be measured.
Where recent work is heading
Recent reviews are organised by domain. Medical image analysis (van der Velden et al., 2022), education (Khosravi et al., 2022) and industry each have dedicated reviews. Shin (2021) studies how explainability and causability affect trust and acceptance. Ali et al. (2023) and Dwivedi et al. (2023) give up-to-date overviews, and Saeed and Omlin (2023) provide a meta-survey of challenges.
Most-cited foundational papers
Published before 2021 and ranked by how often later work cites them. Read the abstract of each and the full text of the three or four closest to your question. Citation count measures attention, not quality, so treat this as a map of what the field has argued about rather than a ranking of what is true.
Arun Rai (2020). Journal of the Academy of Marketing Science.
Cited by 1,278doi:10.1007/s11747-019-00710-5
Most-cited papers since 2021
Primary studies and conceptual papers from 2021 onwards. A paper published in 2024 has had a few years to accumulate citations where the works in the section above have had decades, so compare these counts with each other rather than with the ones above.
Hassan Khosravi and 9 others (2022). Computers and Education: Artificial Intelligence.
Cited by 808Open accessdoi:10.1016/j.caeai.2022.100074
Recent reviews and meta-analyses
The fastest way into a literature. A good review gives you the structure of the field, a reference list to mine and, in its limitations section, the gaps other researchers have already spotted.
Waddah Saeed, Christian Omlin (2023). Knowledge-Based Systems.
Cited by 755Open accessdoi:10.1016/j.knosys.2023.110273
How big the literature is, and where it is published
OpenAlex indexes 36,152 works whose title matches this topic. The chart shows how many were published each year from 2000 to 2025; the current year is left out because it is incomplete.
Show the numbers as a table
Year
Works
2000
1
2002
2
2003
4
2004
2
2005
3
2006
2
2007
3
2008
2
2009
6
2010
3
2011
4
2012
9
2013
11
2014
12
2015
8
2016
14
2017
27
2018
99
2019
258
2020
649
2021
1,148
2022
1,776
2023
2,905
2024
4,894
2025
9,796
Journals behind the most-cited work
Counted across the 173 most-cited works on the topic, not across everything published. Browsing recent issues of the first two or three is a reliable way to find current work that has not yet been cited much.
IEEE Access14 papers
Information Fusion10 papers
Artificial Intelligence4 papers
Applied Sciences3 papers
Artificial Intelligence Review3 papers
Computers in Biology and Medicine3 papers
Sensors3 papers
The Science of The Total Environment3 papers
Search strings to copy
Written for databases that accept Boolean operators (Scopus, Web of Science, ERIC, PubMed, EBSCO). Quotation marks keep a phrase together, an asterisk stands in for the end of a word, and OR groups go in brackets. British and American spellings are written out with OR rather than covered by a single-character wildcard, because those wildcards differ between databases: PubMed’s help page documents the asterisk only, and asks for at least four characters before it. Limit the search to title and abstract first; widen it only if you get too little. Our guide to starting a literature review covers how to record what you searched.
Methods
("explainable AI" OR XAI OR "explainable artificial intelligence" OR "interpretable machine learning") AND (SHAP OR LIME OR "feature attribution" OR saliency OR counterfactual*)
Human factors and trust
("explainable AI" OR XAI OR explainability) AND (trust OR "user study" OR "human-centered" OR stakeholder* OR understanding)
Health care
("explainable AI" OR XAI OR explainab* OR interpretab*) AND ("medical imag*" OR clinical OR healthcare OR diagnos*)
Explore explainable AI in Wonders →Opens Wonders on this question. Once you are signed in it becomes a project with suggested sub-topics and keywords you can edit before searching. The trial is 14 days.
Sub-topics to narrow into
A thesis-sized question usually sits inside one of these, combined with a population or a setting.
Taxonomies and surveys
Adadi and Berrada (2018), Barredo Arrieta et al. (2020) and Linardatos, Papastefanopoulos and Kotsiantis (2020).
Feature attribution with SHAP
Lundberg et al. (2020) extend SHAP to tree models; Salih et al. (2025) compare SHAP and LIME.
Interpretable models by design
Rudin et al. (2022) set out principles and ten grand challenges.
Human-centred evaluation
Shin (2021) on explainability, causability, trust and acceptance; Murdoch et al. (2019) stress the role of the human audience.
Tjoa and Guan (2021), van der Velden et al. (2022) and the critique by Ghassemi, Oakden-Rayner and Beam (2021).
How to cite these papers
Every paper above has a DOI, a part of the reference that is easy to leave out. Here is one of them, “Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI)” (2018), in the two styles students ask about most:
APA 7th edition
Adadi, A., & Berrada, M. (2018). Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access, 6, 52138–52160. https://doi.org/10.1109/access.2018.2870052
MLA 9th edition
Adadi, Amina, and Mohammed Berrada. “Peeking inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI).” IEEE Access, vol. 6, 2018, pp. 52138–60, https://doi.org/10.1109/access.2018.2870052.
Check the details against the article itself before you submit: databases, including the one behind this page, sometimes carry the online-first year rather than the volume year. Full rules and more examples are in our guides to APA, MLA, Chicago, Harvard, Vancouver and ABNT, with the rest in the citation guides. You can also format a reference from its DOI with our free citation tools.
Frequently asked questions
What is the difference between interpretability and explainability?
Usage varies. Doshi-Velez and Kim (2017) observed that there was very little consensus on what interpretable machine learning is or how it should be measured, and Murdoch et al. (2019) wrote that the surge of research had led to confusion about what it means to be interpretable. Both papers propose definitions; Murdoch et al. add a framework for selecting and evaluating interpretation methods. Pick one source, quote its definition and use the terms consistently.
Where are the LIME and SHAP papers?
The titles of the original LIME and SHAP papers do not contain the search terms used here, so they are not in this list. Lundberg et al. (2020) in Nature Machine Intelligence, which is listed, covers explanations for tree-based models, and Salih et al. (2025) discuss both methods. Follow their reference lists to the originals and cite those directly.
Is explainable AI required by law?
That is a legal question that the papers on this page cannot settle, and the answer varies by jurisdiction and use. Treat legal claims in computer science papers as pointers and check them against the legislation and legal scholarship for your jurisdiction.
How this page was made
The lists come from OpenAlex, an open index of scholarly works whose data are published under a CC0 licence, queried on September 21, 2026 for works whose title matches ("explainable artificial intelligence" OR "explainable ai" OR xai OR "interpretable machine learning" OR "explainable machine learning"). Only works with a DOI are listed. Each one was checked against the publisher’s own record at Crossref or DataCite (title, year, first author, journal, volume and pages), and in three cases, where the publisher deposited no byline, against PubMed; anything OpenAlex or Crossref flags as retracted was left out, and an editor took out results that matched the words but not the subject. Citation counts are OpenAlex’s on that date and are usually lower than Google Scholar’s, which counts more kinds of document. Ranking by citations tells you what a field has relied on, not what is correct; several heavily cited papers on any topic are cited because later work disputes them. Books without a DOI are missing, which matters in fields where the founding text is a book.