How the literature is organised
Everything in this literature traces back to Albert Bandura. Start with the 1977 Psychological Review paper for the original statement. Bandura (1982) describes self-percepts of efficacy influencing thought patterns, actions and emotional arousal, and reports that in causal tests higher induced self-efficacy produced higher performance and lower emotional arousal. Bandura (1993) sets out the four processes through which perceived self-efficacy works — cognitive, motivational, affective and selection — and applies them at three levels in education: students' beliefs about regulating their own learning, teachers' beliefs about motivating students, and a faculty's collective sense of instructional efficacy. Bandura (1986) and the 1995 edited volume extend the scope further.
Because self-efficacy is defined against a task, the research that followed is organised by domain. In education, Pajares (1996) reviewed the construct in academic settings and showed that particularised measures matching the criterial task outperform global ones in explaining outcomes; Zimmerman (2000) treats it as an essential motive to learn. In information systems, Compeau and Higgins (1995) surveyed Canadian managers and professionals to build and validate a computer self-efficacy measure, and found it influencing outcome expectations, affect and anxiety, and actual computer use. In health, Tan et al. (2021) review self-efficacy and self-care in hypertension and Wong, Mou and Chien (2021) review interventions for breastfeeding self-efficacy.
Main debates
Three arguments recur. First, measurement. Sherer et al. (1982) built a deliberately general scale, with General and Social Self-efficacy subscales, on the argument that past experiences produce differing levels of generalised expectancy; Pajares (1996) found that task-matched measures surpass global ones in explanation and prediction. Both positions are defensible, so say which you have taken.
Second, conceptual overlap. Ajzen (2002) worked through the relationship between self-efficacy, controllability and perceived behavioural control, concluding that perceived behavioural control is a superordinate construct comprising the first two, and that measures of it should contain items for both. If your framework is the theory of planned behaviour, this is the paper that settles what to measure.
Third, direction. Sherer et al. (1982) treat past experience and the attribution of success as sources of self-efficacy, which is the difficulty: performance both predicts and is predicted by it. The reviews are frank about what that does to the evidence. Fang et al. (2021) note that most of the 89 factors they identified for parenting self-efficacy come from one or two cross-sectional studies, and Tan et al. (2021) rate the hypertension evidence low to medium quality, limited by heterogeneity and self-reported behaviour.
Where recent work is heading
The most-cited papers since 2021 apply the construct to new technology. Yilmaz and Karaoglan Yilmaz (2023) randomised 45 undergraduates on a programming course to use ChatGPT or not, and report higher computational thinking, programming self-efficacy and motivation in the experimental group. Zhang et al. (2024) turn the question round, finding with 300 university students that academic self-efficacy relates to AI dependency through academic stress and performance expectations, and listing laziness, misinformation and reduced critical thinking among the consequences students report. Wang and Chuang (2024) develop and validate an artificial intelligence self-efficacy scale, and Wang et al. (2022) study online learning self-efficacy as a mediator of engagement.
Recent reviews are domain-specific: parenting (Fang et al., 2021), teacher self-efficacy for inclusive education (Wray et al., 2022; Yada et al., 2022), communication skills training for health professionals (Mata et al., 2021), hypertension self-care (Tan et al., 2021) and breastfeeding (Wong, Mou and Chien, 2021). If you are writing a thesis, that is the pattern to follow: a known construct, a validated domain scale, and a population or setting that has not been studied with it.