Instructional design · Learning sciences · Review

The Bandwidth Constraint

Cognitive Load, Generative AI, and the Design Conditions for Project-Based Learning at Scale

Abstract

Kirschner, Sweller, and Clark (2006) argued that minimally guided instruction fails because it disregards the limits of working memory, and the exchange that followed relocated rather than refuted that argument: the defense of problem-based learning depends entirely on whether the scaffolding it claims is present and specified. This review accepts that burden and extends it to a condition the original exchange could not anticipate, in which instructional design is performed by a teacher working with a generative model. Six literatures are read together. The cognitive load literature specifies what guidance does, through the worked example effect, the completion effect, the expertise reversal effect, and the contingency and fading that define scaffolding. The experimental literature on generative artificial intelligence converges independently on the same variable: unconstrained conversational access to a model can improve assisted performance while degrading unassisted learning, whereas the same model under pedagogical constraint does not, making interface architecture rather than model capability the operative design decision. The transfer literature indicates that generalizable knowledge is built through context-embedded schema construction within disciplinary content rather than through generic skills instruction, which supplies a principled rather than rhetorical rationale for project-based designs and locates intrinsic motivation as a mechanism rather than a byproduct. The labor economics literature establishes that the capabilities such designs target are rising in market value, while offering no warrant for claims that any school intervention insulates students against labor market change. The professional learning literature indicates that teacher-facing support is weakly effective on average and effective only under identifiable conditions, that implementation quality exhibits a threshold below which intervention is counterproductive, and that teacher and collective efficacy beliefs are associated with instructional quality, job satisfaction, and retention, though the magnitudes most often cited for collective efficacy rest on a contested evidence base. From these the review derives design conditions for AI-mediated project-based learning, including a route around the diagnosis problem that has limited software scaffolding since the term entered instructional design, and specifies what remains unproven.

Keywords: cognitive load theory · project-based learning · instructional scaffolding · generative artificial intelligence · transfer of learning · collective teacher efficacy · teacher retention · culturally relevant education

Declaration of interest. The author is the founder of TeachCraft, an organization building tooling and professional community for project-based instructional design. Two sections warrant explicit weighting. Section 12 states a position attributable to that organization and is labeled as such. Section 7 derives design conditions that the author’s organization builds against, which means a reader is entitled to ask whether those conditions were derived from the evidence or reverse-engineered from an existing product. Each condition traces to work published before that product existed, and the citations permit the reader to check this, but the conflict is real and is better weighed than concealed. No efficacy data for any platform are presented and none are claimed.

Note on evidence tags. Empirical claims carry a tag describing the design of the underlying study, not the strength of the inference drawn from it. Every source has been checked against the journal or publisher record; where a figure could not be verified, it is described qualitatively. One prominent meta-analysis in this area was retracted in April 2026 and is discussed in Section 4 rather than cited as evidence. One widely circulated effect size is examined in Section 9 and found to rest on an unpublished source.

  • Meta-analysis
  • Randomized trial
  • Quasi-experimental
  • Correlational
  • Qualitative
  • Theoretical
  • Descriptive / survey
MechanismPrincipal evidenceDesign conditionStatus
Working memory limits on novel informationKirschner et al. (2006); Alfieri et al. (2011), d = −0.38 unassisted vs. +0.30 enhancedGuidance is a requirement, not a preferenceEstablished
Worked example and completion effectsSweller & Cooper (1985); van Merriënboer (1990)Supply a defensible draft to complete rather than a blank fieldEstablished for learners; untested for design tools
Production improves without comprehensionvan Merriënboer (1990); Fan et al. (2025)Treat artifact gain without understanding as the governing hazardConvergent across two literatures
Expertise reversal; contingency and fadingKalyuga et al. (2003); Wood et al. (1976); Belland et al. (2017), null on fadingSupport must be contingent on a diagnosisTheory strong; meta-analytic evidence mixed
Diagnosis absent from software scaffoldingPuntambekar & Hübscher (2005)Source the diagnosis from the practitionerUntested proposal
Interface determines the sign of the AI effectBastani et al. (2024, preprint); Kestin et al. (2025)Constrain generation; no open conversationOne preprint; replication needed
Transfer is domain-specificPellegrino & Hilton (2012); Barnett & Ceci (2002); Sala & Gobet (2017)Build competencies inside disciplinary content; specify goals before generatingEstablished
Split attention across fragmented sourcesChandler & Sweller (1992)Consolidate the design environmentTheory strong; untested at tool level
Professional development works only when it embedsSims et al. (2025), d = 0.05 overall, 0.14 balanced; Kraft et al. (2018)Training, coaching, and a recurring peer structure togetherComponents evidenced; combination untested
Implementation threshold effectBalfanz & Byrnes (2026)Support schools past the threshold or do not beginSingle program; pre-post design
Table 1. Mechanism, evidence, design condition, and evidentiary status. Effect sizes are reported within their own design context and are not comparable across rows.

1. Introduction

Project-based learning occupies an unusual position in education research. It is simultaneously one of the most widely advocated instructional approaches in K-12 practice and one of the most seriously challenged in the learning sciences. The challenge is not that projects are unpopular. It is that a substantial body of work in cognitive psychology holds that instruction organized around student-directed problem solving is, for novices, an inefficient way to build the long-term memory structures that constitute learning.

That challenge was stated forcefully two decades ago and has not been retired. What has changed is the surrounding condition. Instructional design, historically a slow and expensive act performed by curriculum publishers or by individual teachers working outside contract hours, can now be performed rapidly by a teacher in collaboration with a generative model. This alters the economics of project-based learning and introduces a new object into the analysis, which is the interface through which design occurs.

This review argues that the interface is not incidental. The emerging experimental literature on generative artificial intelligence in learning settings reports effects that reverse sign depending on how access to the model is structured. That finding, arrived at independently, is what cognitive load theory has maintained since the 1980s. Both literatures point at the same variable, which is guidance, and both treat it as a design decision rather than a property of the technology.

The review proceeds as follows. Sections 2 and 3 state the case against minimal guidance and specify what the cognitive load literature establishes about guidance. Section 4 reads the contemporary generative AI evidence against that specification. Sections 5 and 6 address why context-embedded project work has a principled claim on transfer and motivation, and what the labor economics literature does and does not support about the durability of the resulting capabilities. Section 7 derives design conditions. Sections 8 through 10 turn to the teacher side of the problem, covering efficacy beliefs and retention, implementation through peer-supported networks, and community context. Section 11 states what remains unproven. Section 12 states a position.

1.1 Contribution

Three claims in this review are not available by reading its constituent literatures separately. The first is that the dissociation van Merriënboer reported in 1990, in which completing a supplied draft improved what learners produced without improving what they understood, is the same dissociation Fan and colleagues reported for generative assistance in 2025. Two literatures thirty-five years apart identify one failure mode, and the older one specifies its mechanism.

The second is that teacher-facing and student-facing AI designs are different objects with different risk profiles, a distinction the empirical literature has not yet drawn because it has studied only students. The third is that the diagnosis problem which has limited software scaffolding since Puntambekar and Hübscher identified it admits a solution the field has largely ignored, which is to obtain the diagnosis from the practitioner rather than infer it algorithmically.

1.2 Method and its limits

This is a narrative review rather than a systematic one. Sources were identified through citation tracing from the 2006 and 2007 exchange forward, targeted database searching in education and cognitive psychology for the named effects and constructs, and hand-searching of recent generative AI work in learning settings. No formal inclusion protocol was applied and no screening was double-coded. Every citation was verified against the journal or publisher record, and figures that could not be so verified are described qualitatively rather than quantified.

Three limitations follow. Selection is not reproducible, so a different reviewer could assemble a different and equally defensible set. Single authorship means no independent check on interpretation. And the review was written by someone with a commercial interest in one of its conclusions, which is declared in the front matter and discussed in Section 12. Readers should weight the sections deriving design conditions accordingly, and the review has attempted to compensate by applying the same scrutiny to evidence supporting its position as to evidence against it, most visibly in Section 2.3.

Effect sizes throughout are reported within the design context of the study producing them and are not compared across contexts. This matters because the distance between an intervention and its measured outcome largely determines the magnitude available: a scaffolding manipulation measured on student cognition and a professional development program measured on student achievement are not on a common scale, and treating them as though they were is the error identified in Section 8.2.

2. The case against minimal guidance

The 2006 argument is frequently paraphrased into a weaker form that is easy to dismiss. Stated properly it is harder to answer, and it deserves answering rather than dismissal.

TheoreticalKirschner, Sweller, and Clark (2006) argue from human cognitive architecture rather than instructional preference. Working memory is severely limited when processing novel information and effectively unlimited when processing organized schemas retrieved from long-term memory. The function of instruction is therefore to build those schemas. A method that occupies working memory with activity not contributing to schema construction expends the learner’s only scarce resource without return.

Their sharpest claim is epistemological. The manner in which a discipline generates knowledge is not automatically the manner in which a novice acquires it. Practicing scientists operate from expertise permitting productive search through a problem space. A novice placed in the same position performs means-ends analysis, which consumes working memory on goal attainment rather than on problem structure. The learner can solve the problem and acquire little that generalizes.

TheoreticalThe accompanying empirical claim is the stronger of the two. Reviewing roughly half a century of advocacy, the authors report no body of controlled research supporting minimally guided instruction, and report that controlled comparisons consistently favor guided instruction, with the advantage increasing as learner expertise decreases.

2.1 What is being pooled

A definitional problem runs through this literature and should be stated before evidence is cited. Project-based learning and problem-based learning are distinct traditions. Problem-based learning originated in medical education in the 1960s and is typically organized around an ill-structured case addressed over days or weeks by a small tutorial group. Project-based learning developed in K-12 settings around the production of an artifact over a longer arc, frequently for an audience beyond the classroom. The 2006 critique was aimed at both, together with discovery and inquiry learning, on the grounds that all four share minimal guidance as a design property.

This review follows the same convention where the underlying mechanism is shared, and marks the distinction where it matters. Evidence from problem-based learning is used for claims about the effects of structured problem solving on schema construction, which is the mechanism common to both. It is not used for claims about extended production, community audience, or the logistical demands specific to projects. Where a cited study is problem-based rather than project-based, the text says so.

A second and less tractable problem is that the meta-analyses pooled below do not define project-based learning identically, so the comparability of their headline estimates is limited. Readers should treat the convergence of those estimates in the same direction as more informative than any single value among them.

2.2 The 2007 exchange

Educational Psychologist published three commentaries and a reply. Hmelo-Silver, Duncan, and Chinn (2007) argued that problem-based and inquiry learning are not minimally guided but extensively scaffolded, and that evidence for well-scaffolded implementations is stronger than the original article allowed. Schmidt, Loyens, van Gog, and Paas (2007) argued that problem-based learning as practiced in medical education is compatible with human cognitive architecture. Kuhn (2007) argued that the guided-versus-unguided framing was itself the wrong question.

TheoreticalSweller, Kirschner, and Clark (2007) conceded little, and the core of their reply is difficult to refute. If a program is heavily scaffolded, sequenced, and guided, it is not the instruction under critique, and applying a constructivist label does not make the critique apply. The burden falls on the advocate to specify the guidance rather than assert it.

This review adopts that position. Hmelo-Silver and colleagues offered a strong defense of a version of project-based learning that exists but is uncommon. That defense is available only to programs that can name their guidance, locate it in the instructional sequence, and demonstrate that it is present in classrooms rather than in the philosophy. The critique therefore lands on prevailing practice even where it misses the ideal, and the appropriate response is a design obligation rather than a rebuttal.

2.3 What the efficacy evidence shows

The debate was conducted in 2006 and 2007 largely on theoretical grounds. It has since been settled empirically, in a direction that vindicates both parties on the points each was right about.

Meta-analysisAlfieri, Brooks, Aldrich, and Tenenbaum (2011) meta-analyzed 164 studies and ran the comparison the exchange required. Across 580 comparisons, unassisted discovery was inferior to explicit instruction, d = −0.38. Across 360 comparisons, discovery enhanced with feedback, worked examples, or scaffolding was superior to other methods, d = 0.30. The sign of the effect reverses with the presence of guidance.

Meta-analysisLazonder and Harmsen (2016) meta-analyzed 72 studies of guidance within inquiry learning and found positive effects on learning activities (d = 0.66), performance success (d = 0.71), and learning outcomes (d = 0.50), with guidance type moderating performance success but not the other outcomes.

These two results are the empirical resolution of the 2006 exchange, and they support the reading advanced in Section 2.2 rather than either polemical position. Kirschner and colleagues were right that unguided discovery fails. Hmelo-Silver and colleagues were right that scaffolded versions are a different intervention and perform well. The disagreement was in part about which object each party had in view, and the meta-analytic evidence distinguishes the objects cleanly.

Meta-analysisChen and Yang (2019) synthesized 30 studies covering 12,585 students across 189 schools in nine countries and found a mean weighted effect of d+ = 0.71 favoring project-based over traditional instruction. Zhang and Ma (2023) pooled 190 effect sizes from 66 studies and reported an overall SMD of 0.441, with academic achievement at 0.650, and identified project duration, group size, and subject area as significant moderators.

Those pooled estimates require more scrutiny than they usually receive, including from advocates who benefit from them. Zhang and Ma report a Begg rank test of Z = 5.082, p < 0.05, which indicates possible publication bias, and set it aside on the strength of a fail-safe N of 2,546. Fail-safe N is a weak instrument for this purpose, since it assumes unpublished studies average a null effect and takes no account of their likely distribution, and its use to dismiss a significant asymmetry test should not be treated as resolving the question. The same analysis reports I² = 87.4 percent with Q = 1496.2, p < 0.001, which is very high heterogeneity and means the pooled mean is a poor summary of a dispersed distribution. The moderator analysis is therefore the informative part of that paper and the headline SMD is the least informative part of it.

This does not overturn the finding that project-based learning outperforms traditional instruction on average, which is supported by convergent estimates and by randomized evidence reviewed below. It does mean that any specific number quoted from this literature carries more uncertainty than its decimal places suggest, and that claims should be made about direction and about the conditions identified by moderators rather than about magnitude.

Randomized trialCluster randomized trials support the pooled estimates in specific contexts. Krajcik et al. (2023) found an effect of 0.277 SD on an independent NGSS-aligned science assessment across 46 schools and 2,371 third graders. Saavedra et al. (2022) found positive and significant effects on Advanced Placement examination performance across two courses, holding for both low- and high-income students.

Randomized trialThe picture is not uniform. Duke et al. (2021) found higher growth in social studies and informational reading among second graders in low-SES settings but no effect on writing or motivation, and reported that consistency of implementation with session plans predicted growth. The moderator is fidelity rather than curriculum.

Randomized trialWirkala and Kuhn (2011) compared lecture, small-group problem-based learning, and solitary problem-based learning on identical material and found both problem-based conditions outperformed lecture at a nine-week delayed assessment, with group and solitary conditions equivalent. The problem-solving structure carried the effect; collaboration did not add to it.

Read together, these results support a reconciliation the 2007 exchange gestured at without reaching. Structured, scaffolded project-based learning produces moderate positive effects, and the variance is governed by design and implementation quality rather than by the label.

3. What the guidance literature specifies

An instruction to provide support is not a design principle. The cognitive load literature is considerably more specific, and the specificity is what makes it operational.

3.1 Worked examples and completion

Quasi-experimentalSweller and Cooper (1985) found that students who studied worked algebra examples outperformed students who solved equivalent problems, and did so in less time. Studying a completed solution directs attention to solution structure; solving an unfamiliar problem directs attention to goal attainment. These are different cognitive acts.

Quasi-experimentalvan Merriënboer (1990) compared program completion with program generation in high school programming. Learners who modified and completed existing programs outperformed learners generating programs from scratch on construction and retention. The completion strategy occupies the space between a worked example and an open task.

Quasi-experimentalThe same study found no difference between conditions in the ability to interpret programs. Completion improved what learners could produce without improving what they understood about underlying structure. This dissociation between production and comprehension recurs in the generative AI literature reviewed in Section 4, and it is the characteristic risk of any interface supplying a draft for a human to modify.

3.2 Expertise reversal and fading

TheoreticalKalyuga, Ayres, Chandler, and Sweller (2003) established that guidance benefiting novices can impair more expert learners, because processing redundant support consumes working memory the expert would otherwise allocate to the task. No level of guidance is universally correct, and any fixed level becomes more misaligned as competence develops.

TheoreticalRenkl and Atkinson (2003) proposed the resolution, which is to fade worked steps systematically across a planned sequence, moving learners from example study toward independent problem solving in graded increments rather than in a single transition.

QualitativeThis is what scaffolding denoted when Wood, Bruner, and Ross (1976) introduced the term. Their tutor diagnoses the learner continuously, supplies only what the learner cannot yet supply, and withdraws support as the learner assumes it. Contingency and removal are not features of scaffolding but constitutive of it.

TheoreticalPuntambekar and Hübscher (2005) documented what was lost as scaffolding migrated from human tutors into software. Tools deliver fixed supports and describe them as scaffolds, but absent ongoing diagnosis there is nothing for support to be contingent upon, and absent contingency there is nothing to fade. Much of what educational technology terms scaffolding is more accurately termed structure.

3.3 The meta-analytic evidence and what complicates it

Meta-analysisBelland, Walker, Kim, and Lefler (2017) synthesized 144 experimental studies and 333 outcomes on computer-based scaffolding in STEM and found a consistently positive effect on cognitive outcomes, g = 0.46, robust across grade levels, assessment levels, and contexts of use.

Meta-analysisThe same synthesis found that the effect did not vary with the presence or absence of scaffolding change. Fading did not outperform static support. This complicates the account in Section 3.2 and is reported here rather than omitted. The tension between a strong theoretical rationale for fading and a null meta-analytic result remains unresolved.

3.4 Collaboration as a conditional mechanism

TheoreticalKirschner, Sweller, Kirschner, and Zambrano (2018) extended cognitive load theory to collaborative settings. A group can function as a single expanded working memory, distributing a task exceeding individual capacity across members. The extension is explicitly conditional: collaboration returns more than it costs only when task complexity is high enough to offset transaction costs of communication and coordination. Below that threshold, group structures impose load without corresponding benefit.

This condition, rather than a general preference for collaborative learning, is what justifies the peer-network arguments in Section 9, and it is why the Wirkala and Kuhn null on student collaboration is not paradoxical.

4. Generative AI and the reappearance of the guidance variable

Adoption has moved faster than evidence. Survey data establish that the condition under analysis is general rather than prospective.

Descriptive / surveyDoss et al. (2025) report that in spring 2025, 53 percent of English language arts, mathematics, and science teachers and 54 percent of students reported using AI for school, each an increase exceeding 15 percentage points in a single year. Only 35 percent of district leaders reported providing student training on AI, and 34 percent of teachers reported policies addressing AI and academic integrity.

Descriptive / surveySurvey work by the Walton Family Foundation and Gallup (2025) with 2,232 U.S. public school teachers found 60 percent using AI tools for work and 32 percent using them weekly, with weekly users estimating savings of 5.9 hours per week. The most frequent applications were preparation, creating materials, and modifying materials for student needs.

4.1 The sign of the effect depends on the interface

Randomized trialBastani et al. (2024), in a pre-registered randomized controlled trial with nearly 1,000 students across three grade levels in a Turkish high school, compared a standard conversational interface with a version constrained by pedagogical safeguards. Both improved performance during assisted practice, the constrained version substantially more. On subsequent unassisted examinations, students who had used the unconstrained interface performed 17 percent worse than control, while students who had used the constrained version showed no significant difference from control. This work is a preprint at the time of writing and is weighted accordingly, though its design is stronger than much of the published literature in this area.

The same underlying model produced harm or neutrality depending on whether the interface permitted learners to obtain answers without productive struggle. Capability was constant; architecture was the variable.

That result is load-bearing for this review and is a preprint, so its status should be stated. Tan and Rajaratnam have circulated a critique arguing that baseline proficiency, motivation, teacher effectiveness, outside tutoring, socioeconomic background, and technology familiarity were inadequately controlled. Randomization addresses most of these in expectation at this sample size, which blunts the force of the objection, but the study remains unreplicated in a single school system with a single subject, and the argument built on it in Section 7 should be read as contingent on replication.

Randomized trialKestin et al. (2025) offer a second result from the constrained end of the design space. In a crossover randomized trial with 194 undergraduates in introductory physics, an AI tutor built on research-based pedagogical principles and deliberately constrained to withhold direct answers and elicit student reasoning produced learning gains of approximately 0.63 SD over in-class active learning, with higher engagement and lower median time on task.

This study is frequently presented as the mirror image of the preceding one, and it is not. Its comparison condition is active teaching, not unconstrained model access, so it contains no arm that isolates constraint as the active ingredient. What it establishes is that a carefully designed tutor can outperform a strong comparison condition, which is a meaningful result and a different one. Its setting is also a single course at a highly selective institution, and crossover designs carry order effects that the paper addresses but cannot eliminate. It corroborates that constrained designs can work; it does not demonstrate that constraint is why.

Randomized trialFan et al. (2025) found the dissociation predicted by the completion literature. Among 117 university students, those supported by ChatGPT produced greater improvement in essay scores than those supported by human experts, writing analytics, or no support, yet showed no corresponding advantage in knowledge gain or transfer. The authors characterize the risk as metacognitive laziness, in which learners offload regulation of the task rather than performing it.

Fan and colleagues report for a generative interface exactly what van Merriënboer reported for completion problems: improvement in the artifact without corresponding improvement in understanding. Two literatures separated by thirty-five years converge on the same failure mode.

4.2 Teacher-facing and student-facing designs are different objects

A distinction the literature has not yet drawn cleanly is worth drawing here, because it governs which of the above risks apply. Every study reviewed in Section 4.1 concerns a student interacting with a model. The harm Bastani and colleagues measured is a student obtaining answers without productive struggle. The offloading Fan and colleagues documented is a student ceding regulation of a writing task.

A system in which the generative model is available only to the teacher, for the purpose of designing instruction, does not place students in that interaction at all. The documented student-facing risks are absent by construction rather than mitigated by policy, which is a stronger guarantee than a usage rule. The corresponding risk does not disappear; it relocates to the practitioner, where the same dissociation between produced artifact and understood structure applies, and Section 7.5 treats it as the governing design problem for teacher-facing tools.

The distinction matters for evaluation as well as design. A teacher-facing tool should be evaluated on the quality of instruction it produces and on what happens to teacher expertise over time, not on measures imported from the student-tutoring literature, and the outcome measures in the two cases are not interchangeable.

4.3 A caution about the pooled evidence

Meta-analysisWu et al. (2026) synthesized 35 experimental and quasi-experimental studies comprising 134 effect sizes and 4,193 participants and reported an overall effect of g = 0.670 for ChatGPT on student learning outcomes, with cognitive outcomes at g = 0.872 and non-cognitive outcomes at g = 0.539, and no significant publication bias detected.

That estimate warrants two qualifications. First, the underlying studies are short in duration and frequently measure assisted rather than unassisted performance, which is precisely the distinction on which the Bastani result turns. A pooled positive effect on measures taken with the tool available is not evidence about learning that persists when it is withdrawn.

Second, the field has already produced one prominent failure. A widely cited meta-analysis of ChatGPT effects published in 2025 was retracted in April 2026 following identification of discrepancies that, in the editor’s judgment, undermined confidence in its conclusions. It is listed in the references, marked as retracted, and is not treated as evidence anywhere in this review. Its brief career as a citation is itself informative about the maturity of this literature.

5. Transfer, schema, and motivation

The preceding sections establish a constraint on how project-based learning must be designed. They do not establish why one would choose it. That case rests on transfer, and the transfer literature is more demanding than its usual invocation in advocacy suggests.

5.1 What the transfer literature actually permits

TheoreticalBarnett and Ceci (2002) demonstrated that transfer is not a unitary construct, offering a taxonomy separating what transfers from the contexts across which it transfers along six dimensions, and observed that far transfer claims routinely fail to specify which form of transfer is asserted, rendering many of them untestable.

TheoreticalSala and Gobet (2017) reviewed chess instruction, music training, and working memory training and found that apparent far transfer to general cognitive ability attenuates toward zero as design quality improves. The consistency of this pattern across three unrelated domains constrains any claim that a school experience builds a general capacity.

Taken alone, these findings are frequently read as an argument against project-based learning. They are better read as an argument against a specific competing proposition, which is that general capacities can be trained directly and will then apply anywhere. The consensus position is more precise.

TheoreticalThe National Research Council consensus report (Pellegrino & Hilton, 2012) concluded that transferable competencies are real, that transfer is largely domain-specific, and that competencies develop through deeper learning of disciplinary content rather than through generic skills instruction detached from content.

This is the load-bearing finding for project-based designs, and its implication runs opposite to the way such designs are often marketed. It does not support teaching collaboration, communication, or problem solving as standalone strands. It supports building them inside demanding disciplinary work, because the schema that transfers is a schema about something.

The stronger formulation is this review’s inference rather than a conclusion of the cited reports, and is offered as an argument to be tested rather than a finding to be relied on: context is not a motivational wrapper around content but the condition under which knowledge becomes organized enough to be retrieved and applied elsewhere. The National Research Council committee concluded that transfer is domain-specific and develops through deeper disciplinary learning. It did not make the further claim about contextual richness advanced here, and the studies that would adjudicate it, comparing schema construction in rich against impoverished contexts with transfer as the outcome, have not been run at classroom scale.

Read against Sections 2 and 3, this yields a coherent position rather than a contradiction. Cognitive load theory specifies that schemas must be constructed under managed load, which requires guidance. The transfer literature specifies that schemas generalize to the extent that they are built within rich, meaningful context. A project that is contextually rich but cognitively unguided fails the first requirement. A sequence of decontextualized exercises under tight guidance can satisfy the first and fail the second. The design problem is to satisfy both, and neither literature alone states it.

5.2 Motivation as mechanism rather than byproduct

TheoreticalRyan and Deci (2000) established that intrinsic motivation is supported by conditions satisfying autonomy, competence, and relatedness, and that extrinsic contingencies can undermine it. Motivation on this account is not an affective bonus but a determinant of the quality and persistence of cognitive engagement.

Meta-analysisThe empirical picture for project-based learning is consistent but not uniform. Zhang and Ma (2023) found affective effects at SMD = 0.389 overall, with learning interest the strongest component at 0.713, though self-efficacy gains were weaker.

Randomized trialDuke et al. (2021) found no motivational effect in the primary comparison, while reporting that greater consistency with the session plans was associated with higher motivation. As with achievement, the motivational effect appears contingent on implementation rather than guaranteed by design.

QualitativeTwo qualitative studies identify the mechanism by which extended project work produces the relatedness condition. Pieratt (2011), studying a school built around project-based pedagogy, found that the sustained teacher-student contact across a project arc generated relational knowledge that teachers then used to personalize and differentiate instruction. Fitzgerald (2020), analyzing a third-grade science project unit, found literacy and social-emotional opportunities arising together within disciplinary work rather than as separate strands, while concluding that the teacher’s facilitation skill determined whether those opportunities were realized at all.

The second finding is the more important one, and it recurs throughout this review. The same design in a less prepared teacher’s hands produces weaker outcomes, which makes teacher capacity a moderator of the motivational mechanism rather than a separate concern.

The mechanism worth stating precisely is that personal and community connection to the work supplies the relatedness and autonomy conditions that self-determination theory identifies, and that sustained engagement is what permits the extended, effortful schema construction that transfer requires. Motivation matters here instrumentally, because the cognitive work is long and voluntary attention is the resource that funds it. Claims that projects are valuable because students enjoy them do not survive contact with this literature. Claims that engagement is a precondition for the duration of effort that deep schema construction demands are defensible.

6. Labor market value and the limits of durability claims

A frequent justification for project-based designs is that they build capabilities holding value as labor markets change. The demand-side evidence is strong. The durability claim requires more care.

CorrelationalDeming (2017) found that employment and wage growth since 1980 have been strongest in occupations requiring high levels of social skill, that social skill intensity predicts employment and wage growth even controlling for cognitive skill requirements, and that social and cognitive skills function as complements rather than substitutes. Tasks requiring flexible interpersonal coordination have resisted automation because they are difficult to specify as rules.

CorrelationalHeckman and Kautz (2012) established that measured non-cognitive skills predict labor market and life outcomes at magnitudes comparable to cognitive test scores, and that they remain malleable later in development than cognitive skills, which is the principal economic argument for adolescent intervention.

Descriptive / surveyAmerica Succeeds and Lightcast (2025), analyzing approximately 76 million U.S. job postings from 2023 and 2024, found 76 percent requesting at least one durable skill, up from 64 percent in 2021, and 47 percent requesting three or more, with demand rising fastest in technical rather than human-facing occupations.

CorrelationalEcton and Dougherty (2023) found substantial heterogeneity in career and technical education outcomes across fields and student groups, with returns differing sharply between areas such as health care, information technology, and construction. Applied learning yields returns conditional on its application.

What this evidence supports is a conditional statement: the capabilities that well-designed project work targets are rising in market value, are complementary to technical skill rather than substitutable for it, and remain malleable during the years schools have access to students. That is a strong basis for prioritization.

What it does not support is the stronger claim, frequently made, that such instruction insulates students against labor market disruption. No study demonstrates that a school-based intervention produces capabilities that remain valuable through an unspecified future reorganization of work. The honest formulation is that these capabilities are currently undervalued relative to demand and appear more resistant to automation than routine task proficiency, which is a defensible reason to teach them and not a guarantee about anyone’s future.

7. Design conditions

The literature reviewed above converges on a small number of conditions under which AI-mediated project-based learning could be expected to work. They are stated as conditions rather than findings, because each is an inference from adjacent evidence rather than a directly tested proposition.

7.1 Constrained generation rather than open conversation

An open conversational interface positions the user as a search agent in an unbounded problem space. For a student this reproduces the means-ends search identified as load without learning, and the Bastani result indicates the cost empirically. For a teacher performing instructional design, the same structure imposes the same search cost on a different actor. A constrained sequence substitutes ordered decisions for open search, which is the move worked examples make.

The corollary is that such a system should supply a defensible draft for the human to complete and correct rather than a blank field, since completion is the closest studied analogue to generative assistance. That corollary carries the risk documented in Sections 3.1 and 4.1, which is improvement in the artifact without improvement in understanding, and it should be treated as the principal hazard of the approach rather than an incidental caveat.

7.2 Goal specification prior to generation

The failure mode cognitive load theory predicts for project-based learning is an activity that is elaborate, engaging, and substantively empty. Specifying learning goals as an input constraint rather than as post-hoc justification is a structural rather than exhortative remedy, since it makes intrinsic load a property of the target content rather than of project logistics.

Descriptive / surveyThe magnitude of the problem this addresses is documented outside the cognitive load tradition. TNTP (2018), following nearly 4,000 students, found that students spent more than 500 hours per year on assignments not aligned to grade level and on weak instruction, equivalent to six months of lost class time per core subject, with students of color, students from low-income backgrounds, students with mild to moderate disabilities, and English learners disproportionately affected.

Quasi-experimentalTNTP and Zearn (2021), analyzing more than 6,000 elementary classrooms, found that students given grade-level content with just-in-time support completed 27 percent more grade-level lessons than students routed into remediation, with larger differences in majority-minority schools. Holding the standard constant while adjusting support is the operative pattern.

7.3 Iterative formative evidence rather than terminal assessment

Contingency requires visibility. Guidance can be made contingent only on something observable, and in an extended project the observable signal must be generated during the work rather than after it. Process documentation, intermediate artifacts, and structured reflection are therefore not supplementary to scaffolding but the precondition for it, since without them a teacher has no basis on which to calibrate support until the opportunity to adjust has passed.

7.4 Practitioner-supplied diagnosis as the contingency signal

Section 3.2 identified the unresolved problem in software scaffolding: without ongoing diagnosis there is nothing for support to be contingent upon, and Puntambekar and Hübscher’s critique is that tools deliver fixed structure and call it scaffolding. The usual proposed remedy is algorithmic, inferring learner state from interaction data. There is a second route that the literature has largely overlooked, which is to obtain the diagnosis from the practitioner.

A design of that kind treats the project as a live object with a health state, and takes teacher-reported experience as the signal that updates it. The reportable conditions are ordinary and specific: a grouping arrangement that is not functioning, a misconception that has spread across the class rather than sitting with individuals, a pace that has outrun comprehension. Each implies a different instructional response, and each is observable to the teacher well before it appears in any assessment data. A system that accepts these reports and adapts the project accordingly, while retaining a record of the instructional changes made and the pedagogical techniques applied, satisfies both halves of the Wood, Bruner, and Ross definition: support becomes contingent on a diagnosis, and the record of adjustment makes fading deliberate rather than accidental.

If practitioner reports are accurate enough to drive useful adaptation, this would resolve a limitation that otherwise applies to any generative design tool, and would do so without requiring the system to infer competence from usage traces. It would also make the expertise reversal problem less acute in this configuration than the account in Section 3.2 predicts, since the person best positioned to detect that support has become redundant is the practitioner, who can say so. Each clause of that sentence is conditional, and the condition has not been tested. No study known to the author has compared practitioner-reported adaptation against either fixed structure or algorithmic inference, which makes this the most speculative proposal in the review and the one most in need of the trial specified in Section 11.

Two further limitations qualify the claim. The diagnosis is only as good as the teacher’s reading of the room, and the measurement literature reviewed in the Appendix gives reason for caution about practitioner self-report generally, particularly the recalibration effect in which developing expertise produces lower self-assessment. And a longitudinal profile of instructional adjustments is a research instrument as well as a design feature, which raises questions about who may see it that are governed by the evaluation arrangements discussed in Section 8.4 rather than by the design itself. Neither limitation is disqualifying, and neither has been studied.

7.5 Automating logistics, not judgment

A distinction that the cognitive load framework makes available, and that discussions of AI in education frequently miss, is between load arising from the intellectual substance of a task and load arising from its administration. Alignment documentation, scheduling, resource assembly, and format compliance impose extraneous load on teachers without contributing to instructional quality. These are the appropriate targets for automation, and the time-savings findings in Section 4 suggest the available margin is substantial.

Quasi-experimentalA related and more specific mechanism concerns the number of environments a practitioner must hold simultaneously. Chandler and Sweller (1992) established the split-attention effect, in which learners required to integrate information from separate sources incur load from the integration itself rather than from the material. A teacher assembling a project across a planning document, a standards database, a resource repository, an assessment tool, and a reporting system is performing exactly that integration, and the load is a property of the fragmentation rather than of the work.

Consolidating those functions into a single environment is therefore a cognitive load intervention rather than a convenience, and it is the clearest case in this review of a design decision with an unambiguous theoretical warrant. What it does not have is direct evidence at the level of instructional design tools, where the relevant comparison has not been run.

The inverse case is equally important. Instructional judgment, including what is worth teaching, what a particular group of students needs, and what constitutes acceptable evidence of learning, is the substance of the professional act. Automating it does not reduce extraneous load but removes the germane cognitive work through which expertise develops, which is the teacher-side analogue of the metacognitive offloading documented in Section 4. A defensible system therefore keeps the practitioner in the loop continuously by design rather than by permission, automating the administrative and preserving the deliberative.

There is a second argument for that boundary, which concerns professional conditions rather than cognition, and it is taken up in Section 8.

8. Teacher efficacy, satisfaction, and retention

The preceding sections treat the teacher as an instructional designer. The professional literature treats the teacher as an employee in an organization, and the two framings intersect at efficacy beliefs.

8.1 What the efficacy evidence supports

Meta-analysisZee and Koomen (2016), synthesizing forty years of research across 165 articles, found teacher self-efficacy positively associated with instructional quality, classroom processes, student academic adjustment, and dimensions of teacher well-being including job satisfaction and reduced burnout.

CorrelationalGoddard, Hoy, and Woolfolk Hoy (2000) established collective teacher efficacy as a measurable school-level construct and found it positively associated with student achievement across schools, with the association surviving controls for prior achievement and socioeconomic status.

CorrelationalIngersoll (2001), analyzing national teacher turnover data, found that attrition is better explained by organizational conditions than by demographic or compensation factors alone, with the degree of teacher influence over instructional decisions among the conditions predicting whether teachers remain.

Read together, these results describe a plausible causal chain in which conditions that increase a teacher’s sense of instructional capability are associated with better instruction, greater satisfaction, and lower attrition. That chain is consistent with the design boundary drawn in Section 7.4: a system that removes administrative burden while preserving instructional judgment supports efficacy, while one that automates the judgment may reduce workload and erode the sense of professional agency that predicts retention.

8.2 A caution about a widely cited magnitude

Collective teacher efficacy is frequently cited in practitioner literature as the single largest influence on student achievement, on the strength of an effect size of 1.57 reported in Hattie’s Visible Learning syntheses. That figure should not be relied upon, for two reasons that are worth stating plainly in a review that expects to be checked.

TheoreticalThe 1.57 estimate originates in an unpublished doctoral dissertation (Eells, 2011) rather than in a peer-reviewed meta-analysis, and it is derived largely from correlational studies, which does not support the causal reading it typically receives in practitioner materials.

TheoreticalThe broader synthesis in which it is popularized has been subject to sustained methodological criticism. Bergeron and Rivard (2017), writing from a statistical perspective, identify problems in the pooling and comparison of effect sizes drawn from heterogeneous designs and metrics, and argue the resulting rankings cannot bear the interpretive weight placed on them.

The defensible position is that the direction of the collective efficacy finding is supported by peer-reviewed work including Goddard and colleagues, that the construct is worth attending to, and that the specific magnitude circulating in professional development materials rests on an evidence base that will not survive scrutiny. Practitioners citing 1.57 to justify investment are making an argument weaker than the one available to them.

8.3 Compliance, autonomy, and the conditions of practice

The accountability environment documented by TNTP (2018) has a professional as well as an instructional cost. When the observable unit of accountability is coverage and compliance, instructional decisions migrate toward what is defensible rather than what is warranted, which is a description of extrinsic regulation in the terms Ryan and Deci set out and a plausible contributor to the organizational conditions Ingersoll identifies.

This suggests an evaluative criterion for instructional tooling that is separate from learning outcomes. A system that makes alignment and evidence automatic rather than effortful lowers the cost of defensible practice, which is a route to reducing compliance pressure without reducing accountability. That proposition is coherent with the literature reviewed here and has not, to the author’s knowledge, been tested directly.

8.4 Evaluation as feedback rather than audit

The preceding argument has a specific application to teacher evaluation, where the evidence for the prevailing approach is unusually discouraging.

Descriptive / surveyTNTP (2009) documented what it termed the widget effect, in which evaluation systems failed to distinguish among teachers in any meaningful way, with the overwhelming majority rated at the top of scales that in practice had one category. Evaluation existed without differentiation, and therefore without information.

Quasi-experimentalThe most substantial attempt to correct this produced a sobering result. Stecher et al. (2018) evaluated the Intensive Partnerships for Effective Teaching initiative across seven sites over six years and more than $200 million. The sites succeeded in implementing measures of teaching effectiveness and using them in human resource decisions. The authors report that the initiative nonetheless “did not achieve its goals for student achievement or graduation,” with outcomes not dramatically better than comparison sites, and offer incomplete implementation, external policy change, insufficient time, or a flawed theory of action as candidate explanations.

The natural reading is that measuring teachers more precisely does not by itself improve instruction. What the evaluation apparatus produced was a judgment about a teacher rather than information a teacher could act on, and it imposed the cost of producing evidence on the person being judged.

That diagnosis implies a different arrangement rather than abandonment of evaluation. If the artifacts an evaluator needs are generated as a byproduct of instructional design rather than assembled afterward by the teacher, the burden of proof shifts off the practitioner. If the same artifact is legible to peers as well as evaluators, the encounter can carry instructional information in both directions rather than terminating in a rating. Read against Sims and colleagues, an evaluation conversation grounded in a shared, specific plan is a plausible embedding mechanism, which is the purpose their meta-analysis identifies as most often missing. Read against Kraft and colleagues, it is a low-cost approximation of the contingent diagnosis that coaching supplies and that cannot be staffed at scale.

This is a hypothesis with a clear failure mode worth naming. An artifact designed to satisfy an evaluator can become the object of optimization, reproducing the compliance orientation it was meant to relieve. Whether shared planning artifacts function as feedback or as a more efficient audit is an empirical question about implementation and local culture, and it is not settled by the design.

9. Implementation through peer-supported networks

The fidelity findings in Section 2.2 locate substantial outcome variance in enactment rather than design. The professional learning literature indicates both that this variance is movable and that most attempts to move it fail.

Meta-analysisSims et al. (2025) meta-analyzed 104 randomized controlled trials of teacher professional development and found an overall average effect of d = 0.05. Programs addressing all four of their proposed purposes, comprising insight, motivation, technique, and embedding, averaged d = 0.14 with greater consistency, and programs incorporating more of the identified mechanisms showed modestly greater impact.

That result inverts a common assumption. Teacher professional development in the aggregate is close to ineffective. What distinguishes the exceptions is not duration or topic but whether the design addresses the full path from understanding a practice to habitually performing it. A program delivering insight without embedding, or technique without motivation, is predicted to fail, and the meta-analytic average suggests most do.

Descriptive / surveyField evidence on implementation quality is starker than the professional development literature alone conveys. Balfanz and Byrnes (2026), reporting three years of data from more than one hundred middle and high schools implementing student success systems, found that teams reaching solid or strong implementation reduced course failure and chronic absenteeism by roughly 16 percent in year one, with effects accelerating to a 44 percent reduction in course failure by year three. Teams achieving only partial implementation, 24 percent of the sample in the first year, saw chronic absenteeism worsen by four percentage points.

Partial implementation did not produce a smaller benefit. It produced a worse outcome than the baseline, which means the relationship between implementation quality and results is not merely monotonic but has a threshold below which the intervention is counterproductive.

Any program that cannot support schools past that threshold is better not attempted, and any efficacy estimate averaging across implementation levels understates both the upside and the risk.

Meta-analysisKraft, Blazar, and Hogan (2018) meta-analyzed 60 causal studies of teacher coaching and found pooled effects of 0.49 SD on instruction and 0.18 SD on student achievement, among the stronger results available for any professional learning mechanism.

Meta-analysisThe same analysis found that effects from effectiveness trials of larger programs are a fraction of those from efficacy trials of smaller programs. Coaching does not scale linearly, and evidence generated at small scale should not be extrapolated to system-wide implementation.

The scaling constraint is the practical argument for distributing support across a network rather than concentrating it in expert coaches, whose supply is the binding limit. The evidence for that substitution is real but weaker than the coaching evidence it would partially replace.

Meta-analysisLomos, Hofman, and Bosker (2011) meta-analyzed the relationship between professional community and student achievement and found a small but significant summary effect of d = 0.25, while noting persistent conceptual and methodological inconsistency in how professional community is defined and measured.

Descriptive / surveyVescio, Ross, and Adams (2008) reviewed research on professional learning communities and concluded that participation is associated with changes in teaching practice and with improvements in student learning, while observing that the empirical base consists largely of small qualitative studies rather than controlled designs.

Meta-analysisRohrbeck et al. (2003) provide meta-analytic support for peer-assisted learning producing positive achievement outcomes. That evidence concerns elementary students rather than adult professionals, and its extension to teacher learning is an assumption rather than a finding.

The theoretical warrant for peer structures is specific rather than sentimental. Collaborative cognitive load theory predicts benefit precisely when task complexity exceeds individual working memory capacity, which describes redesigning a course while teaching it. Below that threshold the same structure is a standing meeting consuming time without return, and the condition should be stated whenever the claim is made.

Reading these results against the Sims framework yields a defensible composition. Initial in-person training can supply insight and technique. Ongoing coaching supplies the contingent, individualized diagnosis that Wood, Bruner, and Ross identified as constitutive of scaffolding, applied to adult learners. A recurring peer working structure supplies embedding, which the meta-analytic evidence suggests is most often omitted and most consequential. Such a structure is also the natural site for the collective efficacy beliefs discussed in Section 8, since those beliefs are school-level properties that individual coaching is poorly positioned to build. No study has evaluated that combination as a unit, and its components carry unequal evidentiary weight.

10. Community context and the stakeholder network

Section 5 established that context is a condition for transfer rather than a motivational wrapper. That raises a design question about which context, and the answer has consequences beyond the cognitive.

TheoreticalAronson and Laughter (2016), synthesizing research across mathematics, science, social studies, language arts, and English language learning, found culturally relevant education associated with improved student outcomes across content areas, while noting that standardized curricula and testing regimes have marginalized such approaches in reform discourse.

Meta-analysisJeynes (2007), meta-analyzing 52 studies of urban secondary students, found parental involvement associated with academic outcomes at approximately 0.5 to 0.55 standard deviations across achievement measures, with effects holding across racial groups.

These literatures are usually kept separate from the cognitive load literature, and the separation is unhelpful. If schema construction depends on context, and if the contexts a student can reason about most richly are those they inhabit, then local and community-grounded problem contexts are not a matter of representation alone. They are a matter of the prior knowledge available for new schemas to attach to, which is a cognitive argument for cultural responsiveness rather than only an ethical one.

Public and local data sources make this operationally tractable in a way it was not previously. A project grounded in local economic, environmental, or civic data is anchored in a context about which students already possess prior knowledge, and it produces work with a genuine external audience. The stakeholder relationships that follow, connecting schools with community organizations and families, are consistent with the engagement literature above, though no study has evaluated the specific composition of AI-assisted design, locally grounded context, and community audience as a unit. The claim available here is that each component has independent support and that their combination is plausible, not that the combination has been tested.

11. What remains unproven

Six gaps are identifiable, and each is addressable.

First, the fading question is unresolved. Expertise reversal predicts that constant guidance imposes a cost on developing expertise; the largest relevant meta-analysis finds no difference between faded and static scaffolding. Resolving this requires trials manipulating fading directly over sufficient duration for expertise to develop.

Second, the generative AI evidence measures the wrong thing too often. The distinction between assisted and unassisted performance is the distinction on which the Bastani result turns, and much of the pooled literature does not preserve it. Studies should report unassisted post-tests as a primary outcome and treat assisted performance as a manipulation check.

Third, teacher AI literacy lacks causal evidence. Frameworks and instruments substantially outpace demonstration that training on them alters practice or student outcomes. Evidence on this point is reviewed in the Appendix; the summary position is that cluster-randomized designs with observed practice as the outcome are the necessary next step.

Fourth, the design-to-enactment gap is unmeasured. Specifying learning goals before generation constrains the artifact. It does not establish that the enacted lesson carries grade-appropriate cognitive demand. Instruments separately assessing designed and enacted demand within the same classrooms would isolate the contribution of design tools from that of the teachers using them.

Fifth, the components of implementation support have not been separated. Training, coaching, and peer structures are typically deployed and evaluated as a bundle, and providers of such bundles have an interest in attributing effects to the component they supply. Designs varying support dose independently of tool access are necessary before any component can be credited.

Sixth, the retention argument is untested. The chain running from reduced administrative burden through preserved instructional agency to efficacy, satisfaction, and retention is assembled here from separate literatures. Each link has support. The chain as a whole has not been evaluated, and it is the kind of proposition that could be tested with existing instruments in a multi-year district study.

Seventh, practitioner-supplied diagnosis has not been compared with the alternatives. Section 7.4 proposes that teacher-reported project health can serve as the contingency signal that software scaffolding otherwise lacks. Whether such reports are accurate enough to drive useful adaptation, how they compare with algorithmic inference from interaction data, and whether a longitudinal record of instructional adjustment improves subsequent design are all open. The relevant study would compare adaptation driven by practitioner report against fixed structure, with enacted instructional quality as the outcome.

Eighth, the threshold effect deserves direct study. Balfanz and Byrnes found partial implementation producing worse outcomes than baseline. If that threshold is general rather than specific to on-track systems, it has consequences for how any program should be sold, staged, and withdrawn, and it is not currently characterized well enough to act on.

12. Conclusion and position

The 2006 critique of minimally guided instruction was correct in its central claim and has not been superseded. Its practical implication was never that project-based learning should be abandoned, but that its guidance must be specified rather than assumed, and the subsequent efficacy literature is consistent with that reading: structured, well-implemented project-based learning produces moderate positive effects, and variance is governed by design and fidelity.

The generative AI evidence arriving now recapitulates the same lesson in a new setting. The direction of effect depends on whether the interface preserves productive struggle or removes it, and the difference between an instrument that harms unassisted learning and one that does not lies in constraints imposed by designers rather than in the capability of the underlying model. This is a design finding, and it places responsibility on those building instructional tools rather than on the technology.

The transfer literature supplies the reason to accept the additional design burden that project-based learning imposes. Generalizable knowledge is built within rich disciplinary context rather than through generic skills instruction, and the capabilities that such work develops are rising in labor market value. The professional learning literature supplies the constraint on implementation: most teacher-facing support fails, and what distinguishes the exceptions is whether the design carries a practice from understanding through to habit.

12.1 A position, stated as such

What follows is the author’s position rather than a finding of this review, and is separated from the preceding sections for that reason.

The constraint that historically prevented project-based learning from scaling was never a shortage of evidence or of teacher willingness. It was that designing a rigorous, standards-anchored, locally grounded project is expert work requiring time that school systems do not allocate, which confined high-quality implementation to individuals willing to absorb the cost personally. Generative tooling changes that cost structure, and the evidence reviewed here indicates it changes it in either direction depending on how the tooling is built.

The position taken here is that scaling project-based learning responsibly requires both halves of this review at once: AI-native tooling constrained along the lines of Section 7, and the peer-supported professional community described in Section 9, with community and family stakeholders in the arrangement described in Section 10. Tooling without community produces designs nobody enacts well, which the fidelity findings predict. Community without tooling leaves the original time constraint in place. The combination is infrastructure rather than product, and it is the thesis on which the author’s organization operates.

That position is a hypothesis. It is consistent with the literature reviewed here and is not established by it, and the studies specified in Section 11 are the ones that would determine whether it is correct.

Appendix: Teacher AI literacy

If interface architecture governs whether generative tools help or harm, the competence of the professional selecting and operating those tools becomes consequential. The conceptual literature is well developed.

TheoreticalLong and Magerko (2020) provided the first widely adopted definition of AI literacy as a set of competencies enabling individuals to evaluate AI technologies critically, communicate and collaborate with them, and use them as tools. Ng, Leung, Chu, and Qiao (2021) organized the exploratory literature into four dimensions, comprising knowing and understanding AI, using and applying AI, evaluating and creating with AI, and AI ethics.

Descriptive / surveyUNESCO (2024) published an AI competency framework for teachers structured as fifteen competencies across five dimensions, comprising a human-centred mindset, ethics of AI, AI foundations and applications, AI pedagogy, and AI for professional learning, each specified at three progression levels.

The empirical position is considerably weaker than the conceptual one.

Quasi-experimentalTan, Cheng, and Ling (2025) evaluated a six-month professional development intervention with 125 university teachers using an intelligent-TPACK framework and found significant aggregate gains in AI competency. Some participants recorded negative gain scores, which the authors attribute to metacognitive recalibration as participants moved from unconscious to conscious incompetence.

Descriptive / surveyA scoping review by Baizhanov et al. (2026) concluded that the corpus provides stronger support for measurement-structure claims than for causal claims about professional development effectiveness, and characterized the field as one in which pedagogical innovation is advancing faster than evidence on validity, transferability, and long-term impact.

The recalibration finding has a methodological consequence beyond this literature. If competent instruction in AI use causes practitioners to revise their self-assessment downward, pre-post self-report will understate the effects of effective programs and overstate the effects of superficial ones. This is structurally identical to the reference bias problem in student non-cognitive measurement, where West et al. (2016) found students at higher-performing schools rating themselves lower on conscientiousness and self-control because respondents judge themselves against locally available comparisons. In both cases the implication is that evaluation should privilege observed practice over self-report.

Two further measurement cautions apply to any evaluation of the designs discussed in this review. Duckworth and Yeager (2015) catalogued the limitations of self-report measures of personal qualities used for accountability or program evaluation, concluding that current instruments are inadequate for high-stakes comparison. And Murphy (2021) argues that universal design for learning, which is embedded in many instructional design tools, has been adopted at scale on a thin causal evidence base, observing that no rigorous published research demonstrates improvement attributable to an intervention designed on its principles. It is defensible as a design heuristic and is not defensible as an evidence-based intervention. Taylor, Oberle, Durlak, and Weissberg (2017), by contrast, found school-based social and emotional learning interventions produced effects persisting well beyond the intervention period, which establishes that such outcomes are movable even where their measurement remains contested.

References

  • Alfieri, L., Brooks, P. J., Aldrich, N. J., & Tenenbaum, H. R. (2011). Does discovery-based instruction enhance learning? Journal of Educational Psychology, 103(1), 1–18. doi.org/10.1037/a0021017
  • America Succeeds & Lightcast. (2025, July). Durable by design: An update on the high demand for durable skills.
  • Aronson, B., & Laughter, J. (2016). The theory and practice of culturally relevant education: A synthesis of research across content areas. Review of Educational Research, 86(1), 163–206. doi.org/10.3102/0034654315582066
  • Baizhanov, N., Sharimbayev, B., Churbanova, Z., Khaidarova, A., & Shinetova, L. (2026). Artificial intelligence and teacher competence: A scoping review of assessment, analytics, and professional development. Frontiers in Artificial Intelligence, 9, 1901449. doi.org/10.3389/frai.2026.1901449
  • Balfanz, R., & Byrnes, V. (2026). GRAD Partnership year three impact results. Everyone Graduates Center, Johns Hopkins University.
  • Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612–637. doi.org/10.1037/0033-2909.128.4.612
  • Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. Working paper. [Preprint; not peer reviewed at time of writing.] doi.org/10.2139/ssrn.4895486
  • Belland, B. R., Walker, A. E., Kim, N. J., & Lefler, M. (2017). Synthesizing results from empirical research on computer-based scaffolding in STEM education: A meta-analysis. Review of Educational Research, 87(2), 309–344. doi.org/10.3102/0034654316670999
  • Bergeron, P.-J., & Rivard, L. (2017). How to engage in pseudoscience with real data: A criticism of John Hattie’s arguments in Visible Learning from the perspective of a statistician. McGill Journal of Education, 52(1).
  • Chandler, P., & Sweller, J. (1992). The split-attention effect as a factor in the design of instruction. British Journal of Educational Psychology, 62(2), 233–246. doi.org/10.1111/j.2044-8279.1992.tb01017.x
  • Chen, C.-H., & Yang, Y.-C. (2019). Revisiting the effects of project-based learning on students’ academic achievement: A meta-analysis investigating moderators. Educational Research Review, 26, 71–81. doi.org/10.1016/j.edurev.2018.11.001
  • Deming, D. J. (2017). The growing importance of social skills in the labor market. The Quarterly Journal of Economics, 132(4), 1593–1640. doi.org/10.1093/qje/qjx022
  • Doss, C. J., Bozick, R., Schwartz, H. L., Chu, L., Rainey, L. R., Woo, A., Reich, J., & Dukes, J. (2025). AI use in schools is quickly increasing but guidance lags behind: Findings from the RAND survey panels (RR-A4180-1). RAND Corporation.
  • Duckworth, A. L., & Yeager, D. S. (2015). Measurement matters: Assessing personal qualities other than cognitive ability for educational purposes. Educational Researcher, 44(4), 237–251. doi.org/10.3102/0013189X15584327
  • Duke, N. K., Halvorsen, A.-L., Strachan, S. L., Kim, J., & Konstantopoulos, S. (2021). Putting PjBL to the test: The impact of project-based learning on second graders’ social studies and literacy learning and motivation in low-SES school settings. American Educational Research Journal, 58(1), 160–200. doi.org/10.3102/0002831220929638
  • Ecton, W. G., & Dougherty, S. M. (2023). Heterogeneity in high school career and technical education outcomes. Educational Evaluation and Policy Analysis, 45(1), 157–181. doi.org/10.3102/01623737221103842
  • Eells, R. J. (2011). Meta-analysis of the relationship between collective teacher efficacy and student achievement [Unpublished doctoral dissertation]. Loyola University Chicago.
  • Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gaševic, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530. doi.org/10.1111/bjet.13544
  • Fitzgerald, M. S. (2020). Overlapping opportunities for social-emotional and literacy learning in elementary-grade project-based instruction. American Journal of Education, 126(4). doi.org/10.1086/709545
  • Goddard, R. D., Hoy, W. K., & Woolfolk Hoy, A. (2000). Collective teacher efficacy: Its meaning, measure, and impact on student achievement. American Educational Research Journal, 37(2), 479–507. doi.org/10.3102/00028312037002479
  • Heckman, J. J., & Kautz, T. (2012). Hard evidence on soft skills. Labour Economics, 19(4), 451–464. doi.org/10.1016/j.labeco.2012.05.014
  • Hmelo-Silver, C. E., Duncan, R. G., & Chinn, C. A. (2007). Scaffolding and achievement in problem-based and inquiry learning: A response to Kirschner, Sweller, and Clark (2006). Educational Psychologist, 42(2), 99–107. doi.org/10.1080/00461520701263368
  • Ingersoll, R. M. (2001). Teacher turnover and teacher shortages: An organizational analysis. American Educational Research Journal, 38(3), 499–534. doi.org/10.3102/00028312038003499
  • Jeynes, W. H. (2007). The relationship between parental involvement and urban secondary school student academic achievement: A meta-analysis. Urban Education, 42(1), 82–110. doi.org/10.1177/0042085906293818
  • Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31. doi.org/10.1207/S15326985EP3801_4
  • Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. doi.org/10.1038/s41598-025-97652-6
  • Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why minimal guidance during instruction does not work: An analysis of the failure of constructivist, discovery, problem-based, experiential, and inquiry-based teaching. Educational Psychologist, 41(2), 75–86. doi.org/10.1207/s15326985ep4102_1
  • Kirschner, P. A., Sweller, J., Kirschner, F., & Zambrano R., J. (2018). From cognitive load theory to collaborative cognitive load theory. International Journal of Computer-Supported Collaborative Learning, 13, 213–233. doi.org/10.1007/s11412-018-9277-y
  • Kraft, M. A., Blazar, D., & Hogan, D. (2018). The effect of teacher coaching on instruction and achievement: A meta-analysis of the causal evidence. Review of Educational Research, 88(4), 547–588. doi.org/10.3102/0034654318759268
  • Krajcik, J., Schneider, B., Miller, E. A., Chen, I.-C., Bradford, L., Baker, Q., Bartz, K., Miller, C., Li, T., Codere, S., & Peek-Brown, D. (2023). Assessing the effect of project-based learning on science learning in elementary schools. American Educational Research Journal, 60(1), 70–102. doi.org/10.3102/00028312221129247
  • Kuhn, D. (2007). Is direct instruction an answer to the right question? Educational Psychologist, 42(2), 109–113. doi.org/10.1080/00461520701263376
  • Lazonder, A. W., & Harmsen, R. (2016). Meta-analysis of inquiry-based learning: Effects of guidance. Review of Educational Research, 86(3), 681–718. doi.org/10.3102/0034654315627366
  • Lomos, C., Hofman, R. H., & Bosker, R. J. (2011). Professional communities and student achievement: A meta-analysis. School Effectiveness and School Improvement, 22(2), 121–148. doi.org/10.1080/09243453.2010.550467
  • Long, D., & Magerko, B. (2020). What is AI literacy? Competencies and design considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (pp. 1–16). ACM. doi.org/10.1145/3313831.3376727
  • Murphy, M. P. A. (2021). Belief without evidence? A policy research note on Universal Design for Learning. Policy Futures in Education, 19(1), 7–12. doi.org/10.1177/1478210320940206
  • Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, 100041. doi.org/10.1016/j.caeai.2021.100041
  • Pellegrino, J. W., & Hilton, M. L. (Eds.). (2012). Education for life and work: Developing transferable knowledge and skills in the 21st century. National Research Council; National Academies Press. doi.org/10.17226/13398
  • Pieratt, J. R. (2011). Teacher-student relationships in project based learning: A case study of High Tech Middle North County [Doctoral dissertation, Claremont Graduate University]. doi.org/10.5642/cguetd/13
  • Puntambekar, S., & Hübscher, R. (2005). Tools for scaffolding students in a complex learning environment: What have we gained and what have we missed? Educational Psychologist, 40(1), 1–12. doi.org/10.1207/s15326985ep4001_1
  • Renkl, A., & Atkinson, R. K. (2003). Structuring the transition from example study to problem solving in cognitive skill acquisition: A cognitive load perspective. Educational Psychologist, 38(1), 15–22. doi.org/10.1207/S15326985EP3801_3
  • Rohrbeck, C. A., Ginsburg-Block, M. D., Fantuzzo, J. W., & Miller, T. R. (2003). Peer-assisted learning interventions with elementary school students: A meta-analytic review. Journal of Educational Psychology, 95(2), 240–257. doi.org/10.1037/0022-0663.95.2.240
  • Ryan, R. M., & Deci, E. L. (2000). Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being. American Psychologist, 55(1), 68–78. doi.org/10.1037/0003-066X.55.1.68
  • Saavedra, A. R., Lock Morgan, K., Liu, Y., Garland, M. W., Rapaport, A., Hu, A., Hoepfner, D., & Haderlein, S. K. (2022). The impact of project-based learning on AP exam performance. Educational Evaluation and Policy Analysis, 44(4), 638–666. doi.org/10.3102/01623737221084355
  • Sala, G., & Gobet, F. (2017). Does far transfer exist? Negative evidence from chess, music, and working memory training. Current Directions in Psychological Science. doi.org/10.1177/0963721417712760
  • Schmidt, H. G., Loyens, S. M. M., van Gog, T., & Paas, F. (2007). Problem-based learning is compatible with human cognitive architecture: Commentary on Kirschner, Sweller, and Clark (2006). Educational Psychologist, 42(2), 91–97. doi.org/10.1080/00461520701263350
  • Sims, S., Fletcher-Wood, H., O’Mara-Eves, A., Cottingham, S., Stansfield, C., Goodrich, J., Van Herwegen, J., & Anders, J. (2025). Effective teacher professional development: New theory and a meta-analytic test. Review of Educational Research, 95(2). doi.org/10.3102/00346543231217480
  • Stecher, B. M., Holtzman, D. J., Garet, M. S., Hamilton, L. S., Engberg, J., Steiner, E. D., Robyn, A., Baird, M. D., Gutierrez, I. A., Peet, E. D., Brodziak de los Reyes, I., Fronberg, K., Weinberger, G., Hunter, G. P., & Chambers, J. (2018). Improving teaching effectiveness: Final report. The Intensive Partnerships for Effective Teaching through 2015–2016 (RR-2242-BMGF). RAND Corporation.
  • Sweller, J., & Cooper, G. A. (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction, 2(1), 59–89. doi.org/10.1207/s1532690xci0201_3
  • Sweller, J., Kirschner, P. A., & Clark, R. E. (2007). Why minimally guided teaching techniques do not work: A reply to commentaries. Educational Psychologist, 42(2), 115–121. doi.org/10.1080/00461520701263426
  • Tan, X., Cheng, G., & Ling, M. H. (2025). Enhancing teachers’ AI competency: A professional development intervention study based on the intelligent-TPACK framework. Computers and Education: Artificial Intelligence, 9, 100521. doi.org/10.1016/j.caeai.2025.100521
  • Tan, S., & Rajaratnam, V. (2024). Critique of Generative AI Can Harm Learning study design. Working paper. [Preprint; not peer reviewed.] doi.org/10.2139/ssrn.4898213
  • Taylor, R. D., Oberle, E., Durlak, J. A., & Weissberg, R. P. (2017). Promoting positive youth development through school-based social and emotional learning interventions: A meta-analysis of follow-up effects. Child Development, 88(4), 1156–1171. doi.org/10.1111/cdev.12864
  • TNTP. (2009). The widget effect: Our national failure to acknowledge and act on differences in teacher effectiveness (2nd ed.).
  • TNTP. (2018). The opportunity myth: What students can show us about how school is letting them down and how to fix it.
  • TNTP & Zearn. (2021). Accelerate, don’t remediate: New evidence from elementary math classrooms.
  • UNESCO. (2024). AI competency framework for teachers. United Nations Educational, Scientific and Cultural Organization.
  • van Merriënboer, J. J. G. (1990). Strategies for programming instruction in high school: Program completion vs. program generation. Journal of Educational Computing Research, 6(3), 265–285. doi.org/10.2190/4NK5-17L7-TWQV-1EHL
  • Vescio, V., Ross, D., & Adams, A. (2008). A review of research on the impact of professional learning communities on teaching practice and student learning. Teaching and Teacher Education, 24(1), 80–91.
  • Walton Family Foundation & Gallup. (2025). Teachers and AI: Findings from a national survey of 2,232 U.S. public school teachers.
  • Wang, J., & Fan, W. (2025). The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis. Humanities and Social Sciences Communications, 12, 621. RETRACTED 22 April 2026; referenced in Section 4.3 as an object of discussion, not as evidence.
  • West, M. R., Kraft, M. A., Finn, A. S., Martin, R. E., Duckworth, A. L., Gabrieli, C. F. O., & Gabrieli, J. D. E. (2016). Promise and paradox: Measuring students’ non-cognitive skills and the impact of schooling. Educational Evaluation and Policy Analysis, 38(1), 148–170. doi.org/10.3102/0162373715597298
  • Wirkala, C., & Kuhn, D. (2011). Problem-based learning in K–12 education: Is it effective and how does it achieve its effects? American Educational Research Journal, 48(5), 1157–1186. doi.org/10.3102/0002831211419491
  • Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry, 17(2), 89–100. doi.org/10.1111/j.1469-7610.1976.tb00381.x
  • Wu, X., Zhu, P., Zhang, J., Yin, M., & Wang, Y. (2026). ChatGPT’s impact on student learning outcomes: A meta-analysis of 35 experimental studies. Humanities and Social Sciences Communications, 13, 684. doi.org/10.1057/s41599-026-07019-z
  • Zee, M., & Koomen, H. M. Y. (2016). Teacher self-efficacy and its effects on classroom processes, student academic adjustment, and teacher well-being: A synthesis of 40 years of research. Review of Educational Research, 86(4), 981–1015. doi.org/10.3102/0034654315626801
  • Zhang, L., & Ma, Y. (2023). A study of the impact of project-based learning on student learning effects: A meta-analysis. Frontiers in Psychology, 14, 1202728. doi.org/10.3389/fpsyg.2023.1202728

Join the TeachCraft teacher community

Plan projects with other educators, share what works, and get feedback before your next unit launches.