← Back to all topics
Teaching · Ch. 14–15 (The Path Forward) · AI-Literacy Protocol

The Instrument vs. the Reasoner: A Classroom AI-Literacy Protocol

DRAFT — review before publishing

Ch. 15 draws the line this whole page works from. AlphaFold, the AI system that solved a fifty-year protein-folding problem, doesn't just output a shape — it outputs a confidence score for every part of that shape, called pLDDT, colored dark blue for a region it trusts and orange or red for one it doesn't. It's built to tell you how much to trust its own answer. A fluent AI reasoner — the kind writing an essay, not folding a protein — has no pLDDT. It produces the same tone of confident authority whether it's right or making something up. Everybody's Science puts it plainly: “AI as a scientific instrument is extraordinary. But AI as a reasoner is a different matter: it produces fluency, not argument.” The skill that matters is evaluating the answer, not generating it — and Ch. 14 is honest about where that skill actually breaks down in a real classroom: “My students are told to use AI constantly… and most of them are bad at it, not because the tool is bad, but because they don't yet know how to frame a question that asks for more than the first paragraph a search engine would have handed them anyway.”

The naive-vs-expert AI test

Ch. 14 ran this test directly and it's worth reproducing live, in front of a class, in about ten minutes — the point isn't the answer, it's watching the same tool respond differently to a better question.

From Ch. 14Same tool, same day, only the question changed

Naive: “Do vaccines cause autism?” — a correct, generic answer that could have been written before the question was asked.

Evidence-literate: a question built from Wakefield's actual mechanistic claim (the “leaky gut” hypothesis) and the developmental-timing confound — which pulled a specific, sourced answer citing two Danish national cohort studies (Madsen et al., NEJM, 2002, 537,303 children; Hviid et al., Annals of Internal Medicine, 2019, 657,461 children) that directly ruled out the proposed mechanism.

Naive: “Does a snowstorm in Texas mean climate change isn't real?” — correct, generic, requires no atmospheric science.

Evidence-literate: a question naming the actual live scientific debate (Arctic amplification, polar vortex stretching, Judah Cohen's research) — which pulled an honest “this is genuinely unresolved, here's specifically what's disputed and what isn't” answer instead of a one-liner.

Nothing changed about the AI. What changed was prior knowledge, applied as inquiry — the same CER practice this site's Article #5 builds, aimed at a new tool.

What the current research says

Practice What the evidence shows
Generating your own questions, not just answering someone else's The “generation effect” (Slamecka & Graf, 1978) is one of the most replicated findings in learning research: people remember material they generated far better than material they only read.
Students writing and trading their own quiz questions A real classroom study using PeerWise — students write multiple-choice questions, share them, answer each other's — found reliable generation and retrieval-practice gains, with measurable exam-score improvement (Kelley et al., 2019).
Expecting to explain your work to someone else The “protégé effect”: even just expecting to teach material improves how well you learn it, and students who actually teach it outperform those who only expected to (Nestojko et al., 2014).
Using an AI chatbot to study Recent research (2024–2025) finds a moderately positive effect on learning outcomes overall — but with a real caveat: unprompted reliance on chatbot output risks surface-level understanding, and teacher presence measurably increases engagement with AI-assisted learning compared to the same tool used without it.
Metacognitive prompting with gradual release Dr. Erin Peters-Burton's Metacognitive Prompting Intervention (Peters & Kitsantas, 2010), built on Barry Zimmerman's self-regulated-learning model: the teacher models reasoning aloud, then the student tries with support, then with less support, then alone.

Put together, the research points at exactly the gap Ch. 14 names: AI-generated content helps most when a student has to generate, defend, and refine it — not when they simply accept it.

The protocol: Generate, Defend, Refine

This is a direct adaptation of Peters-Burton's gradual-release model, built around the exact practice this article grew out of: students building their own concept inventory and study questions for each chapter, using AI as a drafting partner rather than an answer key.

  1. Generate. Before opening an AI chat, the student names two or three concepts from the chapter they personally find hardest — that list is the seed, not the AI's own guess at what's important. Then they ask the AI to help draft study questions built from those specific concepts, the same move as the evidence-literate prompts above.
    Example seed prompt

    I'm studying [CHAPTER/TOPIC]. The two ideas I'm least sure I understand are [CONCEPT A] and [CONCEPT B]. Write me three study questions that would actually test whether I understand the difference between them, not just whether I can define each one separately.

  2. Defend. Before the question goes into the shared class study guide, the student explains it out loud — to the teacher, or to a partner — and answers one question about it: what does a correct answer to this actually require someone to understand? A question that only requires recall gets sent back to Generate.
  3. Refine. Using that conversation, the student revises the question — sharpening it, fixing something the AI got subtly wrong, or replacing a recall question with one that requires reasoning. The revised version, not the AI's first draft, goes into the shared concept inventory the whole class studies from.

Release the support gradually across a course, the way Peters-Burton's model prescribes: model the whole cycle aloud yourself on the first chapter, run it in pairs with you circulating for the next few, then let students run it independently — which is what a DE-level class can do for every chapter, as long as the Defend step never gets skipped. That conversation is where the actual learning happens, not in the AI chat window.

For individual research projects

The same three-step cycle extends directly to a student's own research question. When a student brings back an AI-generated summary of their topic, don't treat it as a finished answer — treat it the way Ch. 14's CER framework treats any claim: what's the claim here, what evidence would actually support it, and what would change it. A student who can defend an AI's answer under those three questions has learned something. A student who can only repeat it hasn't. This is the same interrogation this site's CER article (Article #5) already builds toward — use that page's AI prompts directly for the longer-form version of this same habit.

Questions for students

  1. Explain what a pLDDT confidence score is and what it tells a scientist using AlphaFold. Remember
  2. Explain why a fluent AI-written paragraph doesn't come with anything like a pLDDT score. Understand
  3. Run the Generate step yourself: name a concept you're unsure about and draft an AI prompt built from your specific confusion, not a generic request. Apply
  4. Compare the naive and evidence-literate questions in the vaccine example above — what specific prior knowledge made the second question possible? Analyze
  5. Evaluate a study question you generated with AI — does answering it correctly require recall, or real understanding? Evaluate
  6. Write a Refine version of a study question that started as a recall-only question. Create

Ch. 14–15 material (the naive-vs-expert AI test, AlphaFold/pLDDT, Peters-Burton citation) lifted and adapted directly from Everybody's Science, First Draft v3. Exact citations there include: Madsen K.M. et al., New England Journal of Medicine, 2002; Hviid A. et al., Annals of Internal Medicine, 2019; Theo Baker, “What A.I. Did to My College Class,” The New York Times, May 17, 2026; Jumper J. et al. on AlphaFold, 2024 Nobel Prize in Chemistry (Hassabis, Jumper, and Baker); Peters, E. and Kitsantas, A., “Self-regulation of student epistemic thinking in science: The role of metacognitive prompts,” Educational Psychology, 2010; Barry Zimmerman's self-regulated-learning model. New citations for this page: Slamecka, N.J. and Graf, P., “The generation effect: Delineation of a phenomenon,” Journal of Experimental Psychology, 1978; Kelley, M.R., Chapman-Orr, E.K., Calkins, S., and Lemke, R.J., “Generation and Retrieval Practice Effects in the Classroom Using PeerWise,” Teaching of Psychology, 2019; Nestojko, J.F. et al., “Expecting to teach enhances learning and organization of knowledge in free recall of text passages,” Memory & Cognition, 2014. AI-chatbot research summary reflects 2024–2025 literature broadly, not one single study — verify current sources before publishing.