AI Fluency Learning · Research

The Research Behind This

The courses on this hub are not opinion. They rest on empirical research: Anthropic's AI Fluency Index, which measured concrete fluency behaviors across tens of thousands of real, anonymized conversations. Here is what it found, what it means for training, and how to discuss it with your team. We summarize the findings and cite the source.

3 readings~30 minDiscussion guidePoints
0% Complete
Not started
What you'll learn
  • What 50,000 real conversations reveal about fluency
  • Why iteration tracks every other fluency behaviour
  • Why scrutiny falls where output looks finished: the artifact effect
  • How to run the team discussion and the three activities
Reading 01

The AI Fluency Index

Method (all figures in this reading are from Anthropic's AI Fluency Index, February 2026). Using privacy-preserving analysis, researchers examined 9,830 multi-turn Claude conversations from one week in January 2026, testing for 11 of the 24 fluency behaviors, the ones directly observable in chat. (The other 13, like honesty about AI's role in submitted work, happen outside the conversation and are arguably the most consequential; qualitative study of those is planned.) Results were stable across days of the week and across six languages. The dataset has since grown past 50,000 conversations across chat and agentic coding tools.

Finding one: iteration tracks every other behaviour

85.7% of the conversations involved iteration and refinement, and those conversations showed roughly double the fluency behaviors of conversations without it (2.67 additional behaviors versus 1.33). Conversations with iteration were 5.6 times more likely to include questioning the model's reasoning and 4 times more likely to include identifying missing context. Iteration is the behavior most strongly associated with every other one, which makes staying in the conversation the first habit worth building.

Finding two: the artifact paradox

In the 12.3% of conversations that produced artifacts (documents, code, tools), the conversations were more directive: clarifying goals (+14.7 percentage points), specifying format (+14.5), providing examples (+13.4), iterating (+9.7). and less evaluative: identifying missing context dropped 5.2 points, fact-checking dropped 3.7, questioning reasoning dropped 3.1. Where the output looks finished, critical evaluation falls away, and that is precisely where complex tasks make models weakest.

Finding three: almost nobody sets the terms

Only about 30% of conversations included any instruction about how the AI should interact. Instructions like "push back if my assumptions are wrong," "walk me through your reasoning before answering," and "tell me what you're uncertain about" change the dynamic of everything that follows, and they are almost always absent.

Honest limits, stated by the researchers themselves: the sample skews toward early adopters, covers one platform and one week, uses binary classification, cannot see evaluation that happens in the user's head or outside the chat, and the findings are correlational, not causal.

What to do with it
  • Treat the first response as a starting point, always.
  • The moment output looks finished is the moment to ask what is missing.
  • Set the terms of the collaboration up front.
Reading 02

A research-backed curriculum

Extending the Index across products revealed that fluency develops along two tracks that behave very differently, which has direct consequences for anyone building AI training.

Each product has a signature move

The gateway behavior that lifts all other fluency indicators differs by surface. In conversational chat, it is iterating: conversations that refine through follow-ups score higher on every other dimension, while single-message conversations show almost no critical evaluation. In agentic tools, where the AI works semi-independently, it is clarifying the goal before the work starts. Teach the signature move first; a curriculum that skips it builds on nothing.

Description grows with practice; Discernment must be taught

Description skills, from in-the-moment shaping (add context, give examples) to durable configuration (standing instructions, project setups), grow organically with exposure. People find their way to them. Discernment does not: it fails to grow with tenure and fails to transfer from feature familiarity. Watching an AI work (reading its output, running its code) only catches errors you can see; it misses wrong assumptions, missing context, and plausible-but-false claims. A result that looks right can still encode the wrong approach.

The training implication

If your training time is limited, concentrate it on Discernment, and close every learning session with a "now question it" step: is this usable, or does it need another round? Does the reasoning hold, or did it just sound confident? What would make this wrong?

The curriculum model in three lines: teach the signature move first; let Description advance along its natural spectrum from momentary to durable; and revisit Discernment at every single step.

Reading 03

A team discussion guide

For teams of any size: leadership groups, faculties, or any group that works with AI. Allow 45 to 60 minutes, have people skim reading one in advance, and pick two or three topics plus one activity.

Topic 1: what your team already does well

Most conversations already show real fluency behaviors: 86% include iteration, about half include goal clarification, about 41% include examples. Which behaviors do you think are strongest on your team, and which weakest? Do people iterate and push back, or accept the first output, and what drives that?

Topic 2: the iteration effect

Iteration correlates with everything, but correlation is not causation. What might explain the link? What barriers stop people iterating: time pressure, perceived effort, not knowing what to ask next? If you designed a team norm to encourage iteration, what would a good follow-up message look like?

Topic 3: the artifact effect

Have you ever accepted an AI output because it looked polished, and discovered problems later? Where does your team's real evaluation happen: in the conversation, or after? What would help people keep a critical eye when outputs look finished?

Topic 4: cultivating fluency in your organization

Of the three recommendations (iterate more, question polished outputs, set the terms), which is most actionable for your team now? What does your current AI skills development actually look like: formal training, peer learning, trial and error? If you could build one fluency skill next quarter, which, and what is the first concrete step?

Three activities

  • The iteration challenge (15 min). Everyone runs a real task, and must send at least three follow-up messages that refine, push back, or redirect before calling the output complete. Debrief: what would you have missed by stopping at the first response?
  • Spot the gap (15 min). Have AI generate an artifact on a topic the group knows well. Together, find what looks right but may be wrong, what is missing, and what assumptions went unchecked.
  • Set the terms (10 min). Everyone writes a short reusable collaboration preamble: how they want the model to interact, what pushback they want, what it should flag. Share and compare.

Sources. Anthropic Education Report: The AI Fluency Index (2026), its research-backed curriculum companion, and the discussion guide by the AI Fluency Program at Anthropic. The findings and figures are theirs; this page summarizes and cites rather than reproduces. The 4D framework is by Rick Dakan and Joseph Feller, CC BY-NC-SA 4.0.