Skip to main content
Research

How to cite

For how conversations were classified — the coding rules on this page:

Keith, M. J. (2026). AIEM mode coding standard (Protocol version 10). https://aimodes.ai/research/rubric-theory

For the typology itself, its measurement properties and its empirical associations:

Keith, M. J., Wood, D. A., & Posey, C. (2026). The Eight-Mode AI Engagement Typology: Differential Cognitive Signatures and a Self-Report–Behavior Gap. PsyArXiv. https://doi.org/10.31234/osf.io/53pwv_v1

The preprint is under review at Scientific Reports and has not been peer reviewed; anyone citing it should say so. The two references license different claims, and neither stands in for the other.

Rubric Derivation Methodology

How the grading rubrics used in aimodes.ai are derived, validated, and revised.

This document is the theoretical anchor for every task-type rubric in the system. The rubrics are not instructor intuition cast as numbers — they are falsifiable theoretical claims derived from cognitive science, education research, and the AI Engagement Typology (AIT). Each rubric's per-mode target proportion and recommended mode path can be traced back through this methodology to a specific theoretical justification.

Primary author: Mark Keith (BYU) with David Wood and Clay Posey. Version: 0.1 (draft, 2026-04-22) Status: Methodology sections (1–3, 5, 6) drafted. Section 4 (per-task derivations) pending author review of methodology before proceeding. Downstream uses: grading spec (/6-Apps/aimodes-grading-spec/), MyEducator integration, the research program, aimodes.ai public /research/rubric-theory page.


Executive summary

The AI Engagement Typology classifies user-AI interactions into eight modes grouped across three epistemic tiers (Passivity, Partnership, Agency). Each task a student completes in aimodes.ai is graded against a target mode distribution — the proportion of the conversation that should be spent in each mode — and a recommended mode path — the order in which modes should first appear.

The central methodological claim of this document is that target distributions and mode paths are not arbitrary. They are derived from a four-step procedure combining:

  1. Task decomposition — breaking the task into its constituent cognitive operations (Anderson & Krathwohl, 2001);
  2. Mode mapping — matching each cognitive operation to the mode whose behavioral signature serves that operation;
  3. Cognitive-load weighting — proportioning modes by the relative effort and time each operation demands (Sweller, 1988);
  4. Expertise calibration — adjusting the distribution by the student's expertise level using scaffolding theory (Vygotsky, 1978) and the expertise reversal effect (Kalyuga, Ayres, Chandler, & Sweller, 2003).

The resulting rubrics are theory-driven in v0.1, with no empirical weighting from student data. Section 6 describes the ML pipeline that will convert them into empirically-calibrated rubrics once sufficient data is collected — moving from theoretical_v1 to empirical_v1 over roughly two semesters of operation.


1. Theoretical foundation

1.1 Why modes are cognitive operations, not behaviors

The eight AIT modes are not behaviors in the operant sense. Each mode specifies a cognitive relationship between user and AI — who is holding the reasoning, who is holding the knowledge, who is holding the evaluation — and the user's observable messages are downstream evidence of that relationship. Treating modes as cognitive operations rather than surface behaviors matters for rubric derivation: a "target distribution" is a claim about what cognitive work the task requires, not what words the student should type.

Theoretical grounding here rests on three traditions:

  • Hammond's Cognitive Continuum Theory (CCT) (Hammond, 1988, 1996). Cognition operates on a continuum from intuitive (fast, pattern-based, low-effort) to analytic (deliberate, rule-based, high-effort). AIT modes differ systematically in where they position the user on this continuum: Oracle and Production Assistant anchor the intuitive pole; Verification, Critical Challenger, and Problem Setter anchor the analytic pole; Tutor and Collaborative Problem-Solver are quasi-rational middle positions that can drift either way depending on execution.
  • Dual-process theories of cognition (Evans & Stanovich, 2013; Kahneman, 2011). System 1 and System 2 thinking correspond approximately to the Passivity and Agency tiers. Partnership tier engagement recruits both systems, which is why it is neither always better nor always worse than either pole — it is appropriate for tasks where the student is actively learning structure they do not yet possess.
  • Bloom's revised taxonomy (Anderson & Krathwohl, 2001). Each mode maps to a Bloom's level, summarized in AI_ENGAGEMENT_MODEL.md. The mapping is not bijective — some modes span two levels — but it provides the cognitive-operation vocabulary that step 1 of the derivation procedure (task decomposition) uses.

1.2 Why tier structure matters for rubrics

The three tiers are not a hierarchy. "More Agency" is not "better." The framework measures fit-to-task, not elevation. A factual lookup is best served by Oracle; a concept you already understand is best served by Verification; a poorly-posed research question is best served by Problem Setter. The target distribution for a task is the theoretically-correct mix for that task, not a student's developmental target.

This is the core reason a task rubric cannot be derived by asking "what's the best way to use AI?" It must be derived by asking "what cognitive operations does this task require, and which modes serve them?"

1.3 Why expertise shifts the distribution

Expertise is the single largest moderator of what mode-mix best serves a given task. Three traditions support this:

  • Chase & Simon (1973) and later work on expertise (Ericsson, Krampe, & Tesch-Römer, 1993; Ericsson, 2006) show that experts have chunked, structured knowledge in their domain that novices lack. Novices learning the same material benefit from scaffolded, explanation-heavy engagement (Tutor); experts derive little from the same engagement and benefit instead from challenge and reframing (Critical Challenger, Problem Setter).
  • Dreyfus & Dreyfus (1980) five-stage model of skill acquisition predicts qualitatively different interaction patterns at each stage. Novices operate on explicit rules; experts operate on situational pattern recognition. The rubric framework compresses Dreyfus's five stages to three tiers (Beginner / Intermediate / Expert) for tractability.
  • Expertise reversal effect (Kalyuga, Ayres, Chandler, & Sweller, 2003; Kalyuga, 2007) demonstrates that instructional techniques optimal for novices actively harm expert learners, and vice versa. Applied to AI engagement: a heavy Tutor/Oracle pattern that supports a novice's understanding may undermine an expert's analytical processing by offloading cognitive work the expert needs to do themselves to stay calibrated.

The practical implication: every target distribution must be tier-adjusted. A single universal distribution per task is theoretically indefensible.


2. Derivation methodology

The procedure for deriving a task-type rubric has five steps. This procedure will be applied to all 26 task types in Section 4; the output is a structured rubric file conforming to the schema in aimodes-grading-spec/task-types/README.md.

Step 1 — Decompose the task into cognitive operations

Using Bloom's revised taxonomy as the operation vocabulary, list every cognitive operation the task requires, at a grain size where each operation is served by 1–2 modes. Operations are verbs: remember, understand, apply, analyze, evaluate, create, plus meta-cognitive operations (reframe, verify, expand).

Example (task: "Research and fact-check a historical claim"):

  • Understand the claim's context and scope (Bloom: Understand)
  • Locate sources (Bloom: Remember/Apply)
  • Evaluate source credibility (Bloom: Evaluate)
  • Integrate evidence into a position (Bloom: Analyze)
  • Stress-test the position against counter-evidence (Bloom: Evaluate)
  • Reframe if evidence contradicts (Bloom: Create, meta-level)

Step 2 — Map operations to modes

For each operation, identify the mode(s) whose behavioral signature (per the canonical classification guide) serves that operation. Most operations map to one dominant mode and one secondary mode.

Example (continuing):

  • Understand → Tutor (3), Oracle (1) secondary
  • Locate sources → Oracle (1), Collaborative (4) secondary
  • Evaluate credibility → Verification (5), Critical Challenger (7) secondary
  • Integrate evidence → Collaborative (4), Creative Expander (6) secondary
  • Stress-test → Critical Challenger (7)
  • Reframe → Problem Setter (8)

Step 3 — Weight by cognitive load and time

Each operation has a relative cognitive load (how much thinking it demands) and a relative time cost (how long it takes in the conversation). These two dimensions combine into a weight. Cognitive load is estimated from Cognitive Load Theory (Sweller, 1988, 1994) using the intrinsic / extraneous / germane tripartite framework (Sweller, van Merriënboer, & Paas, 1998) — intrinsic load (inherent to the task), extraneous load (imposed by presentation), and germane load (schema construction). Time cost is estimated from task-family averages observed in pilot data and adjusted downward when instructional scaffolding compresses time.

The weight w_i for operation i is (load_i × time_i) / Σ(load_k × time_k). The base target distribution for mode m is the sum of w_i over all operations whose dominant mode is m, plus half the weight for operations whose secondary mode is m.

This is the most contested step in the derivation. Estimates of cognitive load and time are not directly observable; they are expert judgments informed by the literature and pilot data. Section 5 (validation) and Section 6 (ML pipeline) are how these estimates are checked and revised.

Step 4 — Apply expertise calibration

The base distribution from Step 3 is adjusted by tier-specific multipliers. The current multipliers in src/lib/engagement/theoretical-distributions.ts are stated below for reference; Section 3 derives them.

ModeBeginnerIntermediateExpert
1 — Oracle1.40.80.3
2 — Production1.21.00.6
3 — Tutor1.51.00.5
4 — Collaborative1.01.21.0
5 — Verification0.71.11.4
6 — Creative0.61.21.3
7 — Challenge0.41.01.6
8 — Problem Setter0.30.91.8

The tier-adjusted distribution is base_m × multiplier_m, normalized to sum to 1.0. The mode path is also tier-adjusted: Beginners get Tutor (3) prepended if not already first; Experts get Problem Setter (8) prepended if not already first.

Step 5 — Normalize and sanity-check

The tier-adjusted distribution is normalized to sum to 1.0. Three sanity checks are applied before the rubric is finalized:

  1. Floor check. No mode's proportion is below 0.02 unless the operation set for the task genuinely excludes that mode. A proportion below 0.02 signals the mode is absent rather than rare, and should be represented as 0 explicitly.
  2. Ceiling check. No mode exceeds 0.50 unless the task is single-mode by design (e.g., a drill that is intentionally Tutor-heavy). A ceiling violation usually means the operation set is too narrow.
  3. Tier sanity. The Passivity / Partnership / Agency tier proportions are computed. For a Beginner version of any non-drill task, Partnership tier should be ≥ 0.40; for an Expert version, Agency tier should be ≥ 0.50. Violations are evidence the multipliers haven't fully propagated.

3. Expertise calibration framework

3.1 Why three levels, not five

Dreyfus & Dreyfus (1980) propose five stages: novice, advanced beginner, competent, proficient, expert. Our framework compresses this to three. The compression is justified by two considerations:

  • Measurement granularity. With currently-available data (N=91 pilot, N=331 Wave 1 survey), five levels produce sparse cells and unstable tier-specific distributions. Three levels produce distributions that are estimable with current sample sizes.
  • Pedagogical interpretability. Instructors and students can meaningfully self-assess into three levels with relatively high agreement; five levels produce much more measurement error at the boundaries. This matters because current aimodes assignments ask students to self-report their expertise level per task.

A future version of the framework (v2.0) may move to five levels once Wave 2 and Study 2+3 data collection enables it. That transition would require re-deriving every task rubric because tier boundaries shift.

3.2 Why these multipliers

The multipliers in the table in Step 4 are derived from four constraints:

  1. Monotonicity constraint. Within each mode, the multiplier moves monotonically across tiers in a direction consistent with the cognitive-relationship theory. Passivity modes (1, 2) decrease with expertise; Agency modes (5–8) increase. Partnership modes (3, 4) do not follow a strict monotone: Tutor (3) declines with expertise (experts need less explanation); Collaborative (4) stays approximately flat because collaboration is useful across levels, differing only in who contributes what.

  2. Tier-sum constraint. Applied to a task whose base distribution has 30% Passivity, 40% Partnership, 30% Agency, the Beginner multipliers produce approximately 42% Passivity, 46% Partnership, 22% Agency (pulling toward Passivity + Partnership). The Expert multipliers produce approximately 10% Passivity, 31% Partnership, 75% Agency before normalization — normalization brings this to 9% / 27% / 64%. This is theoretically consistent with the expertise reversal literature (Kalyuga et al., 2003).

  3. Zero-avoidance constraint. No multiplier is zero; the minimum is 0.3 (Problem Setter for Beginners). This preserves the possibility that a Beginner could use Problem Setter productively — the framework does not forbid any mode at any level, it only specifies what's typical.

  4. Anchor-matching constraint. The multipliers were set so that applying them to a baseline task where the target is uniform (0.125 per mode) produces approximately the Wave 1 observed distributions by self-reported expertise level (N=331). This is a weak constraint — Wave 1 data is a self-report survey, not behavioral data — but it keeps the multipliers within plausibility.

3.3 Limitations acknowledged in v0.1

  • Self-report expertise is noisy. Students' self-reported tier is only approximately correct. Future versions will infer tier from conversation behavior using the predictor.ts module, with the self-report as a prior.
  • Task-independent multipliers are a simplification. Some tasks may warrant task-specific multipliers (e.g., creative writing may shift less with expertise than mathematical proof construction does). v0.1 assumes task-independent multipliers; Section 6 (ML pipeline) is how we test that assumption.
  • Three-level compression loses information. Students on boundary between two levels are miscalibrated by construction. The model is robust to this in aggregate but biased for individuals.
  • Path rules are theory-driven, not yet validated. §3.4 specifies tier-specific path transformations (anchor prepend, side-cycle interleave, tail collapse). The transformations are theoretically motivated but not yet tested against observed student paths at scale. Wave 2 / Study 2+3 path data will validate or revise.

3.4 Path derivation by tier

The base mode path captures the canonical sequence of cognitive operations a task requires. Distributions tell us how often each mode appears; paths tell us in what order. Two students who land on identical mode distributions but in different sequences are doing different work — order encodes dependency between operations. Tier-adjusted paths therefore need their own derivation rules, not just inherited multipliers from the distribution side. Three rules apply at the path level.

Rule 1 — Tier-anchor prepend. A Beginner path prepends Tutor (Mode 3) at position 0 if not already first. An Expert path prepends Problem Setter (Mode 8) at position 0 if not already first. Intermediate paths inherit the base path unchanged.

The asymmetry has theoretical grounding. Beginners lack the schemas needed to act productively in the task domain (Sweller, 1994); Tutor at the start delivers the missing instructional scaffolding before any other operation can succeed. Experts already hold the schemas; for them, the binding constraint is whether the framing of the task is correct. Problem Setter at the start lets them interrogate the frame before committing effort. Reversing the prepend (Tutor for Experts, Problem Setter for Beginners) would harm both — the expertise-reversal effect (Kalyuga et al., 2003) for Experts, and meta-cognitive overload of premature framing for Beginners.

Rule 2 — Tier-characteristic side-cycles. Beginners interleave Tutor as a recurring side-cycle throughout the path, not only at position 0. After every two consecutive non-Tutor operations, the Beginner path returns to Mode 3 to consolidate before continuing. The grounding is cognitive load theory (Sweller, 1988): Beginners spend working memory encoding task structure itself, leaving little capacity for new content unless it is periodically chunked. Schnotz & Kürschner (2007) describe this as the germane-load offload pattern. A Beginner path of [3, 8, 6, 4, 7, 2, 5] therefore renders in practice as [3, 8, 6, 3, 4, 7, 3, 2, 5] — Tutor as connective tissue, not just opening move.

Experts have an analogous side-cycle in Verification (Mode 5). After every two consecutive non-Verification operations, the Expert path returns to Mode 5. Expert paths do not end with verification; they verify continuously, in tight loops with each generative step. This pattern is already noted inline in §4.10 (research) and §4.6 (exam-prep) but generalizes: an Expert path of [8, 6, 4, 7, 2, 5] is more accurately [8, 6, 5, 4, 7, 5, 2, 5]. Intermediates have no characteristic side-cycle; they execute the base path with one anchor prepend (if needed) and minimal interruption.

Side-cycle suppression rules. Two suppressions keep the rule from over-firing: (a) no insertion at the final position of the path (a side-cycle after the last operation has no consolidating effect), and (b) no insertion when the next operation in the base path is already the side-cycle mode (the consolidation will happen anyway). The counter resets to zero whenever the side-cycle mode appears, whether inserted or native.

Rule 3 — Beginner cycle collapse. When the base path ends with a repeated cycle of length k appearing two or more times consecutively (e.g., A → B → A → B), the Beginner path truncates to a single instance of that cycle (A → B). Beginners are still consolidating the first iteration when the task time-box closes; a second iteration is unrealistic under the working-memory budget that defines the tier. Intermediate and Expert paths preserve the repetition. This rule generates the truncated tier paths visible in §4.5 (study), §4.6 (exam-prep), and §4.10 (research).

Empty-path tasks. All three rules are suppressed when the base path is empty. Baseline (§4.1), fitness-check (§4.12), and the two catch-all tasks (§4.25, §4.26) are deliberately free-form; prescribing a tier-specific sequence on top of an empty base would manufacture structure that doesn't exist in the task design.

Application order. When a task's base path is non-empty: (1) collapse repeated tail cycles for Beginner only, (2) prepend the tier anchor (3 for Beginner, 8 for Expert) if not already first, (3) interleave the tier side-cycle (3 for Beginner, 5 for Expert) after every two consecutive non-side-cycle operations. Intermediate paths apply only step 2 implicitly (no anchor needed because Intermediate inherits the base unchanged).

Validation implications. Path rules are checked alongside distribution rules. Specifically: any non-empty Beginner path must start with Mode 3, must contain at least one additional Mode 3 entry at position ≥ 2 if the post-collapse path length is ≥ 4, and must not contain a tail repetition. Any non-empty Expert path must start with Mode 8 and should contain Mode 5 at multiple positions if the path length is ≥ 4. The spec package's task-types/README.md codifies these as machine-checkable rules.


4. Per-task derivations

This section derives the target mode distribution, base mode path, and tier-adjusted variants for each of the 26 task types defined in src/lib/engagement/task-types.ts. Each derivation follows the five-step procedure from Section 2 and applies the expertise multipliers from Section 3.

4.0 How to read these tables

Each derivation presents:

  • Task description — a one-sentence characterization of what the student is doing.
  • Cognitive operations — the discrete operations the task requires, mapped to Bloom's revised taxonomy.
  • Mode mapping and base distribution — the proportion of the task-relevant cognitive work served by each of the 8 modes, summing to 1.00.
  • Tier-adjusted distributions — the base distribution multiplied by the Section 3 expertise multipliers and renormalized. Values shown in the per-task tables are rounded to the nearest whole percent for readability; per-row sums may differ from 100% by ≤ 2 percentage points due to rounding. The precise normalized values (sum exactly 1.000) live in aimodes-grading-spec/task-types/*.yaml — that spec package is the authoritative numerical source for downstream implementers. Methodology tables here are the human-readable presentation.
  • Tier-adjusted mode paths — the recommended sequence of modes per the path derivation rules in §3.4. Beginner paths prepend Tutor (Mode 3), interleave Tutor side-cycles after every two non-Tutor operations, and collapse repeated tail cycles. Expert paths prepend Problem Setter (Mode 8) and interleave Verification (Mode 5) side-cycles after every two non-Verification operations. Intermediate paths inherit the base path with no modification. Paths shown in the per-task tables are post-transformation; the base path is shown only on the Base row.

Distribution notation in text is {M1, M2, M3, M4, M5, M6, M7, M8}. Mode paths are written A → B → C.


4.1 task-baseline — Baseline Assessment (Book 1, Ch 1)

Task. The student has a free-form conversation with AI on any topic they choose, so the framework can measure their natural engagement pattern before any instruction.

Cognitive operations. None prescribed — the measurement goal is diagnostic. Students bring whatever operations they would bring outside the course.

Base distribution and rationale. Because this task measures natural behavior rather than prescribes it, the base target mirrors the modal pattern we would expect from a population of students approximating untrained novices. Drawing on Wave 1 pilot data and the broader literature on novice AI use (cognitive offloading, Risko & Gilbert, 2016), Passivity modes dominate: {.30, .20, .15, .10, .05, .10, .05, .05}. The student is not graded against this target — the baseline is used only as a reference pattern for comparison across subsequent tasks.

Mode12345678Path
Base3020151051055(none)
Beginner38222093521(none)
Intermediate2520151261255(none)
Expert121610139171112(none)

Notes. Because baseline has no prescribed path, the path-adjustment rules (prepend Tutor for Beginner, Problem Setter for Expert) do not apply. The table is still shown to illustrate what pattern a student of that tier would be expected to produce naturally, for comparison to their actual submission. Scoring for the baseline task is pass/fail on completion; the tier-adjusted distributions are diagnostic, not evaluative.


4.2 task-how-ai-works — Understanding AI (Book 1, Ch 2)

Task. The student runs a calibration test to identify where AI output is reliable versus unreliable, producing a short reflection on what they learned.

Cognitive operations.

  • Understand a claim or explanation AI provides (Bloom: Understand)
  • Apply a test or verification procedure to that claim (Apply, Evaluate)
  • Identify discrepancy or error (Evaluate)
  • Synthesize a personal heuristic for when to trust AI output (Analyze, Create)

Base distribution and rationale. This is a Tutor-plus-Verification task. The student must understand what AI is claiming before they can evaluate it; verification dominates once the claim is understood. Minor Collaborative (Mode 4) work happens when the student reasons together with AI about edge cases. Problem Setter is not useful here — the frame is given ("test AI's reliability"), not to be questioned. Base: {.15, .05, .30, .15, .25, .05, .03, .02}. Path: 3 → 5.

Mode12345678Path
Base1553015255323 → 5
Beginner1964114163103 → 5
Intermediate1252917276323 → 5
Expert531717407648 → 3 → 5

Notes. Expert path prepends Problem Setter even though the base task gives the frame, because expert students should be interrogating which calibration claims are worth testing before spending time testing them. This is the expertise reversal effect applied to meta-cognition (Kalyuga, 2007).


4.3 task-quick-answers — Oracle Audit (Book 1, Ch 3)

Task. The student completes two rounds: Round 1 is natural Oracle-heavy behavior; Round 2 applies an "Oracle Filter" (three-question check before accepting any AI answer). Students produce both transcripts plus a comparison reflection.

Cognitive operations.

  • Ask for factual information (Remember)
  • Evaluate whether the answer is plausible (Evaluate)
  • Verify against another source or own reasoning (Evaluate)
  • Reflect on the shift between rounds (Analyze)

Base distribution and rationale. The task is intentionally split-personality: Round 1 expects heavy Oracle, Round 2 expects heavy Verification. The aggregate base distribution balances across both rounds, with Verification (M5) dominant because the learning happens in Round 2. Light Collaborative (M4) and Critical Challenger (M7) capture students who go beyond the minimum to push back on AI answers. Base: {.15, .10, .15, .15, .25, .05, .10, .05}. Path: 1 → 5 (reflecting the round-to-round shift).

Mode12345678Path
Base151015152551051 → 5
Beginner22122316183423 → 1 → 5
Intermediate121015182761041 → 5
Expert468153561698 → 1 → 5

Notes. Beginner path prepends Tutor because a novice benefits from understanding why the Oracle Filter works before applying it. Expert path prepends Problem Setter because experts should question whether the filter's three canonical questions are even the right ones for their domain.


4.4 task-learning — Concept Deep Dive (Book 1, Ch 4)

Task. The student picks a concept they don't understand and uses AI to build understanding through Socratic dialogue, then produces a one-paragraph teach-back written from memory.

Cognitive operations.

  • Frame what is unknown about the concept (Analyze, meta-level)
  • Receive scaffolded explanation (Understand)
  • Restate concept in own words (Understand, Apply)
  • Self-test at multiple Bloom's levels (Evaluate)
  • Challenge own first interpretation (Evaluate)

Base distribution and rationale. Tutor (M3) dominates because the core work is learning. Collaborative (M4) is heavy because the student must contribute reasoning to avoid passive reception. Verification (M5) is substantial because self-testing is required. Problem Setter (M8) appears modestly because the student is asked to question their framing of the concept before asking AI. Base: {.10, .02, .35, .25, .20, .03, .02, .03}. Path: 8 → 3 → 7 → 5.

Mode12345678Path
Base1023525203238 → 3 → 7 → 5
Beginner1324722132113 → 8 → 3 → 7 → 5
Intermediate823328213238 → 3 → 7 → 5
Expert312029324468 → 3 → 5 → 7 → 5

Notes. Beginner path places Tutor first (default for beginners); Problem Setter becomes accessible only after initial understanding. Expert path is unchanged because Problem Setter was already first. This rubric is re-derivation of the existing concept-deep-dive file in the codebase; the base vector matches.


4.5 task-studying — Study Sprint (Book 1, Ch 5)

Task. Across three spaced sessions, the student uses AI for retrieval practice, spaced review, and interleaving, producing three transcripts plus a memory-change reflection.

Cognitive operations.

  • Attempt recall before receiving information (Remember)
  • Check accuracy of recall (Evaluate)
  • Generate practice problems for self-study (Apply — Tutor per canonical definition)
  • Interleave concepts across sessions (Analyze)
  • Notice and describe forgetting (Evaluate, meta-level)

Base distribution and rationale. Tutor (M3) dominates because retrieval and practice-problem generation both classify as Tutor under the canonical rules. Verification (M5) is substantial because students check their own answers. Problem Setter (M8) is non-trivial because the student must interrogate what they need to study. Minor Production Assistant (M2) for formatting flashcards. Base: {.10, .08, .40, .07, .20, .02, .02, .11}. Path: 3 → 5 → 3 → 5 (retrieval-verification cycles across sessions).

Mode12345678Path
Base1084072022113 → 5 → 3 → 5
Beginner139556131133 → 5
Intermediate884082222103 → 5 → 3 → 5
Expert352383234228 → 3 → 5 → 3 → 5

Notes. Paths with repeated elements are collapsed for tier display (the underlying repetition is preserved in the base path). Expert path prepends Problem Setter because experts should question their knowledge gaps before generating practice material. Re-derivation of existing study-sprint rubric; base vector matches.


4.6 task-exam-prep — Exam Prep Simulation (Book 1, Ch 6)

Task. The student runs a four-stage exam-prep workflow: (1) map territory + self-assess, (2) practice exam, (3) reasoning review, (4) stress test with Critical Challenger. Plus a reflection comparing stress-test insights to practice-exam insights.

Cognitive operations.

  • Self-assess confidence on topics (Evaluate, meta-level)
  • Receive scaffolded practice problems (Apply — Tutor)
  • Attempt problems without answers (Apply)
  • Review reasoning on right and wrong answers (Evaluate)
  • Invite adversarial questions targeting weakest areas (Evaluate)

Base distribution and rationale. Tutor (M3) is dominant for the practice and review stages; Critical Challenger (M7) is heavily weighted for the stress-test stage — this is what distinguishes this task from plain studying. Verification (M5) captures the reasoning-review stage. Low Production Assistant because exam prep should be student-generated reasoning, not AI-generated answers. Base: {.05, .03, .30, .07, .20, .03, .25, .07}. Path: 3 → 5 → 7 → 5.

Mode12345678Path
Base533072032573 → 5 → 7 → 5
Beginner845081621123 → 5 → 7 → 3 → 5
Intermediate432982242463 → 5 → 7 → 5
Expert1214626436128 → 3 → 5 → 7 → 5

Notes. Expert path prepends Problem Setter because experts should interrogate what kind of exam this is before deciding what stress test is appropriate. Re-derivation of existing exam-prep-simulation rubric.


4.7 task-writing — The AI Writing Partner (Book 1, Ch 7)

Task. Write an analytical essay using AI across the full writing workflow: understand the assignment, build the argument, draft with ownership boundaries, critically challenge the draft, and verify citations.

Cognitive operations.

  • Interrogate the assignment frame (Analyze, meta-level)
  • Construct an argument (Create)
  • Draft prose paragraphs (Apply, Create)
  • Revise under adversarial challenge (Evaluate)
  • Verify factual claims and citations (Evaluate)

Base distribution and rationale. Writing is the most mode-diverse task in the book — this is why it is the canonical mode-fluency task. Production Assistant (M2) is substantial because students are supposed to use AI for craft-level revision within boundaries, but it is not dominant. Collaborative (M4) and Critical Challenger (M7) carry equal weight because argument-building and revision are co-dominant phases. Creative Expander (M6) supports divergent thinking about approaches. Verification (M5) is essential given AI's hallucination risk on citations. Problem Setter (M8) is non-trivial because the assignment-framing step is load-bearing. Base: {.05, .15, .05, .20, .15, .15, .15, .10}. Path: 8 → 6 → 4 → 7 → 2 → 5.

Mode12345678Path
Base515520151515108 → 6 → 4 → 7 → 2 → 5
Beginner9229251311743 → 8 → 6 → 3 → 4 → 7 → 3 → 2 → 5
Intermediate41452216171488 → 6 → 4 → 7 → 2 → 5
Expert18217181721168 → 6 → 5 → 4 → 7 → 5 → 2 → 5

Sum note: Expert row sums to 129% before normalization in illustration; the displayed values are post-normalization rounded to whole percents.

Notes. Beginner path prepends Tutor because a first-time essay writer needs to understand the assignment type (synthesis vs. evaluation vs. application) before framing can even happen. Expert distribution heavily shifts away from Production Assistant — expert writers write their own prose and use AI for challenge, not generation. This supports the Five Principles section of the book chapter (writing ownership spectrum).


4.8 task-problem-solving — Problem Sets (Book 1, Ch 8)

Task. The student works through 3+ analytical problems with AI as thinking partner, identifies patterns across problems, and articulates the underlying principle.

Cognitive operations.

  • Attempt problem before asking AI (Apply)
  • Receive guided discovery hints, not answers (Understand, Apply)
  • Iterate reasoning collaboratively (Analyze)
  • Verify solution through alternative method (Evaluate)
  • Identify generalizable principle across problems (Analyze, Create)

Base distribution and rationale. Collaborative Problem-Solver (M4) is dominant because the task is explicitly "think with AI" rather than "get answer from AI." Tutor (M3) is substantial because guided-discovery interactions classify as Tutor. Verification (M5) is essential to check reasoning, not just answers. Low Production because problem work should show the student's reasoning, not AI's solution. Base: {.08, .05, .25, .30, .20, .02, .05, .05}. Path: 8 → 4 → 5 → 7.

Mode12345678Path
Base852530202558 → 4 → 5 → 7
Beginner1163629141223 → 8 → 4 → 3 → 5 → 7
Intermediate652434212548 → 4 → 5 → 7
Expert231332293898 → 4 → 5 → 7

Notes. Beginner path correctly prepends Tutor because a student who does not yet have the problem-type schemas needs instruction before guided discovery works (expertise reversal — Kalyuga et al., 2003).


4.9 task-brainstorming — Reframing & Ideation (Book 1, Ch 9)

Task. Starting from a vague problem, the student generates three distinct problem framings, chooses one with rationale, and produces a range of candidate solutions within the chosen frame.

Cognitive operations.

  • Question the framing of the problem (Analyze, meta-level)
  • Generate divergent framings (Create)
  • Generate divergent solutions within a frame (Create)
  • Stress-test favored ideas (Evaluate)
  • Converge on a direction with reasoning (Evaluate)

Base distribution and rationale. Creative Expander (M6) is dominant because divergent generation across options is the defining operation. Problem Setter (M8) is heavy because three-framing exercises are meta-level work. Critical Challenger (M7) supports the stress-test phase. Base: {.03, .03, .05, .12, .05, .35, .20, .17}. Path: 8 → 6 → 7 → 6 (reframe → generate → challenge → re-expand).

Mode12345678Path
Base3351253520178 → 6 → 7 → 6
Beginner6612185321283 → 8 → 6 → 3 → 7 → 6
Intermediate2351353919148 → 6 → 7 → 6
Expert112953424238 → 6 → 5 → 7 → 6

Notes. Beginner distribution is more balanced than other tasks because a novice brainstormer needs some Tutor-style scaffolding on what divergent thinking looks like before they can produce it independently. Expert version is Problem Setter + Creative + Challenger dominant — this is the shape of expert ideation.


4.10 task-research — Research & Fact-Checking (Book 1, Ch 10)

Task. The student researches a claim or question using AI for source discovery and synthesis, producing a verification log and an annotated research synthesis.

Cognitive operations.

  • Frame research question (Analyze, meta-level)
  • Generate candidate search terms and source types (Create)
  • Verify every source exists and says what AI claims (Evaluate)
  • Stress-test synthesis against counter-evidence (Evaluate)

Base distribution and rationale. Verification (M5) is overwhelmingly dominant because the canonical research workflow (canonical-skill §Mode 5, combined with hallucination risk from canonical-skill §Mode 1) makes verification the load-bearing component. Critical Challenger (M7) is substantial because counter-evidence matters. Creative Expander (M6) supports candidate-source generation. Base: {.05, .03, .05, .10, .35, .10, .25, .07}. Path: 8 → 6 → 5 → 7 → 5.

Mode12345678Path
Base5351035102578 → 6 → 5 → 7 → 5
Beginner10511143581433 → 8 → 6 → 3 → 5 → 7 → 3 → 5
Intermediate4351136112468 → 6 → 5 → 7 → 5
Expert1128371031108 → 6 → 5 → 7 → 5

Notes. Under §3.4 the Expert path now retains both Verifications rather than collapsing them — continuous verification is represented by Mode 5 appearing multiple times in the path, not by truncation. Research task rubric emphasizes citation-existence verification explicitly in the artifact criteria.


4.11 task-decisions — Making Decisions (Book 1, Ch 11)

Task. Apply a 6-stage decision model (Frame → Options → Evaluate → Decide → Commit → Review) with AI, producing a decision matrix and a decision journal entry.

Cognitive operations.

  • Frame decision (Analyze, meta-level)
  • Generate options (Create)
  • Evaluate options against criteria (Evaluate)
  • Stress-test the favored option (Evaluate)
  • Verify critical assumptions (Evaluate)

Base distribution and rationale. Decisions are balanced-mode tasks: Problem Setter (framing) + Creative (options) + Verification (assumption checking) + Critical Challenger (stress-test) all carry weight. Collaborative (M4) captures iterative reasoning on criteria. Base: {.05, .03, .08, .15, .20, .15, .20, .14}. Path: 8 → 6 → 4 → 7 → 5.

Mode12345678Path
Base53815201520148 → 6 → 4 → 7 → 5
Beginner105162119121163 → 8 → 6 → 3 → 4 → 7 → 3 → 5
Intermediate43817211719128 → 6 → 4 → 7 → 5
Expert11312221525208 → 6 → 5 → 4 → 7 → 5

Notes. Grounded in Paul-Elder's framework for critical thinking (intellectual standards) and Kahneman's (2011) dual-system work on decision biases. The Beginner version includes more Tutor because novice decision-makers benefit from explicit frameworks before attempting free-form application.


4.12 task-fitness-check — Capstone Fitness Check (Book 1, Ch 12)

Task. End-of-course free-form assessment where students demonstrate mode fluency across a task of their choosing, producing an AI Practice Statement.

Cognitive operations.

  • Select a task that exercises multiple modes (meta-level)
  • Demonstrate mode-switching in a single conversation (Analyze, Evaluate, Create)
  • Reflect on which modes felt natural and which required effort (Analyze, meta-level)

Base distribution and rationale. This is the capstone equivalent of the baseline. Unlike baseline (which measures untrained behavior), fitness check measures trained behavior — students should now show a balanced, Agency-rich profile. The base distribution is deliberately flat across Partnership and Agency modes to reward breadth over any single concentration. Low but non-zero Passivity because Oracle can still be the right mode for factual look-ups. Base: {.08, .10, .15, .15, .15, .12, .15, .10}. Path: (none — student chooses).

Mode12345678Path
Base810151515121510(none)
Beginner1314261712873(none)
Intermediate61014171614149(none)
Expert2671419142216(none)

Notes. Because this is the capstone, the tier-adjusted targets are not the grading bar — students are graded on breadth (≥ 5 modes represented) and Agency-tier proportion (≥ 40% for pass). The tier-adjusted targets are shown as developmental references.


4.13 task-communication — Communication & Email (Book 2, Ch 13)

Task. Draft, tone-tune, and triage professional communications (emails, messages, meeting notes) using AI while maintaining authentic voice.

Cognitive operations.

  • Understand communication goal (Understand)
  • Generate draft (Create)
  • Tune tone for audience (Apply, Evaluate)
  • Verify content accuracy (Evaluate)

Base distribution and rationale. Production Assistant (M2) is dominant because AI drafting is the productive core. Tutor (M3) appears when students ask for communication-style explanations. Critical Challenger is low — you stress-test decisions, not emails. Base: {.08, .35, .08, .15, .15, .10, .05, .04}. Path: 8 → 2 → 5.

Mode12345678Path
Base8358151510548 → 2 → 5
Beginner11421215106213 → 8 → 2 → 3 → 5
Intermediate6348171612538 → 2 → 5
Expert3234162314988 → 2 → 5

Notes. The Production Assistant weight reflects that communication tasks are legitimate AI-drafting domains provided the student has authentic input and voice-check discipline.


4.14 task-data — Data & Spreadsheets (Book 2, Ch 14)

Task. Clean, analyze, and communicate findings from a dataset using AI for code, formula, and visualization assistance.

Cognitive operations.

  • Understand dataset and question (Understand)
  • Plan analysis approach (Analyze)
  • Generate code/formulas (Create — Production)
  • Debug and verify (Evaluate)
  • Interpret and communicate findings (Analyze, Create)

Base distribution and rationale. Balanced across Tutor (M3, for understanding statistical concepts), Production (M2, for code/formula generation), Collaborative (M4, for iterative analysis), and Verification (M5, for checking results). Base: {.08, .20, .20, .18, .20, .05, .05, .04}. Path: 8 → 3 → 4 → 2 → 5.

Mode12345678Path
Base8202018205548 → 3 → 4 → 2 → 5
Beginner11232917143213 → 8 → 3 → 4 → 2 → 3 → 5
Intermediate6191921216538 → 3 → 4 → 2 → 5
Expert3131120307988 → 3 → 5 → 4 → 2 → 5

Notes. Unlike writing, data work legitimately involves heavy Production because code-generation is a well-defined task. Verification is critical because AI-generated data code routinely has subtle bugs (wrong join logic, off-by-one indexing).


4.15 task-presentations — Presentations & Speaking (Book 2, Ch 15)

Task. Design a presentation with outline, speaker notes, slides, and Q&A prep. Use AI for structure, visual concepts, and rehearsal.

Cognitive operations.

  • Frame presentation goal and audience (Analyze, meta-level)
  • Generate slide content and visual ideas (Create)
  • Draft speaker notes (Create — Production)
  • Stress-test likely questions (Evaluate)

Base distribution and rationale. Production (M2) and Creative (M6) are both substantial — presentations benefit from AI-generated drafts AND from diverse visual/framing options. Critical Challenger appears for Q&A prep. Base: {.05, .25, .10, .15, .15, .15, .10, .05}. Path: 8 → 6 → 4 → 2 → 7 → 5.

Mode12345678Path
Base525101515151058 → 6 → 4 → 2 → 7 → 5
Beginner83316161110423 → 8 → 6 → 3 → 4 → 2 → 3 → 7 → 5
Intermediate4249171617948 → 6 → 4 → 2 → 7 → 5
Expert21551521191698 → 6 → 5 → 4 → 2 → 5 → 7 → 5

4.16 task-career — Career Acceleration (Book 2, Ch 16)

Task. Build or revise resume, cover letter, and interview prep document with AI.

Cognitive operations.

  • Frame target role and differentiators (Analyze, meta-level)
  • Iterate on self-narrative with AI (Analyze, Evaluate)
  • Generate drafts (Create — Production)
  • Stress-test answers to likely interview questions (Evaluate)

Base distribution and rationale. Production (M2) is substantial because resume-drafting is a legitimate Production use. Collaborative (M4) carries the self-narrative work. Critical Challenger for interview prep. Verification ensures job-search facts (company info, salary data) are correct. Base: {.10, .25, .10, .15, .15, .10, .08, .07}. Path: 8 → 4 → 2 → 7 → 5.

Mode12345678Path
Base102510151510878 → 4 → 2 → 7 → 5
Beginner15311616116323 → 8 → 4 → 3 → 2 → 7 → 3 → 5
Intermediate82410171612868 → 4 → 2 → 7 → 5
Expert315515221313138 → 4 → 5 → 2 → 7 → 5

4.17 task-finance — Personal Finance & Life Optimization (Book 2, Ch 17)

Task. Build a 90-day financial plan using AI for research, calculation, and decision-journal entries.

Cognitive operations.

  • Frame financial goal (Analyze, meta-level)
  • Understand relevant concepts (Understand — Tutor)
  • Generate options (Create)
  • Calculate scenarios (Apply)
  • Verify claims and numbers (Evaluate)

Base distribution and rationale. Tutor dominates because personal-finance tasks usually require learning (what a Roth IRA is, how compound interest works) before acting. Verification is essential because AI financial advice has hallucination risk on interest rates, tax rules, etc. Base: {.12, .08, .20, .15, .20, .10, .10, .05}. Path: 8 → 3 → 4 → 5 → 7.

Mode12345678Path
Base128201520101058 → 3 → 4 → 5 → 7
Beginner17103116146423 → 8 → 3 → 4 → 5 → 3 → 7
Intermediate98191721121048 → 3 → 4 → 5 → 7
Expert45101528131698 → 3 → 5 → 4 → 5 → 7

Notes. Domain-specific warning: financial information from AI must be verified against authoritative sources (IRS, SEC, institution documentation). The verification weight is conservative.


4.18 task-collaboration — Team Collaboration (Book 2, Ch 18)

Task. Build a team charter and meeting workflow document using AI for structure and meeting-facilitation support.

Cognitive operations.

  • Frame team context (Analyze)
  • Generate charter options (Create)
  • Iterate collaboratively on meeting structure (Analyze, Create)
  • Draft templates (Create — Production)

Base distribution and rationale. Collaborative (M4) is dominant because team-related tasks inherently involve bidirectional iteration. Creative Expander supports option generation. Production for templates. Base: {.05, .15, .10, .25, .15, .15, .10, .05}. Path: 8 → 4 → 6 → 2 → 5.

Mode12345678Path
Base515102515151058 → 4 → 6 → 2 → 5
Beginner82017281210423 → 8 → 4 → 3 → 6 → 2 → 3 → 5
Intermediate4149281517948 → 4 → 6 → 2 → 5
Expert1852420181588 → 4 → 5 → 6 → 2 → 5

4.19 task-social-media — Social Media & Content (Book 2, Ch 19)

Task. Build brand voice guide and content calendar using AI for ideation and draft generation.

Cognitive operations.

  • Frame audience and brand voice (Analyze)
  • Generate post ideas (Create)
  • Draft content (Create — Production)
  • Verify factual claims in content (Evaluate)

Base distribution and rationale. Production (M2) and Creative (M6) both substantial. Creative reflects ideation load; Production reflects draft-generation load. Verification is essential because public-facing content is high-stakes if wrong. Base: {.05, .30, .08, .12, .15, .20, .05, .05}. Path: 8 → 6 → 2 → 5.

Mode12345678Path
Base5308121520558 → 6 → 2 → 5
Beginner83913131113223 → 8 → 6 → 3 → 2 → 5
Intermediate4288141623548 → 6 → 2 → 5
Expert2184122126898 → 6 → 5 → 2 → 5

4.20 task-coding — Coding & Technical (Book 2, Ch 20)

Task. Build or modify code artifacts using AI for design, writing, debugging, and reasoning-log documentation.

Cognitive operations.

  • Design approach (Analyze, Create)
  • Understand unfamiliar libraries/patterns (Understand — Tutor)
  • Generate code (Create — Production)
  • Debug with AI assistance (Analyze, Evaluate)
  • Verify behavior (Evaluate)

Base distribution and rationale. Production (M2) is substantial because code generation is a legitimate use. Tutor is heavy because coding tasks often require learning APIs. Verification is essential because AI-generated code has known hallucination patterns (non-existent functions, wrong API signatures). Base: {.08, .25, .18, .15, .20, .05, .05, .04}. Path: 8 → 3 → 4 → 2 → 5.

Mode12345678Path
Base8251815205548 → 3 → 4 → 2 → 5
Beginner11292614143213 → 8 → 3 → 4 → 2 → 3 → 5
Intermediate6241717216548 → 3 → 4 → 2 → 5
Expert3161016317988 → 3 → 5 → 4 → 2 → 5

Notes. Verification weight is deliberately conservative; mature coding practice with AI requires continuous testing, not periodic checks. The canonical-skill classifier correctly routes code-debugging messages to Collaborative or Verification depending on artifact presence.


4.21 task-entrepreneurship — Entrepreneurship & Side Projects (Book 2, Ch 21)

Task. Build business model canvas, validation plan, and iterate on venture concept using AI.

Cognitive operations.

  • Frame opportunity (Analyze, meta-level)
  • Generate business-model options (Create)
  • Iterate via customer-feedback reasoning (Analyze)
  • Stress-test assumptions (Evaluate)
  • Verify market/competitor claims (Evaluate)

Base distribution and rationale. Balanced across Creative (options), Collaborative (iteration), Critical Challenger (assumption testing), and Problem Setter (opportunity framing). Production is modest because the canvas itself is small-format. Base: {.05, .10, .12, .15, .13, .15, .15, .15}. Path: 8 → 6 → 4 → 7 → 5.

Mode12345678Path
Base5101215131515158 → 6 → 4 → 7 → 5
Beginner91522191111763 → 8 → 6 → 3 → 4 → 7 → 3 → 5
Intermediate4101217141714138 → 6 → 4 → 7 → 5
Expert15513161720238 → 6 → 5 → 4 → 7 → 5

4.22 task-health — Health Research & Decisions (Book 2, Ch 22)

Task. Research a health-related decision using AI for literature synthesis and personal-context reasoning. Produce decision document + verification log.

Cognitive operations.

  • Frame health question (Analyze)
  • Understand relevant concepts (Understand — Tutor)
  • Verify claims against authoritative sources (Evaluate)
  • Stress-test advice (Evaluate)

Base distribution and rationale. Verification (M5) is dominant and disproportionately weighted because health stakes are high and AI hallucinations carry real-world harm. Tutor supports understanding. Critical Challenger supports stress-testing. Base: {.08, .05, .20, .15, .30, .08, .08, .06}. Path: 8 → 3 → 5 → 7.

Mode12345678Path
Base852015308868 → 3 → 5 → 7
Beginner1263216235323 → 8 → 3 → 5 → 7
Intermediate651917319858 → 3 → 5 → 7
Expert23914391012108 → 3 → 5 → 7

Notes. This task carries an explicit academic/medical caveat: AI is not a substitute for licensed medical advice, and the rubric reinforces verification against authoritative health sources (peer-reviewed literature, CDC/NIH, etc.). Artifact criteria require evidence of source-verification.


4.23 task-ethics — Ethics & AI Thinking (Book 2, Ch 23)

Task. Analyze an ethics case using AI as a thinking partner, applying an ethics framework and stress-testing conclusions.

Cognitive operations.

  • Frame the ethical question (Analyze, meta-level)
  • Apply an ethical framework (Apply, Analyze)
  • Collaborate to weigh competing considerations (Analyze, Evaluate)
  • Stress-test position (Evaluate)

Base distribution and rationale. Collaborative (M4) and Critical Challenger (M7) are co-dominant because ethical reasoning benefits from both iteration and adversarial check. Problem Setter is heavy because framing an ethics case correctly is half the work. Base: {.03, .05, .15, .20, .15, .10, .20, .12}. Path: 8 → 4 → 7 → 5.

Mode12345678Path
Base351520151020128 → 4 → 7 → 5
Beginner5728251371043 → 8 → 4 → 3 → 7 → 5
Intermediate251423161119108 → 4 → 7 → 5
Expert12617181127188 → 4 → 5 → 7 → 5

4.24 task-ai-agents — AI Agents & Automation (Book 2, Ch 24)

Task. Design and specify an AI agent or automation workflow, with documentation and monitoring plan.

Cognitive operations.

  • Frame automation goal (Analyze, meta-level)
  • Understand agent capabilities (Understand — Tutor)
  • Iterate design collaboratively (Analyze, Create)
  • Draft specification (Create — Production)
  • Verify agent behavior (Evaluate)

Base distribution and rationale. Production (for specifications), Tutor (for learning agent concepts), Collaborative (for design iteration), Verification (for validating agent behavior). Problem Setter to frame the right automation opportunity. Base: {.05, .20, .15, .20, .15, .10, .08, .07}. Path: 8 → 3 → 4 → 2 → 5.

Mode12345678Path
Base52015201510878 → 3 → 4 → 2 → 5
Beginner7252421116323 → 8 → 3 → 4 → 2 → 3 → 5
Intermediate41914231611868 → 3 → 4 → 2 → 5
Expert212820211313128 → 3 → 5 → 4 → 2 → 5

4.25 task-casual — Casual & Personal (catch-all)

Task. Open-ended personal use (recipes, travel planning, hobby help, daily-life questions).

Cognitive operations. Whatever the user brings — no task-prescribed operations.

Base distribution and rationale. Like baseline, this task's "target" is a reference pattern, not a goal. Casual use is Oracle-dominant because daily questions are mostly factual ("what should I cook with these ingredients"). Low Agency-tier because stakes are low and cognitive load is low. Base: {.40, .25, .10, .10, .05, .05, .03, .02}. Path: (none).

Mode12345678Path
Base402510105532(none)
Beginner47251383210(none)
Intermediate342610136632(none)
Expert1924816111086(none)

Notes. Casual and work tasks are not graded against targets — the distribution shown is descriptive. Students may see how their casual-use pattern compares to others', but the comparison is informational.


4.26 task-work — Work & Professional (catch-all)

Task. Open-ended professional use outside the 24 structured task types (meeting prep, internal comms, project scoping, etc.).

Cognitive operations. Variable across professional domains.

Base distribution and rationale. Professional use is Production-dominant (drafting, summarizing) but with more Collaborative and Verification than casual because work has higher accuracy stakes. Base: {.20, .30, .10, .15, .10, .05, .05, .05}. Path: (none).

Mode12345678Path
Base2030101510555(none)
Beginner263314146321(none)
Intermediate1630101811654(none)
Expert7226181781011(none)

Notes. Like casual, work is descriptive rather than prescriptive. Students may still submit work tasks for pattern analysis and track their distribution over time.


4.27 task-path-design — Path Design Practice (Module 3 capstone)

Task. After learning the eight modes in Module 3, students pick a real task they're facing — academic, professional, or personal — and design a mode path for it before opening the conversation. They specify which modes they plan to use, in what order, and why. They run the conversation following their designed path and submit transcript + a short reflection on whether the path survived contact with the actual task.

Cognitive operations. Frame the task → Select modes → Sequence modes into a coherent path → Execute the path → Reflect on path-vs-reality divergence. The dominant Bloom level is Evaluate on the front and back ends and Create in the middle.

Why this task is different. Every other task in the spec has a prescribed mode_path for at least one tier — the rubric scores execution against the prescribed sequence. Path-design is the only task where mode_path is intentionally empty across all three tiers because the student is the path designer. The rubric grades whether their self-designed path was coherent and whether their conversation followed it.

Base distribution and rationale. Module 3 introduces the eight modes; the path-design exercise is meant to demonstrate deliberate, varied mode use. Base distribution: {.10, .15, .10, .10, .15, .15, .15, .10} — flat-ish but with mid-Partnership/Agency tilt (M5/6/7 emphasized over M1/3/4/8) because path design rewards reflective, synthetic, and collaborative engagement. Path: (student-designed).

Mode12345678Path
Base1015101015151510(student-designed)
Beginner16211812121174(student-designed)
Intermediate81510121617159(student-designed)
Expert385919182216(student-designed)

Artifact. Required. The artifact is the design document plus the reflection — not the transcript. Sonnet-graded. Three criteria:

  • Design coherence (40%) — does the planned path fit the chosen task? Are mode choices justified? Is the sequence intentional?
  • Execution fidelity (30%) — did the student follow their designed path? Deviations are not penalized but must be noticed and accounted for in reflection.
  • Reflection quality (30%) — how thoughtful is the post-hoc analysis of where the path held vs. broke?

Notes. Path-design is the only task whose rubric grading weight is dominated by the artifact rather than the engagement distribution; the engagement distribution is descriptive (what the student naturally produced) rather than the primary scoring axis. Beginners tend toward more M1/2/3 (information-seeking, verifying, tutoring) — they apply the modes most legible to a novice. Experts tend toward more M5/7/8 (challenging, collaborating, problem-setting) — modes that require established mode fluency. Both can score well if their design matches their task; tier patterns describe natural production, not quality.

This task corresponds to the new Module 3 ("The AI Engagement Toolkit") inserted into the Learn-with-AI architecture per decisions.md 2026-04-25 §A. When the broader chapter renumber sweep happens (Module 3 chapter slot is now occupied by the toolkit; old Module 3 → new Module 4, etc.), this section may move to §4.3 with the rest of the §4.x sequence shifting; the rubric content here is stable regardless of section-number assignment.


4.28 Cross-task observations

Reading across all 26 task-tier rubrics, a few systematic patterns emerge:

  • Beginners concentrate on 3–4 modes per task. Across tasks, Beginner distributions rarely have more than 4 modes above 10%. This is consistent with Dreyfus & Dreyfus (1980) — novices operate on explicit scaffolds, not full repertoires.
  • Experts are mode-diverse. Expert distributions show more modes above 10%, often 6–7, because experts adaptively deploy whichever mode fits the moment. This is the empirical signature of the mode-fluency construct.
  • Verification (M5) grows with expertise on every single task. The multiplier (1.4 for Expert) combines with the fact that Verification's base proportion is already substantial on most tasks. This is the single most robust pattern in the rubric system.
  • Production Assistant (M2) is expertise-agnostic for tasks that legitimately require it (data, coding, communication, career). For tasks where AI drafting undermines the learning objective (writing, research), Production declines sharply with expertise.
  • Problem Setter (M8) grows fastest with expertise. The multiplier is 1.8 — the steepest in the framework. This reflects the theoretical claim that meta-level problem reframing is the defining expert cognitive move.

These patterns are predictions, not facts. The ML pipeline (§6) is designed to test whether observed distributions from high-performing students actually match these tier-adjusted targets. Persistent mismatches are the trigger for methodology revision.


5. Validation protocol

The rubrics produced by the methodology above are theoretical. Validation proceeds at three levels.

5.1 Construct validity of the modes themselves

The eight modes are a taxonomy of user-AI interaction. Their construct validity rests on factor structure: do survey items claiming to measure different modes cluster into the expected factors? Do observed behavioral classifications form the hypothesized tier structure?

Wave 1 survey data (N=331, collected 2026-01 through 2026-03 from student samples only) shows:

  • Eight-factor structure for the AIT self-report items holds at exploratory factor analysis (details in 3-Research/AI engagement/).
  • Tier structure validates: Oracle correlates negatively with Mode 5–8 items (r = −0.16 to −0.35); Modes 5–8 cluster (inter-correlations r = 0.58–0.70).

This establishes that the modes and tiers are empirically distinguishable constructs. It does not yet establish that any particular target distribution is optimal — that requires conversation-level data.

5.2 Classifier reliability

The classifier is an LLM (Haiku 4.5) applying the canonical mode definitions to messages. Its reliability is measured by Cohen's kappa against human coders who have calibrated to the canonical guide.

Target: κ ≥ 0.80 on a held-out calibration set of 200 messages spanning all 8 modes. As of 2026-04-22, this calibration set does not yet exist; it is the critical deliverable blocking the MyEducator integration (see Grading Spec CHANGELOG §0.2.0).

Procedure: Two independent coders code the set after reading the canonical guide. Disagreements are adjudicated by a third coder. The classifier is then run on the same set; κ against the adjudicated labels is the reliability estimate. Classifications where the classifier's confidence is in [0.50, 0.70] are re-run through a Sonnet-class model; both labels are stored.

5.3 Rubric validity (the hard part)

A rubric is valid if students who perform well on the task — as judged by an independent grader — also score well on the rubric, and vice versa. This is a predictive-validity claim that can only be tested with real student data.

Validation will be staged:

  • Stage A (Months 1–3): Collect conversations, classifications, rubric scores, and independent artifact grades. Compute the correlation between track-adjusted rubric scores and artifact grades. Expected: r > 0.40 at the task level, indicating the rubric captures task-relevant variance. Lower correlations indicate the rubric is miscalibrated for that task.

  • Stage B (Months 3–6): For each task where Stage A correlation is below 0.40, examine which rubric components underperform. Revise the methodology (Section 2 Step 3 weights, or Section 3 multipliers) for those tasks. Re-derive the rubric, bump version, and re-run Stage A.

  • Stage C (Months 6–12): For each task where the revised rubric still underperforms, treat the theoretical target as provisional and move the task to the empirical calibration track described in Section 6. This is the fallback for tasks where theory is insufficient and data must dominate.

  • Stage D (Ongoing): Wave 2 of the AAEV instrument (planned 2026 Q3) collects self-report plus behavioral alignment from the same students. This enables a cross-method validity check: do students who report high Agency-tier self-efficacy actually exhibit it behaviorally? Disagreements are evidence of measurement problems in one or both channels.

5.4 Research paper tie-in

The research program reports these validations as part of the Methods and Results sections. Each paper carries the rubric version it used; a rubric revision triggers a re-analysis for any paper whose data predates the revision.


6. Continuous revision — the ML pipeline

Theory produces the first rubric. Data produces every rubric after that. This section specifies the ML pipeline that converts aimodes's operational data stream into periodic, rigorous rubric revisions.

6.1 Data sources

Every submission writes the following to the engagement_reports and submissions tables in Supabase:

  • Raw transcript
  • Per-message mode classifications with confidence scores
  • Mode-distribution vector (8 values, summing to 1.0)
  • Tier-distribution vector (3 values)
  • 6-component objective score breakdown
  • Track-adjusted score
  • Task type, declared expertise level, user archetype
  • Artifact grade (where applicable; currently null for most assignments)
  • Instructor grade (from MyEducator gradebook where it integrates; currently unavailable)
  • Timestamps, user ID, assignment ID

The pipeline augments this with derived quantities: transition matrices per conversation, convergence of the user's archetype over time, and mode-path adherence per assignment.

6.2 Pipeline stages

  1. Extraction job (daily). Pulls newly-inserted engagement_reports rows from the last 24 hours into a read-only analytics schema. Excludes rows flagged artifact_needs_review = true or where the submission is still status = pending.

  2. Per-task aggregation (weekly). For each task type × self-reported expertise level, computes:

    • Empirical mean mode distribution across all submissions
    • Variance of each mode's proportion
    • Empirical modal mode path (most common first-appearance order across the top-quartile of submissions by artifact grade, where artifact grades exist)
    • Sample size n
  3. Divergence detection (weekly). For each task × tier cell with n ≥ 50, computes the L1 distance between the current theoretical target and the empirical distribution of top-quartile students (by artifact grade where available, by track score as fallback). Cells where L1 > 0.30 are flagged for review.

  4. Candidate revision (monthly). For flagged cells, a candidate revised target is proposed as 0.6 × theoretical + 0.4 × empirical_top_quartile. The weight mix (60/40) is conservative in v0.1; as sample sizes grow and the ML pipeline gains confidence, the weighting can shift toward empirical. Candidates are never auto-applied — they are written to a review queue.

  5. Human review (monthly). Keith, Wood, Posey review candidate revisions. Each candidate must pass three checks:

    • Theoretical plausibility. Does the revision make sense under the Section 2 methodology? If not, the methodology is flagged for revision rather than the rubric.
    • Sample size adequacy. Is n ≥ 50 per cell robust enough for the variance observed?
    • Stability across cohorts. Do cohorts from different institutions produce consistent divergences, or is the divergence specific to one cohort (possibly reflecting cohort composition rather than true rubric miscalibration)?
  6. Version bump and rollout. Approved revisions bump the rubric's minor version. New submissions score against the new rubric; historical submissions retain their original version tag (per the canonical skill file's versioning rule — no retroactive rescoring). The new version propagates to aimodes-grading-spec/ and, via the diff tool, to MyEducator.

  7. Longitudinal re-validation (per semester). Every semester, the full set of rubrics is re-run against a freshly-collected test set of ≥ 500 conversations. Cohen's kappa against human-adjudicated labels is recomputed. If kappa drops below 0.80 on any task, the classifier and/or canonical mode definitions are investigated — this signals either drift in the classifier model (new Haiku version) or drift in student behavior beyond the rubric's parameters.

6.3 ML models in the pipeline

The pipeline uses four ML components:

  • Top-quartile identification (supervised). When artifact grades are available, a gradient-boosted regressor predicts artifact grade from the 6-component rubric score. Top quartile is defined by predicted grade (robust to grader inconsistency) rather than raw artifact grade. When artifact grades are absent, top quartile falls back to raw track score.

  • Expertise inference (supervised). The predictor.ts module already predicts the user's expertise level from their mode distribution. This prediction is compared to self-report to identify mis-calibrated users whose self-reported tier diverges from their behavioral pattern. Such users are excluded from per-tier aggregation until the prediction stabilizes.

  • Archetype evolution (unsupervised). A Markov model fits transitions between the 6 archetypes (Delegator → Partner → Verifier → Creator → Challenger → Architect) across a user's sequence of submissions. Stable archetype estimates inform the per-user tier prior. Outside the rubric-revision pipeline, this model also feeds user-facing longitudinal trajectory visualization.

  • Drift detection (unsupervised). Changes in task-level distributions over time are monitored with a CUSUM control chart. Sudden shifts signal either instructor-driven changes in how the task is framed in the book, classifier drift (new model version), or cohort composition change. CUSUM alerts go to the monthly review queue.

6.4 Research integrity

The ML pipeline is not an auto-tuning system that optimizes rubrics for student grade correlation. It is a candidate-proposal system that surfaces divergences; every revision is human-approved with explicit theoretical justification. This matters because auto-tuned rubrics drift toward post-hoc fitting — rewarding whatever students actually do rather than what theory says they should do. The pipeline's role is to surface evidence that theory is miscalibrated, not to replace theory with curve-fitting.

This distinction is critical for the research program. Papers that cite aimodes rubrics must be able to say which version was in effect during data collection and what theoretical justification underwrote that version. Auto-tuned rubrics break that traceability.


7. References

Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom's taxonomy of educational objectives. Longman.

Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55–81. https://doi.org/10.1016/0010-0285(73)90004-2

Dreyfus, S. E., & Dreyfus, H. L. (1980). A five-stage model of the mental activities involved in directed skill acquisition (Tech. Rep. ORC-80-2). University of California, Berkeley Operations Research Center.

Edwards, J. R. (1991). Person-job fit: A conceptual integration, literature review, and methodological critique. In C. L. Cooper & I. T. Robertson (Eds.), International review of industrial and organizational psychology (Vol. 6, pp. 283–357). Wiley.

Ericsson, K. A. (2006). The influence of experience and deliberate practice on the development of superior expert performance. In K. A. Ericsson, N. Charness, P. J. Feltovich, & R. R. Hoffman (Eds.), The Cambridge handbook of expertise and expert performance (pp. 683–703). Cambridge University Press.

Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3), 363–406. https://doi.org/10.1037/0033-295X.100.3.363

Evans, J. S. B. T., & Stanovich, K. E. (2013). Dual-process theories of higher cognition: Advancing the debate. Perspectives on Psychological Science, 8(3), 223–241. https://doi.org/10.1177/1745691612460685

Hammond, K. R. (1988). Judgment and decision making in dynamic tasks. Information and Decision Technologies, 14(1), 3–14.

Hammond, K. R. (1996). Human judgment and social policy: Irreducible uncertainty, inevitable error, unavoidable injustice. Oxford University Press.

Kahneman, D. (2011). Thinking, fast and slow. Farrar, Straus and Giroux.

Kalyuga, S. (2007). Expertise reversal effect and its implications for learner-tailored instruction. Educational Psychology Review, 19(4), 509–539. https://doi.org/10.1007/s10648-007-9054-3

Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31. https://doi.org/10.1207/S15326985EP3801_4

Keith, M., Wood, D., & Posey, C. (working paper). Two cognitive channels in human–AI engagement: calibration governs behavior, metacognition governs perception.

Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055

Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. https://doi.org/10.1016/j.tics.2016.07.002

Schoenfeld, A. H. (1985). Mathematical problem solving. Academic Press.

Sperber, D., Clément, F., Heintz, C., Mascaro, O., Mercier, H., Origgi, G., & Wilson, D. (2010). Epistemic vigilance. Mind & Language, 25(4), 359–393. https://doi.org/10.1111/j.1468-0017.2010.01394.x

Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4

Sweller, J. (1994). Cognitive load theory, learning difficulty, and instructional design. Learning and Instruction, 4(4), 295–312. https://doi.org/10.1016/0959-4752(94)90003-5

Sweller, J., van Merriënboer, J. J. G., & Paas, F. G. W. C. (1998). Cognitive architecture and instructional design. Educational Psychology Review, 10(3), 251–296. https://doi.org/10.1023/A:1022193728205

Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes (M. Cole, V. John-Steiner, S. Scribner, & E. Souberman, Eds.). Harvard University Press.

Zimmerman, B. J. (2002). Becoming a self-regulated learner: An overview. Theory into Practice, 41(2), 64–70. https://doi.org/10.1207/s15430421tip4102_2


Version history

0.1 (2026-04-22) — initial draft

  • Sections 1–3, 5, 6, 7 drafted as methodology anchor for the grading spec package.
  • Section 4 (per-task derivations for 26 task types × 3 tiers = 78 rubrics) deferred pending author review of the methodology.
  • ML pipeline described at stage level; implementation-level specs (job schedules, data schemas, alert thresholds) deferred to a separate engineering document once the pipeline is prioritized for build.
  • Publishes at /research/rubric-theory on aimodes.ai; source-of-truth for the grading spec package at /6-Apps/aimodes-grading-spec/.

AIEM Mode Coding Standard

Protocol version 10. The operative rule set applied when aimodes.ai assigns AI Engagement Model modes to conversational turns. Published in full so that any score it produces can be checked against the rules that produced it.

How to cite this document

Keith, M. J. (2026). AIEM mode coding standard (Protocol version 10). https://aimodes.ai/research/rubric-theory

Cite the coding standard for how conversations were classified. Cite the preprint below for the typology itself, its measurement properties, and its empirical associations. The two license different claims, and one does not stand in for the other.

Keith, M. J., Wood, D. A., & Posey, C. (2026). The Eight-Mode AI Engagement Typology: Differential Cognitive Signatures and a Self-Report–Behavior Gap. PsyArXiv. https://doi.org/10.31234/osf.io/53pwv_v1

The preprint is under review at Scientific Reports and has not been peer reviewed. Anyone citing it should say so.

Reading this document

Three things about its status, because they change how the rules should be read and none of them is obvious from the text.

"Operative" and "canon" are not the same thing here, and the gap is deliberate. Several rules below are marked PROPOSED. That does not mean rejected, provisional in quality, or awaiting evidence. It means they were implemented in the instrument before the canonical package was formally amended, because production was actively producing measurements known to be wrong and waiting would have meant knowingly collecting more of them. The governance record in §8 exists so the canon leads rather than trails. A rule marked PROPOSED is a rule that ran.

§8 is not an appendix. Rulings dated 2026-08-17 and later are recorded in the §8 table but their body text has not yet been written into §§1–7. Until that pass runs, §8 is the authority for those rules and the body is silent on them. A coder who reads §§1–7 straight through and stops will code against an incomplete rule set. The document says this about itself; it is repeated here because it is the single most consequential thing to know before using it.

Three sections are adopted but not yet implemented: §4.3b (sequence-dependent coding), §4.3c (adjudicated boundary cases), and §5a (confidence and escalation). Hand-coding a corpus from this document will therefore not exactly reproduce what the automated coder produces. For work that compares human coding against aimodes.ai output, that difference is the first thing to control for.

On version numbers. Three schemes run in parallel and are easy to confuse. The protocol version (10, above) is stamped onto every stored classification and is what makes a score traceable. The instrument version (CS-2.x, see §0a) tracks the rule set's own lineage. The canonical package version (2.0.0) tracks the wider document set. When reporting results, the protocol version is the one that matters: it identifies the exact rules a given score was produced under.


Protocol version: 10 (§4.3b, §4.3c and §5a are ADOPTED but NOT YET IMPLEMENTED — see §8) Instrument version: CS-2.1 (see §0a). Canon package release 2.0.0; 2.0.1 / 2.1.0 pending.

Status: OPERATIVE in the aimodes.ai instrument. Rules 1–4 and the initiation property are PROPOSED to the canon and awaiting Mark Keith's decision (IMP-2026-007, IMP-2026-008). Sequence dependence, the C1–C5 adjudications and the confidence/escalation rule are APPROVED (IMP-2026-015, IMP-2026-016) and land in releases 2.0.1 / 2.1.0, which have not yet been cut. See §8. Implementing package version: 2.0.0 Last revised: 2026-09-02


0. What this document is, and which document to read

This is the operative coding standard: the complete set of rules actually applied when assigning AIEM modes to conversational turns. It is written to be self-contained. An analyst or agent rescoring a corpus should be able to work from this file alone, without reconstructing any prior discussion.

If you needRead
To code a corpusThis file. It is sufficient on its own.
What the eight modes mean in theoryTHEORY_CONTRACT.md
The full specification, decision trees, exemplarsMODE_CLASSIFICATION_GUIDE.md
Why a rule says what it saysevidence/intake/2026-08-15-*.md
How the rules were derived and what is still unprovenevidence/reviews/2026-08-15-instrument-development-research-record.md

This file and src/lib/engagement/coding-protocol.ts in aimodes-site are two renderings of one protocol. If they disagree, that is a defect in itself — report it rather than choosing one.

Why a separate operative file exists

Between 2026-08-14 and 2026-08-15 the protocol lived only as prose pasted into model prompts. Four copies existed and all four had drifted. One consumer — the five-model reliability panel — had no copy at all and ran on the bare mode definitions, which invalidated every agreement statistic computed against it (§7.2). Nothing detected this for as long as no single artifact was authoritative. This file is that artifact.


0a. Version scheme — CS-x.y, and why it is not the package version

Decided by Mark Keith, 2026-08-31. The instrument now carries its own human-facing version, because "protocol 8" is meaningful only inside this repository and is the wrong thing to quote to a collaborator.

The label is prefixed CS- and it is NOT the AIEM Canonical Package release number. The package is separately at 2.0.0 with 2.0.1 / 2.1.0 pending. Both would otherwise read "2", describing different things, and this programme has already paid for one silent numbering collision (migration 159 renumbered to 161 after losing a race; three impact IDs claimed twice on 2026-08-19/21). The prefix makes the collision impossible rather than unlikely.

LabelWhat it isData
CS-0Pre-versioning. Not a version — the protocol lived as four hand-maintained prose copies that had drifted (IMP-2026-010), so it was not constant across the period.2,036 moves / 204 submissions, 2026-03-25 → 2026-08-11. Not comparable to anything later; mode 0 did not exist, so all eight mode shares are inflated.
CS-1.1 … CS-1.8The stamped protocol_version integers 1–8, one sub-version each. CS-1.1 is the bare mode definitions the original panel ran; CS-1.8 is the current operative prompt.Panel corpus is CS-1.4 (484 turns). Current instrument rows are CS-1.8 (351 moves / 51 submissions).
CS-2.0The current instrument, stamped 2026-08-31. CS-1.8's prompt, plus every ruling through 2026-08-31 recorded in §8, the corrected attribution rule in the guide, the accuracy standard at §8a, and the falsification re-run paid and reported at §8 item 5.Panel re-run: 497 turns at protocol 8.

The stored protocol_version integers are never rewritten. CS-1.n is a label laid over protocol_version = n, not a replacement for it. That field exists so a score can always be traced to the exact rules that produced it — the one thing that made the three conclusion reversals reconstructable after the fact — and renumbering stored data would destroy exactly that.

What this means for anyone quoting a version today: the instrument is CS-2.0, and the underlying prompt is protocol_version = 8 (CS-1.8's block, unchanged — CS-2.0 is a documentation and governance milestone, not a prompt change, so no stored row needs rewriting).

One honest limit on the CS-2.0 stamp: eleven rulings from 2026-08-17 onward are recorded in §8 but their body text is not yet written into §§1–7. §8 is the authority for those rules until the propagation pass writes them through. An implementer must read §8, not only the body.


1. Units of analysis

There are two units, and keeping them apart is what lets a long multi-request turn be measured correctly (Krippendorff, 2013, on the distinction between context and recording units).

UnitDefinitionCarries
Turn (context unit)One uninterrupted contribution by the participant, however longattribution (§2), initiation (§5), response_to_offer (§3), sequence position
Engagement move (coding unit)One distinct act directed at the AI — a request, question, challenge, correction, or checkthe mode

Most turns contain exactly one move, and for those the two coincide. A turn carrying three distinct requests carries three moves and receives three codes. Segmenting rules, worked example, and the rationale are in §4.3a.

Mode shares are computed over MOVES, so they still sum to one and remain a valid composition. A turn that yields no move at all is recorded as mode 0 (§3), keeping the turn in turn-level counts while contributing nothing to the mode distribution.

Turns are coded in conversational order with the preceding assistant turn visible. Several rules below are undefined without it. Where a corpus does not preserve assistant turns, say so explicitly in the method section; do not silently code as though it did.

Assistant turns are never coded. They are retained solely as context for reading the participant's next turn.


2. Step 1 — Attribution: whose turn is this?

Do this before anything else. A turn assigned to the wrong speaker measures the AI and reports it as the human, and no later rule can repair it.

Judge by REGISTER, not by length. This overrides any impression from formatting or length.

  • Whoever is offering to do the work is the assistant. Whoever is asking for it is the participant.

Assistant turns are recognisable by:

  • Offers of service: "Would you like…", "I can…", "Happy to…", "Say the word and I'll…", "Standing offer:"
  • Speaking about the participant's materials from outside: "the file you attached", "the data you gave me"
  • Delivering a finished product: structured output with headings, tables, bulleted recommendations
  • Asking the participant what they want next, or explaining the exercise back to them
  • Openers like "Sure!", "Certainly!", "Great question!", or apologetic corrections

Participant turns ask for something, supply material, react, decide, or push back.

Explicit role markers win outright. "User:", "Human:", "Me:" → human. "Assistant:", "AI:", "ChatGPT:", "Claude:", "GPT:", "Bot:" → AI.

⚠ The commonest attribution failure

A short, friendly, question-shaped turn is often the AI. Length-based heuristics miss the assistant's single most frequent turn: a brief offer to do more work.

  • "Would you like this as a one-page board slide?"
  • "Is there a specific piece you want me to go deeper on?"
  • "Standing offer: I can add the profit-aware alternative — say which and I'll do it."

All three are short and conversational. All three are the AI. Every misattributed turn found in validation was of this shape.

This corrects a prior canonical instruction. MODE_CLASSIFICATION_GUIDE.md tells the coder to "err toward classifying long structured content as AI and short conversational content as human." That instruction is empirically wrong and is superseded here. See §8.

When you cannot tell, OMIT the turn. A missing turn costs one observation; a misattributed turn corrupts the measure. If a transcript opens on what reads as an AI response, the participant's prompt was probably lost in copying — do not code the AI's response as theirs.

⚠ AN ASSISTANT TURN IS NOT MODE 0.

Mode 0 means the participant said something carrying no codeable move. An assistant turn is not the participant speaking at all. Where a coder works one turn at a time and cannot omit, it must record speaker: "assistant" and leave the mode null.

The difference is not cosmetic. Mode 0 turns stay in the denominator as participant turns; assistant turns must leave it entirely. Collapsing the two inflates the uncodeable rate and deflates every mode share.

Measured: of 95 mode-0 turns in the validation corpus, 22 were assistant output — several at UNANIMOUS agreement, five models confidently coding the AI as an uncodeable human. The reported "19.6% of turns are unclassifiable" was therefore wrong as a statement about participants; the true rate is nearer 13–14%. This is the same conflation as counting a non-responding model as a dissenting vote, and it has now been found in four separate places.


3. Step 2 — Is the turn codeable? (mode 0)

Some turns carry no codeable intent. Assign mode 0. This is a category, not a failure. Forcing a bare "ok" into Oracle inflates the passivity share with something that is not an engagement move at all.

Mode 0 when the turn is:

  • a bare acknowledgement with no direction: "thanks", "ok", "got it", "nice"
  • a salutation or sign-off
  • a restatement of a previous turn adding no new substantive request
  • a fragment too short to identify intent
  • a file or attachment marker — a bare filename such as cabana_prospects.csv, or a file label with no request attached. These are upload artifacts of pasted transcripts, not turns.
  • an interface artifact captured when the transcript was copied out of a chat product: "Edit", "Copy", "Regenerate", "Retry", "Share", "Show more", "2/2", "You said:", "ChatGPT said:", a bare model name, a timestamp. Chrome, not conversation — and common in pasted transcripts.
  • an emoji-only or reaction-only turn with no request attached.
  • an aside not addressed to the AI — "(submitting this now)", a note to a teammate.

⚠ Two look-alikes that are NOT mode 0.

A bare continuation is a move. "Continue", "go on", "keep going", "next" ask for more of the same output and therefore direct the assistant's work → Production Assistant (2), or whatever was underway. They resemble acknowledgements and are the opposite: an acknowledgement closes, a continuation asks for more. This follows directly from the INITIATES-or-CLOSES test.

A self-repair attaches to the move it repairs. "Sorry, I meant CAC not LTV" corrects a request already made rather than making a new one. If the repair carries a fresh request, code that request; if it only fixes the prior one, it is mode 0.

A PROCESS DIRECTIVE is a move — code it by what it demands. Participants routinely direct HOW the AI should work rather than asking it to do a task, and these were falling into mode 0 for want of a rule:

DirectiveCode
Constrains the evidence base or resets what counts as an acceptable answer — "only rely on what I put in as a prompt, don't use external sources"Problem Setter (8) where it changes what the problem is; Verification (5) where it controls answer reliability
Demands counterargument — "always present both sides, good and bad"Critical Challenger (7)
Sets format, length or housekeeping — "use bullet points", "keep the full transcript"Production Assistant (2), or mode 0 if purely administrative and demanding no work

These are often HIGH-AGENCY moves. Someone constraining the AI's evidence base is doing something much closer to framing than to nothing.

⚠ EXCEPTION — an affirmation that answers an assistant question is NOT mode 0.

"Yes", "yes please do that", "sure, go ahead" following an assistant offer are codeable. A "yes" is a second pair-part: its meaning is inherited from the question it answers, not carried in its own words (Schegloff & Sacks, 1973). "Yes" to "shall I compile this?" is delegation; "yes" to "shall we ask whether this is the right goal?" is assent to reframing. The strings are identical.

Code by what was offered, and set initiation to co.

It is only a bare acknowledgement when nothing was asked.

⚠ A bare DECLINE is mode 0 — and the mode never looks forward.

"No", "no thanks", "not that" answering an assistant offer is mode 0. Do not code it by what the participant asks for in their next turn. That next turn carries its own code, in its own right.

Forward-looking codes break turn-as-unit: the code for turn n would depend on turn n+1, which makes every code provisional until the conversation ends and gives two coders two chances to diverge on one label. Merging a decline with the redirect that follows it is no better — a coder must then first judge whether to merge, which is a new discretionary decision with its own unreliability, and it produces variable-width units that turn-level statistics cannot use.

But a decline is not the absence of an engagement move, and it must not vanish. It is the participant refusing an AI-proposed direction, which is precisely the co-regulation construct. So every turn that answers an assistant offer carries a separate structured property:

response_to_offerWhen
acceptedParticipant assents. Code by what was offered; initiation: co.
declinedParticipant refuses. Mode 0. The following turn is coded independently.
ignoredParticipant replies but does not engage the offer at all. Code their turn normally.

ignored is deliberately separate from declined: collapsing them would inflate the decline rate with inattention.

Decline rate is then a variable in its own right — arguably a better agency signal than any mode share, since it measures the participant overriding the AI's direction rather than following it.

Warrant: declines are dispreferred second pair-parts and are normally packaged with accounts or alternatives (Pomerantz, 1984), so where a decline carries a substantive move in the same turn, that move is codeable and the turn is not mode 0. Mode 0 applies to the bare decline.

Ruled 2026-08-15 by Mark Keith, on Fable's objection that coding a decline by the following turn introduces forward-looking codes. The objection was accepted; the structured-property addition preserves the information the simple version would have lost.

⚠ A bare APPROVAL that closes an exchange is mode 0 — the INITIATES/CLOSES test.

"Looks good." "Perfect." "Looks good — one page, board-ready." "Nice, that works."

These are sequence-closing thirds (Schegloff, 2007, Sequence Organization in Interaction). After a base adjacency pair — request, then compliance — a third-position turn can close the sequence rather than launch a new one, and the canonical forms are exactly these: "oh", "okay", and assessments. They close; they do not act.

The test that governs this whole family of turns: does the turn INITIATE something, or CLOSE something?

TurnInitiates or closesCode
"Looks good, now add the retention slide"initiates — a new directioncode the direction (Production Assistant)
"Excellent, this is very professional. Continue."initiates — "Continue" directs further buildingProduction Assistant (§4.3, unchanged)
"Looks good — one page, board-ready."closes — assesses and stopsmode 0, closes_sequence: true
"Perfect, thanks."closesmode 0, closes_sequence: true

Why not code it as a repeat of the prior turn's mode: that counts one engagement move twice. The request was already coded when it was made; the approval adds no second move.

Why not merge it with the prior turn: merging makes whether to merge a discretionary coder judgement and produces variable-width units that turn-level statistics cannot use — the same objection that settled the decline case above.

What this buys: the UNCHECKED-ACCEPTANCE RATE — how often a participant closes a sequence with bare approval rather than testing the output. This is a better passivity signal than any mode share, because it measures what the participant did NOT do, and it is the behaviour the Incentive Trap case is built to expose: accepting the AI's plan without checking it.

Evidence this was needed. In the 484-turn corpus at protocol v4, approval-closers ran mean agreement 0.64 against a corpus mean of 0.817, and split 2 mode-0 / 4 Passivity-tier / 3 Verification. Approval was still being read as verification in a third of cases — the exact inversion §4.3 exists to prevent — because §3 and §4.3 disagreed about this turn class until now.

Ruled 2026-08-15 by Mark Keith.

Pipeline consequence. Whenever transcript text is shortened for storage or context, the assistant's closing move — its final question or offer — must be preserved. Head-truncation destroys exactly the pair-part the next participant turn depends on.

Mode 0 turns are excluded from mode-distribution statistics and included in the total turn count.


4. Step 3 — Assign the mode

The eight modes are defined in THEORY_CONTRACT.md. Assign by dominant intent where a turn spans several. The rules below settle the boundaries that the definitions alone do not.

4.1 Mode 8 — Problem Setter is Schön's problem setting

Not "asking a meta question". Schön (1983), The Reflective Practitioner: problems are not given; they are constructed from situations that are puzzling and uncertain. Problem setting is how we name the things we will attend to and frame the context in which we attend to them.

Two qualifying moves, either sufficient:

  • NAMING — introducing something as a relevant thing of the situation that the task did not treat as material: an unstated cost, an affected party, a time horizon, a downstream consequence. Selecting what counts is problem setting even when no goal is contradicted.
  • FRAMING — setting or resetting the boundaries of attention, or imposing a coherence that says what is wrong here and in what direction things should change. Challenging the stated objective is one instance of this, not the whole of it.

THE DECISIVE TEST: is the frame being USED or CONSTRUCTED?

Schön contrasts problem setting with Technical Rationality, which solves problems whose ends are already given and cannot perform problem setting precisely because it depends on knowing the end.

A participant directing expert, multi-step analysis toward the goal they were handed is Collaborative, however sophisticated that analysis is. Depth does not promote a turn.

A participant who changes what the problem is — especially after the data contradicts the brief, which is reflection-in-action, where the situation "talks back" — is Problem Setter.

Note on scope: Schön describes an ongoing practitioner stance across a reflective conversation. Coding discrete turns is a coarser object and is an extension of the theory, not the theory itself. Record this as a limitation in any write-up.

4.2 Pasted task material is coded by the ASK, never by its content

Participants routinely paste an assignment, case brief, prompt sheet or dataset description into the chat. That text is context, not the participant's reasoning. A brief that itself discusses goals, mandates or "is this the right target" does not make the forwarder a Problem Setter — they forwarded those words, they did not write them.

Pasted withCode as
an explicit or implicit "do this for me" (make a plan, write the recommendation, find what drives conversion)Production Assistant (2)
a bounded question ("what does CAC mean here?")Oracle (1)
the participant's own framing or challenge, in their own wordscode their sentence — may well be Problem Setter

The test: is the meta-level move in the participant's words, or merely present in the material they forwarded?

Handing over a task wholesale is the canonical delegation move and one of the most common real behaviours there is. It is not a deficiency to be hidden or stripped from the data; it is a mode, and it is scored as that mode. Grounded in cognitive offloading (Risko & Gilbert, 2016).

4.3 Acceptance is not verification

"Looks great, continue", "Excellent, this is very professional", "Perfect, keep going" are not Verification.

  • Where the participant directs further building → Production Assistant (2)
  • Where they simply consume → Oracle (1)

Verification (5) requires testing a claim against something outside the claim: naming a number and asking where it came from, requesting a recomputation, checking an assertion against the data. Approval is the opposite of verification.

Verification also requires a participant-produced artifact — their own complete or near-complete work, presented to be checked. Without one, it is not Verification.

4.3a Turns that span several modes — the ENGAGEMENT MOVE is the coding unit

A long turn carrying three distinct requests is performing three engagement moves, and it is coded as three. Forcing one code onto it is measurement error, not parsimony.

Two units, not one (Krippendorff, 2013, on sampling vs recording vs context units):

UnitWhat it isWhat it carries
Turn — the context unitOne uninterrupted participant contributionattribution (§2), initiation (§5), response_to_offer (§3), sequence position
Engagement move — the coding unitOne distinct act directed at the AIthe mode

Most turns contain exactly one move, and for those the two units coincide.

Segmenting rule. A new engagement move begins where the participant performs a distinct act toward the AI — a new request, question, challenge, correction, or check. Operationally:

  • A new imperative or interrogative addressed to the AI starts a new move.
  • Supporting material — context, data, constraints, pasted briefs, "here's the file" — attaches to the move it serves. It is not a move of its own.
  • Elaboration, restatement, or hedging of the same request stays in the move it belongs to.
  • Politeness wrappers, greetings, sign-offs and bare assessments are not moves. A turn made only of these yields zero moves and is recorded as mode 0 (§3), which keeps the turn in the denominator of turn-level counts while contributing nothing to the mode distribution.

Worked example. "Can you summarise what drives conversion? Also I don't think sign-ups is the right target given these economics — push back on that. And check the CAC figure against the data dictionary." → three moves: Oracle (1), Problem Setter (8), Verification (5).

Why this, and not the alternatives

Multi-labelling one turn (double counting) is rejected. Mode shares are a composition and the ILR transform requires them to sum to one. One unit carrying two codes breaks the simplex and silently invalidates every compositional statistic downstream. Segmentation does not have this problem — three moves of one code each still sum to one, over moves rather than turns. That is the whole reason to segment rather than multi-label.

Turn-as-sole-unit is rejected on construct validity. The AIEM measures engagement moves. If one turn contains three, a one-code-per-turn instrument systematically under-counts exactly the participants who engage most densely — and it does so worst in the long-session, high-volume use that the framework most wants to describe. The composition is a property of the ANALYSIS; the move is a property of the PHENOMENON. When they conflict, the phenomenon wins and the analysis adapts.

Theoretical grounding. Turns are built from smaller complete units, each a possible contribution in its own right — the turn-constructional unit of conversation analysis (Sacks, Schegloff & Jefferson, 1974). A single utterance can perform several illocutionary acts (Austin, 1962; Searle, 1969), and discourse analysis has always coded at the level of the act or move rather than the speaking turn (Sinclair & Coulthard, 1975). Content-analysis methodology says the same: the recording unit should be chosen to match the construct, not the transcript's formatting (Chi, 1997; Rourke, Anderson, Garrison & Archer, 2001; De Wever, Schellens, Valcke & Van Keer, 2006).

The TCU is defined partly by prosody and transition-relevance, which text chat does not have. The engagement move is its written-CMC analogue: a segment that could stand alone as a complete request.

The cost, stated plainly

Segmenting introduces a second reliability problem: coders must agree on where the boundaries are, not only on what the codes are. Unitizing disagreement is frequently larger than coding disagreement and is usually left unreported (Strijbos, Martens, Prins & Jochems, 2006).

This is measurable and must be measured. Report Krippendorff's α_U, the unitizing reliability coefficient (Krippendorff, 1995, 2013), alongside the coding reliability. An unreported unitizing step is a hidden researcher degree of freedom; a reported one is a method.

Also report the move-per-turn distribution. It is a substantive finding in its own right — dense multi-move turns are a real usage pattern in long working sessions, and turn-level instruments cannot see them at all.

Migration

Turn-level and move-level codes are not interchangeable, so a corpus recoded under this rule is not directly comparable to one coded before it. In the 484-turn validation corpus roughly 6–10% of turns carry more than one move, so the divergence is bounded but real.

Report both while any longitudinal comparison is live: the move-level distribution as primary, and a turn-level distribution (dominant move per turn) for continuity with prior waves. Stamp the protocol version on both.

4.3b A move is coded IN SEQUENCE, never in isolation

Where a turn's text underdetermines the mode, the preceding turn — and the code that turn received — decide it. A continuation inherits the move it continues.

The exception, which is the whole reason this is a rule and not a shortcut: a turn that CHANGES DIRECTION overrides the inherited code. Inheritance is the default, not the answer. Read the turn for a new direction first; inherit only if there is none.

Why this is not optional. A turn like "Excellent. Let's check 3-4." has no readable mode on its own face. Coding it in isolation forces the coder to invent an intent, and the five-model panel shows exactly what that produces — see the evidence in §4.3c. The mode is recoverable, but only from the sequence.

Three consequences, each of which costs something:

  1. Isolated-turn coding is invalid, including for a reference standard. A gold set of shuffled turns is not a harder version of this task; it is a different task, and one the instrument is not performing. Any expanded gold set must present turns in their conversational order with prior turns visible. This supersedes the "unit mismatch" framing in §7.5, which treated the problem as a limitation to disclose rather than a defect to fix.
  2. Human coders code in conversation order and never shuffle. Randomising turn presentation to suppress order effects — ordinarily good practice — destroys the information this rule depends on.
  3. It is the empirical case for the sequence-level pass. If a code depends on prior codes, a coder that reads one turn at a time is structurally unable to produce it reliably.

Theoretical grounding. Turns are constructed in relation to what precedes them, and a turn's action is not recoverable from its text alone (Sacks, Schegloff & Jefferson, 1974). Sequences, not utterances, are the unit in which actions become interpretable (Schegloff, 2007) — the same source this standard already relies on for the sequence-closing third in §3. Schön's (1983) reflective conversation is an ongoing stance across an exchange rather than a property of any single utterance. And co-regulation is inherently sequential: who originated a move is defined by what came before it (Hadwin, Järvelä & Miller, 2011), which is why initiation was already a sequential property being recorded at turn level.

4.3c Adjudicated boundary cases

Mark Keith's rulings on contested turns from the 484-turn validation corpus. Each was contested by the five-model panel; each is now settled. Apply these directly — they are the operative reading of §4.1 through §4.3b, not illustrations of it.

C1 / C2 — Reviewing produced output against a standard is Verification (5)

"Let's QA all 7 slides." → "Excellent. Let's check 3-4." "Now let's convert to PDF and visually review." → "7 pages. Let's review each."

Both are Verification. QA and review test an artifact against something outside it — here a submission rubric. The continuation turns inherit it under §4.3b.

Evidence, and it is the reason this ruling matters more than the individual turns. In one conversation the same move recurs four times and the panel codes it two different ways:

#TurnPanel codeAgreement
34"Let's QA all 7 slides."20.6
35"Excellent. Let's check 3-4."20.4
36"This is a very clean, clear pivot table. Let's check slide 5."50.6
38"This is very strong — clean waterfall, transparent methodology. Let's check 6 and 7."51.0

A second conversation runs the same move five times and draws codes 2, 0, 0, 2, 2. Nine instances of one move type, three different codes. The Production and mode-0 readings cluster on the THINNEST instances; at the richest wording the panel is unanimous on Verification. This is not a boundary between two modes — it is the instrument degrading where context is sparse, which is precisely what §4.3b exists to repair.

⚠ Turn alternation does NOT make a move Collaborative (4). Back-and-forth is a property of the sequence, not of the move. Were iteration sufficient, Collaborative would become the default code for every extended session and lose its discriminant value entirely — and it is already the most-assigned code in the corpus. Collaborative requires the participant to contribute substantive content, to co-construct. Directing and approving in increments is not that.

"Now let's convert to PDF and visually review" is also a clean two-move turn under §4.3a: convert (Production, 2) then review (Verification, 5).

C3 — Introducing an unstated cost is Problem Setter (8)

Preceding turn: "Give me an example of focusing on the right demographic…" (Production). Contested turn: "Now consider the conversion costs and reanalyze the proposal."

Problem Setter. §4.1 defines NAMING as "introducing something as a relevant thing of the situation that the task did not treat as material: an unstated cost, an affected party, a time horizon" — this turn is the literal case in the rule.

"Reanalyze" is the delivery mechanism, not the move. Strip the cost clause and "reanalyze the proposal" inherits Production under §4.3b. The engagement move is the introduction of cost as material.

Measurement consequence, which decides it. In the buried-cost case family the unstated cost IS the hidden dimension the activity exists to measure. A participant who raises it unprompted is the participant catching the trap. Coding that as Collaborative blurs the exact signal the activity was built to collect.

C4 — Testing an enumerated objection is Verification (5)

"Let me test the referral-dependency objection, which is the other one that could matter."

Verification. Testing a claim against something outside it (§4.3). One move — "which is the other one that could matter" is an account, not a second move. Note that "the other one" implies the objections were named in an earlier turn; that earlier turn is the Problem Setter and this one is the test.

C5 — Two moves in one turn: naming, then checking

"What are potential confounders that may have affected our analysis? How have we accounted for the impact of these confounders?"

Two moves. What are potential confounders → Problem Setter (8), naming factors the analysis did not treat as material. How have we accounted for them → Verification (5), checking the analysis against those factors.

This is the reference example for §4.3a. A turn-level instrument must discard one of the two codes, and either choice misreports the turn.

4.4 Remaining boundaries

  • Practice-problem / quiz / flashcard / self-test generation is Tutor (3), not Problem Setter and not Production. Pedagogical intent dominates even when the surface form ("generate 10 practice problems") resembles Production.
  • Creative Expander (6) requires explicit multiplicity — multiple options, alternatives, or divergent perspectives. Extending a single idea is Collaborative or Tutor.
  • Context can outrank form. A bounded data question is Oracle standing alone, and may be Collaborative inside a sequence where the participant is directing which cuts to examine. Judge a turn within its conversation. See §4.3b, which makes this binding rather than advisory.

4.5 ⚠ Appropriateness does not promote a mode

Modes are observable functions, not identities or moral labels (THEORY_CONTRACT.md). Any mode can be wise or unwise in context. Oracle used well is still Oracle. A bounded question does not become a higher tier because it was the right move.

Scoring appropriateness would make the instrument normative — "passivity is bad" would become an assumption of the measure rather than something data could disconfirm. One cohort in the validation corpus showed the Agency tier negatively associated with catching a buried cost. A normative instrument could not have observed that.


5. Step 4 — Record initiation

Every coded turn carries an initiation value. This never changes the mode.

ValueMeaning
selfThe participant originated the cognitive move in their own words.
coThe AI proposed the move; the participant assented.
unknownThe preceding assistant turn is unavailable. Recorded, never guessed.

How to tell: a short affirmation ("yes", "yes please", "sure, do that") answering an assistant question is always co — its content is inherited from the question. A turn that introduces the move in the participant's own words is self.

Assenting to "shall we step back and ask whether this is the right goal?" is problem setting, and the mode is Problem Setter. It is co-regulated problem setting rather than self-regulated.

Why initiation is partitioned and never weighted

Regulation research distinguishes self-regulated, co-regulated, and socially shared regulation as three distinct levels, not three magnitudes of one thing (Hadwin, Järvelä & Miller, 2011, 2018). An AI-proposed move a participant accepts is co-regulation precisely, not by analogy.

Half credit was proposed and rejected:

  1. Co-regulation is a level, not a discount. In the Vygotskian lineage the model draws on, development runs through co-regulation toward self-regulation. A 0.5 multiplier misrepresents a developmental stage as a deficiency.
  2. Any coefficient would be arbitrary and unfalsifiable. Nothing in the literature supplies the number, and once it is inside a composite the score no longer contains the information needed to test it.
  3. It destroys the more interesting variable — whether a participant originates agency moves or merely accepts them.

5a. Step 5 — Confidence, and when to escalate

Confidence is PANEL AGREEMENT — the proportion of independent coders on the modal code. It is never the coder's own self-reported certainty.

Why self-report cannot do this job

MODE_CLASSIFICATION_GUIDE.md Steps 3–4 specify a self-reported 0.00–1.00 confidence with escalation of the 0.50–0.70 band. Measured on the 484-turn corpus, that design does not work. Self-reported confidence barely moves as coders go from perfect agreement to near-total disagreement:

Panel agreementTurnsMean self-reported confidence
Unanimous 5/52310.954
Strong 4/51060.929
Split 3/5890.908
Badly split ≤2/5580.896

A spread of 0.058 across the entire range. On seven badly-split turns every coder reported ≥0.95 confidence while disagreeing with each other. The specified escalation band would fire on 2 of 484 turns (0.4%) and catch 1 of the 147 contested ones; nothing falls below 0.50, so nothing would ever be marked unclassifiable by confidence.

Implemented as specified it is a quality gate that catches nothing while making every output look vetted. That is worse than having none, because it converts an unmeasured risk into a false assurance. Self-report is a judgement by the same process that made the call; independence is what makes agreement informative, and only a panel supplies it.

Retain the self-reported value, uninterpreted. It is the evidence for this finding. Do not overwrite it with the agreement figure.

The escalation rule

Escalate a move to an independent second pass when EITHER holds:

TriggerThreshold
Modal proportion≤ 0.60 (3 of 5 or fewer on the modal code)
Panel completenessfewer than 5 valid votes

On the validation corpus that is ~180 of 484 moves, about 37%.

Justification 1 — marginal value (primary; derived, not borrowed). Escalation buys one additional independent vote, and is worth buying only where it can change the answer. With five valid votes, 5/5 and 4/5 cannot be overturned by a sixth vote — 4/5 becomes 4/6 against 2/6 and the modal code stands. 3/5 can: a sixth vote makes it 3–3. So the set where escalation can possibly change the modal code is exactly modal count ≤ 3. Above that threshold escalation is pure cost with no reachable effect on the outcome.

Justification 2 — convergence with the reliability literature. Krippendorff (2004, 2013) sets α ≥ .800 for firm conclusions and .667 as the floor below which not even tentative conclusions are warranted. A five-rater panel's achievable proportions bracket these cleanly: 4/5 = .800 meets the firm standard, 3/5 = .600 falls below the .667 floor. The same cut is where the panel drops below the literature's minimum for any conclusion at all.

Justification 3 — method. Ambiguous segments should be adjudicated rather than assigned by fiat (Chi, 1997). Escalation is the automated form of adjudication; the human-coder analogue is consensus review.

⚠ Declare this, do not bury it: a raw modal proportion is NOT Krippendorff's α. α corrects for chance agreement and a modal proportion does not. With nine categories (modes 1–8 plus 0) chance agreement is roughly .11, so the inflation is modest — but Justification 2 is an analogy, not a derivation, and must be reported as one. Justification 1 carries the rule on its own. The correct long-run move is to compute α on the panel and set the threshold on α itself.

Why panel completeness is a separate trigger and not a refinement. 66 of 484 turns (13.6%) drew fewer than five valid votes, and ten drew exactly one. A single vote yields agreement = 1.00, so those ten turns present as unanimous and maximally reliable on the strength of one coder — the exact inverse of the truth. Any threshold on agreement alone admits them as the highest-confidence turns in the corpus.

Escalation fails closed. If the second pass errors or does not return, the move is flagged and excluded from aggregates. It is never silently coded at the original confidence.

What escalation does to the numbers

  • Denominator. Excluded moves leave the mode-share denominator, which still sums to one over retained moves, so the ILR composition survives. The denominator's meaning changes, so report retained-move count alongside n and protocol version.
  • Flag rate becomes a reported quantity (§6). A participant whose moves are systematically hard to code may be engaging in a genuinely distinct way, or the instrument may be failing on them. Both are findings, and neither is visible if flagged moves are silently dropped.

6. Reporting

Report six quantities, not one:

MetricComputed overReads as
Mode profile (total)all coded turns, self + coWhat happened in this conversation.
Mode profile (self-initiated)self onlyWhat the participant originated.
Co-regulation ratioco / (self + co), overall and per modeHow much engagement was AI-scaffolded.
Decline ratedeclined / all turns answering an offerHow often the participant overrode the AI's direction.
Unchecked-acceptance ratecloses_sequence turns / all sequence closingsHow often output was approved without being tested.
Flag rateescalated-and-unresolved moves / all moves (§5a)How much of this profile the instrument could not code confidently.
  • The engagement composite continues to be computed over all coded turns, unchanged, so historical scores stay comparable.
  • Mode 0 turns are excluded from all mode-distribution denominators and included in total turn count.
  • unknown turns are included in the total profile and excluded from the self-initiated profile and from the co-regulation denominator, so missing context depresses neither.

The co-regulation ratio is the new research variable. High Agency-tier share with a high co-regulation ratio means the AI is doing the metacognitive work — a pattern the instrument previously could not see.


7. Mandatory disclosures when publishing

Any paper, report, or claim resting on codes produced by this standard must state the following. None of it is optional, and none of it has been resolved by adopting this standard.

7.1 The measure has moved before

On one corpus of 484 turns, the substantive conclusion changed three times as coding defects were repaired, on identical data:

Coder versionWhat the data then appeared to show
v1Partnership predicts catching the buried cost
v2 (forwarded material coded by the ask)Verification predicts it; Problem Setter does not
v3 (assistant turns omitted; acceptance ≠ verification)Both prior readings were partly artifacts

No mode-to-outcome claim from that corpus was stable until v3. Always report the protocol version alongside any finding.

7.1a The protocol has been falsification-tested — and it moves the headline down

The five-model panel was re-run on the same 484 turns under this protocol on 2026-08-15. Agreement rose: mean 0.724 → 0.817, unanimity 31.0% → 47.7%, contested share 20.2% → 9.5%.

But the modal code changed on 48% of turns and the tier on 43%, and all four defects pushed the same direction — toward overstating engagement. Corrected tier shares over codeable turns:

beforeafter
Passivity49.4%68.9%
Partnership21.7%9.5%
Agency28.9%21.6%

Additionally, 19.6% of turns are unanimously unclassifiable and belong in no mode distribution at all.

Report both halves. Agreement rising while half the labels move is the signature of a measure that was previously unconstrained. And agreement is not accuracy: a sharper rule raises convergence whether it is right or arbitrary. The rules earn their standing from theory and from adjudicated real cases; the re-run shows only that independent coders given them converge.

⚠ These figures do not cover §2, and §2 has the largest blast radius of any rule here

The re-run performed no attribution step. Verified 2026-08-16 by reading scripts/rerun-panel-protocol-v4.ts: its write block records modal_mode, agreement, valid_votes and per_model and contains no speaker field at all. assistant_votes is 0 on all 484 rows of the snapshot, against a maximum of 3 on the live table, because the column was never populated rather than because no assistant turns were found.

Consequence: assistant turns entered these agreement statistics as participant turns with modes assigned to them. Two adjacent turns in one conversation — "You're right, and I should have said this before building the tab…" and "Want me to replace the Q3 tab with the profit-aware version…?", the second being the canonical assistant offer §2 exists to catch — drew modal codes 5 and 8 on five valid votes each. A conservative lexical scan finds 16 such turns; the true count is higher, and establishing it requires an attribution pass rather than a regex.

This is the same defect as §7.2, one level up, and it is now the second occurrence: a panel that does not run the instrument it is measuring. §7.2 caught it for the protocol as a whole. This is the same error confined to a single rule, and it survived precisely because the rule was assumed rather than exercised.

What may still be cited from this re-run: the direction of movement, and the fact that the four repaired defects all pushed toward overstating engagement. What may not: any agreement figure as evidence that §2 is applicable or reliable, and any tier share as a corrected value, since misattributed turns are in the denominator. Both await a re-run with attribution in the loop.

7.2 Do not cite the 2026-08-14 panel-vs-coder agreement figures

An earlier five-model panel produced "the deployed coder differs from the panel on 47–49% of turns". That figure does not measure coder reliability and must not be reported as though it does. The panel ran on the bare mode definitions with none of §2–§5; the deployed coder ran the full protocol. The comparison measured protocol vs no protocol.

The direction is also the opposite of what it appears: in the single largest contested cluster — 27 copies of the same forwarded case brief — the panel's modal answer was Problem Setter and the deployed coder said Oracle, and §4.2 says the deployed coder is right.

Corollary, which generalises: panel modal ≠ ground truth. Model agreement bounds reliability, not validity. Five models sharing a prompt and overlapping training data can agree and be jointly wrong.

7.2a Only ONE rule has been empirically validated individually

A leave-one-out ablation of all eleven rules (2026-08-16) confirmed exactly one:

  • §4.2, pasted material coded by the ask — agreement +0.066 (p = 0.001), gold accuracy +10.5 pp. Consistent on both criteria, and the only significant result.

The other ten are not validated, and are also not refuted. The 19-item gold set has no power to adjudicate them: with 0–3 discordant pairs, McNemar's floor is p = 0.5, so no result was reachable whatever the truth. They stand on theory and on adjudicated real cases, which is a legitimate warrant and a weaker one than measurement.

Say this plainly in any write-up. Do not describe the protocol as "validated"; describe one rule as validated and the rest as theory-grounded and untested. The distinction is the difference between a defensible instrument paper and an overclaim a reviewer will find.

7.3 Inter-rater reliability HAS been measured — and the result is worse than not having one

This section previously read "no inter-rater reliability has been earned." That was false when written. A two-coder human validation was completed on 2026-08-06 by two trained research assistants on a 178-message stratified sample, independent round followed by a recorded consensus resolution. It lives in the project validation archive and in the manuscript sources, and was referenced by neither this standard nor the canonical package. Same failure class as the AIT_mode_classification_v1.0.2.md finding in §8: decisive evidence outside the canon, invisible for as long as nothing pointed at it.

MeasureValue
Human–human, independent roundκ = .49 (60.0% exact, n = 165, 95% CI [.40, .58])
Human consensus vs. primary classifierκ = .31 (40.2% exact, n = 164)
Model–model (Fleiss, three models)κ = .61

The three models agree with one another substantially more than any of them agrees with human consensus. That is correlated model error, demonstrated rather than hypothesised — and it is the empirical case against ever citing panel agreement as validity evidence. §5a uses panel agreement as a confidence signal, which remains defensible; it is not, and must never be presented as, accuracy.

Per-label precision — P(human consensus = label | classifier label):

LabelPrecisionLabelPrecision
Tutor (3)96% (24/25)Oracle (1)28%
Production (2)78% (18/23)Critical Challenger (7)11%
Creative Expander (6)44%Verification (5)6% (1/18)
Collaborative (4)43%Problem Setter (8)0% (0/22)

⚠ These figures describe PREDECESSOR constructs, not the ones defined in this standard

Read analysis/validation_2026-05-30/RA_codebook.txt beside classify_validation.py before citing any number above. The RA codebook and the classifier prompt are byte-identical definitions, so the validation is internally sound — humans and models were reading the same rulebook. But that rulebook is not this one, and on exactly the two labels in question the definitions are not narrower versions of the current ones. They are different constructs:

LabelDefinition validated in 2026-08Definition in this standard
Problem Setter (8)"asks the AI to help define/reframe the problem before solving"§4.1 — the participant themselves names or frames something the task did not treat as material
Verification (5)"asks the AI to check/critique the user's own work or claims"§4.3 — testing a specific claim against something outside it, on any artifact regardless of who produced it

The Problem Setter definitions differ on who does the framing — nearly opposite. The Verification definition carries the participant-artifact requirement that was deleted from §4.3 on 2026-08-17.

What follows, stated precisely:

  • The 0% and 6% figures are real and correctly reported for the constructs they tested, which are the constructs the SR manuscript's classifier used. Nothing here impugns that paper's validation.
  • They do not transfer to §4.1 or §4.3 as written. Those constructs have never been human-validated. The honest statement is untested, not invalid — and the two must not be conflated in either direction.
  • So the §4.3c rulings are neither vindicated nor undermined by this validation. They are unmeasured, and a re-validation under the current definitions is what would settle them.

Three further defects would depress modes 5 and 8 independently of any of the above, and all three are fixed by rules adopted since: RAs coded isolated messages with no prior turns (§4.3b now forbids this); they were given no task brief, though §4.1's definition is explicitly task-relative — "that the task did not treat as material" cannot be applied without knowing what the task treated; and the template supplied assistant text under a message column labelled as the user's (§2), which the coders caught twice unaided.

Consequence for the RA re-run: an instrument carrying the current definitions, the preceding turns, the task brief, mode 0, and move-level segmentation. Re-running the old sheet would reproduce the old numbers and answer nothing.

Two independent corroborations of rules in this standard, from the coders themselves:

  1. Coders identified two messages as AI responses misattributed as user messages — independent confirmation of §2 from humans reading the same corpus.
  2. Coders rated isolated messages while the classifier saw full transcripts, and several coder notes state the code "depends on prior context." Two RAs, working without knowledge of this standard, discovered §4.3b. That is the strongest available evidence for the sequence-dependence rule, and it also means κ = .49 is a floor: part of the human–human disagreement is context deprivation, not construct ambiguity.

What must be reported. κ = .49 for human–human is modest — this instrument is hard for humans too, and any claim of coder reliability must cite this figure rather than a model-agreement statistic. The corpus-weighted consensus–classifier agreement is 58%, and the classifier is trustworthy for the two mass modes that dominate the corpus and untrustworthy for the rare agentic ones.

7.4 Self-consistency — the original finding has a mundane explanation, and it is not self-inconsistency

This entry previously read as a threat to validity. It is now most likely a database defect, and the correction matters because the two have opposite implications.

The identical string "What is the churn risk of the final recommendation?" was found coded Tutor once and Verification once, and recorded here as evidence that the coder contradicts itself on identical input.

Investigation on 2026-08-16 found a sufficient alternative cause. Two endpoints each wrote a complete coding of the same submission — /api/v1/classify wrote one and /api/v1/analyze re-coded the same turns and inserted a second set without deleting the first. 311 turns across 30 of 204 submissions carried two codings, 135 of them conflicting. One turn holding two codings from two independent passes is not a coder contradicting itself; it is the expected disagreement between two passes, made to look like self-contradiction by a storage bug. Fixed at source, and a unique constraint now prevents recurrence.

What this does and does not license. It does not establish that the coder is self-consistent — that has still never been measured, and remains a precondition for any reliability figure. It establishes only that the single piece of evidence previously offered for self-inconsistency does not support the claim. Re-run the test deliberately on deduplicated data before either asserting or denying self-consistency.

7.5 Other standing limitations

  • One case world. All validation turns come from a single case (Cabana "Incentive Trap"). Boundary rules may not transfer.
  • Unit mismatch — a limitation to disclose SUPERSEDED by §4.3b, which makes it a defect to fix. The hand-labelled set codes isolated turns; the deployed coder codes whole conversations. Under §4.3b a move is coded in sequence, so an isolated-turn reference standard is not a harder version of the task but a different one. The expanded gold set must present turns in conversational order.
  • Turn-level Schön is an extension, not Schön (§4.1).
  • Initiation depends on the assistant turn surviving preprocessing. Where participants paste pre-truncated transcripts, unknown will be common and the ratio correspondingly weaker.
  • The self/co distinction has not been tested for reliability. It looks cleaner than mode assignment, but "looks cleaner" is not a measurement.

8. Governance status of each rule

Per ADOPTION_AND_CHANGE_CONTROL.md, nothing here alters canon until Mark Keith decides. Rules were implemented in the instrument ahead of approval because production was actively producing wrong measurements; this table exists so the canon leads rather than trails.

⚠ Rows carrying a ruling date of 2026-08-17 or later are RULED but their BODY TEXT IS NOT YET WRITTEN into §§1–7. They were added to this table on 2026-08-31 so the governance record is complete and no ruling is invisible; writing them into the body is the single propagation pass that cuts 2.0.1 / 2.1.0. Until that pass runs, this table is the authority for those rules and the body is silent on them — a coder must read §8, not only §§1–7. That is a defect of sequencing, recorded here rather than hidden.

§RuleQueue itemStatus
2Attribution by register; omit rather than guessIMP-2026-007PROPOSED (patch 2.0.1)
2Supersedes the guide's long=AI / short=human tie-breakerIMP-2026-007PROPOSED — canon currently states the opposite
3Mode 0 unclassifiableIMP-2026-009PROPOSED — specified in the guide, never implemented until 2026-08-15
3Assistant closing move preserved through preprocessingIMP-2026-008Adopted in the pipeline 2026-08-15
4.1Mode 8 = Schön naming AND framingIMP-2026-007PROPOSED
4.2Pasted material coded by the askIMP-2026-007Mark's ruling, 2026-08-14
4.3Acceptance is not verificationIMP-2026-007Mark's ruling, 2026-08-15
4.5Appropriateness does not promoteIMP-2026-007Restates THEORY_CONTRACT.md; no change
5–6Initiation recorded, partitioned not weightedIMP-2026-008PROPOSED (minor 2.1.0)
3Bare decline is mode 0; response_to_offer accepted/declined/ignored; decline rate reportedIMP-2026-011Mark's ruling, 2026-08-15 — CANON
3Bare closing approval is mode 0 (INITIATES-or-CLOSES test)IMP-2026-012Mark's ruling, 2026-08-15
1, 4.3aEngagement move is the coding unit; turn is the context unit; α_U reportedIMP-2026-013Mark's ruling, 2026-08-15 (minor 2.1.0)
4.3bA move is coded in sequence; continuations inherit; a change of direction overridesIMP-2026-015Mark's ruling, 2026-08-16 (minor 2.1.0)
4.3cFive adjudicated boundary cases (C1–C5)IMP-2026-015Mark's rulings, 2026-08-16/17
5a, 6Confidence is panel agreement, not self-report; escalate at ≤0.60 or <5 valid votes; flag rate reportedIMP-2026-016Mark's ruling, 2026-08-17 (minor 2.1.0)
4.3Delete the "Verification requires a participant-produced artifact" clause — verification can occur on AI-generated content and the provenance of the checked artifact is irrelevantIMP-2026-017Mark's ruling, 2026-08-17
4.3B1 cluster: "any other ideas besides X" after rejecting a suggestion = Creative Expander; "does this include X? let's add Y" = two moves, Oracle then Collaborative; rejecting an AI idea outright = Collaborative. None are VerificationIMP-2026-017Mark's ruling, 2026-08-17
4.1, 5Problem Setter has two branches, both valid — the participant frames (initiation=self) or asks the AI to frame (initiation=co). §4.1 states only the first and must be widened; the existing initiation field carries the difference, so no new field is neededIMP-2026-018Mark's ruling, 2026-08-17
4.3Verification is about the claim, not its ownerIMP-2026-018Mark's ruling, 2026-08-17
4.3The verification-independence discriminator. Verification (5) requires the check to CONSULT OR CREATE EVIDENCE INDEPENDENT of the checker's own knowledge — an external record, an independent re-derivation or execution, or elicited counter-evidence. Where the AI can answer from what it already knows, no verification act occurred: code Tutor (3) if building understanding, Oracle (1) if retrieving a fact. Independence is about the KIND of evidence, not its cost. Replaces the discriminator lost when the participant-artifact clause was deletedIMP-2026-019Mark's ruling, 2026-08-17
4.3aDual-mode turns do not get two scored codes. Resolve in order: (a) segment if distinct spans exist; (b) code ONE move as Problem Setter with initiation=co; (c) secondary_mode UNSCORED as the fallback. True multi-labelling remains rejected — two scored codes on one unit breaks the compositional simplex and silently invalidates the ILR transformIMP-2026-020Mark's ruling, 2026-08-17
5reframe_direction as metadata, never a mode discriminator: self / requested / adopted / imposed. initiation cannot distinguish "the AI reframed me" from "I reframed the AI", and those are opposite events wearing the same codeIMP-2026-020Mark's ruling, 2026-08-17
4.1Calibration is not problem setting. Telling the AI how to view the problem is calibrating the AI. Problem setting proper requires the PARTICIPANT'S OWN frame to be constructed or changed. Operational test: did the participant's frame change or get constructed in this move, or was it already held and merely transmitted? This NARROWS mode 8IMP-2026-022Mark's ruling, 2026-08-17
4.1The abstraction reference point is RELATIVE, not absolute. A move is Problem Setter when it steps UP a level from the frame the conversation has already established — questioning or resetting an objective prior turns treated as fixed. A move operating AT the established level is not. Absolute level is coder-relative and unusable. Where a task brief exists the brief sets the level (external, fixed, identical across coders); where none exists, the reference is the conversation prefix, turns 1..n−1IMP-2026-027Mark's ruling, 2026-08-21
4.3dThe forward window is ADJACENT-ONLY (n±1), with provisional-then-final as one mechanism: a code is provisional when assigned and final when either n+1 arrives or the conversation ends. Intent is read FROM the move; n+1 is evidence about what the move already was, never a consequence it produced. Mode 0 is a determination, not a tie — the uptake check can never convert a mode 0 into a modeIMP-2026-021Mark's ruling, 2026-08-31
4.3The artifact-existence test, written explicitly rather than left to be derived: where the move is silent on intent, a referenced artifact that ALREADY EXISTS in the conversation makes "hand me / give me / show me X" retrieval → Oracle; one that does not exist makes it a build instruction → ProductionIMP-2026-021Mark's ruling, 2026-08-17

An agent applying this standard should apply all of it, and record in its method section that Rules marked PROPOSED were pending ratification at the time of coding.

Specification/implementation gap — CLOSED 2026-08-15 (protocol v7)

The engagement-move unit (§1, §4.3a) is now implemented in the aimodes.ai instrument. All four coordinated changes shipped together:

  1. Prompt — the shared protocol carries segmenting rules; both classifier schemas emit moves[] per turn.
  2. Schema and parser — ParsedMove / ParsedMessage, with flattenMoves() as the single path any distribution must go through. A flat one-mode-per-turn response is still accepted and synthesised into a single move, so a prompt regression degrades to the old behaviour instead of failing every conversation.
  3. Storage — migration 155 adds move_index, move_text, initiation, response_to_offer and protocol_version to message_classifications; every historical row is a valid single-move turn at move_index = 0.
  4. Scoring — mode and tier shares computed over moves; turn counts still over turns.

Duplicate classifications — CLOSED 2026-08-16

The unique index on (submission_id, message_index, move_index) is applied (migration 157). Getting there took three steps, in this order, and the order was load-bearing:

  1. Cause. /api/v1/classify and /api/v1/analyze each wrote a complete coding of the same submission; analyze inserted without deleting. 311 duplicated turns across 30 of 204 submissions, 135 conflicting, 2026-04-08 to 2026-08-11. See §7.4 — this is also the explanation for the recorded self-inconsistency.
  2. Data. All 622 rows for the affected turn slots were copied to message_classifications_dupes_20260815 — a full snapshot including survivors, so the deletion is reversible in both directions — and the newest coding per slot kept. That table is evidence; it must not be dropped.
  3. Code, then constraint. analyze now deletes before inserting and reapplies instructor overrides. The unique index was applied only after that fix was confirmed live, because applying it first converts silent duplication into a hard failure on every submission.

7.6 Measurement reliability of the instrument itself (measured 2026-09-02)

These are properties of the CODER, not of the rules, and they bound every figure in §7.1a. They were found by coding the same conversation twice under identical conditions — same prompt, same model, same protocol — which isolates instrument instability from the coder-difference confound that cross-model disagreement carries.

1. temperature was never set, so the classifier ran at the SDK default of 1.0 for the life of the instrument. Pinned to 0 on 2026-09-01. The multi-model panel was unaffected — its router already defaulted to 0 — so panel figures are more trustworthy than single-model ones, and the panel's disagreement is genuine boundary ambiguity rather than sampling noise.

2. Test-retest agreement is ~75%, and pinning temperature only moved it from 68.4% to 74.6%. The instability is mostly NOT sampling. Every disagreement observed crossed a TIER boundary.

3. Unitizing is less stable than categorising. 23 moves out of ~59 compared were segmented differently between two identical runs. α_U therefore cannot be assumed better than the categorical figure and must be reported separately, per §6.

4. Score-level instability is bimodal, not uniform. Across 11 submissions coded five times each: mean spread 1.8 points, but 9 of 11 were perfectly deterministic while 2 swung 3 and 11 points, both crossing a 10-point band. Most conversations are stable; a minority are volatile, and the volatile ones are not predictable from length — the volatile cases had MORE moves (12, 17) than the stable ones (7, 9), which refutes the obvious "short conversations are noisier" hypothesis.

5. One score component caused the largest observed swing, and it was the formula, not the rules. scoreVerification was a step function (0% → 0, any → 40, ≥20% → 100). Two turns moving out of Verification took the component 100 → 0 and the student's overall score 10 points, changing their band, while every other component was identical. Cliff edges convert one-move classification noise into a discontinuous score change. Replaced with a piecewise-linear curve through the same anchor points 2026-09-02; intent and breakpoint values unchanged, discontinuities gone.

Consequence for practice: a single classification pass is not sufficient for a graded score. Code twice; where the two disagree, code a third time and take the mode. This is automatic and needs no human review.

7.7 Self-disagreement is the better boundary-case detector

When five models disagree, "this turn is ambiguous" cannot be separated from "these models differ." When ONE model disagrees with itself under identical conditions, the model is held constant, so the disagreement is attributable to the turn and the rules alone. Coding the corpus twice and collecting the disagreements yields a pre-filtered adjudication queue with no coder-difference confound — a cleaner instrument-development method than the cross-model disagreement this programme has used, and cheaper than a panel.

Its first application found §4.6 below, which cross-model disagreement had not surfaced.

4.6 Supplying information the assistant asked for is not an engagement move

Ruled 2026-09-02. Distinguish two assistant acts that produce opposite codes:

  • The assistant offers to act ("shall I model that?") and the participant accepts → the participant has authorised a move. Code by what was offered, initiation=co.
  • The assistant asks for information ("what are your costs?") and the participant supplies it → the participant has made no move: mode 0, initiation=co, response_to_offer=accepted.

Warrant. An answer's action is constrained by the question that made it relevant (Schegloff, adjacency pairs — already relied on in §3 for the sequence-closing third and in §4.3d for the next-turn proof procedure). The second part of the pair does what the first made relevant; it does not perform a new act. The move belongs to whoever initiated it, and here that is the assistant.

Why it was needed. The prior text said only "code them by WHAT WAS OFFERED", which defers the decision without supplying a procedure. On one measured conversation the coder split nine such turns three ways across Oracle, Tutor and Collaborative between identical runs, swinging the score 19 points.

The discriminator, so the rule cannot swallow every response: mode 0 applies where the participant supplies the requested fact and nothing more. The turn becomes codeable the moment they do something with it — draw a conclusion, question the premise, redirect the enquiry, or volunteer a consideration the assistant did not ask about.

8a. The revert standard — ACCURACY, not agreement (Mark's ruling, 2026-08-31)

This SUPERSEDES the previous standard — "a rule that does not raise agreement is not a clarification and will be reverted" — which is replaced, not amended.

If a rule lowers agreement but has a theoretical reason to justify greater accuracy, the rule is kept. Mark Keith, 2026-08-31: "The goal is accuracy."

The old standard would have reverted the protocol 5–8 rules for being agreement-neutral, and doing so would have restored a known misattribution — coding assistant turns as participant turns. Agreement measures whether coders converge, not whether they converge on the truth. A rule can raise agreement by making a hard distinction disappear, which is the failure this replacement exists to prevent.

⚠ Unresolved, and it is the reason this clause is dangerous. "Theory says it is more accurate" is unfalsifiable if it may be asserted after the agreement figure is seen. The analyst proposed, and Mark has NOT yet ruled on, three conditions for keeping a rule under this clause: (a) the accuracy claim and its mechanism are stated before the agreement figure is known; (b) the rule names what observation would falsify the accuracy claim; (c) both are recorded, so the exemption is auditable. Without them this clause removes the falsification discipline the research record calls the only thing separating this programme from post-hoc rationalisation. Tracked as IMP-2026-030.

Still open — not settled by this document

  1. Confidence scoring with escalation — the rule is now specified in §5a; not yet implemented.

  2. Sequence-level pass (guide Step 5) — specified, not implemented. §4.3b is the theoretical and empirical case for it; until it exists, inherited codes rest on a coder that reads whole conversations but has no explicit sequence representation.

  3. Provisional-then-final coding (§4.3d, ruled 2026-08-31) — specified, not implemented. Needs a provisional code, a final code, a finalisation event, post_context_available on every move, and retention of superseded provisional codes. That makes three specified-unimplemented rules, not two.

  4. Remaining contested axes. Five cases are adjudicated in §4.3c. Recount on the current corpus, at panel agreement < 0.8, puts the largest remaining disagreements at Production↔Problem Setter (45 turns) and Collaborative↔Problem Setter (40) — not the Oracle↔Production axis previously recorded here as largest. Mode 8 is where the real ambiguity sits.

  5. The re-run HAS now happened — 2026-08-31, and the result is agreement-neutral. Full paired comparison, 478 turns, protocol 8 against protocol 4:

    protocol 4protocol 8
    mean agreement0.8160.809
    low agreement (≤0.60)23.8%24.5%
    few valid votes (<5)13.6%19.0%
    clear (≥0.80)69.2%65.7%

    Against no protocol at all the same rules are a large gain (mean agreement 0.732 → 0.798, unanimous 31.0% → 41.8%, contested 15.2% → 7.6%), which is the comparison the 2026-08-15 falsification step was pre-committed to. Against protocol 4 they are flat on disagreement.

    The apparent rise in escalation is not degradation: it is entirely the vote-count clause, and exactly 19.0% of turns have both <5 valid votes and at least one assistant vote — no residue. See IMP-2026-031; §5a's threshold is miscalibrated under protocol 8 and currently overstates the flag rate.

    These rules are KEPT, under the standard ruled at §8a. They are agreement-neutral but deliver a correctness gain the agreement metric cannot see: the speaker field identified 91 turns carrying assistant votes and 11 where the panel majority says the turn is the assistant's — all of which protocol 4 was coding as participant turns with a mode.


9. References

Austin, J. L. (1962). How to Do Things with Words. Oxford University Press.

Chi, M. T. H. (1997). Quantifying qualitative analyses of verbal data: a practical guide. Journal of the Learning Sciences, 6(3), 271–315.

De Wever, B., Schellens, T., Valcke, M., & Van Keer, H. (2006). Content analysis schemes to analyze transcripts of online asynchronous discussion groups: a review. Computers & Education, 46(1), 6–28.

Hadwin, A. F., Järvelä, S., & Miller, M. (2011). Self-regulated, co-regulated, and socially shared regulation of learning. In B. J. Zimmerman & D. H. Schunk (Eds.), Handbook of Self-Regulation of Learning and Performance (pp. 65–84). Routledge.

Hadwin, A. F., Järvelä, S., & Miller, M. (2018). Self-regulation, co-regulation, and shared regulation in collaborative learning environments. In D. H. Schunk & J. A. Greene (Eds.), Handbook of Self-Regulation of Learning and Performance (2nd ed., pp. 83–106). Routledge.

Krippendorff, K. (1995). On the reliability of unitizing continuous data. Sociological Methodology, 25, 47–76.

Krippendorff, K. (2004). Reliability in content analysis: some common misconceptions and recommendations. Human Communication Research, 30(3), 411–433.

Krippendorff, K. (2013). Content Analysis: An Introduction to Its Methodology (3rd ed.). Sage.

Pomerantz, A. (1984). Agreeing and disagreeing with assessments: some features of preferred/dispreferred turn shapes. In J. M. Atkinson & J. Heritage (Eds.), Structures of Social Action: Studies in Conversation Analysis (pp. 57–101). Cambridge University Press.

Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688.

Rourke, L., Anderson, T., Garrison, D. R., & Archer, W. (2001). Methodological issues in the content analysis of computer conference transcripts. International Journal of Artificial Intelligence in Education, 12, 8–22.

Sacks, H., Schegloff, E. A., & Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4), 696–735.

Schegloff, E. A. (2007). Sequence Organization in Interaction: A Primer in Conversation Analysis. Cambridge University Press.

Schegloff, E. A., & Sacks, H. (1973). Opening up closings. Semiotica, 8(4), 289–327.

Schön, D. A. (1983). The Reflective Practitioner: How Professionals Think in Action. Basic Books.

Schön, D. A. (1987). Educating the Reflective Practitioner. Jossey-Bass.

Searle, J. R. (1969). Speech Acts: An Essay in the Philosophy of Language. Cambridge University Press.

Sinclair, J. McH., & Coulthard, R. M. (1975). Towards an Analysis of Discourse: The English Used by Teachers and Pupils. Oxford University Press.

Strijbos, J.-W., Martens, R. L., Prins, F. J., & Jochems, W. M. G. (2006). Content analysis: what are they talking about? Computers & Education, 46(1), 29–48.

Questions or feedback on the methodology? Get in touch. Both documents are versioned and under active revision; the derivation methodology keeps its changelog in §7, and the coding standard records the governance status of every rule in its own §8.

Also in the research library: Pedagogy