Grading audit

Independent verifier output, side-by-side with the brain's curriculum state. Every score is from Claude (model family separate from the brain's training engine). Rationale and next-attempt instructions reproduced verbatim, never paraphrased. No mock data — if a field cannot be populated, it is shown empty with a written note explaining why.

Source: graded_compositions.jsonl Compositions on disk: Showing latest: Refreshed:

Dual Score — last 50 compositions

Each row: the brain's articulation of a concept, with the external Claude verifier's score, the verbatim rationale for that score, and the next-attempt instruction the grader returned to the brain. Sorted by recency.

When Persona / Concept Layer Claude Brain self Δ Rationale
Loading…

Calibration

Distribution of Claude's scores across the loaded compositions. A calibrated external verifier produces a spread, not a single peak. Mean score and per-bucket counts are derived from the same JSONL the dual-score panel reads.

Mean Claude score (rows shown)
Compositions in panel
Total compositions on disk
Brain self-scores available
Delta column is null because the rehearsal loop does not yet emit a per-composition self_score alongside Claude's grade. When self_score begins populating, this panel will surface honest brain-vs-grader deltas. Until then, only Claude's grades are shown — verifiably non-self-reported.

Blocked gates — L4 concepts in_progress

Mastery gate: three independent compositions at score 1.0 are required to flip a concept from in_progress to mastered. A single 1.0 is not enough. The brain must reproduce that score across separate attempts before the gate unlocks. These are the concepts currently sitting at the L4 gate.

Loading…