Every score we publish is checked by independent judges: trained humans, and a model held to their standard. When they disagree, the human wins and the model learns. Try your ear first.




Every score beginsSustained effort toward a goal, held through setbacks — one of the eight durable skills we measure. Below: what Grit sounds like at each level. Click a level to hear it in a real sentence.
Grit is emerging: trying is mentioned, but no setback and no follow-through appear yet.
Grit is developing: one sustained effort at something difficult, without the full arc.
Grit is present: a setback, sustained effort, and an outcome, causally linked.
Grit is strong: the full arc, and the lesson transfers to new situations.

“At first I kept messing up the routine and I wanted to quit, but I kept practicing it every night until it finally clicked. Because I stuck with it, I made the team, and now when something is hard I remind myself it just takes reps.”
The working interface our raters use, doors open. Set the Human Score on each row — where you differ, yours publishes.
Sample responses, students, and counts · Illustrative
You just ran the scoring room. Now slow one disagreement down and take the call yourself: do you hear L2, or L3?

“My auntie always says quitters never win, so I guess I just don’t quit. Like last semester with algebra, everybody said drop it, and I stayed.”
Detected a persistence verb, but classified “quitters never win” as a borrowed phrase, not first-person evidence.

“My grandpa always tells me slow is smooth, and honestly he’s right. When my science fair project kept failing, I rebuilt it three times until it worked.”
Same pattern — a borrowed phrase carrying lived persistence — now credited correctly, because a human corrected it.
A system that never disagrees with itself is a system nobody is checking. Every divergence is surfaced, resolved by a human, and turned into training data.
Rubrics define the skill. Humans apply the rubric. The model is trained on their judgment and overruled by it — every time, by design.
Experts agree with each other at κ 0.87 — the model sits inside the human range.
Within one rubric level of the expert, on responses the model has never seen.
Identical to the expert. Everything further apart is overruled by a human.
Source · Published model reliability data · impacterpathway.com/model-card