Under The Hood
In Your Hands
The Big Picture
The Solutions
The Methods
The Foundations
The Case
The Science
The Proof
The Arc
The People
The Promise
Construct Validity

These Scores Mean
What We Say They Mean

You're measuring something — but how do you know the thing you're measuring is actually curiosity? Fair question. The answer rests on two ideas — precise language, and rubrics that hold the measurement accountable — plus the empirical checks that test whether the scores behave the way those definitions say they should. This page walks through all three.

The Objection

"A model read a student's reflection and called it curiosity. How are you sure the thing it detected actually is curiosity — and not fluency, confidence, or vocabulary?"

This is the right question to ask of any instrument that claims to measure a psychological construct — and it's exactly the question construct validity answers. Our answer isn't "trust the model." It's the opposite: the model is never the authority on what curiosity is. The definition comes first, the rubric operationalizes it, trained human raters apply it — and the machine learning system is held accountable to all three.

The Answer, in Two Parts

Precise Language.
Accountable Rubrics.

Anchor 01 · Defining the Skill

Precision of Language Leads to Precision of Behavior

Dr. Marc Brackett
Founding Director, Yale Center for Emotional Intelligence

Marc Brackett's research at Yale begins from a simple observation: you cannot develop — or measure — what you cannot name precisely. When people acquire exact words for what they're experiencing, their behavior becomes correspondingly exact. Vague language produces vague understanding; precise language produces precise action.

Precision of language leads to precision of behavior.

Measurement inherits the same rule. A construct like curiosity is only measurable to the degree it is defined in precise language first — what it is, what it is not, and what it looks like at each stage of development. That's why every durable skill we measure starts as a written definition grounded in published developmental research, anchored by Harvard's Making Caring Common project, before a single response is ever scored. The word comes before the number.

Brackett, M. — research on emotional granularity and the RULER approach, Yale Center for Emotional Intelligence. See also Permission to Feel (2019).

Anchor 02 · Holding the Score Accountable

Rubrics Are What Make Machine Scoring Accountable

Dr. Teresa Ober & colleagues
ETS · Journal of Educational Data Mining

In a 2026 study in the Journal of Educational Data Mining, Teresa Ober (ETS) and colleagues tested whether large language models could reliably code constructs like persistence and self-efficacy from more than 10,000 turns of authentic student dialogue. The finding that matters most: reliability depended on how precisely the construct was defined. Well-defined constructs with clear behavioral indicators could be coded consistently; loosely defined ones wobbled with every configuration change — same student words, different judgment.

A general-purpose chat model, asked open-endedly whether a student "seems curious," is not a measurement instrument. A defined construct, operationalized in a rubric and checked against expert human coders, is.

This is the line between Human Skills Analytics and a chat-model wrapper. We never ask a model for its opinion of a student. Each durable skill is defined in a rubric, trained educators apply that rubric to tens of thousands of authentic responses, and the machine learning system is trained and validated against their judgment — inside those lanes, inside those definitions, consistently over time. The rubric is what the score is accountable to.

Ober, T., Zhang, S., Zapata-Rivera, D., Schroeder, N., & Botelho, A. (2026). Using LLMs to Identify Indicators of Persistence from Students' Dialogues with a Pedagogical Agent. Journal of Educational Data Mining, 18(1), 208–243. Read the paper →

The Practical Difference

Opinion vs. Measurement

A Chat-Model Wrapper

"This Response Seems Curious"

  • No definition of the construct — the model improvises one per response.
  • Judgment shifts with settings, phrasing, and model version.
  • No human raters to be accountable to; nothing to audit against.
  • Produces commentary. Can't produce a defensible score, so it never claims one.
Rubric-Aligned Measurement

"This Response Scores a 3 on the Curiosity Rubric"

  • The construct is defined in precise language before anything is scored.
  • A rubric operationalizes that definition into observable evidence.
  • Trained educators set the standard; ML is validated against them and versioned.
  • Every score is auditable back through rubric, rater, and definition.
Ruling Out the Alternatives

Not Fluency. Not Vocabulary.
Not Length.

A precise definition and an accountable rubric establish what we intend to measure. On their own they don't prove the scores behave that way — a rubric can still be applied in a way that quietly rewards polish. So the definitions are only the beginning of the argument. Four kinds of evidence test whether a curiosity score is actually about curiosity.

Discriminant · Check 01

Is It Just Rewarding Students Who Talk Well?

If a curiosity score were really a fluency score, it would rise and fall with response length, polish, and vocabulary sophistication. So scores are examined against exactly those features — and against the other seven anchors — using multitrait–multimethod analysis, which separates how much of a score reflects the skill being measured from how much reflects the way the evidence was collected. A score that mostly tracked eloquence would show up there.

Structural · Check 02

Are the Eight Anchors Really Eight Things?

If durable skills collapsed into one undifferentiated "strong student" impression, confirmatory factor analysis would reveal a single factor rather than eight. The rubrics are built so the anchors stay distinguishable in students' own evidence — a student can score high on curiosity and low on perseverance in the same response — and the measurement structure has to reproduce that separation, not smooth it away.

Convergent · Check 03

Does It Agree With What Educators Already See?

A valid score shouldn't be a surprise to the adults who know the student. Scores are checked against independent judgments of the same skill — calibrated rater consensus first, then district-held indicators the scoring system never sees. Agreement is evidence the construct is real; disagreement is a finding we investigate rather than tune away.

Fairness · Check 04

Does a Level 3 Mean the Same Thing for Every Student?

A rubric that rewards standard academic English isn't measuring a durable skill; it's measuring proximity to a dialect. Scores are examined for measurement invariance across student groups, home language, and response length, so the same evidence earns the same score whoever is speaking. Authentic voice is scored for what the student says — not for how closely it resembles an essay.

Models are versioned and benchmarked against human raters using standard agreement metrics; exact figures are reported per assessment and model version on the ML Model Card, with independent psychometric validation detailed in Case Studies.

The Construct Chain

From Definition to Every Score

Construct validity is an unbroken chain from what the research says a skill is to the score a student receives. Nothing in the chain is invented by the technology.

Step 1 · Language

Precise Definition

Each durable skill is defined in exact language, grounded in published developmental research.

Step 2 · Rubric

Rubric-Aligned Criteria

The definition becomes an asset-based rubric describing observable evidence at each level.

Step 3 · Raters

Trained Human Raters

Calibrated educators apply the rubric to tens of thousands of authentic responses. Their judgment is the standard.

Step 4 · Model

Versioned ML

The system is trained on rater judgment, validated against it, then frozen — the same evidence always earns the same score.