Under The Hood
In Your Hands
The Big Picture
The Solutions
The Methods
The Foundations
The Case
The Science
The Proof
The Arc
The People
The Promise
Transparency · Model Card v6.0

How Our ML Scores
Student Voice

Every IMPACTER model release ships with a public Model Card: architecture, training data, human-agreement benchmarks, and fairness checks. Because a score you can't audit is a score you can't trust.

0.000
Avg QWK
Human–ML agreement (Quadratic Weighted Kappa). Industry threshold: ≥ 0.70.
0.0%
Avg Exact Accuracy
ML scores exactly matching trained human raters.
0.0%
Avg Adjacent Accuracy
Within ±1 point of human scores — the human–human benchmark.
01 · The Science

Validated Rubrics.
Not Black-Box LLMs.

Eight competencies, each benchmarked against trained human raters. Click any competency to see the linguistic markers the model actually reads.

Agreement with Human Raters,
Trait by Trait.

1.0 = perfect agreement · 0.70 = production threshold

Validated by
Dr. Andrew Arnold
Senior Scientific Advisor
  • Principal ML Engineer, Shopify
  • Chief Scientist, AWS CodeWhisperer
  • PhD in Machine Learning, Carnegie Mellon
  • Architect of IMPACTER's scoring pipeline
Dashed line = 0.70 industry threshold
≥ 0.70
00.250.500.751.0

    SOURCE: HARVARD MCC DUAL-AXIS RUBRIC · IMPACTER MODEL CARD v6.0
    02 · Overview & Scope

    What These Models Do —
    And What They Don't.

    The v6.0 suite scores K-12 open-ended responses on eight human-skills competencies, calibrated to human expert scoring of student reasoning, reflection, and narrative evidence.

    Intended Use

    • Scoring student reflections for Portrait of a Graduate durable-skills competencies
    • Behavioral health and social-emotional reflection scoring
    • Workforce and career readiness / CTE assessments
    • Short-form constructed-response tasks

    Out of Scope

    • High-stakes summative academic content scoring
    • Individual diagnostic or clinical decisions without human review
    • Any use outside rubric-anchored prompts and approved task types
    03 · Training & Governance

    Trained on Authentic Voice.
    Frozen for Audit.

    Models train on district-provided responses — written and transcribed speech — hand-scored by human experts on a 0–4 scale, then evaluated on held-out sets of newly collected responses.

    1

    Registered

    Every new version is registered in MLflow with full lineage.

    2

    Validated

    Promoted to Production only if QWK improves over the current model — with SMD and subgroup fairness checks on every release.

    3

    Frozen

    Production models are frozen for deterministic scoring — identical responses receive identical scores, every administration.

    Model Configuration

    Architecture
    DeBERTa-v3-base · 184M params
    Version
    MODEL SUITE v6.0
    Loss Function
    CORAL ORDINAL REGRESSION
    Platform
    DATABRICKS + MLFLOW REGISTRY
    Score Scale
    0–4 RUBRIC-ALIGNED ORDINAL
    Promotion Gates
    QWK ≥ 0.70 · DEGRADATION ≤ 0.10 · |SMD| < 0.15
    04 · Fairness

    Checked for Drift
    Before Every Promotion.

    Distributional SMD checks — by score band and response length — show no systematic drift. Full subgroup analysis (EL status, race/ethnicity) runs on partner-specific human-scored validation sets.

    Structural SMD — all bands well under threshold

    THRESHOLD: |SMD| < 0.15 · OBSERVED: 0.03–0.08
    |SMD| < 0.15

    Version note: All statistics report Model Suite v6.0 performance on held-out validation data. Performance varies by assessment, prompt design, and cohort; each release publishes updated figures. Emotional inference and biometric capture are disabled for K-12 deployments.

    Measurement Your Review Committee
    Can Interrogate.

    Get the full Model Card, scoring rubric, and technical architecture documentation — or walk through it live with our data science team.