How Our ML Scores
Student Voice
Every IMPACTER model release ships with a public Model Card: architecture, training data, human-agreement benchmarks, and fairness checks. Because a score you can't audit is a score you can't trust.
Validated Rubrics.
Not Black-Box LLMs.
Eight competencies, each benchmarked against trained human raters. Click any competency to see the linguistic markers the model actually reads.
Agreement with Human Raters,
Trait by Trait.
1.0 = perfect agreement · 0.70 = production threshold
- Principal ML Engineer, Shopify
- Chief Scientist, AWS CodeWhisperer
- PhD in Machine Learning, Carnegie Mellon
- Architect of IMPACTER's scoring pipeline
What These Models Do —
And What They Don't.
The v6.0 suite scores K-12 open-ended responses on eight human-skills competencies, calibrated to human expert scoring of student reasoning, reflection, and narrative evidence.
Intended Use
- Scoring student reflections for Portrait of a Graduate durable-skills competencies
- Behavioral health and social-emotional reflection scoring
- Workforce and career readiness / CTE assessments
- Short-form constructed-response tasks
Out of Scope
- High-stakes summative academic content scoring
- Individual diagnostic or clinical decisions without human review
- Any use outside rubric-anchored prompts and approved task types
Trained on Authentic Voice.
Frozen for Audit.
Models train on district-provided responses — written and transcribed speech — hand-scored by human experts on a 0–4 scale, then evaluated on held-out sets of newly collected responses.
Registered
Every new version is registered in MLflow with full lineage.
Validated
Promoted to Production only if QWK improves over the current model — with SMD and subgroup fairness checks on every release.
Frozen
Production models are frozen for deterministic scoring — identical responses receive identical scores, every administration.
Model Configuration
- Architecture
- DeBERTa-v3-base · 184M params
- Version
- MODEL SUITE v6.0
- Loss Function
- CORAL ORDINAL REGRESSION
- Platform
- DATABRICKS + MLFLOW REGISTRY
- Score Scale
- 0–4 RUBRIC-ALIGNED ORDINAL
- Promotion Gates
- QWK ≥ 0.70 · DEGRADATION ≤ 0.10 · |SMD| < 0.15
Checked for Drift
Before Every Promotion.
Distributional SMD checks — by score band and response length — show no systematic drift. Full subgroup analysis (EL status, race/ethnicity) runs on partner-specific human-scored validation sets.
Structural SMD — all bands well under threshold
THRESHOLD: |SMD| < 0.15 · OBSERVED: 0.03–0.08Version note: All statistics report Model Suite v6.0 performance on held-out validation data. Performance varies by assessment, prompt design, and cohort; each release publishes updated figures. Emotional inference and biometric capture are disabled for K-12 deployments.
Measurement Your Review Committee
Can Interrogate.
Get the full Model Card, scoring rubric, and technical architecture documentation — or walk through it live with our data science team.