Under The Hood
In Your Hands
The Big Picture
The Solutions
The Methods
The Foundations
The Case
The Science
The Proof
The Arc
The People
The Promise
The Evidence

Research Library

Three kinds of evidence, kept honestly separate: independent research we didn't write, our own technical documentation of how the scoring engine actually works today, and the applied field guides we build alongside district partners.

Independent Research

We Didn't Write These

Peer-reviewed and academic work that informs our approach — cited on their own merits, not as marketing copy.

Diagram of a seven-turn student-tutor dialogue with an inferred persistence trajectory rising across the conversation
Journal of Educational Data Mining · Vol. 18, No. 1 · 2026

Using LLMs to Identify Indicators of Persistence from Students’ Dialogues with a Pedagogical Agent

Ober, Zhang, Zapata-Rivera, Schroeder & Botelho — ETS & University of Florida

Directly on point: a study of how well large language models can code psychological constructs like persistence and self-efficacy from real student dialogue. Its finding — that well-defined constructs code reliably while ambiguous ones don't — is a design input for our own scoring pipeline, cited directly in our v6.0 technical note below.

View PDF ↗
Diagram of NCS-1 signaling at a dentate gyrus synapse promoting long-term potentiation
Neuron (Cell Press) · Vol. 63 · 2009

NCS-1 in the Dentate Gyrus Promotes Exploration, Synaptic Plasticity, and Rapid Acquisition of Spatial Memory

Saab, Georgiou, Nath, Lee, Wang, et al. — Mount Sinai Hospital, University of Toronto

Foundational neuroscience on how exploratory behavior forms in the brain, from a genetically engineered mouse model. We include it as background on the biology of curiosity — not as evidence about our own product, which it doesn't test or mention.

View PDF ↗
Diagram of directed causal influence between four time-series variables strengthening across sequential windows
Carnegie Mellon University

Temporal Causal Modeling with Graphical Granger Methods

Andrew Arnold — Machine Learning Department, CMU

The causal-discovery technique behind how we think about upstream drivers of growth (e.g., which shifts in mindset actually precede gains elsewhere) rather than just correlated outcomes.

View PDF ↗
Our Technical Methodology

How Our Scoring Engine Has Evolved

We're publishing this progression on purpose. Early concept papers explored a broader, more experimental scoring model; the system running in production today is narrower, deterministic, and independently benchmarked. Read them as a record of how the thinking sharpened, not as three descriptions of the same system.

Current — Model Suite v6.0

A Stratified, Human-in-the-Loop Sampling Layer for Continuous Calibration of a Production Rubric-Aligned Ordinal Scoring Pipeline

The technical note describing what's actually running today: a fine-tuned DeBERTa-v3-base encoder with a CORAL ordinal regression head, scoring against a combined Making Caring Common / CASEL-aligned rubric, with a weekly stratified human-review loop that feeds continuous recalibration.

View PDF ↗
Cover of the v6.0 Ordinal Scoring Pipeline technical note
0.918
Average QWK across the eight character competencies
77.3%
Exact-match agreement with human raters
98.0%
Adjacent (±1) agreement with human raters
184M
Parameters, DeBERTa-v3-base encoder, fine-tuned per competency
Earlier Concept Papers
A student and a teacher shaking hands in a school hallway, overlaid with a circuit-pattern graphic
Early Vision Paper

Emotional Intelligence Meets Machine Learning: IMPACTER's Blueprint for Elevating Student Voice

An early framing of the case for machine learning in reading student language for social-emotional signal — the vision that eventually became the rubric-first, deterministic architecture running today.

View PDF ↗
A teacher helping two students working on laptops in a classroom
Early Concept Paper

Multi-Dimensional Scoring Methodology: A Holistic Approach to Assessing Social-Emotional Growth

An earlier scoring model (DistilBERT-based, with peer-relative and engagement adjustments) exploring how much context should shape a score. The current system takes a narrower, more defensible path: one deterministic rubric-aligned score per response, no peer-relative adjustment.

View PDF ↗
Two children shaking hands over a chessboard at a tournament
Early Concept Paper

The Impact Potential Score: An Elo-Inspired Growth Model

An exploration of chess-style relative rating systems as a growth metaphor. The current system uses a fixed 0–4 ordinal rubric scale instead — easier for a teacher or parent to read, and directly comparable to standard constructed-response scoring (AP, NAEP, ETS performance items).

View PDF ↗
Applied With Partners

Field Guides, Co-Developed With Districts

Not peer-reviewed studies — practitioner-facing guides that translate the methodology above for the people who have to act on it.

Cover of A Human-Friendly Guide to IMPACTER's Machine Learning Model, showing a neon hexagon over a collage of student photos
A Resource For Our Partners

A Human-Friendly Guide to IMPACTER's Machine Learning Model

Written for school and district leaders, not engineers: how the scoring system works, why it's trustworthy, and what makes it different from traditional assessment approaches, in plain language throughout.

Training Data Human-in-the-Loop Review Model Versioning Glossary + FAQ
View PDF ↗
Cover of the Behavioral Health Screening Guide, co-developed with Santa Clara County Office of Education, showing the title and the six behavioral health domains wheel
With Santa Clara County Office of Education

From Self-Report to Student Voice: A Guide to Performance-Based Behavioral Health Screening

Written for district behavioral health leaders, school psychologists, and county MTSS teams evaluating screener options. Walks through where Likert-scale self-report screeners (PHQ-9, GAD-7, BASC-3 BESS) fall short in K–12 settings, and how a performance-based, rubric-validated approach fits alongside them within a Multi-Tiered System of Support.

MTSS-Aligned Screener Comparison Implementation Models
View PDF ↗