
Overview01
Start here: grading error down ~23% against official College Board scores.
Most test-prep sites hand every student the same practice set. Academy of Testers tries to work out what this student knows and ask the question that will teach us the most.
The platform covers AP and SAT: curated practice exams, subject resources, unit overviews, an AI study assistant, and two ML systems I designed and built. The first is an adaptive SAT math engine that tracks mastery of 8 skills. The second is a retrieval-augmented essay grader for AP free-response questions.
- 01
Architected an adaptive learning engine that models per-student skill mastery by combining Bayesian Knowledge Tracing with Item Response Theory, selecting questions to maximize information gain at the learner's estimated ability. I extended the 8-skill model with forgetting-curve decay and prerequisite-graph penalty propagation, and it powers the mastery radar.
- 02
Improved AP essay-grading consistency, measured as a ~23% reduction in scoring error (QWK 0.62) against official College Board scores, by building a RAG pipeline. It embeds each free response, retrieves rubric clauses and score-banded exemplars from a Postgres vector store by cosine similarity, and returns calibrated results as constrained JSON.
System architecture02
A monorepo with a Spring Boot 3.2 / Java 17 API, a React 18 + TypeScript + Vite frontend, and PostgreSQL 16. Flyway migrations are the only way the schema changes. The frontend deploys to Vercel and the API and database run on Render. Three subsystems hang off the student: the study experiences, the AI services (chat + FRQ grading over a shared RAG layer) and the SAT adaptive engine.

The mastery radar03
Each student gets a vector of eight mastery weights, one per SAT math skill. Every weight is a probability that the skill is learned. The radar shows that vector directly. The dashed ring is the 0.85 mastery threshold, and the engine aims practice at the gaps inside it.

Solid = where you are now. Hollow = your best. The gap is forgetting.
After every adaptive session the same radar comes back with the session's result, so students see exactly which skills moved.

The radar is hand-rolled SVG plus the motion package, themed through CSS variables so it picks up
every site theme. There's no charting library. The client does no model math: weights arrive
already decayed from the server, so a stray Math.exp in the frontend would be a bug.
Adaptive engine, piece by piece04
The engine lives in com.aot.sat.engine as pure functions: no Spring, no repositories,
no now(). Clocks and data are passed in, so the whole model can be tested without a database.
Five pieces fit together:
1 · Bayesian Knowledge Tracing
Standard four-parameter BKT (P(T)=0.10, P(S)=0.10, P(G)=0.25 for 4-option multiple choice). Each answer gives a Bayesian posterior and then a learning transition. Then I damp the step: only 25% of an upward move is applied, but 50% of a downward one, so mastery is easier to lose than to earn. One lucky guess can't push a student over the threshold.
| Starting at w = 0.40 | BKT posterior | + learning | Damped step | New weight |
|---|---|---|---|---|
| Correct answer | 0.706 | 0.735 | 25% of +0.335 | 0.484 |
| Wrong answer | 0.082 | 0.173 | 50% of −0.227 | 0.287 |
Worked from the engine's constants. A wrong answer drops the weight by 0.113, which is about 1.4× the gain from a right one.
Why does a wrong answer move mastery more than a right one?
Gains are damped to 25% of the step and losses to 50%, so one lucky guess can't fake mastery.
So mastery is easier to lose than to earn. Got it.
2 · Prerequisite-graph propagation
Skills form a directed graph (e.g. linear functions depend on algebra). When a student misses a question, part of that drop flows one level upstream to each prerequisite. Correct answers don't propagate.
3 · Forgetting curve
Weights decay exponentially toward a floor of 0.20, and the decay is applied lazily on read.
MasteryService is the only place weights are read, so a stale value can't leak out. Decay can
only pull a weight down. A weight already below the floor never drifts up toward it.
Skills fade without practice. This is the exact curve the engine uses.
4 · IRT question selection
Each mastery weight maps to an ability θ on the logit scale. Every item has 3PL parameters (a, b, c), and its Fisher information at the student's θ measures how much the answer will tell us. Each candidate's score blends three terms:
5 · Data hygiene
- The client never sees answers:
correctIndex, explanation,irt_band difficulty are withheld until after submission. Knowing an item's difficulty changes how students answer, which would contaminate the data used to recalibrateirt_b. - Question IDs are permanent, because the response history references them forever.
- The question bank's source of truth is reviewable JSON. A deterministic, idempotent normalizer turns raw practice tests into that JSON, and a generator emits the Flyway seed migration from it.
What a session looks like
Each question is tagged with the skill it tests. After answering, the student sees right or wrong and a worked explanation; the weight updates happen on the server. The session bar has a timer, pause, end, and the same tools as the real test.


RAG essay grader05
AP free-response essays are graded on detailed rubrics, and an LLM that's simply asked to grade an essay 0–6 drifts. The grader grounds every call in the rubric, the matching rubric clauses, and real scored student exemplars, then checks itself against official College Board scores.
Vectors in Postgres, no vector DB
The corpus is small (AP rubrics, a handful of exemplars per prompt, per-skill curriculum). So embeddings
live in Postgres as JSON float arrays, and cosine ranking runs in a small pure-Java VectorMath engine
over rows pre-filtered by metadata (corpus, subject, prompt, score band). No pgvector extension is needed,
and swapping to pgvector later would only change the repository layer. Every chunk carries a unique
ref_id that the model must quote, and each retrieval is logged with rank and similarity for weekly spot checks.
Why diversify by score band?
Nearest-neighbor retrieval naturally returns mid-band essays, the ones most similar to a typical response. The grader then compresses scores toward the middle. Pulling a wider pool and stratifying it so the prompt always sees low, mid and high anchors fixed most of that compression.
Evaluation: a real A/B against official scores
The harness extracts verbatim student responses and their official scores from College Board "Sample Responses + Scoring Commentary" packets. It keeps only typed responses and drops scanned handwriting and garbled OCR. Then it grades each one twice: a baseline arm with no rubric and no retrieval, and the grounded arm. Test essays (2023/2024) never touch the RAG corpus (2025), so nothing leaks.
APUSH is the honest test: never tuned on, and error still fell 22%.
| Set | n | Arm | MAE | Exact | Within-1 | QWK |
|---|---|---|---|---|---|---|
| English Lang + Lit | 39 | baseline | 0.974 | 33% | 77% | 0.590 |
| grounded | 0.744 | 38% | 87% | 0.624 | ||
| APUSH LEQ (held out) | 8 | baseline | 1.125 | 25% | 75% | 0.521 |
| grounded | 0.875 | 38% | 75% | 0.569 |
Reported honestly: the samples are small (39 + 8), the English set was tuned over about four passes, and APUSH is the clean, untuned check. Packets don't include source passages, so both arms grade without them. That's fine for the A/B difference, but these aren't production accuracy numbers. After this run the grader config was frozen.
Where students use it

Real 2025 prompt on the left, rubric rows and AI grading on the right.
The rest of the platform06
Around the two ML systems sits the rest of a real product, which I built too:
Subject hubs

My AP Planner

Testy, the AI study assistant
