Welcome to Braedyn Thompson's portfolio!
See the roles I'm targeting →

Now playing
Berkeley Lab SeismicSoCal BearLM
PROJECT 01 · ED-TECH · ML

Academy of
Testers

A full-stack AP/SAT study platform with an adaptive learning engine that models what each student knows, and a RAG essay grader calibrated against official College Board scores.

Jan 2023 – Jun 2026ML & AI Software DeveloperBerkeley, CA
Academy of Testers homepage
Live siteacademyoftesters.com: pick AP (29 subjects: unit reviews, real 2025 free-response questions, timed mocks) or SAT (adaptive practice, topic lessons, full-length tests).

Overview01

Start here: grading error down ~23% against official College Board scores.

-23%AP grading error (MAE)vs. naive LLM grader
0.62QWK vs. official scoresEnglish Lang + Lit, n=39
8skills in the mastery modelSAT math radar axes
4models fusedBKT · IRT · decay · prereq graph

Most test-prep sites hand every student the same practice set. Academy of Testers tries to work out what this student knows and ask the question that will teach us the most.

The platform covers AP and SAT: curated practice exams, subject resources, unit overviews, an AI study assistant, and two ML systems I designed and built. The first is an adaptive SAT math engine that tracks mastery of 8 skills. The second is a retrieval-augmented essay grader for AP free-response questions.

  • 01

    Architected an adaptive learning engine that models per-student skill mastery by combining Bayesian Knowledge Tracing with Item Response Theory, selecting questions to maximize information gain at the learner's estimated ability. I extended the 8-skill model with forgetting-curve decay and prerequisite-graph penalty propagation, and it powers the mastery radar.

  • 02

    Improved AP essay-grading consistency, measured as a ~23% reduction in scoring error (QWK 0.62) against official College Board scores, by building a RAG pipeline. It embeds each free response, retrieves rubric clauses and score-banded exemplars from a Postgres vector store by cosine similarity, and returns calibrated results as constrained JSON.

System architecture02

A monorepo with a Spring Boot 3.2 / Java 17 API, a React 18 + TypeScript + Vite frontend, and PostgreSQL 16. Flyway migrations are the only way the schema changes. The frontend deploys to Vercel and the API and database run on Render. Three subsystems hang off the student: the study experiences, the AI services (chat + FRQ grading over a shared RAG layer) and the SAT adaptive engine.

Academy of Testers architecture diagram
ArchitectureStudent-facing web routes call the platform API. AI chat and FRQ grading share RAG retrieval over a chunk store, and the SAT adaptive engine (diagnostic, sessions, mastery tracking) persists learning state to PostgreSQL.

The mastery radar03

Each student gets a vector of eight mastery weights, one per SAT math skill. Every weight is a probability that the skill is learned. The radar shows that vector directly. The dashed ring is the 0.85 mastery threshold, and the engine aims practice at the gaps inside it.

SAT mastery radar
ScreenshotThe mastery map on the SAT dashboard. Solid dots are current mastery; hollow outer dots are each skill's peak, so the gap between them is what the forgetting curve has taken back. The dashed outer ring is the mastery target.

Solid = where you are now. Hollow = your best. The gap is forgetting.

After every adaptive session the same radar comes back with the session's result, so students see exactly which skills moved.

Session complete screen with radar
ScreenshotEnd of an adaptive session: the score for the session and the updated eight-skill radar, with options to practice again or return to the dashboard.

The radar is hand-rolled SVG plus the motion package, themed through CSS variables so it picks up every site theme. There's no charting library. The client does no model math: weights arrive already decayed from the server, so a stray Math.exp in the frontend would be a bug.

Adaptive engine, piece by piece04

The engine lives in com.aot.sat.engine as pure functions: no Spring, no repositories, no now(). Clocks and data are passed in, so the whole model can be tested without a database. Five pieces fit together:

01Diagnostic24 items: 8 skills × easy/med/hard → prior-blended starting weights
02BKT updateBayesian posterior on each answer + learning transition
03Prereq penaltya miss bleeds into upstream skills (κ = 0.35)
04Forgettinglazy exponential decay toward a 0.20 floor
05IRT selectiongap + Fisher info − recency, softmax top-5

1 · Bayesian Knowledge Tracing

Standard four-parameter BKT (P(T)=0.10, P(S)=0.10, P(G)=0.25 for 4-option multiple choice). Each answer gives a Bayesian posterior and then a learning transition. Then I damp the step: only 25% of an upward move is applied, but 50% of a downward one, so mastery is easier to lose than to earn. One lucky guess can't push a student over the threshold.

Starting at w = 0.40BKT posterior+ learningDamped stepNew weight
Correct answer0.7060.73525% of +0.3350.484
Wrong answer0.0820.17350% of −0.2270.287

Worked from the engine's constants. A wrong answer drops the weight by 0.113, which is about 1.4× the gain from a right one.

Why does a wrong answer move mastery more than a right one?

Gains are damped to 25% of the step and losses to 50%, so one lucky guess can't fake mastery.

So mastery is easier to lose than to earn. Got it.

2 · Prerequisite-graph propagation

Skills form a directed graph (e.g. linear functions depend on algebra). When a student misses a question, part of that drop flows one level upstream to each prerequisite. Correct answers don't propagate.

wprereq ← clamp( wprereq − κ · σedge · Δskill ) // κ = 0.35, depth 1, no recursion // the miss above (Δ = 0.113) costs a full-strength prerequisite 0.040

3 · Forgetting curve

Weights decay exponentially toward a floor of 0.20, and the decay is applied lazily on read. MasteryService is the only place weights are read, so a stale value can't leak out. Decay can only pull a weight down. A weight already below the floor never drifts up toward it.

Forgetting curve: w(t) = 0.20 + (w₀ − 0.20) · e^(−0.0112 t)00.250.50.7510306090120150180days since the skill was last practicedmastery weightmastered (0.90)developing (0.60)floor 0.20
FigureWith λ = 0.0112/day, the part of a weight above the floor halves roughly every 62 days (marked point), so a skill mastered in the fall is flagged for review well before spring.

Skills fade without practice. This is the exact curve the engine uses.

4 · IRT question selection

Each mastery weight maps to an ability θ on the logit scale. Every item has 3PL parameters (a, b, c), and its Fisher information at the student's θ measures how much the answer will tell us. Each candidate's score blends three terms:

score = 0.45 · skill_gap(w) + 0.40 · info(θ; a,b,c) / infomax − 0.15 · e^(−days_since_seen / 7) // then sample from the top 5 with a softmax (τ = 0.15), so identical states don't repeat questions
3PL item information: which question tells us the most about this student?0.00.10.20.3-4-2024student ability θ (logit of mastery weight)Fisher informationeasy (b = −1.5)medium (b = 0)hard (b = +1.5)
FigureFisher information under the 3PL model (a = 1, c = 0.25). The marked student (w = 0.45, θ ≈ −0.20) learns the most from medium items. Easy items tell us almost nothing about them.

5 · Data hygiene

  • The client never sees answers: correctIndex, explanation, irt_b and difficulty are withheld until after submission. Knowing an item's difficulty changes how students answer, which would contaminate the data used to recalibrate irt_b.
  • Question IDs are permanent, because the response history references them forever.
  • The question bank's source of truth is reviewable JSON. A deterministic, idempotent normalizer turns raw practice tests into that JSON, and a generator emits the Flyway seed migration from it.

What a session looks like

Each question is tagged with the skill it tests. After answering, the student sees right or wrong and a worked explanation; the weight updates happen on the server. The session bar has a timer, pause, end, and the same tools as the real test.

Adaptive session question
ScreenshotA question from an adaptive SAT Math session (question 5 of 10, tagged Exponential Functions), answered correctly with its worked explanation. Math renders as LaTeX.
Calculator and reference sheet in a session
ScreenshotTest-day tools inside a session: a graphing calculator and the official-style reference sheet, open alongside the question.

RAG essay grader05

AP free-response essays are graded on detailed rubrics, and an LLM that's simply asked to grade an essay 0–6 drifts. The grader grounds every call in the rubric, the matching rubric clauses, and real scored student exemplars, then checks itself against official College Board scores.

01Embedthe student response + rubric as the retrieval query
02Retrievepool of 16 chunks by cosine similarity, filtered by subject/prompt
03Diversifystratify by score band → inject 9 (low / mid / high anchors)
04GradeGPT-4o, temperature 0, calibrated anti-compression prompt
05ConstrainJSON response format → per-row points + cited ref_ids

Vectors in Postgres, no vector DB

The corpus is small (AP rubrics, a handful of exemplars per prompt, per-skill curriculum). So embeddings live in Postgres as JSON float arrays, and cosine ranking runs in a small pure-Java VectorMath engine over rows pre-filtered by metadata (corpus, subject, prompt, score band). No pgvector extension is needed, and swapping to pgvector later would only change the repository layer. Every chunk carries a unique ref_id that the model must quote, and each retrieval is logged with rank and similarity for weekly spot checks.

Why diversify by score band?

Nearest-neighbor retrieval naturally returns mid-band essays, the ones most similar to a typical response. The grader then compresses scores toward the middle. Pulling a wider pool and stratifying it so the prompt always sees low, mid and high anchors fixed most of that compression.

Evaluation: a real A/B against official scores

The harness extracts verbatim student responses and their official scores from College Board "Sample Responses + Scoring Commentary" packets. It keeps only typed responses and drops scanned handwriting and garbled OCR. Then it grades each one twice: a baseline arm with no rubric and no retrieval, and the grounded arm. Test essays (2023/2024) never touch the RAG corpus (2025), so nothing leaks.

Mean absolute error vs. official College Board scores (lower is better)0.0000.5001.000Baseline (no rubric, no retrieval)RAG-grounded graderEnglish Lang + LitEnglish Lang + Lit — Baseline (no rubric, no retrieval): 0.9740.974English Lang + Lit — RAG-grounded grader: 0.7440.744APUSH LEQ (held out)APUSH LEQ (held out) — Baseline (no rubric, no retrieval): 1.1251.125APUSH LEQ (held out) — RAG-grounded grader: 0.8750.875
FigureMAE fell 24% on English and 22% on APUSH, a subject and rubric the grader was never tuned on.
Agreement with official scores (higher is better)0.000.501.00BaselineGroundedEnglish, QWKEnglish, QWK — Baseline: 0.590.59English, QWK — Grounded: 0.620.62APUSH, QWKAPUSH, QWK — Baseline: 0.520.52APUSH, QWK — Grounded: 0.570.57English, within-1English, within-1 — Baseline: 0.770.77English, within-1 — Grounded: 0.870.87
FigureQuadratic weighted kappa rose on both sets. Within-one-point agreement on English went from 77% to 87%.

APUSH is the honest test: never tuned on, and error still fell 22%.

SetnArmMAEExactWithin-1QWK
English Lang + Lit39baseline0.97433%77%0.590
grounded0.74438%87%0.624
APUSH LEQ (held out)8baseline1.12525%75%0.521
grounded0.87538%75%0.569

Reported honestly: the samples are small (39 + 8), the English set was tuned over about four passes, and APUSH is the clean, untuned check. Packets don't include source passages, so both arms grade without them. That's fine for the A/B difference, but these aren't production accuracy numbers. After this run the grader config was frozen.

Where students use it

AP FRQ practice and grading screen
ScreenshotThe FRQ practice screen for a real 2025 AP English Language synthesis question: the official prompt PDF on the left, and on the right the rubric rows (thesis, evidence & commentary, sophistication), a suggested 55-minute timer, the response box and Grade my answer. Tabs open the scoring guide and scored samples, and drafts save as you type.

Real 2025 prompt on the left, rubric rows and AI grading on the right.

The rest of the platform06

Around the two ML systems sits the rest of a real product, which I built too:

Exam hubsBrowse AP and SAT, search and filter subjects, and drill into subject pages.
ResourcesPractice exams, unit overviews, topical review and video resources, with PDFs served inline.
AI study chatCurriculum-grounded assistant on the same RAG layer, with per-user rate limiting.
FlashcardsCards, stacks and per-card progress tracking.
AuthJWT access/refresh tokens, email verification, account lockout.
StreaksSAT streaks, streak repair and a focus mode with user preferences.

Subject hubs

AP English Language subject hub
ScreenshotAn AP subject hub (English Language): unit overviews, videos and reference sheets to learn the material, then practice questions, past exams, flash cards, AI-graded FRQ practice, mixed review and a timed mock exam that predicts a 1–5 score.

My AP Planner

My AP Planner mastery map
ScreenshotThe AP planner: a mastery map for each of the student's classes, built from Unit Practice. A question counts as mastered after 3 correct answers (2 in a row), and mastery fades: every 2 weeks without a correct answer it slips a tier.

Testy, the AI study assistant

Testy AI study chat
ScreenshotTesty answering a history question. Questions can be scoped to a subject, and usage is rate-limited per user (the counter shows messages left this hour).