Welcome to Braedyn Thompson's portfolio!
See the roles I'm targeting →

Now playing
Berkeley Lab SeismicSoCal BearLM
PROJECT 03 · RAG · LOCAL LLM

BearLM

A fully local, zero-cost Q&A assistant over UC Berkeley CS and Data Science courses. Every answer is cited to the exact PDF page, and it refuses when the material doesn't cover the question.

Aug 2026 – presentData & AI Engineer$0 / query
recall@1 · retrieval ablation0.56dense0.60hybrid0.82+rerank

Overview01

0.82recall@1from 0.56 dense-only
0.83RAGAS faithfulnessfrom 0.55
0.97Answer relevancyfrom 0.80
10Kchunks indexed10 courses · 768-dim

Course Q&A bots fail in two ways: they retrieve the wrong passage, or they make up an answer when there isn't a right one. BearLM is built to fix both.

BearLM runs entirely on one machine with LangChain, Ollama, Llama 3.1 8B, nomic-embed-text and Chroma, so it needs no API keys and has no cloud bill. It covers CS61A, CS61B, CS70, CS188, CS189, Data 8, Data 100, Data 101, Data 144 and generative ML. Beyond the shared corpus, Projects let you bring your own PDFs and hold multi-turn, grounded conversations over them.

  • 01

    Built a fully local, zero-cost RAG Q&A assistant (LangChain, Ollama, Llama 3.1 8B, nomic-embed-text, Chroma) over ~10K chunks spanning UC Berkeley CS/DS courses. I blended keyword and meaning-based search (BM25 + vector embeddings) with reranking to pull the most relevant sources, lifting recall@1 from 56% to 82% and recall@5 to 86%.

  • 02

    Reduced hallucinated answers by grounding every response strictly in its retrieved sources. That combines a relevance filter that scores and drops weak matches, query rewriting for sharper retrieval, strict context-only prompting, and clickable page-level citations, raising RAGAS faithfulness from 0.55 to 0.83 (0.97 relevance).

  • 03

    Engineered an ETL pipeline that ingests mixed file types (PDF, Markdown, Jupyter), token-chunks them with metadata, and batch-embeds ~10K records with retry/backoff into Chroma, BM25 and SQLite for hybrid search and analytics.

BearLM Ask page with cited answer
ScreenshotThe Ask page: a grounded answer to "What is an Index Scan?" with its sources as clickable course · section tags (Data 101, GenML). Above, the Improve prompt button and the query-rewrite toggle; below, the live course list in the corpus.

Every answer lists exactly where it came from.

Hybrid retrieval02

01Rewriteoptional LLM query rewrite for a sharper search string
02BM25keyword recall: exact terms, function names, notation
03Densenomic-embed-text vectors in Chroma: meaning and paraphrase
04RRF fuseReciprocal Rank Fusion (k = 60) merges both lists
05Rerankbge-reranker-base cross-encoder: 15 candidates → top 4

Why hybrid?

Course questions mix two kinds of language. Some are exact: "__init__", "Dijkstra", a specific CS61B lab name. Others are conceptual: "why does this recursion blow up?" Dense vectors are good with paraphrase and bad with rare exact terms. BM25 is the reverse. Reciprocal Rank Fusion combines them without having to calibrate their scores against each other, and a cross-encoder reranker then reads each question/passage pair together to pick the best four.

The reranker is the big jump: recall@1 from 0.60 to 0.82.

Retrieval ablation · 50-question benchmark generated from real corpus chunks0.000.501.00Dense onlyHybrid BM25 + denseHybrid + rerankerrecall@1recall@1 — Dense only: 0.560.56recall@1 — Hybrid BM25 + dense: 0.600.60recall@1 — Hybrid + reranker: 0.820.82recall@3recall@3 — Dense only: 0.680.68recall@3 — Hybrid BM25 + dense: 0.780.78recall@3 — Hybrid + reranker: 0.860.86recall@5recall@5 — Dense only: 0.760.76recall@5 — Hybrid BM25 + dense: 0.840.84recall@5 — Hybrid + reranker: 0.860.86MRRMRR — Dense only: 0.640.64MRR — Hybrid BM25 + dense: 0.700.70MRR — Hybrid + reranker: 0.840.84
FigureEach stage measurably helps. Hybrid search lifts recall, and the reranker sharpens the top rank (recall@1 0.60 → 0.82, MRR 0.70 → 0.84).

A benchmark with exact gold labels

Instead of hand-labelling relevance, the eval set is generated from real corpus chunks. Each question is written from one specific chunk, so that chunk's (course, section, chunk_id) is the exact answer key. recall@k then asks a precise question: did retrieval surface the true source? It's a synthetic benchmark with real gold labels, 50 questions across the courses.

Anti-hallucination03

Better retrieval isn't enough on its own. The generator also has to stay inside what was retrieved. Four defenses work together:

Confidence gateThe cross-encoder scores every chunk. Below 0.02 it's dropped, and if nothing clears the bar, BearLM refuses instead of guessing.
Strict promptAnswer only from the context, otherwise return an exact refusal string. Temperature 0, 8K context.
Query rewritingOptionally rewrites vague questions into sharper retrieval queries, and the UI shows which query was used.
Page citationsDeduplicated (course, section) citations that open the source PDF at the cited page.

What happens when the course materials don't cover the question?

BearLM refuses instead of guessing, and every answer it does give cites its sources.

No making things up, then. Respect.

RAGAS answer quality · local Llama 3.1 judge0.000.501.00Dense onlyHybrid + reranker + groundingFaithfulnessFaithfulness — Dense only: 0.550.55Faithfulness — Hybrid + reranker + grounding: 0.830.83Answer relevancyAnswer relevancy — Dense only: 0.800.80Answer relevancy — Hybrid + reranker + grounding: 0.970.97
FigureRAGAS with a local Llama 3.1 judge. Better context in gives better-grounded answers out: faithfulness 0.55 → 0.83, relevancy 0.80 → 0.97.

Reported honestly: the RAGAS runs are small (n = 6) with a small local judge, and on one hybrid sample the judge couldn't produce parseable output, so RAGAS dropped it. The retrieval ablation (n = 50) is the stronger evidence. The gate threshold is calibrated to bge's sigmoid scores (off-topic ≈ 0.0, real matches ≥ ~0.04). "Summarize this doc" questions skip the gate, since they don't match any single passage.

Ingestion pipeline & architecture04

01ExtractPDF (pdfplumber), Markdown, Jupyter .ipynb, some sources are cloned course repos
02Chunktoken-length RecursiveCharacterTextSplitter (tiktoken) + {course, section, source, chunk_id}
03Embedbatched nomic-embed-text with retry and exponential backoff
04LoadChroma vectors · BM25 index · SQLite analytics

Re-ingesting rebuilds the whole store, about 10.4K chunks. On Windows, firing thousands of embedding requests at a local server runs into socket exhaustion (WSAENOBUFS), so embedding is batched with retry and backoff. The pipeline then finishes reliably instead of dying halfway through. The BM25 index (~40 s to build over 10K docs) is pre-warmed on a background thread at startup and cached.

A BACKEND=local|openai flag keeps the original OpenAI + MongoDB Atlas path working for side-by-side comparison. Local costs $0.0000/query against roughly $0.0006 for OpenAI, with a median answer time of about 6 s on an 8 GB RTX 4060.

BearLM architecture diagram
ArchitectureThe learner uses the web and Projects interfaces through a FastAPI HTTP API. Answer orchestration (rag.py) combines hybrid retrieval, confidence reranking and the answer model over Ollama. Corpus ingestion extracts, splits and embeds course materials into vector collections.

The product05

BearLM is a full product, not just a script: a FastAPI backend and a hand-built vanilla-JS frontend (marked, DOMPurify, KaTeX, all vendored locally) that renders Markdown tables, code blocks and LaTeX.

ProjectsBring your own PDFs, keep multiple persistent chats, and get project-first answers that fall back to the course corpus.
Corpus scopingPer-project checkboxes choose which courses an answer may fall back on.
Chat with a PDFAd-hoc single-document sessions, strictly grounded with no fallback.
AnalyticsSQLite usage stats, most-cited courses, unanswered questions, and auto-generated practice questions for the weakest area.

Projects

BearLM Projects workspace
ScreenshotA Projects workspace (Data101): multiple saved chats in the sidebar, Reset context, and corpus-fallback scoping. With no PDFs in the project, the answer is marked From course corpus and cites its sections.

Analytics

BearLM analytics page
ScreenshotThe analytics page: questions asked, courses cited, answered vs. no-source counts, the course that needs the most help (with its hot-spot topics and a Generate practice questions button), and most-cited courses.

It spots the weakest course and writes practice questions for it.