Overview01
Course Q&A bots fail in two ways: they retrieve the wrong passage, or they make up an answer when there isn't a right one. BearLM is built to fix both.
BearLM runs entirely on one machine with LangChain, Ollama, Llama 3.1 8B, nomic-embed-text and Chroma, so it needs no API keys and has no cloud bill. It covers CS61A, CS61B, CS70, CS188, CS189, Data 8, Data 100, Data 101, Data 144 and generative ML. Beyond the shared corpus, Projects let you bring your own PDFs and hold multi-turn, grounded conversations over them.
- 01
Built a fully local, zero-cost RAG Q&A assistant (LangChain, Ollama, Llama 3.1 8B, nomic-embed-text, Chroma) over ~10K chunks spanning UC Berkeley CS/DS courses. I blended keyword and meaning-based search (BM25 + vector embeddings) with reranking to pull the most relevant sources, lifting recall@1 from 56% to 82% and recall@5 to 86%.
- 02
Reduced hallucinated answers by grounding every response strictly in its retrieved sources. That combines a relevance filter that scores and drops weak matches, query rewriting for sharper retrieval, strict context-only prompting, and clickable page-level citations, raising RAGAS faithfulness from 0.55 to 0.83 (0.97 relevance).
- 03
Engineered an ETL pipeline that ingests mixed file types (PDF, Markdown, Jupyter), token-chunks them with metadata, and batch-embeds ~10K records with retry/backoff into Chroma, BM25 and SQLite for hybrid search and analytics.

Every answer lists exactly where it came from.
Hybrid retrieval02
Why hybrid?
Course questions mix two kinds of language. Some are exact: "__init__", "Dijkstra", a
specific CS61B lab name. Others are conceptual: "why does this recursion blow up?" Dense vectors are good
with paraphrase and bad with rare exact terms. BM25 is the reverse. Reciprocal Rank Fusion combines them
without having to calibrate their scores against each other, and a cross-encoder reranker then reads each
question/passage pair together to pick the best four.
The reranker is the big jump: recall@1 from 0.60 to 0.82.
A benchmark with exact gold labels
Instead of hand-labelling relevance, the eval set is generated from real corpus chunks. Each
question is written from one specific chunk, so that chunk's (course, section, chunk_id) is the exact
answer key. recall@k then asks a precise question: did retrieval surface the true source? It's a synthetic
benchmark with real gold labels, 50 questions across the courses.
Anti-hallucination03
Better retrieval isn't enough on its own. The generator also has to stay inside what was retrieved. Four defenses work together:
What happens when the course materials don't cover the question?
BearLM refuses instead of guessing, and every answer it does give cites its sources.
No making things up, then. Respect.
Reported honestly: the RAGAS runs are small (n = 6) with a small local judge, and on one hybrid sample the judge couldn't produce parseable output, so RAGAS dropped it. The retrieval ablation (n = 50) is the stronger evidence. The gate threshold is calibrated to bge's sigmoid scores (off-topic ≈ 0.0, real matches ≥ ~0.04). "Summarize this doc" questions skip the gate, since they don't match any single passage.
Ingestion pipeline & architecture04
Re-ingesting rebuilds the whole store, about 10.4K chunks. On Windows, firing thousands of embedding
requests at a local server runs into socket exhaustion (WSAENOBUFS), so embedding is batched with
retry and backoff. The pipeline then finishes reliably instead of dying halfway through. The BM25 index (~40 s
to build over 10K docs) is pre-warmed on a background thread at startup and cached.
A BACKEND=local|openai flag keeps the original OpenAI + MongoDB Atlas path working for
side-by-side comparison. Local costs $0.0000/query against roughly $0.0006 for OpenAI, with a median
answer time of about 6 s on an 8 GB RTX 4060.

The product05
BearLM is a full product, not just a script: a FastAPI backend and a hand-built vanilla-JS frontend (marked, DOMPurify, KaTeX, all vendored locally) that renders Markdown tables, code blocks and LaTeX.
Projects

Analytics

It spots the weakest course and writes practice questions for it.