Systems
Multi-Tenant Legal RAG Assistant
A legal-document retrieval system with two ingestion engines, hybrid search and per-session document isolation: a FastAPI backend behind a Cloudflare tunnel serving a static React SPA.
- Built
- 2026
- Role
- Project lead
- Unit tests, no external deps
- 210Unit tests, no external deps
- Full unit suite runtime
- ~6 sFull unit suite runtime
- Isolated Python environments
- 3Isolated Python environments
- VRAM ceiling designed against
- 8 GBVRAM ceiling designed against
Overview
Legal questions are unusually hostile to pure vector search. The thing that makes a question specific is almost always a rare identifier: a section number, a case citation, a statute short-name. That is precisely the token an embedding model averages away. Ask about “Section 302” and a dense retriever will happily hand back everything about criminal procedure in general.
So this system runs two retrievers and fuses them. BM25 catches the literal identifier; dense vectors catch the paraphrase. The interesting decision is how they are combined: by rank, using reciprocal rank fusion, rather than by score. A cosine distance and a BM25 score share no scale and no distribution, so any normalisation you pick between them is really a hyperparameter fitted to one corpus, and silently wrong on the next one.
Around that sits the part nobody demos: two incompatible document parsers running as isolated subprocesses, a Celery queue that has to respect an 8 GB VRAM ceiling, and a tenancy model that fails closed on both the write and the read path.
Architecture
React SPA (GitHub Pages)
│ HTTPS
▼
Cloudflare Tunnel
│
▼
FastAPI ──────────► SQLite (WAL) document + job state
│
├──────────────► ChromaDB :8001 3 collections
│
└──► Redis :6379 ──► Celery (solo pool)
│
├─► subprocess: marker-pdf (.venv-marker)
└─► subprocess: Unlimited-OCR (.venv-ocr)What it does
- Hybrid retrieval: BM25 lexical search and dense vector search fused with reciprocal rank fusion.
- Two selectable ingestion engines. marker-pdf for layout-faithful extraction, Unlimited-OCR for scanned and handwritten documents.
- Per-session tenancy: a UUID4 acts as a capability token, enforced independently on the embedder and the query router.
- Practice-area tagging combining metadata filtering, gentle centroid steering, and a widened top-k to compensate.
- Background ingestion through Celery with live job state persisted to SQLite in WAL mode.
- Static React SPA on GitHub Pages talking to a self-hosted backend through a Cloudflare tunnel, with no cloud bill.
How it works
- Reciprocal rank fusion over ranks rather than scores, so the two retrievers never need a shared scale.
- Three-virtualenv architecture: marker-pdf and Unlimited-OCR pin irreconcilable transformers and pillow versions, so both run as subprocesses and the application environment stays entirely torch-free.
- Contextual chunk embedding: each chunk is vectorised as previous-tail plus chunk plus next-head (±220 characters) while the stored and cited text remains the chunk itself. The vector sees the neighbourhood; the citation stays exact.
- VRAM budgeting for a 3.3B-parameter bf16 OCR model plus SAM-ViT-B and CLIP-L: max length 32768 down to 8192, image size 1024 down to 640, page batch 4, Celery on a solo pool.
- Fail-closed tenancy: the embedder refuses unscoped writes and the router raises on an unfiltered query, so a missing session id is an error rather than a data leak.
- 210 unit tests running with no external dependencies, plus a live smoke test asserting twenty properties against a real Chroma server.
Problems and answers
Dense retrieval loses the rare identifier that makes a legal question specific.
Run BM25 alongside it and fuse by rank, so the lexical leg can rescue a query the embedder blurred.
The two best document parsers pin mutually incompatible versions of transformers and pillow.
Give each its own virtualenv and invoke it as a subprocess; the app environment never imports torch at all.
Upstream OCR defaults OOM immediately on an 8 GB consumer card.
Cut max sequence length and image resolution, batch pages in fours, and pin Celery to a solo pool, documented as a real coherence tradeoff rather than hidden.
Multi-tenant document isolation is the kind of thing that fails silently and catastrophically.
Enforce the session scope on both sides of the store, and make the absence of a scope raise rather than default.