Skip to content

Systems

Multi-Tenant Legal RAG Assistant

A legal-document retrieval system with two ingestion engines, hybrid search and per-session document isolation: a FastAPI backend behind a Cloudflare tunnel serving a static React SPA.

Built
2026
Role
Project lead
Unit tests, no external deps
210Unit tests, no external deps
Full unit suite runtime
~6 sFull unit suite runtime
Isolated Python environments
3Isolated Python environments
VRAM ceiling designed against
8 GBVRAM ceiling designed against

Overview

Legal questions are unusually hostile to pure vector search. The thing that makes a question specific is almost always a rare identifier: a section number, a case citation, a statute short-name. That is precisely the token an embedding model averages away. Ask about “Section 302” and a dense retriever will happily hand back everything about criminal procedure in general.

So this system runs two retrievers and fuses them. BM25 catches the literal identifier; dense vectors catch the paraphrase. The interesting decision is how they are combined: by rank, using reciprocal rank fusion, rather than by score. A cosine distance and a BM25 score share no scale and no distribution, so any normalisation you pick between them is really a hyperparameter fitted to one corpus, and silently wrong on the next one.

Around that sits the part nobody demos: two incompatible document parsers running as isolated subprocesses, a Celery queue that has to respect an 8 GB VRAM ceiling, and a tenancy model that fails closed on both the write and the read path.

Architecture

React SPA (GitHub Pages)
      │  HTTPS
      ▼
Cloudflare Tunnel
      │
      ▼
   FastAPI ──────────► SQLite (WAL)      document + job state
      │
      ├──────────────► ChromaDB :8001    3 collections
      │
      └──► Redis :6379 ──► Celery (solo pool)
                                │
                                ├─► subprocess: marker-pdf    (.venv-marker)
                                └─► subprocess: Unlimited-OCR (.venv-ocr)
The two parsers sit behind a subprocess boundary because their dependency pins cannot coexist in one interpreter.

What it does

  • Hybrid retrieval: BM25 lexical search and dense vector search fused with reciprocal rank fusion.
  • Two selectable ingestion engines. marker-pdf for layout-faithful extraction, Unlimited-OCR for scanned and handwritten documents.
  • Per-session tenancy: a UUID4 acts as a capability token, enforced independently on the embedder and the query router.
  • Practice-area tagging combining metadata filtering, gentle centroid steering, and a widened top-k to compensate.
  • Background ingestion through Celery with live job state persisted to SQLite in WAL mode.
  • Static React SPA on GitHub Pages talking to a self-hosted backend through a Cloudflare tunnel, with no cloud bill.

How it works

  • Reciprocal rank fusion over ranks rather than scores, so the two retrievers never need a shared scale.
  • Three-virtualenv architecture: marker-pdf and Unlimited-OCR pin irreconcilable transformers and pillow versions, so both run as subprocesses and the application environment stays entirely torch-free.
  • Contextual chunk embedding: each chunk is vectorised as previous-tail plus chunk plus next-head (±220 characters) while the stored and cited text remains the chunk itself. The vector sees the neighbourhood; the citation stays exact.
  • VRAM budgeting for a 3.3B-parameter bf16 OCR model plus SAM-ViT-B and CLIP-L: max length 32768 down to 8192, image size 1024 down to 640, page batch 4, Celery on a solo pool.
  • Fail-closed tenancy: the embedder refuses unscoped writes and the router raises on an unfiltered query, so a missing session id is an error rather than a data leak.
  • 210 unit tests running with no external dependencies, plus a live smoke test asserting twenty properties against a real Chroma server.

Problems and answers

  • Dense retrieval loses the rare identifier that makes a legal question specific.

    Run BM25 alongside it and fuse by rank, so the lexical leg can rescue a query the embedder blurred.

  • The two best document parsers pin mutually incompatible versions of transformers and pillow.

    Give each its own virtualenv and invoke it as a subprocess; the app environment never imports torch at all.

  • Upstream OCR defaults OOM immediately on an 8 GB consumer card.

    Cut max sequence length and image resolution, batch pages in fours, and pin Celery to a solo pool, documented as a real coherence tradeoff rather than hidden.

  • Multi-tenant document isolation is the kind of thing that fails silently and catastrophically.

    Enforce the session scope on both sides of the store, and make the absence of a scope raise rather than default.

OlderKurt, an Offline-First AI Tutor