AI/ML
Kurt, an Offline-First AI Tutor
A fully local desktop assistant that transcribes lectures, answers questions grounded in your own documents, and compiles the result into typeset PDFs. No network required.
- Built
- May 2026 to Jul 2026
- Role
- Project lead
- Faster than realtime transcription
- 10-20xFaster than realtime transcription
- Top-k chunks injected as context
- 5Top-k chunks injected as context
- Characters per overlapping chunk
- ~1000Characters per overlapping chunk
- Swappable LLM providers
- 5Swappable LLM providers
Overview
Kurt started as a lecture summariser and turned into a study environment. It records or ingests audio, transcribes it locally, files the result into a per-course knowledge base, and then answers questions against that base, all on the machine in front of you, with no API key and no upload.
The constraint that shaped everything was that it had to work offline. That rules out hosted transcription, hosted embeddings, hosted inference, and least obviously, hosted mathematics rendering. A tutor for a mathematics student that cannot draw an integral sign is not a tutor.
Tkinter has no MathJax and no HTML math support. The fix was to treat matplotlib's mathtext engine as a typesetter: intercept LaTeX blocks in the model's streamed output, render each to a transparent PNG, and embed it inline in the chat transcript. It is a slightly absurd solution that works completely.
Screens

What it does
- Local speech-to-text with voice activity detection, so silence never reaches the model.
- Per-course RAG knowledge bases with document ingestion, chunking and local vector storage.
- Inline mathematical rendering in a Tkinter chat window: genuinely offline LaTeX display.
- Summaries compiled to LaTeX and typeset to PDF without a multi-gigabyte TeX installation.
- Provider hot-swap between Ollama, OpenRouter, Anthropic, OpenAI and Gemini, with the settings UI reconfiguring per provider.
- Theme and UI-scaling controls, because a study tool gets used for hours at a time.
How it works
- faster-whisper, a CTranslate2 reimplementation of Whisper, CUDA-accelerated to 10-20x realtime with VAD skipping silent spans.
- PyMuPDF and LangChain extraction into roughly 1000-character overlapping semantic chunks, embedded with sentence-transformers/all-MiniLM-L6-v2.
- ChromaDB stored on disk and partitioned per course; retrieval pulls the top five chunks as contextual ground truth injected into the system prompt.
- LaTeX interception on $$ and \[ delimiters, rendered through matplotlib's mathtext engine to transparent PNGs, with tkhtmlview handling Markdown and HTML for the surrounding prose.
- Tectonic rather than TeXLive for PDF compilation. It fetches only the packages a document actually uses instead of installing several gigabytes up front.
Problems and answers
Tkinter cannot render mathematics.
Use matplotlib as a typesetting engine and embed the output as transparent inline images.
A full TeX distribution is a multi-gigabyte dependency.
Compile with Tectonic, which resolves packages on demand and keeps the install small.
Whole-lecture audio contains long silent stretches that waste GPU time.
Gate transcription behind voice activity detection.
Different providers expect different credentials, endpoints and parameters.
Drive the settings UI from a provider descriptor so the fields reconfigure themselves rather than being conditionally hidden.