AI/ML
LaTeX Converter
The earlier image-to-LaTeX pipeline, carrying the dataset assembly, training loop and logging that the 2D-attention model above grew out of.
- Built
- 2026
Overview
The first attempt at image-to-LaTeX, and the reason the second one is shaped the way it is. It carries the parts that survived: dataset assembly across several incompatible handwriting corpora, a training loop, log plotting, and its own architecture document.
Kept here rather than deleted because the interesting thing about it is what it got wrong. This is the version without 2D positional encoding, which is precisely the gap ocr-2.0 exists to close.
What it does
- Dataset download and master-dataset assembly across multiple handwriting corpora.
- Training and evaluation entry points with checkpointing.
- Log plotting for loss and metric curves across runs.
How it works
- Separate download and build steps, so corpus acquisition is decoupled from dataset construction.
- Its own architecture.md, written before the model was built rather than after.