AI/ML
Image-to-LaTeX Handwriting OCR
A 2D-attention encoder-decoder that reads handwritten text and mathematical expressions and emits LaTeX autoregressively: ResNet-18 features into a four-layer Transformer decoder over a 1,007-token vocabulary.
- Built
- Jan 2026 to Jun 2026
- Role
- Project lead
- Token vocabulary
- 1,007Token vocabulary
- Decoder layers by heads
- 4 x 8Decoder layers by heads
- VRAM at batch 4
- ~6 GBVRAM at batch 4
- Measured training throughput
- 2.8 it/sMeasured training throughput
Overview
Reading handwritten mathematics is not a text recognition problem with harder glyphs. The meaning of an expression lives in its two-dimensional layout: a numerator is a numerator because of where it sits relative to a bar, and a superscript is a superscript because of where it sits relative to a baseline.
Standard image-to-sequence models flatten the encoder's spatial feature map into a 1D sequence before the decoder sees it, which throws that layout away and asks attention to reconstruct it from scratch. This model instead carries separate sinusoidal positional encodings for the X and Y axes, so every feature vector arrives at the decoder still knowing where on the page it came from.
The rest of the work was fitting the thing on an 8 GB card: mixed precision, gradient accumulation, Flash Attention through PyTorch's fused kernel, and optional encoder checkpointing, documented with a real VRAM-versus-throughput table rather than a claim.
Architecture
Image 256 × 256 × 3 │ ├─► ResNet-18 encoder → 16 × 16 × 512 ├─► 1×1 conv projection → 16 × 16 × 256 ├─► 2D positional encoding → X + Y sinusoidal │ ├─► Transformer decoder → 4 layers, 8 heads, d = 256 ├─► vocabulary projection → 1,007 tokens │ └─► autoregressive LaTeX sequence
What it does
- ResNet-18 encoder producing a 16x16x512 spatial map, projected to 256 channels.
- Separate X and Y sinusoidal positional encodings, preserving 2D layout into the decoder.
- Four-layer, eight-head Transformer decoder emitting LaTeX autoregressively.
- Trained across CASIA, HME100K, CROHME and IAM plus Hugging Face formula sets and a synthetic generator.
- Evaluated on BLEU, character error rate and exact match. Three metrics, because none alone is honest.
How it works
- Mixed precision via torch.amp.autocast, with gradient accumulation of four steps at batch 4 for an effective batch of 16.
- Flash Attention through F.scaled_dot_product_attention, which is the fused path rather than a hand-rolled kernel.
- Optional gradient checkpointing on the encoder, trading recompute for memory when a run needs the headroom.
- AdamW at lr 3e-4 with gradient clipping at 1.0, converging in roughly 50 epochs.
- Documented VRAM and throughput table: approximately 6 GB at 2.8 it/s at batch 4 on an 8 GB RTX 5060 Ti.
Problems and answers
Flattening the feature map destroys the spatial layout of an equation.
Encode X and Y positions separately and add both, so the decoder retains the grid.
An encoder-decoder plus attention does not fit comfortably in 8 GB.
Mixed precision, gradient accumulation, fused attention, and checkpointing when needed.
Handwriting datasets are fragmented across incompatible formats.
Build a unified dataset assembly step over CASIA, HME100K, CROHME and IAM plus synthetic data.
A single metric flatters an OCR model.
Report BLEU, CER and exact match together: sequence quality, character accuracy, and whether it is actually right.