Skip to content

AI/ML

Image-to-LaTeX Handwriting OCR

A 2D-attention encoder-decoder that reads handwritten text and mathematical expressions and emits LaTeX autoregressively: ResNet-18 features into a four-layer Transformer decoder over a 1,007-token vocabulary.

Built
Jan 2026 to Jun 2026
Role
Project lead
Token vocabulary
1,007Token vocabulary
Decoder layers by heads
4 x 8Decoder layers by heads
VRAM at batch 4
~6 GBVRAM at batch 4
Measured training throughput
2.8 it/sMeasured training throughput

Overview

Reading handwritten mathematics is not a text recognition problem with harder glyphs. The meaning of an expression lives in its two-dimensional layout: a numerator is a numerator because of where it sits relative to a bar, and a superscript is a superscript because of where it sits relative to a baseline.

Standard image-to-sequence models flatten the encoder's spatial feature map into a 1D sequence before the decoder sees it, which throws that layout away and asks attention to reconstruct it from scratch. This model instead carries separate sinusoidal positional encodings for the X and Y axes, so every feature vector arrives at the decoder still knowing where on the page it came from.

The rest of the work was fitting the thing on an 8 GB card: mixed precision, gradient accumulation, Flash Attention through PyTorch's fused kernel, and optional encoder checkpointing, documented with a real VRAM-versus-throughput table rather than a claim.

Architecture

Image 256 × 256 × 3
   │
   ├─► ResNet-18 encoder        →  16 × 16 × 512
   ├─► 1×1 conv projection      →  16 × 16 × 256
   ├─► 2D positional encoding   →  X + Y sinusoidal
   │
   ├─► Transformer decoder      →  4 layers, 8 heads, d = 256
   ├─► vocabulary projection    →  1,007 tokens
   │
   └─► autoregressive LaTeX sequence
The 2D positional encoding is the load-bearing step. Remove it and the model can read symbols but not structure.

What it does

  • ResNet-18 encoder producing a 16x16x512 spatial map, projected to 256 channels.
  • Separate X and Y sinusoidal positional encodings, preserving 2D layout into the decoder.
  • Four-layer, eight-head Transformer decoder emitting LaTeX autoregressively.
  • Trained across CASIA, HME100K, CROHME and IAM plus Hugging Face formula sets and a synthetic generator.
  • Evaluated on BLEU, character error rate and exact match. Three metrics, because none alone is honest.

How it works

  • Mixed precision via torch.amp.autocast, with gradient accumulation of four steps at batch 4 for an effective batch of 16.
  • Flash Attention through F.scaled_dot_product_attention, which is the fused path rather than a hand-rolled kernel.
  • Optional gradient checkpointing on the encoder, trading recompute for memory when a run needs the headroom.
  • AdamW at lr 3e-4 with gradient clipping at 1.0, converging in roughly 50 epochs.
  • Documented VRAM and throughput table: approximately 6 GB at 2.8 it/s at batch 4 on an 8 GB RTX 5060 Ti.

Problems and answers

  • Flattening the feature map destroys the spatial layout of an equation.

    Encode X and Y positions separately and add both, so the decoder retains the grid.

  • An encoder-decoder plus attention does not fit comfortably in 8 GB.

    Mixed precision, gradient accumulation, fused attention, and checkpointing when needed.

  • Handwriting datasets are fragmented across incompatible formats.

    Build a unified dataset assembly step over CASIA, HME100K, CROHME and IAM plus synthetic data.

  • A single metric flatters an OCR model.

    Report BLEU, CER and exact match together: sequence quality, character accuracy, and whether it is actually right.

NewerF1 Telemetry DashboardOlderViT Compression Trade-offs