Byrne-Docling (131M)

Document-understanding variant of Byrne-VLM. ~131M from-scratch SpikeWhale / Byrne VLM. Document image → DocTags (the Granite-Docling / SmolDocling markup - <chart>, <formula>, <code>, OTSL <fcel> table cells, <loc_N> boxes, code-language tags like <_Python_>, …).

Research artifact. Clean DocTags structure and coarse content. Not a production document parser.

document image → 448px letterbox → Byrne-VE (frozen, 784 tokens)
              → Connector → Byrne LM + Family-LoRA (+ atomic DocTags tokens) → DocTags

Current architecture

part detail
Byrne-VE 39M ViT-style encoder, 448px / patch16 → 784 tokens, RMSNorm, 2D-axial RoPE, QK-Norm, SwiGLU, HRM. DINOv2-distilled → DINO self-distilled.
Preprocessing letterbox (aspect-preserving pad) - the whole page is visible (center-crop was cutting ~90% off tall documents).
Connector 2-layer MLP (512→640).
Byrne LM ~90M SpikeWhale LM, custom SpikeTokenizer, 4096 ctx.
Family-LoRA HRM + MoE-SwiGLU adapter on the LM decoder.
DocTags tokens 140 DocTags markup tokens added as ATOMIC tokens (vocab 16512→16652). Each tag (<fcel>, <chart>, <_C++_>, <loc_0>…) is one token - emit it right or not, instead of spelling it char-by-char. Only these 140 new embedding rows were trained on the frozen base LM.

How we got here (the honest development story)

Several iterations. Each one fixed a failure the last one showed.

v1 - 224px + AnyRes tiling (char-level DocTags). First Byrne-Docling used a 224px encoder with AnyRes 2×2 tiling (5 tiles → 980 image tokens) to raise effective resolution. Trained on 6,000 ground-truth (image → DocTags) pairs from SmolDocling synthetic sets (SynthChartNet, SynthFormulaNet, SynthCodeNet) via connector + Family-LoRA. It learned the DocTags format and coarse content (chart column headers, C++ copyright, LaTeX), but greedy decoding looped, so inference uses a repetition penalty. That established the pipeline.

Higher-resolution encoder (v2, 448px). Distilled a native 448px / 784-token encoder (DINOv2 distill → DINO self-distill) to read finer detail. Naively swapping it in with center-crop made things worse - malformed tags and cross-modal confusion (code inside <formula>). Diagnosis: center-crop was squaring tall document pages, throwing away ~90% of the page.

Letterbox preprocessing. Aspect-preserving letterbox (whole page visible) fixed the regression. Content and modality came back (formula→LaTeX, code→right language). One leftover: tag syntax was noisy (<fcel> as <fcil> / <fbar>).

Atomic DocTags tokens (this release). Tag noise was tokenization, not training or resolution. DocTags weren't in the vocab, so the model spelled them character-by-character and one wrong subword wrecked the tag. Adding 140 DocTags as atomic tokens (and training only those embedding rows) cleaned the syntax - proper <chart><loc_0><loc_0><loc_500><loc_500> <bar_chart><fcel>…<nl></chart>, correct </formula> closings, clean <_Java_> language tags. Same trick Granite-Docling uses.

Evaluation (honest)

Held-out chart / formula / code: clean DocTags structure (proper <fcel>/<chart>/<loc_N>/<nl> + closings), correct modality, coarse-correct content (real financial labels, LaTeX, license headers). Remaining errors - repeated values, occasional wrong specifics - are capacity (39M encoder + 90M LM on 6k synthetic pairs), not tokenization. Use a repetition penalty at inference.

Usage

import torch
from generate import load_vlm, caption
from spike_tokenizer import SpikeTokenizer

dev = "cuda" if torch.cuda.is_available() else "cpu"
tok = SpikeTokenizer(vocab_file="tokenizer_doctags.json")
# load_vlm reads new_vocab from the checkpoint and resizes the LM automatically
vlm = load_vlm("weights/byrne_docling.pt", "lm", "weights/byrne_ve.pt", dev)
doctags = caption(vlm, tok, "page.png", dev, max_new=256,
                  repetition_penalty=1.2, no_repeat_ngram=3, letterbox=True)
print(doctags)

CLI:

python generate.py --image page.png --ckpt weights/byrne_docling.pt \
  --vision-ckpt weights/byrne_ve.pt --tokenizer tokenizer_doctags.json \
  --letterbox --max-new 256 --repetition-penalty 1.2 --no-repeat-ngram 3

Files

weights/byrne_docling.pt (connector + Family-LoRA + trained DocTags embeddings) · weights/byrne_ve.pt (448px letterbox encoder) · lm/ (Byrne base LM) · tokenizer_doctags.json (vocab 16652) · model code.

Citation

@misc{byrne_docling_2026,
  title  = {Byrne-Docling: A Tiny SpikeWhale VLM for Document DocTags (letterbox + atomic tokens)},
  author = {Quazim0t0},
  year   = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Quazim0t0/Byrne-Docling-131M}}
}

License

Apache-2.0.

Escarda vs Byrne - vision family comparison

Byrne = HRM refine. Escarda = Byrne + JEPA on the vision encoder and the LM trunk. Auxiliary only. Zero inference cost.

Vision encoder (DINOv2 teacher-alignment, n=1024 held-out):

Byrne-VE Escarda-VE
Params 39.34M 39.60M (+JEPA head)
CLS cosine 0.776 0.771
PATCH cosine 0.600 0.584
JEPA self-consistency - 0.040

Docling (same held-out doc images, atomic DocTags): both emit well-formed DocTags. Byrne-Docling is a bit more complete on the hardest samples (closes </formula>, includes the <code> wrapper), which matches the slightly higher teacher-alignment. Escarda-Docling is structurally on par and has the JEPA representation-learning trait.

Pros/cons. Byrne (HRM): higher teacher-alignment, all capacity on distillation fidelity; no self-supervised objective. Escarda (HRM+JEPA): self-supervised neighbour-prediction (richer spatial structure) at zero inference cost, trading ~1-3% teacher-alignment. Same size class.

Family repos: Byrne-VE · Escarda-VE · Byrne-Docling-131M · Escarda-Docling-126M

Engram repair (update)

The n-gram Engram in this checkpoint was degenerate: every token hashed to bucket 0, so only 1 of 4096 rows ever trained and the memory was a single constant (more training could not fix it). This version applies a behavior-preserving repair (rescale the frozen hash compressor + broadcast bucket 0) - outputs unchanged - then a short distill with Engram tables unfrozen so the now-reachable buckets populate.

Verified: hash bucket usage 1/4096 → 4096/4096. Engram tables are a real populated n-gram memory, not a constant. DocTags output stays valid.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Quazim0t0/Byrne-Docling-131M 1

Collection including Quazim0t0/Byrne-Docling-131M