Byrne-Docling (131M)
Document-understanding variant of Byrne-VLM.
~131M from-scratch SpikeWhale / Byrne VLM. Document image → DocTags (the
Granite-Docling / SmolDocling markup - <chart>, <formula>, <code>, OTSL
<fcel> table cells, <loc_N> boxes, code-language tags like <_Python_>, …).
Research artifact. Clean DocTags structure and coarse content. Not a production document parser.
document image → 448px letterbox → Byrne-VE (frozen, 784 tokens)
→ Connector → Byrne LM + Family-LoRA (+ atomic DocTags tokens) → DocTags
Current architecture
| part | detail |
|---|---|
| Byrne-VE | 39M ViT-style encoder, 448px / patch16 → 784 tokens, RMSNorm, 2D-axial RoPE, QK-Norm, SwiGLU, HRM. DINOv2-distilled → DINO self-distilled. |
| Preprocessing | letterbox (aspect-preserving pad) - the whole page is visible (center-crop was cutting ~90% off tall documents). |
| Connector | 2-layer MLP (512→640). |
| Byrne LM | ~90M SpikeWhale LM, custom SpikeTokenizer, 4096 ctx. |
| Family-LoRA | HRM + MoE-SwiGLU adapter on the LM decoder. |
| DocTags tokens | 140 DocTags markup tokens added as ATOMIC tokens (vocab 16512→16652). Each tag (<fcel>, <chart>, <_C++_>, <loc_0>…) is one token - emit it right or not, instead of spelling it char-by-char. Only these 140 new embedding rows were trained on the frozen base LM. |
How we got here (the honest development story)
Several iterations. Each one fixed a failure the last one showed.
v1 - 224px + AnyRes tiling (char-level DocTags). First Byrne-Docling used a
224px encoder with AnyRes 2×2 tiling (5 tiles → 980 image tokens) to raise
effective resolution. Trained on 6,000 ground-truth (image → DocTags) pairs
from SmolDocling synthetic sets (SynthChartNet, SynthFormulaNet,
SynthCodeNet) via connector + Family-LoRA. It learned the DocTags format
and coarse content (chart column headers, C++ copyright, LaTeX), but greedy
decoding looped, so inference uses a repetition penalty. That established
the pipeline.
Higher-resolution encoder (v2, 448px). Distilled a native 448px /
784-token encoder (DINOv2 distill → DINO self-distill) to read finer detail.
Naively swapping it in with center-crop made things worse - malformed tags
and cross-modal confusion (code inside <formula>). Diagnosis: center-crop
was squaring tall document pages, throwing away ~90% of the page.
Letterbox preprocessing. Aspect-preserving letterbox (whole page visible)
fixed the regression. Content and modality came back (formula→LaTeX,
code→right language). One leftover: tag syntax was noisy (<fcel> as
<fcil> / <fbar>).
Atomic DocTags tokens (this release). Tag noise was tokenization, not
training or resolution. DocTags weren't in the vocab, so the model spelled
them character-by-character and one wrong subword wrecked the tag. Adding 140
DocTags as atomic tokens (and training only those embedding rows)
cleaned the syntax - proper <chart><loc_0><loc_0><loc_500><loc_500> <bar_chart><fcel>…<nl></chart>, correct </formula> closings, clean
<_Java_> language tags. Same trick Granite-Docling uses.
Evaluation (honest)
Held-out chart / formula / code: clean DocTags structure (proper
<fcel>/<chart>/<loc_N>/<nl> + closings), correct modality,
coarse-correct content (real financial labels, LaTeX, license headers).
Remaining errors - repeated values, occasional wrong specifics - are
capacity (39M encoder + 90M LM on 6k synthetic pairs), not tokenization.
Use a repetition penalty at inference.
Usage
import torch
from generate import load_vlm, caption
from spike_tokenizer import SpikeTokenizer
dev = "cuda" if torch.cuda.is_available() else "cpu"
tok = SpikeTokenizer(vocab_file="tokenizer_doctags.json")
# load_vlm reads new_vocab from the checkpoint and resizes the LM automatically
vlm = load_vlm("weights/byrne_docling.pt", "lm", "weights/byrne_ve.pt", dev)
doctags = caption(vlm, tok, "page.png", dev, max_new=256,
repetition_penalty=1.2, no_repeat_ngram=3, letterbox=True)
print(doctags)
CLI:
python generate.py --image page.png --ckpt weights/byrne_docling.pt \
--vision-ckpt weights/byrne_ve.pt --tokenizer tokenizer_doctags.json \
--letterbox --max-new 256 --repetition-penalty 1.2 --no-repeat-ngram 3
Files
weights/byrne_docling.pt (connector + Family-LoRA + trained DocTags embeddings) ·
weights/byrne_ve.pt (448px letterbox encoder) · lm/ (Byrne base LM) ·
tokenizer_doctags.json (vocab 16652) · model code.
Citation
@misc{byrne_docling_2026,
title = {Byrne-Docling: A Tiny SpikeWhale VLM for Document DocTags (letterbox + atomic tokens)},
author = {Quazim0t0},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Quazim0t0/Byrne-Docling-131M}}
}
License
Apache-2.0.
Escarda vs Byrne - vision family comparison
Byrne = HRM refine. Escarda = Byrne + JEPA on the vision encoder and the LM trunk. Auxiliary only. Zero inference cost.
Vision encoder (DINOv2 teacher-alignment, n=1024 held-out):
| Byrne-VE | Escarda-VE | |
|---|---|---|
| Params | 39.34M | 39.60M (+JEPA head) |
| CLS cosine | 0.776 | 0.771 |
| PATCH cosine | 0.600 | 0.584 |
| JEPA self-consistency | - | 0.040 |
Docling (same held-out doc images, atomic DocTags): both emit well-formed
DocTags. Byrne-Docling is a bit more complete on the hardest samples (closes
</formula>, includes the <code> wrapper), which matches the slightly higher
teacher-alignment. Escarda-Docling is structurally on par and has the JEPA
representation-learning trait.
Pros/cons. Byrne (HRM): higher teacher-alignment, all capacity on distillation fidelity; no self-supervised objective. Escarda (HRM+JEPA): self-supervised neighbour-prediction (richer spatial structure) at zero inference cost, trading ~1-3% teacher-alignment. Same size class.
Family repos: Byrne-VE · Escarda-VE · Byrne-Docling-131M · Escarda-Docling-126M
Engram repair (update)
The n-gram Engram in this checkpoint was degenerate: every token hashed to bucket 0, so only 1 of 4096 rows ever trained and the memory was a single constant (more training could not fix it). This version applies a behavior-preserving repair (rescale the frozen hash compressor + broadcast bucket 0) - outputs unchanged - then a short distill with Engram tables unfrozen so the now-reachable buckets populate.
Verified: hash bucket usage 1/4096 → 4096/4096. Engram tables are a real populated n-gram memory, not a constant. DocTags output stays valid.