Yo-ByT5: Byte-Level Diacritic Restoration for Yorùbá

Yo-ByT5 is a fine-tune of google/byt5-small that restores Yorùbá diacritics (tone marks: acute ◌́, grave ◌̀, mid unmarked; underdots: ẹ, ọ, ṣ) to undiacritized text.


Quickstart

import unicodedata
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

model_id = "lazymonster/yobyt5-restoration"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

text = "Eko ni kokoro aseyori."
inputs = tokenizer(text, return_tensors="pt", max_length=1024, truncation=True)
outputs = model.generate(inputs["input_ids"], max_length=1024, num_beams=1)

predicted = tokenizer.decode(outputs[0], skip_special_tokens=True)
predicted = unicodedata.normalize("NFC", predicted)
print(predicted)
# Ẹ̀kọ́ ni kọ́kọ́rọ́ àṣeyọrí.

Benchmark Results

Evaluated on the YAD Test Benchmark (3,330 sentences) using base-word Needleman-Wunsch sequence alignment and Unicode NFC normalization.

The Corrected Reference repairs 1,415 non-standard underdot codepoints (U+0329 to standard U+0323) in the raw YAD release.

Greedy decoding uses num_beams=1. Beam search uses num_beams=5.

Metric Official, Greedy Official, Beam Corrected, Greedy Corrected, Beam
DER (Total) ↓ 10.45% 10.58% 10.00% 10.14%
DER-tone ↓ 8.76% 8.93% 8.76% 8.93%
DER-underdot ↓ 5.66% 6.12% 5.69% 6.15%
WDER ↓ 16.81% 16.88% 15.58% 15.64%
CER ↓ 3.82% 3.75% 3.55% 3.48%
WER ↓ 16.11% 16.03% 14.93% 14.86%
BLEU ↑ 0.6837 0.6841 0.6837 0.6841
ChrF ↑ 0.8431 0.8431 0.8431 0.8431

Training Data

Dataset Sentences License
MENYO-20k (Yorùbá train split) 9,942 CC BY-NC 4.0
Biblica® Open Yorùbá Contemporary Bible 2017 36,371 CC BY-SA
Total Train 46,313
Validation Set 5,305

Training Procedure

  • Architecture: google/byt5-small (299.6M parameters)
  • Compute: Google Cloud TPU v6e-8
  • Optimizer: AdamW, linear decay, weight decay 0.01, gradient clip norm 0.5
  • Batch Size: Per-device batch 4, gradient accumulation 2 (effective global batch size 64)
  • Phase 1: Learning rate 2e-4, 300 warmup steps, 4 epochs
  • Phase 2: Learning rate 1e-4, 0 warmup steps, 4 epochs (released checkpoint)

Recommendations

  1. Unicode normalization: Always normalize input and generated text to NFC before scoring or processing.
  2. Chunking: Chunk inputs longer than 1024 bytes at sentence or clause boundaries.
  3. Decoding: Greedy decoding is fast and accurate. Beam search (num_beams=5) offers slight quality gains on edit distance metrics.

Citation & License

License

  • Model weights: Apache 2.0
  • Training data: Subject to original licenses (MENYO-20k: CC BY-NC 4.0; Biblica: CC BY-SA)

Citation

@misc{gali2026yobyt5efficienthighfidelitydiacritic,
      title={Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yor\`ub\'a}, 
      author={Ahmad Samuel Gali and Shamsuddeen Hassan Muhammad},
      year={2026},
      eprint={2610.01634},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2610.01634}, 
}

Acknowledgments

Trained as a member of the HausaNLP Research Group with compute support from the Google TPU Research Cloud (TRC) program.

Downloads last month
730
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for lazymonster/yobyt5-restoration

Finetuned
(335)
this model

Dataset used to train lazymonster/yobyt5-restoration

Space using lazymonster/yobyt5-restoration 1

Paper for lazymonster/yobyt5-restoration