KLM 1.1

KLM 1.1

KLM 1.1 is an 88M-parameter base language model, the continued training of KeyLM-75M. It scores 17.40 on the Open SLM Intelligence Index, up from KeyLM-75M's 10.44. That moves it from 38th to 9th among models under 100M parameters on the Open SLM leaderboard.

KeyLM is now KLM. This release keeps the KeyLM architecture and continues from its weights.

Why KLM 1.1 exists

KeyLM-75M was the first model we trained, and it was flawed in many ways. Most of its data was casual chat (Reddit, WildChat, LMSYS and similar), its tokenizer was small, and it stopped at about 18B tokens while still at a high learning rate. It is currently 38th on the leaderboard among models under 100M parameters.

The goal of KLM 1.1 was to find out how far that same base model could be pushed by more training alone, without changing its architecture. Over about 120B more tokens, we gave it a larger tokenizer, retrained it on a cleaner, education-, math- and reasoning-heavy mix, and ended with a proper learning-rate decay.

KLM 1.1 on the Open SLM leaderboard, models under 100M parameters

The gain shows up on every benchmark, most on ARC-Easy (+8.6 points) and ArithMark-3 (+5.7).

KeyLM-75M and KLM 1.1 on each benchmark

KeyLM-75M KLM 1.1
Int Index 10.44 17.40
Rank, under 100M #38 #9
HellaSwag 29.66% 33.65%
ARC-Easy 35.73% 44.36%
ARC-Challenge 23.98% 25.17%
PIQA 60.50% 64.80%
ArithMark-3 30.10% 35.80%
Parameters 75.3M 88.1M
Vocabulary 12,020 24,576
Training tokens ~18.7B ~138.6B in total

How it was trained

KLM 1.1 went through three stages after KeyLM-75M, each continuing from the last one's weights:

Stage Tokens What changed Int Index after
KeyLM-75M (start) 18.7B 10.47
Tokenizer expansion 20.0B Vocabulary grown from 12,020 to 24,576 tokens; trained on web, math and code 15.15
Short continuation 5.0B Same mix, learning rate re-warmed 15.52
Long continuation 94.9B New mix, warmup-stable-decay schedule 17.40

Int Index values in this table are from our own evaluation runs. Our harness scores KeyLM-75M at 10.47; the leaderboard lists 10.44.

Int Index across training, and held-out loss during the long continuation

Held-out loss fell at every evaluation of the long continuation, from 2.131 to 1.814 nats per token, averaged over held-out text from each of its 11 data sources. Most of the benchmark gain arrived in the final learning-rate decay.

Tokenizer expansion

The 24,576-token vocabulary keeps all of KeyLM-75M's tokens and adds 12,299 new byte-level BPE merges learned on the training mix. Each new token starts from the average of the old tokens it replaces. The larger vocabulary encodes the same text in 1.31–1.41× fewer tokens, with the biggest savings on code and math. This stage trained on UltraX Ultra-FineWeb web text (80%), UltraData-Math (15%) and Code-Reasoning (5%).

Long continuation

In the long continuation, the learning rate warmed up to 3e-5, held there for the first 80% of the run (75.9B tokens), then decayed linearly to 10% of peak over the final 20% (19.0B tokens). The decay used a mix heavier in math and textbooks.

Source Stable phase Decay phase
DCLM-baseline web text 35% 20%
FineWeb-Edu (dedup) 30% 25%
Cosmopedia v2 synthetic textbooks 15% 20%
UltraData-Math L2 12.5% 11%
UltraData-Math L3 textbook exercises and QA 2.5% 6%
OpenMathInstruct-2 problems with solutions 8%
Pretrain-Behaviors reasoning and format rewrites 3% 8%
Code-Reasoning 2% 2%

All sources were filtered before training. The filters dropped documents the model had already trained on, exact repeats, empty documents, and any document sharing a 13-word sequence with an item from the five evaluation sets.

Field Value
Optimizer AdamW (betas 0.9 / 0.95, weight decay 0.1), state carried over from the previous stage
Precision bf16 autocast over fp32 weights
Gradient clipping 1.0

Evaluation

All scores are zero-shot. ARC, HellaSwag and PIQA use lm-eval 0.4.12 with length-normalized accuracy. ArithMark-3 uses the official ArithMark-3.0 script. Our harness reproduces the leaderboard's published KeyLM numbers to within ±0.17 points per task.

Architecture

KLM 1.1 keeps the KeyLM-75M architecture unchanged. Only the embedding and output layers grew with the vocabulary, from 75.3M to 88.1M parameters.

Field Value
Parameters 88,108,544
Architecture Decoder-only transformer (Llama/Qwen3-style)
Layers 24
Hidden size 512
Attention 8 query heads, 2 KV heads (GQA), per-head QK-RMSNorm
Feed-forward SwiGLU, hidden size 1,280
Position encoding RoPE (theta 10,000)
Context length 2,048 tokens
Vocabulary 24,576 (byte-level BPE)
Embeddings Untied input and output
Weights fp32 safetensors, as trained

The weights load as a standard Transformers Qwen3ForCausalLM, since the network computes exactly what that class computes. No custom code is needed, and tools that support Qwen3 can run it.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "MinimaLabs/KLM-1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16)

inputs = tokenizer("The three primary colors are", return_tensors="pt")
outputs = model.generate(
    **inputs, max_new_tokens=40, do_sample=True,
    temperature=0.7, top_p=0.9, repetition_penalty=1.1,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Limitations

  • Base model only. KLM 1.1 completes text. It is not instruction-tuned and does not hold a conversation. The tokenizer has ChatML tokens (<|im_start|>, <|im_end|>), but the model never saw them in training.
  • Small model. Its world knowledge is thin and its facts are often wrong. Do not rely on it for factual answers, math or code.
  • English only, with a 2,048-token context window.
  • No safety alignment. Add your own filtering before any user-facing use.

License

Apache-2.0. See LICENSE.

Citation

@misc{klm11_2026,
  title  = {KLM 1.1: continued training of KeyLM-75M},
  author = {Eclipse-Senpai},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/MinimaLabs/KLM-1.1}}
}
Downloads last month
172
Safetensors
Model size
88.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train MinimaLabs/KLM-1.1