Instructions to use MinimaLabs/KLM-1.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MinimaLabs/KLM-1.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="MinimaLabs/KLM-1.1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("MinimaLabs/KLM-1.1") model = AutoModelForCausalLM.from_pretrained("MinimaLabs/KLM-1.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MinimaLabs/KLM-1.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MinimaLabs/KLM-1.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MinimaLabs/KLM-1.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/MinimaLabs/KLM-1.1
- SGLang
How to use MinimaLabs/KLM-1.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MinimaLabs/KLM-1.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MinimaLabs/KLM-1.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MinimaLabs/KLM-1.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MinimaLabs/KLM-1.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use MinimaLabs/KLM-1.1 with Docker Model Runner:
docker model run hf.co/MinimaLabs/KLM-1.1
KLM 1.1
KLM 1.1 is an 88M-parameter base language model, the continued training of KeyLM-75M. It scores 17.40 on the Open SLM Intelligence Index, up from KeyLM-75M's 10.44. That moves it from 38th to 9th among models under 100M parameters on the Open SLM leaderboard.
KeyLM is now KLM. This release keeps the KeyLM architecture and continues from its weights.
Why KLM 1.1 exists
KeyLM-75M was the first model we trained, and it was flawed in many ways. Most of its data was casual chat (Reddit, WildChat, LMSYS and similar), its tokenizer was small, and it stopped at about 18B tokens while still at a high learning rate. It is currently 38th on the leaderboard among models under 100M parameters.
The goal of KLM 1.1 was to find out how far that same base model could be pushed by more training alone, without changing its architecture. Over about 120B more tokens, we gave it a larger tokenizer, retrained it on a cleaner, education-, math- and reasoning-heavy mix, and ended with a proper learning-rate decay.
The gain shows up on every benchmark, most on ARC-Easy (+8.6 points) and ArithMark-3 (+5.7).
| KeyLM-75M | KLM 1.1 | |
|---|---|---|
| Int Index | 10.44 | 17.40 |
| Rank, under 100M | #38 | #9 |
| HellaSwag | 29.66% | 33.65% |
| ARC-Easy | 35.73% | 44.36% |
| ARC-Challenge | 23.98% | 25.17% |
| PIQA | 60.50% | 64.80% |
| ArithMark-3 | 30.10% | 35.80% |
| Parameters | 75.3M | 88.1M |
| Vocabulary | 12,020 | 24,576 |
| Training tokens | ~18.7B | ~138.6B in total |
How it was trained
KLM 1.1 went through three stages after KeyLM-75M, each continuing from the last one's weights:
| Stage | Tokens | What changed | Int Index after |
|---|---|---|---|
| KeyLM-75M (start) | 18.7B | 10.47 | |
| Tokenizer expansion | 20.0B | Vocabulary grown from 12,020 to 24,576 tokens; trained on web, math and code | 15.15 |
| Short continuation | 5.0B | Same mix, learning rate re-warmed | 15.52 |
| Long continuation | 94.9B | New mix, warmup-stable-decay schedule | 17.40 |
Int Index values in this table are from our own evaluation runs. Our harness scores KeyLM-75M at 10.47; the leaderboard lists 10.44.
Held-out loss fell at every evaluation of the long continuation, from 2.131 to 1.814 nats per token, averaged over held-out text from each of its 11 data sources. Most of the benchmark gain arrived in the final learning-rate decay.
Tokenizer expansion
The 24,576-token vocabulary keeps all of KeyLM-75M's tokens and adds 12,299 new byte-level BPE merges learned on the training mix. Each new token starts from the average of the old tokens it replaces. The larger vocabulary encodes the same text in 1.31–1.41× fewer tokens, with the biggest savings on code and math. This stage trained on UltraX Ultra-FineWeb web text (80%), UltraData-Math (15%) and Code-Reasoning (5%).
Long continuation
In the long continuation, the learning rate warmed up to 3e-5, held there for the first 80% of the run (75.9B tokens), then decayed linearly to 10% of peak over the final 20% (19.0B tokens). The decay used a mix heavier in math and textbooks.
| Source | Stable phase | Decay phase |
|---|---|---|
| DCLM-baseline web text | 35% | 20% |
| FineWeb-Edu (dedup) | 30% | 25% |
| Cosmopedia v2 synthetic textbooks | 15% | 20% |
| UltraData-Math L2 | 12.5% | 11% |
| UltraData-Math L3 textbook exercises and QA | 2.5% | 6% |
| OpenMathInstruct-2 problems with solutions | 8% | |
| Pretrain-Behaviors reasoning and format rewrites | 3% | 8% |
| Code-Reasoning | 2% | 2% |
All sources were filtered before training. The filters dropped documents the model had already trained on, exact repeats, empty documents, and any document sharing a 13-word sequence with an item from the five evaluation sets.
| Field | Value |
|---|---|
| Optimizer | AdamW (betas 0.9 / 0.95, weight decay 0.1), state carried over from the previous stage |
| Precision | bf16 autocast over fp32 weights |
| Gradient clipping | 1.0 |
Evaluation
All scores are zero-shot. ARC, HellaSwag and PIQA use lm-eval 0.4.12 with length-normalized accuracy. ArithMark-3 uses the official ArithMark-3.0 script. Our harness reproduces the leaderboard's published KeyLM numbers to within ±0.17 points per task.
Architecture
KLM 1.1 keeps the KeyLM-75M architecture unchanged. Only the embedding and output layers grew with the vocabulary, from 75.3M to 88.1M parameters.
| Field | Value |
|---|---|
| Parameters | 88,108,544 |
| Architecture | Decoder-only transformer (Llama/Qwen3-style) |
| Layers | 24 |
| Hidden size | 512 |
| Attention | 8 query heads, 2 KV heads (GQA), per-head QK-RMSNorm |
| Feed-forward | SwiGLU, hidden size 1,280 |
| Position encoding | RoPE (theta 10,000) |
| Context length | 2,048 tokens |
| Vocabulary | 24,576 (byte-level BPE) |
| Embeddings | Untied input and output |
| Weights | fp32 safetensors, as trained |
The weights load as a standard Transformers Qwen3ForCausalLM, since the network computes exactly what that class computes. No custom code is needed, and tools that support Qwen3 can run it.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "MinimaLabs/KLM-1.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16)
inputs = tokenizer("The three primary colors are", return_tensors="pt")
outputs = model.generate(
**inputs, max_new_tokens=40, do_sample=True,
temperature=0.7, top_p=0.9, repetition_penalty=1.1,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Limitations
- Base model only. KLM 1.1 completes text. It is not instruction-tuned and does not hold a conversation. The tokenizer has ChatML tokens (
<|im_start|>,<|im_end|>), but the model never saw them in training. - Small model. Its world knowledge is thin and its facts are often wrong. Do not rely on it for factual answers, math or code.
- English only, with a 2,048-token context window.
- No safety alignment. Add your own filtering before any user-facing use.
License
Apache-2.0. See LICENSE.
Citation
@misc{klm11_2026,
title = {KLM 1.1: continued training of KeyLM-75M},
author = {Eclipse-Senpai},
year = {2026},
howpublished = {\url{https://huggingface.co/MinimaLabs/KLM-1.1}}
}
- Downloads last month
- 172