kenning-xl-v0.6 (experimental)

Kenning-XL is a System One decision model built on a decoder. It answers the same wire format as Kenning (POST /v1/systemone; typed questions โ†’ calibrated probabilities, no text), but instead of a cross-encoder it uses a small decoder (Qwen3-1.7B) with a constrained answer-token readout: it renders (state, question) and reads the next-token distribution over the answer tokens in one pass, so the full language model reasons over the whole state. An optional deliberate mode generates a short reasoning trace, then reads the answer โ€” for numeric, policy and table decisions a single pass can't do.

This is the experimental v0.6 engine. The published cross-encoder (kenning-large-v0.5) is still the fast default; Kenning-XL is complementary โ€” strongest on agent and structured-record decisions.

How it's served

It is served by SystemOne Builder's kenning-xl engine (systemone_builder.kenning.xl_serve), which answers one-pass and escalates to deliberate mode only for low-confidence, short/structured questions โ€” so the extra latency is spent only where reasoning helps. See the design notes. It is not a plain transformers classifier; load it through the builder.

Training

  • Base: Qwen/Qwen3-1.7B (Apache-2.0), LoRA fine-tuned (merged), bf16, 1,536-token context.
  • Readout distillation on the v0.5 structured corpus (records, tables, agent steps, logs, text, safety; exact or public labels), plus gold-trace deliberate training: the builder's rule generators compute each label, so they emit the exact arithmetic/threshold/membership reasoning chain for free (kenning/traces.py), and the model learns to reason-then-answer on those.
  • Not trained on any TypeSafe output.

Benchmark (general suite, per-family mean, 1,328 held-out items)

As served (confidence + size-gated escalation), vs the cross-encoder v0.5 and โ€” for comparison only โ€” Clef (clef-flash) and Jev (jev-latest):

Family kenning-xl v0.6 kenning v0.5 Clef Jev
Macro 0.688 0.653 0.791 0.830
Agent (tool calls, task completion) 0.860 0.727 0.793 0.900
Conversation 1.000 0.979 1.000 1.000
Records (rules over JSON) 0.663 0.587 0.857 0.921
Tables 0.660 0.520 0.860 0.940
Text 0.752 0.756 0.841 0.834
Quality (helpful, correct) 0.439 0.479 0.447 0.498
Logs 0.440 0.520 0.740 0.720

Beats Clef on agent decisions, ties conversation, and the gold-trace deliberate reasoning more than doubles the hardest arithmetic task (daily-limit checks 0.42 โ†’ 0.81, transferring to held-out tasks). It does not match Clef overall, and currently trails even the cross-encoder on logs, quality and text โ€” deliberate reasoning isn't trace-trained for log-counting yet. Pick the engine for your task: Kenning-XL for agent/structured decisions, the v0.5 cross-encoder for text/safety/logs.

Licence

Weights under the Apache License 2.0 (merged LoRA adapter on Apache-2.0 Qwen3-1.7B). Training data is permissively licensed or generated; see NOTICE.md. Kenning implements a wire format compatible with TypeSafe AI's System One API; not affiliated with or endorsed by TypeSafe AI, and not trained on its outputs.

Downloads last month
319
Safetensors
Model size
2B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for systemonedev/kenning-xl-v0.6

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1275)
this model