Sieve-9B-Plus

Sieve-9B-Plus is a decision model from the Sieve family. You give it a piece of text or JSON (the state) and typed questions, and it returns a calibrated probability for every option of every question. It reads the state once, scores the options directly, and never generates text.

It is a LoRA adapter (rank 32) and a pointer head on Qwen/Qwen3.5-9B (revision c2022362), and it stores one calibration temperature per question category. The code is at github.com/sthanika-ai/Sieve.

How it works

  • Backbone. Only the text part of Qwen3.5-9B is used; the model cannot produce text.
  • Readout. A pointer head scores the end of each option against the decision point. A calibrated softmax gives the probabilities.
  • Isolation. Each question runs as its own row from the state's cache, so questions never see each other.
type options returns
choice the caller's keys, optionally described (up to 255) a probability per key
noul yes / no P(yes)
score ordered levels a probability per level

Use

pip install "sieve-decisions[cuda,serve] @ git+https://github.com/sthanika-ai/Sieve"
pip install --no-build-isolation causal-conv1d==1.7.0   # recommended: the fused kernel our runs used
from sieve import load_sieve, decide

m = load_sieve("sthanika-ai/Sieve-9B-Plus")
decide(m, {"subject": "Duplicate charge on invoice 4411", "body": "Billed twice. Refund today or we cancel."},
       {"department": {"type": "choice", "instructions": "Which team handles this?",
                       "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs and outages"}},
        "churn_risk": {"type": "noul", "instructions": "Does the customer threaten to cancel?"}})
  • Loading. Load it with sieve 0.2.0 or later; the adapter alone is not a text-generation model. The backbone is downloaded at its pinned revision on first use.
  • GPU memory. About 20 GB in bf16. Longer states need more.
  • Token limits. States and questions up to 16,384 tokens each, as in training. sieve 0.2.1 or later uses these limits by default; 0.2.0 stops at 8,192 tokens. Longer requests are refused, never truncated. The Decision Index engine declares larger limits (a 65,536-token state and a 32,768-token question). In the Decision Index run below, 241 requests had a question longer than 16,384 tokens (at most 29,944, from ToolRet and BRIGHT) and were answered in full; no state was longer than 16,384.
  • Server. sieve-serve --model sthanika-ai/Sieve-9B-Plus --graphs.
  • causal-conv1d. Without it, transformers falls back to a slower PyTorch convolution, whose probabilities can differ slightly. Install it to reproduce our numbers.
  • Temperature per category. head.pt stores a global T = 1.109 and 23 temperatures (1.000 to 1.338), fitted by NLL on the calibration split (43,176 questions). A question's category is its type, its number of options (2, 3-4, 5-9, 10-25, 26+) and the kind of state (empty, text or JSON). A question takes the temperature of its most specific category in the table (choice:3-4:text, then choice:3-4, then choice), else the global T. Temperatures never change an answer, and temperature=1.0 gives raw probabilities. sieve versions before 0.2.0 ignore the table and use the global T.

Results

Decision Index 0.3

Sieve-9B-Plus scores 53.35 (raw index 64.39, breadth 51.95) on the public 0.3 suite of the Decision Index, on our own run with the official kit at 62d2f51 (0.3) and its scorer.

  • Coverage. All 140,178 scoreable requests over 43 benchmarks were answered, with no truncation, no unsupported request and no error. The index averages 37 of the benchmarks.
  • How the run was made. One complete run of the 0.3 suite on two H100 80GB GPUs, in four shards.
  • Results. sthanika-ai/Sieve-9B-Plus-decision-index-results.
  • Not on the leaderboard yet. The score is self-reported. On the 0.3 board the ranking uses the Full score, where this public suite counts 20% and the maintainers' private tests count 80%.
Decision Index 0.3 Knowledge & Reasoning Language Understanding Retrieval & Classification Tools & Automation Arts & Human Taste
53.35 38.3 55.9 64.2 69.8 34.0

The index and area scores are chance-corrected: 0 is random guessing and 100 is perfect. Below, the score is in each benchmark's own metric and the skill is chance-corrected.

All 43 benchmarks
benchmark metric score skill (×100) requests
ACOS per-review F1 0.175 14.8 1,565
Amazon ESCI macro-F1 0.499 37.2 5,000
ANLI macro-F1 0.651 47.7 3,200
API-Bank accuracy 0.858 85.6 508
BANKING77 macro-F1 0.860 85.8 3,080
BBH fixed-option tasks accuracy 0.703 57.0 5,507
BFCL case exact accuracy 0.958 94.3 1,694
BPoMP accuracy 0.882 77.0 5,000
BRIGHT nDCG@10 0.451 37.9 220
cfcolor accuracy 0.623 26.4 5,000
ChessBench accuracy 0.128 5.1 5,000
CLadder accuracy 0.669 33.9 5,000
CLINC150+OOS macro-F1 0.912 91.1 5,500
ContractNLI macro-F1 0.734 61.6 123
CRUXEval accuracy 0.589 34.9 570
FinEntity macro-F1 0.889 83.8 979
GPQA Diamond accuracy 0.464 28.6 196
GSM8K accuracy 0.687 62.4 2,638
Habermas Machine accuracy 0.403 13.4 1,676
HellaSwag accuracy 0.967 95.6 10,042
HLE accuracy 0.100 0.0 501
Home appliance simulator case exact accuracy 0.250 25.0 88
HoVer claim verification accuracy 0.802 60.5 4,000
Humicroedit accuracy 0.614 22.8 2,628
iSarcasmEval Sarcasm F1 · track A, English 0.549 42.1 4,600
MMLU-Pro accuracy 0.594 54.3 12,032
MuSR accuracy 0.560 30.0 752
New Yorker caption matching accuracy 0.659 57.4 528
NLI4CT macro-F1 0.792 59.6 5,500
PhishNChips phishing decisions accuracy 0.841 68.1 2,000
POP909-CL accuracy 0.084 7.1 2,000
RAGTruth response-level hallucination F1 on hallucinated class 0.802 59.0 2,700
SATA-Bench case exact accuracy 0.279 26.9 1,650
ToolRet nDCG@10 0.673 62.3 685
VAST macro-F1 0.554 33.2 3,006
When2Call MCQ accuracy 0.801 73.5 3,652
WinoGrande accuracy 0.908 81.7 1,267
ARC-Challenge (not in the index) accuracy 0.954 93.9 1,172
ARC-Easy (not in the index) accuracy 0.981 97.5 2,376
MMLU (not in the index) accuracy 0.787 71.6 14,033
RouterBench (not in the index) selected quality (quality objective) 0.795 51.9 10,000
SGD/SGD-X (not in the index) macro-F1 0.438 6.5 2,500
SimpleBench (not in the index) accuracy 0.100 0.0 10

Calibration split

The calibration split chose the checkpoint and fitted the temperatures.

split questions accuracy NLL Brier ECE
calibration, per-category temperatures (shipped) 43,176 0.8775 0.3071 0.1676 0.0068
calibration, one global T = 1.109 43,176 0.8775 0.3077 0.1678 0.0074
calibration, raw (T = 1) 43,176 0.8775 0.3091 0.1677 0.0093

In a 2-fold check inside the calibration split (fit on one half, score the other), the per-category table and the single T scored the same: NLL 0.3079 vs 0.3078, ECE 0.0087 vs 0.0084.

Latency

  • Per Decision Index request. Measured on one H100 80GB GPU, single process, eager and in-process, over a random 3,000-request sample of the public 0.3 suite: median 43 ms, mean 138 ms, 80th percentile 58 ms, 95th percentile 144 ms. The answers on that sample matched the full run exactly.

Training

  • Recipe.
    • LoRA r=32, α=64, dropout 0.05 on the attention, MLP and linear-attention projections (12 modules), with the pointer head trained from scratch.
    • AdamW, lr 5e-5 for the adapter and the head, OneCycle with 10% warm-up, 1 epoch, 32 records per step, bf16.
    • States and questions up to 16,384 tokens each; longer records are dropped, never truncated (none were).
    • 5,183 steps, 5.6 hours on two H100 80GB GPUs.
    • Permuted choice options and "none of the above" augmentation, with a cross-entropy loss.
  • Selection. The checkpoint with the highest accuracy on a fixed 2,000-question calibration sample, with NLL breaking ties. This rule was fixed before training, and it chose step 4,500. The benchmarks were never used to choose.
  • Temperatures. Fitted after selection on the whole calibration split: the global T, then one T per category with at least 200 calibration questions.
  • Data. 165,832 training records and 30,722 calibration records.
    • Sources: public datasets and generated decision records.
    • Not published. The training files are not part of this release; the SHA-256 of each file is in training_config.json.

Files

file contents
adapter_model.safetensors, adapter_config.json the LoRA adapter (rank 32, PEFT format)
head.pt the pointer head and the calibrated temperatures (global and per category)
tokenizer.json, tokenizer_config.json Qwen3.5-9B's tokenizer, unchanged
training_config.json, training_metrics.json, train.log the recipe, the run and its log
result.json the calibration, Decision Index and latency numbers above, with the calibration metrics per category
provenance.json SHA-256 of every file and of the training data, and library versions

License

Apache-2.0, see LICENSE.

Downloads last month
40
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sthanika-ai/Sieve-9B-Plus

Finetuned
Qwen/Qwen3.5-9B
Adapter
(785)
this model

Collection including sthanika-ai/Sieve-9B-Plus

Evaluation results

  • Decision Index 0.3, public index (chance-corrected; self-reported, not yet on the leaderboard) on Decision Index 0.3 public suite (140,178 scored requests; the index averages 37 benchmarks)
    self-reported
    53.350