Instructions to use sthanika-ai/Sieve-9B-Plus with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sthanika-ai/Sieve-9B-Plus with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-9B") model = PeftModel.from_pretrained(base_model, "sthanika-ai/Sieve-9B-Plus") - Notebooks
- Google Colab
- Kaggle
Sieve-9B-Plus
Sieve-9B-Plus is a decision model from the Sieve family. You give it a piece of text or JSON (the state) and typed questions, and it returns a calibrated probability for every option of every question. It reads the state once, scores the options directly, and never generates text.
It is a LoRA adapter (rank 32) and a pointer head on Qwen/Qwen3.5-9B
(revision c2022362), and it stores one calibration temperature per question category. The code is at
github.com/sthanika-ai/Sieve.
How it works
- Backbone. Only the text part of Qwen3.5-9B is used; the model cannot produce text.
- Readout. A pointer head scores the end of each option against the decision point. A calibrated softmax gives the probabilities.
- Isolation. Each question runs as its own row from the state's cache, so questions never see each other.
| type | options | returns |
|---|---|---|
choice |
the caller's keys, optionally described (up to 255) | a probability per key |
noul |
yes / no | P(yes) |
score |
ordered levels | a probability per level |
Use
pip install "sieve-decisions[cuda,serve] @ git+https://github.com/sthanika-ai/Sieve"
pip install --no-build-isolation causal-conv1d==1.7.0 # recommended: the fused kernel our runs used
from sieve import load_sieve, decide
m = load_sieve("sthanika-ai/Sieve-9B-Plus")
decide(m, {"subject": "Duplicate charge on invoice 4411", "body": "Billed twice. Refund today or we cancel."},
{"department": {"type": "choice", "instructions": "Which team handles this?",
"criteria": {"billing": "invoices, payments, refunds", "technical": "bugs and outages"}},
"churn_risk": {"type": "noul", "instructions": "Does the customer threaten to cancel?"}})
- Loading. Load it with
sieve0.2.0 or later; the adapter alone is not a text-generation model. The backbone is downloaded at its pinned revision on first use. - GPU memory. About 20 GB in bf16. Longer states need more.
- Token limits. States and questions up to 16,384 tokens each, as in training.
sieve0.2.1 or later uses these limits by default; 0.2.0 stops at 8,192 tokens. Longer requests are refused, never truncated. The Decision Index engine declares larger limits (a 65,536-token state and a 32,768-token question). In the Decision Index run below, 241 requests had a question longer than 16,384 tokens (at most 29,944, from ToolRet and BRIGHT) and were answered in full; no state was longer than 16,384. - Server.
sieve-serve --model sthanika-ai/Sieve-9B-Plus --graphs. - causal-conv1d. Without it, transformers falls back to a slower PyTorch convolution, whose probabilities can differ slightly. Install it to reproduce our numbers.
- Temperature per category.
head.ptstores a global T = 1.109 and 23 temperatures (1.000 to 1.338), fitted by NLL on the calibration split (43,176 questions). A question's category is its type, its number of options (2, 3-4, 5-9, 10-25, 26+) and the kind of state (empty, text or JSON). A question takes the temperature of its most specific category in the table (choice:3-4:text, thenchoice:3-4, thenchoice), else the global T. Temperatures never change an answer, andtemperature=1.0gives raw probabilities.sieveversions before 0.2.0 ignore the table and use the global T.
Results
Decision Index 0.3
Sieve-9B-Plus scores 53.35 (raw index 64.39, breadth 51.95) on the public 0.3 suite of the
Decision Index, on our own run with the official
kit at 62d2f51 (0.3) and its scorer.
- Coverage. All 140,178 scoreable requests over 43 benchmarks were answered, with no truncation, no unsupported request and no error. The index averages 37 of the benchmarks.
- How the run was made. One complete run of the 0.3 suite on two H100 80GB GPUs, in four shards.
- Results. sthanika-ai/Sieve-9B-Plus-decision-index-results.
- Not on the leaderboard yet. The score is self-reported. On the 0.3 board the ranking uses the Full score, where this public suite counts 20% and the maintainers' private tests count 80%.
| Decision Index 0.3 | Knowledge & Reasoning | Language Understanding | Retrieval & Classification | Tools & Automation | Arts & Human Taste |
|---|---|---|---|---|---|
| 53.35 | 38.3 | 55.9 | 64.2 | 69.8 | 34.0 |
The index and area scores are chance-corrected: 0 is random guessing and 100 is perfect. Below, the score is in each benchmark's own metric and the skill is chance-corrected.
All 43 benchmarks
| benchmark | metric | score | skill (×100) | requests |
|---|---|---|---|---|
| ACOS | per-review F1 | 0.175 | 14.8 | 1,565 |
| Amazon ESCI | macro-F1 | 0.499 | 37.2 | 5,000 |
| ANLI | macro-F1 | 0.651 | 47.7 | 3,200 |
| API-Bank | accuracy | 0.858 | 85.6 | 508 |
| BANKING77 | macro-F1 | 0.860 | 85.8 | 3,080 |
| BBH fixed-option tasks | accuracy | 0.703 | 57.0 | 5,507 |
| BFCL | case exact accuracy | 0.958 | 94.3 | 1,694 |
| BPoMP | accuracy | 0.882 | 77.0 | 5,000 |
| BRIGHT | nDCG@10 | 0.451 | 37.9 | 220 |
| cfcolor | accuracy | 0.623 | 26.4 | 5,000 |
| ChessBench | accuracy | 0.128 | 5.1 | 5,000 |
| CLadder | accuracy | 0.669 | 33.9 | 5,000 |
| CLINC150+OOS | macro-F1 | 0.912 | 91.1 | 5,500 |
| ContractNLI | macro-F1 | 0.734 | 61.6 | 123 |
| CRUXEval | accuracy | 0.589 | 34.9 | 570 |
| FinEntity | macro-F1 | 0.889 | 83.8 | 979 |
| GPQA Diamond | accuracy | 0.464 | 28.6 | 196 |
| GSM8K | accuracy | 0.687 | 62.4 | 2,638 |
| Habermas Machine | accuracy | 0.403 | 13.4 | 1,676 |
| HellaSwag | accuracy | 0.967 | 95.6 | 10,042 |
| HLE | accuracy | 0.100 | 0.0 | 501 |
| Home appliance simulator | case exact accuracy | 0.250 | 25.0 | 88 |
| HoVer claim verification | accuracy | 0.802 | 60.5 | 4,000 |
| Humicroedit | accuracy | 0.614 | 22.8 | 2,628 |
| iSarcasmEval | Sarcasm F1 · track A, English | 0.549 | 42.1 | 4,600 |
| MMLU-Pro | accuracy | 0.594 | 54.3 | 12,032 |
| MuSR | accuracy | 0.560 | 30.0 | 752 |
| New Yorker caption matching | accuracy | 0.659 | 57.4 | 528 |
| NLI4CT | macro-F1 | 0.792 | 59.6 | 5,500 |
| PhishNChips phishing decisions | accuracy | 0.841 | 68.1 | 2,000 |
| POP909-CL | accuracy | 0.084 | 7.1 | 2,000 |
| RAGTruth response-level hallucination | F1 on hallucinated class | 0.802 | 59.0 | 2,700 |
| SATA-Bench | case exact accuracy | 0.279 | 26.9 | 1,650 |
| ToolRet | nDCG@10 | 0.673 | 62.3 | 685 |
| VAST | macro-F1 | 0.554 | 33.2 | 3,006 |
| When2Call MCQ | accuracy | 0.801 | 73.5 | 3,652 |
| WinoGrande | accuracy | 0.908 | 81.7 | 1,267 |
| ARC-Challenge (not in the index) | accuracy | 0.954 | 93.9 | 1,172 |
| ARC-Easy (not in the index) | accuracy | 0.981 | 97.5 | 2,376 |
| MMLU (not in the index) | accuracy | 0.787 | 71.6 | 14,033 |
| RouterBench (not in the index) | selected quality (quality objective) | 0.795 | 51.9 | 10,000 |
| SGD/SGD-X (not in the index) | macro-F1 | 0.438 | 6.5 | 2,500 |
| SimpleBench (not in the index) | accuracy | 0.100 | 0.0 | 10 |
Calibration split
The calibration split chose the checkpoint and fitted the temperatures.
| split | questions | accuracy | NLL | Brier | ECE |
|---|---|---|---|---|---|
| calibration, per-category temperatures (shipped) | 43,176 | 0.8775 | 0.3071 | 0.1676 | 0.0068 |
| calibration, one global T = 1.109 | 43,176 | 0.8775 | 0.3077 | 0.1678 | 0.0074 |
| calibration, raw (T = 1) | 43,176 | 0.8775 | 0.3091 | 0.1677 | 0.0093 |
In a 2-fold check inside the calibration split (fit on one half, score the other), the per-category table and the single T scored the same: NLL 0.3079 vs 0.3078, ECE 0.0087 vs 0.0084.
Latency
- Per Decision Index request. Measured on one H100 80GB GPU, single process, eager and in-process, over a random 3,000-request sample of the public 0.3 suite: median 43 ms, mean 138 ms, 80th percentile 58 ms, 95th percentile 144 ms. The answers on that sample matched the full run exactly.
Training
- Recipe.
- LoRA r=32, α=64, dropout 0.05 on the attention, MLP and linear-attention projections (12 modules), with the pointer head trained from scratch.
- AdamW, lr 5e-5 for the adapter and the head, OneCycle with 10% warm-up, 1 epoch, 32 records per step, bf16.
- States and questions up to 16,384 tokens each; longer records are dropped, never truncated (none were).
- 5,183 steps, 5.6 hours on two H100 80GB GPUs.
- Permuted choice options and "none of the above" augmentation, with a cross-entropy loss.
- Selection. The checkpoint with the highest accuracy on a fixed 2,000-question calibration sample, with NLL breaking ties. This rule was fixed before training, and it chose step 4,500. The benchmarks were never used to choose.
- Temperatures. Fitted after selection on the whole calibration split: the global T, then one T per category with at least 200 calibration questions.
- Data. 165,832 training records and 30,722 calibration records.
- Sources: public datasets and generated decision records.
- Not published. The training files are not part of this release; the SHA-256 of each file is in
training_config.json.
Files
| file | contents |
|---|---|
adapter_model.safetensors, adapter_config.json |
the LoRA adapter (rank 32, PEFT format) |
head.pt |
the pointer head and the calibrated temperatures (global and per category) |
tokenizer.json, tokenizer_config.json |
Qwen3.5-9B's tokenizer, unchanged |
training_config.json, training_metrics.json, train.log |
the recipe, the run and its log |
result.json |
the calibration, Decision Index and latency numbers above, with the calibration metrics per category |
provenance.json |
SHA-256 of every file and of the training data, and library versions |
License
Apache-2.0, see LICENSE.
- Downloads last month
- 40
Model tree for sthanika-ai/Sieve-9B-Plus
Collection including sthanika-ai/Sieve-9B-Plus
Evaluation results
- Decision Index 0.3, public index (chance-corrected; self-reported, not yet on the leaderboard) on Decision Index 0.3 public suite (140,178 scored requests; the index averages 37 benchmarks)self-reported53.350