Nativ (נתיב)

The best open Hebrew decision model for commercial use (as of October 2026). Nativ v2 matches or beats OpenAI's Decisions API on 6 of the 9 tests below and is within 5 points on two more. It beats Laya multilingual on all 9 (the leading commercially licensed open alternative). Runs on a laptop CPU.

Give it any Hebrew text, a question and the options, and it returns a probability for every option in a single forward pass. No generation, no prompt parsing, no GPU.

נתיב (nativ) is Hebrew for "path": the model chooses which way a request should go.

Size 364M parameters (dicta-il/neodictabert + an option scorer)
Speed ~70 ms per decision on a laptop CPU
Output a probability per option, with its calibration error reported for every task
Options any Hebrew text, 2 to 7 per question
License CC BY 4.0, commercial use allowed
Benchmark nativ-bench
Training data hebrew-intent-emotion (translated part)

Usage

decider.py is included in this repo (requires torch, transformers, safetensors).

from decider import HebrewDecider

nativ = HebrewDecider.from_pretrained("path/to/nativ-he-decision")
nativ.decide(
    state="שלום, ההזמנה שלי הייתה אמורה להגיע ביום שלישי ועדיין לא קיבלתי אותה. אפשר לבדוק מה קורה?",
    question="לאיזו מחלקה להעביר את הפנייה?",
    options=["משלוחים", "החזרות והחלפות", "חיובים ותשלומים", "תמיכה טכנית"],
)
# one probability per option, in the order given
  • Options are free Hebrew text. Short descriptions work better than label names ("לקבוע שעון מעורר" rather than "alarm_set").
  • decide_many takes a list of {"state", "question", "options"} dicts.
  • The input is limited to 1,024 tokens. Long states are truncated; the question and options are kept.

Versions

version changes
v2 (October 2026, this one) new training data: topic, inference (MultiNLI), yes / no questions (BoolQ), Belebele-style reading questions, more policy arithmetic. SIB-200 0.755 → 0.858, HebNLI 0.476 → 0.808, Belebele 0.551 → 0.588
v1 (October 2026) first release; still available as revision v1

To keep using v1, download it with huggingface_hub.snapshot_download("yoavipo/nativ-he-decision", revision="v1").

Results

The two nativ-bench tests (nativ-bench) were written for this benchmark: their messages and thresholds never appear in Nativ's training data. The other seven are public Hebrew test sets.

  • 🟢 - marks the tests where Nativ scores highest (or ties).
  • Which test datasets each model trained on: Nativ trained on the train splits of MASSIVE, HeQ, OnlpLab and SIB-200 (the 66 train sentences that are not part of Belebele), and on MultiNLI, which the HebNLI test was translated from (every premise in the HebNLI test was left out); never on a validation or test file. Laya-Hebrew trained on HebNLI, HeQ and Hebrew sentiment data, among others. OpenAI Decisions and Laya multilingual have not published their training data.

Accuracy (higher is better)

The share of questions answered correctly, from 0 to 1: 0.95 means 95 of every 100 answers are right.

test (dataset) Nativ v2 OpenAI Decisions Laya multilingual Laya-Hebrew DeepSeek V4 Flash
what does the user want, 4 options (MASSIVE) 0.965 🟢 0.925 0.644 0.887 0.905
does the paragraph answer the question (HeQ) 0.855 🟢 0.737 0.582 0.694 0.847
sentiment of a comment (OnlpLab) 0.905 🟢 0.865 0.680 0.822 0.838
personal information in a message (nativ-bench) 0.954 🟢 0.593 0.566 0.848 0.890
does one sentence follow from another (HebNLI) 0.808 🟢 0.734 0.480 0.606 0.577
topic of a sentence, 7 options (SIB-200) 0.858 🟢 0.858 0.583 0.789 0.858
request vs policy (nativ-bench) 0.926 0.970 0.298 0.406 0.624
positive / negative / neutral (HebrewSentiment) 0.656 0.699 0.429 0.830 0.611
reading comprehension, 4 options (Belebele) 0.588 0.910 0.318 0.758 0.883

Calibration error (lower is better)

How far the model's confidence is from how often it is actually right, from 0 to 1. When a well-calibrated model says 0.9, it is right about 90% of the time; an error of 0.10 means its confidence is off by about 10 points on average. Low error means the probabilities can be trusted as thresholds, for example "send to a person below 0.8".

test (dataset) Nativ v2 OpenAI Decisions Laya multilingual Laya-Hebrew DeepSeek V4 Flash
what does the user want (MASSIVE) .011 🟢 .011 .085 .028 .050
does the paragraph answer the question (HeQ) .019 🟢 .140 .338 .061 .096
sentiment of a comment (OnlpLab) .009 🟢 .076 .062 .064 .132
personal information in a message (nativ-bench) .078 .285 .356 .028 .069
does one sentence follow from another (HebNLI) .089 .062 .191 .114 .242
topic of a sentence (SIB-200) .072 .058 .112 .096 .109
request vs policy (nativ-bench) .060 .038 .373 .365 .311
positive / negative / neutral (HebrewSentiment) .105 .128 .146 .062 .282
reading comprehension (Belebele) .133 .016 .229 .026 .086

The models compared:

  • OpenAI Decisions API (gpt-6-luna, public beta): closed, API only, $0.10 per million input tokens. OpenAI has not published its size or training data. Run on the full test sets in October 2026.
  • Laya multilingual: 322M, Apache 2.0. Its training data is not published.
  • Laya-Hebrew: 378M, CC BY-NC-SA (non-commercial). Trained on HebNLI, HeQ, Hebrew sentiment data, GoEmotions, CLINC150, Banking77, RACE, BoolQ and other datasets.
  • DeepSeek V4 Flash: a large LLM, zero-shot through an API, scored by the probability of each option's letter. A reference point, not a model that runs on a CPU.

Two MIT-licensed zero-shot classifiers were also tested, mDeBERTa-v3-xnli and bge-m3-zeroshot-v2.0-c, and scored below Nativ on every test they ran (for example, intent: 0.542 and 0.754).

Speed (lower is better), one decision at a time, fp32 on an Apple M2 Pro CPU: Nativ 70 ms for a short message, 130 ms for a paragraph · Laya-Hebrew 77 / 160 ms.

What to use it for

use question options
route a support message לאיזו מחלקה להעביר את הפנייה? משלוחים / החזרות / חיובים / תמיכה טכנית
catch personal details before logging האם ההודעה מכילה מידע מזהה אישי? כן / לא
check a request against a policy האם הבקשה עומדת במדיניות? עומדת / לא עומדת / חסר מידע
check a retrieved passage (RAG) האם הקטע עונה על השאלה? כן / לא
tag a topic or an intent מה הנושא של הטקסט? your own list
sentiment of reviews and comments מה הסנטימנט של התגובה? חיובית / שלילית / לא קשורה

A low top probability is a useful signal to hand the case to a person.

Training

The encoder reads [CLS] state [SEP] question [SEP] option1 [SEP] option2 [SEP] ... in one pass; each option is scored from the [CLS] vector and the mean of its tokens, and a softmax gives the probabilities.

Nativ was distilled (Hinton et al., 2015) from several teachers: a fine-tuned DictaLM-3.0-1.7B-Instruct (Apache-2.0) for routing, sentiment, grounded and personal information; DeepSeek V4 Flash (MIT) for intent, emotion, 3-way sentiment, yes / no questions, topic and the generated reading questions; and DictaLM-3.0-24B-Thinking for the rest of the reading data. Inference and policy are learned from their gold labels (policy labels are computed in code). Part of the data was generated, translated or checked by DictaLM-3.0-24B-Thinking and DeepSeek V4 Flash. During training every item's question and options are reworded at random, so the model learns to read the options.

task data license
routing MASSIVE he-IL train (11.5k), Hebrew descriptions of the 60 intents, some labels corrected by hand CC BY 4.0
sentiment OnlpLab Hebrew-Sentiment-Data, deduplicated (5.9k) MIT
grounded HeQ v1.1 train, answerable and unanswerable questions, balanced (15k) CC BY 4.0
personal information ~5k template messages rewritten by the 24B, ~2k messages written by the 24B generated
policy ~21k rule / request pairs, labels computed in code, each rewritten by the 24B or DeepSeek V4 Flash generated
reading CosmosQA translated to Hebrew (24.8k), questions on HeQ paragraphs (5.1k), questions written by DeepSeek V4 Flash on BoolQ's Hebrew paragraphs (6k) CC BY 4.0 / CC BY-SA 3.0 / generated
inference MultiNLI train translated to Hebrew by DeepSeek V4 Flash, human labels (19.8k) OANC / CC BY 3.0 / CC BY-SA 3.0 / MIT
yes / no questions BoolQ translated to Hebrew by DeepSeek V4 Flash (6.9k) CC BY-SA 3.0
topic SIB-200 heb_Hebr train (66 sentences) and 10.1k sentences from the sources above, labelled by DeepSeek V4 Flash over 18 topics CC BY-SA 4.0 / as the sources
intent CLINC150 (11.9k) and Banking77 (10k), translated to Hebrew; 2 to 7 options per item from 228 intents CC BY 3.0 / CC BY 4.0
emotion GoEmotions single-label comments, translated to Hebrew (5.8k); 2 to 7 options from 28 emotions Apache 2.0
sentiment with neutral 12k texts from the sources above, labelled positive / negative / neutral by DeepSeek V4 Flash as the sources

Generated and translated data were filtered for answer shortcuts, changed facts and broken translations. The Hebrew translations of CLINC150, Banking77 and GoEmotions are published as hebrew-intent-emotion.

Limitations

  • Reading comprehension is the weakest task - this comes partly from the size limit of the model and partly from how little reading data is licensed for commercial use.
  • Whether something counts as personal information depends on the application (order numbers, for example). Adjust the threshold or the options.
  • The current version supports Hebrew only (text with a little English in it is fine).

License

CC BY 4.0, as dicta-il/neodictabert. Training data: MASSIVE (Amazon, CC BY 4.0), HeQ (NNLP-IL, CC BY 4.0), OnlpLab Hebrew-Sentiment-Data (MIT), CosmosQA (AllenAI, CC BY 4.0), CLINC150 (CC BY 3.0), Banking77 (PolyAI, CC BY 4.0), GoEmotions (Google, Apache 2.0), MultiNLI (NYU; OANC, CC BY 3.0, CC BY-SA 3.0 and MIT), BoolQ (Google, CC BY-SA 3.0), SIB-200 train (CC BY-SA 4.0), and data generated with DictaLM-3.0 models (Dicta, Apache-2.0) and DeepSeek V4 Flash (MIT). Belebele (Meta, CC BY-SA 4.0), HebrewSentiment, HebNLI and the SIB-200 test were used for evaluation only.

Downloads last month
123
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yoavipo/nativ-he-decision

Finetuned
(5)
this model

Datasets used to train yoavipo/nativ-he-decision