barga

barga (Nepali बर्ग, "category") is a small system-one decision model for English and Nepali. It reads a state (a conversation, a call transcript, a document, a game position) and one or more questions, each with its own candidate options, and returns a probability for every option in a single forward pass. It does not generate text. It decides.

  • Size: 140.6M parameters (ModernBERT / mmBERT-small encoder + a small option scorer), fp32 safetensors, 562 MB.
  • Languages: English, Nepali (Devanagari and romanized), and code-mixed Nepali–English.
  • Architecture: split-layer. The lower 11 of 22 encoder layers read the state once; the top 11 layers read each question jointly with the state. Every question about the same state shares one state encoding.
  • Speed: 0.90 s per turn (median; 1.45 s p95; measured with about one core of background load) on an Apple M3 Pro CPU with 4 threads, for simulator call turns of about 6 questions over states of up to 1,024 tokens; a short input with one question takes 0.05–0.2 s. Peak process memory 2.6 GB.
  • Licence: Apache-2.0. Built on Julia-1 (Apache-2.0), which is built on mmBERT-small (MIT).
  • Made by Ampixa.
  • Try it: play against barga in the Kirtipur courtyard: Bagh-Chal, chess, Ludo, Marriage, Dhumbal and Call Break, with every option it weighed and the exact request and reply. Launch video: English · नेपाली. (The site still runs the previous checkpoint, g14.)

Versions

revision checkpoint what it is for
main (this card, 2026-10-11) G21, seed 17 General decisions plus streaming turn-end timing. Best overall accuracy. Does not pause its own speech when a caller interrupts (see Turn-taking).
turn-taking G22b, seed 17 Voice-agent floor control: pauses for a real interruption, keeps playing through "हजुर". About 1–2 points lower on general decisions.
g14 G14, seed 17 The first release (2026-10-05), unchanged.
barga.load("ampixa/barga")                          # main (G21)
barga.load("ampixa/barga", revision="turn-taking")  # G22b
barga.load("ampixa/barga", revision="g14")          # the 2026-10-05 release

What changed from g14: the training mix adds 15% turn-taking states, built from the same partner calls. Some states are anchored to the call's timeline; others are streaming states polled every 0.2 s while the call runs. On protected test calls the new checkpoint:

  • cuts in during a caller's mid-turn pause half as often (0.12 vs 0.25);
  • rarely misses a turn end (0.02 vs 0.24);
  • scores 90.1 on timeline floor decisions (g14 75.7);
  • is +0.7 on the mean of 17 general test sets.

Some sets are lower: chess (−6.8), AG News (−6, on 100 questions) and the other games (−0.1 to −1.9).

Quickstart

pip install barga
import barga

model = barga.load("ampixa/barga")          # device="cuda" for a GPU

state = ("Caller: Hello, I want to book a check-up for my dog tomorrow morning.\n"
         "Agent: Sure. We have 9:30 or 11:00 tomorrow.\n"
         "Caller: 9:30 is fine.")

for d in model.decide(state, [
    {"id": "intent", "text": "What does the caller want?",
     "options": [{"id": "book", "description": "Book an appointment"},
                 {"id": "cancel", "description": "Cancel an appointment"},
                 {"id": "info", "description": "Only ask for information"}]},
    {"id": "intent_ne", "text": "कल गर्नेले के चाहनुहुन्छ?",
     "options": [{"id": "book", "description": "अपोइन्टमेन्ट बुक गर्न"},
                 {"id": "cancel", "description": "अपोइन्टमेन्ट रद्द गर्न"},
                 {"id": "info", "description": "जानकारी मात्र सोध्न"}]},
]):
    print(d.question_id, d.choice, d.probs)
# intent    book  {'book': 0.984, 'cancel': 0.015, 'info': 0.001}
# intent_ne book  {'book': 0.866, 'cancel': 0.134, 'info': 0.001}

Questions can be "choice" (default), "boolean" (two options valued True/False) or "ordinal" (numeric option values; the result also carries an expected_value). Each takes 2–20 options and an optional rubric. The model reads only the state, the question, the rubric and each option's description, never the ids. The input format is documented in the barga package.

How to read the output

choice is the most probable option; probs gives every option's probability (they sum to 1 within a question). Probabilities compare options within one question. A flat distribution (say a top probability below 0.6) means the model is unsure, which is often correct when the state does not answer the question. Set your own threshold on your own data before acting on a decision automatically.

Intended uses

  • Turn-by-turn decisions in voice and chat agents: intent, slot values, whether the caller confirmed, what to do next.
  • Turn-end timing in a voice agent: whether the caller has finished and the agent should answer now (see Turn-taking; for pausing on interruptions use the turn-taking revision).
  • Classification, routing and yes/no checks over English or Nepali text, with the label set given at run time.
  • Reading-comprehension style checks (does the text support this statement? is this question answerable?).
  • Rule and strategy questions about Nepali-played card and board games (see the limits below).

Evaluation

All numbers are accuracy (%) on held-out test sets that were never used for training or checkpoint selection. The released checkpoint is seed 17. Two-seed is the mean of seed 17 and an identically trained seed-18 replicate, shown to indicate seed variance. g14 is the previous release's checkpoint on the same questions. Sampled sets use a fixed seed (17).

Calls and agent decisions

test set questions released two-seed g14
Held-out real customer-service calls, set A (English questions) 2,589 76.7 76.7 77.0
same calls, Nepali questions 2,589 76.1 76.2 77.1
set A scored against what the agent in the recording did (English / Nepali questions)¹ 2,380 77.4 / 77.7 78.1 / 78.4 76.7 / 77.7
set A in the output style of a speech recognizer (English / Nepali questions) 2,589 71.1 / 71.8 71.5 / 71.7 72.1 / 71.6
Held-out real calls, set B (English) 135 67.4 66.3 64.4
set B in speech-recognizer style 135 62.2 61.9 62.2
Receptionist simulator, template phrasing (300 calls) 1,779 94.4 94.6 94.4
Receptionist simulator, LLM phrasing (300 calls) 1,678 94.9 95.5 96.1
typed-decisions test split 2,000 76.3 76.8 76.3
Timeline floor decisions on protected calls (94 calls) 4,532 90.1 90.4 75.7

¹ About two thirds of set A's questions ask what the agent should do with the floor at that moment, and an audio teacher labelled them. That teacher agrees with what the human agent actually did only about 70% of the time at turn ends. This row scores those questions against the recorded behaviour instead, where it is defined.

Reading comprehension and classification (300-question samples unless noted)

test set released two-seed g14
MultiNLI 68.3 69.3 66.0
MNLI-Nepali 65.0 66.0 64.3
SQuAD 2.0 62.3 61.5 61.0
BoolQ 61.7 63.5 60.7
ShARC 67.3 64.8 62.7
PAWS 57.0 56.0 56.7
CLINC150 intents 93.3 92.2 90.7
AG News (100) 89.0 90.0 95.0
Emotion (100) 82.0 80.5 81.0
Hard cases (1,500; the hard slice of an Ampixa support-ticket decision set) 36.5 31.7 26.3

Games (rule and strategy questions; positions from engines and self-play)

game questions released two-seed g14
Dhumbal 1,440 61.9 62.4 62.0
Ludo 1,592 61.0 61.7 62.9
Marriage 1,494 46.3 46.3 47.7
Call Break 1,828 47.2 47.3 47.8
Chess 1,674 36.0 38.7 42.8
Bagh-Chal 3,726 23.1 23.2 24.0

Game questions have 2–20 options, so chance differs by set: about 14% on Bagh-Chal, and up to 50% on yes/no families.

Turn-taking (streaming floor control, protected test calls)

The stream test replays 94 protected real calls, never used for training, and asks the floor question every 0.2 s: wait, answer now, play a short acknowledgement, or pause the agent's own speech. The scores are raw argmax choices, two-seed means. Lower is better except for stopping.

error main (G21) turn-taking (G22b) g14
Cut-in: answers in a mid-turn pause after which the caller goes on 0.12 0.16 0.25
Misses a turn end (has not answered by the time the recorded agent spoke) 0.02 0.02 0.24
Answers while the caller is speaking (per tick) 0.00 0.00 0.22
Pauses or cuts into its own ongoing speech (per tick) 0.00 0.00 0.81
Stops or answers when the caller only acknowledges during its speech (share of acknowledgements) 0.00 0.00 0.54
Stops for a real interruption within 1.0 s (482 caller runs of ≥ 3 words and ≥ 1.0 s; higher is better) 0.00 0.82 0.72
Stops for a short run (≤ 2 words) 0.00 0.04 0.51

The stop rows come from a separate probe on the same 94 calls. It plays the agent's reply as if it were still playing when the caller starts, and asks at 0.2–2.0 s into the caller's speech.

  • Threshold: answering only when P(COMMIT_RESPONSE) ≥ 0.65 in caller silence (tuned on 30 separate dev calls) lowers main's cut-in to 0.10 and raises its misses to 0.03.
  • No SOFT_ACK: no checkpoint here chooses the short-acknowledgement option (0 of 379 opportunities on the stream test).

Main never pauses its own speech for a caller. Its streaming training data had no stop examples, so it learned "the caller is talking over playback, keep playing". If your agent needs barge-in handling, use the turn-taking revision, or handle barge-in outside the model (for example, a voice-activity rule). The input format for these states is documented on the turn-taking card.

Against Julia-1 on identical questions

Same questions and the same scoring for both models; a question a model cannot encode counts as wrong.

Julia-1 barga (released)
typed-decisions test 72.5 76.3
CLINC150 55.7 93.3
MultiNLI / MNLI-Nepali 30.7 / 31.0 68.3 / 65.0
ShARC / BoolQ / SQuAD 2.0 31.0 / 45.7 / 51.0 67.3 / 61.7 / 62.3
PAWS 54.3 57.0
AG News (100) 94.0 89.0
Emotion (100) 86.0 82.0
Hard cases 33.7 36.5
Julia-1's own evaluation families (12 Open-Jev families, macro) 87.8 56.3

On the real-call and simulator sets barga scores 76–77 and 94–95, against Julia-1's 22 and 23. Read that gap with care:

  • Encoding: Julia-1 cannot encode some of those questions; on the questions both models answer it scores 17.7–24.8.
  • Training data: barga was trained on the training splits of those same sources, and Julia-1 was not.

Julia-1 leads on its own evaluation families: game-playing agents (tile platformer, ViZDoom, Snake), workflow and customer controls, and reasoning controls. barga was not trained on those families.

Limitations

  • No barge-in. This checkpoint does not pause its own speech when a caller interrupts (above). Use the turn-taking revision or an external rule.
  • Near chance on some game families. Chess move legality (yes/no), tic-tac-toe, Bagh-Chal, Marriage discards, Ludo safe squares and Call Break card play score within a few points of chance. Do not use barga as a game engine; use a rules engine for legality and barga, at most, to rank legal moves.
  • Out of domain. Julia-1's agent-control families (above) and the hard support-ticket cases are weak spots.
  • Real-call accuracy is about 77%. Speech-recognizer errors cost about 5 points (71–72%). Keep a human or a confirmation step in the loop for anything with consequences.
  • Calibration. Probabilities are informative but not calibrated to your data; threshold them on your own data.
  • Context. 1,024 state tokens (when a state is longer, barga keeps its beginning and its most recent end) and 768 tokens per question with its options. Longer questions are rejected, not cut.
  • Languages. Tested on English, Nepali (Devanagari and romanized) and code-mixed text. Other languages inherit some ability from mmBERT but were not evaluated.
  • Seed variance. Two identically trained seeds differ by up to 5 points on small sets (100–300 questions), by 5.4 on chess and by 9.6 on the hard cases (36.5 vs 26.9). Differences smaller than that between models are not meaningful.

Out-of-scope uses

Generating text; open-ended question answering without candidate options; medical, legal, financial or safety-critical decisions without human review; profiling or surveillance of people; any use that violates the licences of the training data sources listed below.

Training data

barga was fine-tuned from the encoder of Julia-1 on a weighted mix of 38 training sets:

source share licence / terms
Nepali/English customer-service calls from partners, used under agreement; not released 13.1% private; used under agreement
Turn-taking states built from the same partner calls: timeline-anchored floor decisions and 0.2 s streaming states (labels from recorded timing; short caller runs audio-checked in two agreeing passes); not released 15.0% private; used under agreement
Receptionist simulator: synthetic calls; template phrasing and LLM-written phrasing 11.4% generated by Ampixa
typed-decisions (train split) 7.2% Apache-2.0
Paraphrases of typed-decisions train cases (same targets) 1.0% as typed-decisions
MultiNLI, SQuAD 2.0, ShARC, CLINC150, BoolQ, PAWS (English) 15.9% CC BY 3.0 / CC BY-SA 3.0 / MIT / other (MultiNLI), CC BY-SA 4.0, CC BY-SA 3.0, CC BY 3.0, CC BY-SA 3.0, "may be freely used for any purpose"
MultiNLI, SQuAD 2.0, ShARC, CLINC150, BoolQ machine-translated to Nepali (Sarvam-Translate) 11.5% as the originals
MNLI-Nepali (IRIIS-RESEARCH) 3.9% see the dataset
Bagh-Chal, Marriage, Dhumbal, Ludo, Call Break positions with rule-checked labels 10.2% generated by Ampixa
Chess positions from the Lichess open database with engine labels 2.5% CC0 (Lichess database)
Ampixa support-ticket decision sets (routing, negation, urgency, emotion, abstention) 3.9% Ampixa
Typed-decision replays with teacher soft labels 1.9% generated by Ampixa
AG News, Emotion (dair-ai), Banking77, MASSIVE (scenario) 2.6% research / non-commercial (AG News), educational and research purposes only (Emotion), CC BY 4.0, CC BY 4.0

Protected test calls and the typed-decisions test split were never used for training. No call audio or transcripts are released.

Files

  • backbone/: ModernBERT encoder (config.json, model.safetensors)
  • heads.pt: option scorer weights
  • tokenizer/: tokenizer (the mmBERT/Gemma-style 256k vocabulary)
  • bundle.json: architecture, serialization config and the sha256 of every file above; barga.load verifies them before loading. Its backbone reference was rewritten from a local path to its public source; every hashed file is byte-identical to the evaluated checkpoint.

barga 0.1.1 from PyPI reproduces the reference evaluation exactly on this checkpoint: same choices, probabilities within 1e-6. That covers 50 records each from the simulator, typed-decisions, chess, streaming, stop-probe and timeline test sets.

Citation

@misc{barga2026,
  title  = {barga: a system-one decision model for English and Nepali},
  author = {Ampixa},
  year   = {2026},
  url    = {https://huggingface.co/ampixa/barga}
}

barga builds on Julia-1 (Supersonic Labs) and mmBERT (Johns Hopkins University CLSP).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ampixa/barga

Finetuned
(8)
this model

Evaluation results

  • Accuracy (released checkpoint) on typed-decisions (test split)
    test set self-reported
    76.300
  • Accuracy (released checkpoint) on MultiNLI (300-question sample)
    self-reported
    68.300
  • Accuracy (released checkpoint) on MNLI-Nepali (300-question sample)
    self-reported
    65.000
  • Accuracy (released checkpoint) on SQuAD 2.0 (300-question sample)
    self-reported
    62.300
  • Accuracy (released checkpoint) on CLINC150 (300-question sample)
    self-reported
    93.300