laya-tf-judge

Laya Multilingual fine-tuned for a single decision, packaged for in-browser inference (onnxruntime-web, WASM). It powers a small demo on MEGAFIXEL.

Given a situation and a short Korean reply (5 to 50 characters), it answers one choice question with two options and returns calibrated probabilities. The question asks one thing: does the reply move toward the other person's feelings?

  • ๋งˆ์Œยท๊ด€๊ณ„(F): it does. Naming or acknowledging feelings, comfort, taking their side, asking how they are, affirming the relationship.
  • ์‚ฌ๊ฑดยทํ•ด๊ฒฐ(T): it does not. This covers two kinds of reply:
    • working on the problem: causes, facts, solutions, next steps;
    • keeping a distance: drawing a line ("that's your problem"), a cold judgment ("that was your fault"), pointing to facts instead of comfort ("crying won't change anything"), or an indifferent reaction ("so what?").

Emotion words alone do not make a reply F, and cold-sounding words alone do not make it T; the label follows what the reply does. The option text shown to the model is unchanged from the earlier version; the wider T definition is learned from the data.

It judges one reply, not a person. It is not a personality test. This model was made for a demo and is not perfect (see Known weaknesses).

Changes from the base model (Apache-2.0 ยง4)

  • Fine-tuned from the base model with the official laya.train.finetune (laya 0.3.28): loss="soft-ce", 4 epochs, encoder lr 2.5e-5, head lr 1e-4, option order shuffled each epoch, seed 0.
  • Training data: 1,457 synthetic Korean replies of 5 to 50 characters to 32 everyday situations, written by an AI model and label-reviewed by a different model. The data is not included.
  • Temperature calibrated with the official routine on 183 replies from 4 held-out situations: choice temperature 2.879 (rl_agent_config.json).
  • Vocabulary pruned from 256,000 to 39,001 tokens after fine-tuning (see Vocabulary pruning). The embedding keeps only the rows of the kept tokens. No weight was retrained or changed.
  • Exported with the official scripts/export_onnx.py and quantized differently from the official --quantize: weight-only 8-bit MatMulNBits (block 32) for linear layers and per-tensor int8 for the embedding. See ONNX.

Evaluation

All test situations are outside training. T first option order, as in the demo.

test_short (141) test_old (72) adv_old (12)
base model, zero-shot 53.2% 59.7% 16.7%
this model 91.5% 100% 100%
  • test_short: new 5 to 50 character replies to 12 held-out situations, label-reviewed. Calibration on this set: Brier 0.391 โ†’ 0.066, ECE 0.364 โ†’ 0.035.
  • test_old: replies of at most 50 characters from the earlier test set (written under the earlier definition and used with their original labels).
  • adv_old: replies of at most 50 characters from the earlier adversarial set (T replies that open with an emotion phrase).

test_short by reply type:

Type Label Rows Accuracy
keeping a distance T 36 86.1%
very short T 34 94.1%
warm but blunt-sounding F 36 97.2%
very short F 35 88.6%

Scores are on synthetic text. Replies written by real people will likely score lower.

Known weaknesses

  • A concession followed by a line-drawing T is often read as F, for example "๋งŽ์ด ๋‹นํ™ฉํ•˜์…จ๊ฒ ์ง€๋งŒ ์ €๋„ ์ง€๊ธˆ์€ ๊ฐ€ ๋ณผ ์ˆ˜๊ฐ€ ์—†์–ด์š”" (I'm sure you're upset, but I can't come right now). Some indifferent reactions ("๊ทธ๋ž˜์š”? ์ „ ๋ณ„ ์ƒ๊ฐ ์—†์–ด์š”") are also read as F.
  • Fact-style comfort (F) is sometimes read as T, for example "๋ง ๋ชป ํ•œ ๊ฑฐ ๋‹ค๋“ค ๊ฒช์–ด ๋ดค์–ด~" (everyone has frozen up like that).
  • Some decisions follow single words: an F reply containing ์–ด์ฉŒ๋ผ๊ณ  was read as T, and a T reply containing ๋„ค ์ž˜๋ชป ๋งž์•„ (it was your fault) was read as F.
  • The same reply can land on either side depending on the situation card. The demo's example reply "๋ญ˜ ์–ด๋–ป๊ฒŒ ํ•˜๋ผ๋Š”๊ฑฐ์ง€? ๋งˆ์Œ์€ ์•„ํ”„๊ฒ ์ง€๋งŒ ๋‚˜๋ž‘์€ ์ƒ๊ด€์—†์–ด." gets P(T) 0.72 with one situation and 0.43 with another.
  • Of 12 wrong test_short rows, 9 are in the two weakest types above. 31 of 34 hand-picked check sentences are right.

Vocabulary pruning

The base tokenizer is the mmBERT tokenizer: a Gemma-style BPE with byte fallback and 256,000 tokens. Its embedding (256,000 ร— 768) holds about 61% of the model's weights, and most of those tokens never appear in Korean input. Pruning cuts the browser download (model, tokenizer, configs and the onnxruntime-web runtime) from about 370 MB to about 189 MB.

What was kept (39,001 tokens)

  • All 249 added and special tokens. <pad> <eos> <bos> <unk> <mask> keep ids 0 to 4.
  • The 256 byte-fallback slots (<0x00> to <0xFF>; the base vocabulary has \t in place of <0x09>).
  • Every single-character token for printable ASCII (94), Hangul syllables (all 1,597 in the base vocabulary) and Hangul compatibility jamo (44).
  • Every token that appears when tokenizing all of our training, test and check text and the exact Laya input sequences (question, options, state), in both option orders.
  • The 2,000 most frequent tokens of a small English sample, and the first 500 emoji tokens in base-vocabulary order.
  • Then tokens ranked by frequency in the Korean corpora below, each corpus weighted equally, up to the size limit.
  • When a token is kept, every piece on its BPE merge path is kept too. A word whose original tokens are all kept therefore tokenizes exactly as before. Any other word splits into smaller kept pieces or, as a last resort, into bytes. <unk> never appears.

The tokenizer was edited directly: removed tokens were deleted from the vocabulary, every merge whose left part, right part or result was removed was deleted, and ids were renumbered in the original order. Token ids differ from the base tokenizer; use this tokenizer.json with this model.

Corpora

Used only on the training machine to count token frequencies and to evaluate. No corpus text is included in this repository.

Corpus Counted Held out for evaluation License
Korean Wikipedia (wikimedia/wikipedia 20231101.ko, file 3 of 3) 49 of every 50 articles (92M tokens) sentences from every 50th article CC BY-SA 3.0, GFDL
NSMC movie reviews train (150k reviews) test file CC0 1.0
Korean chatbot data 9 of every 10 Q/A pairs every 10th pair MIT
Korean HateSpeech Dataset news comments unlabeled comments, file 1 (500k) labeled train and dev comments CC BY-SA 4.0
3i4k conversational utterances train_val test CC BY-SA 4.0
WikiText-103 raw, validation and test (English) top 2,000 tokens only โ€“ CC BY-SA 3.0, GFDL

Size options

"Coverage" is the share of token occurrences in the five Korean corpora (equal weight) whose token is kept. "Same tokenization" is measured on 4,500 held-out Korean sentences of 10 to 200 characters that were not used for counting. "Download" is model + tokenizer + configs + the onnxruntime-web and transformers.js files.

Vocabulary Download Korean coverage English coverage Same tokenization (held out) Tokens added (held out)
24,000 177.0 MB 99.847% 84.9% 97.04% +0.29%
32,000 183.5 MB 99.888% 86.6% 97.58% +0.22%
39,001 (this build) 189.2 MB 99.913% 87.8% 97.73% +0.18%
48,000 196.6 MB 99.935% 89.0% 98.20% +0.13%
64,000 209.7 MB 99.962% 91.3% 98.84% +0.09%
256,000 (unpruned) 369.6 MB 100% 100% 100% 0

39,001 is the largest size that keeps the download under 190 MB. On the 3,626 held-out sentences of at most 50 characters (the demo's input length), 98.6% tokenize exactly as before. Most differences are rare names and foreign words in encyclopedic text.

Verification of this model

  • 259 evaluation items (test_short 141, test_old 72, adv_old 12, and 34 hand-picked check sentences), PyTorch, both option orders: token pieces identical on 259/259, probability difference 0, decisions 259/259.
  • Held-out Korean sentences, each used as the reply to one of the test situation cards and judged T first:
    • 500 random sentences: same decision on 500/500 (largest probability difference 0.22, on a long encyclopedic sentence whose tokenization changed).
    • All 102 sentences whose tokenization changed: same decision on 99/102.
    • Over all 4,500 sentences: 4,497 (99.93%) same decision. <unk> count: 0.

ONNX

  • The 8-bit recipe is applied to the full-vocabulary export first; then the int8 embedding rows of the kept tokens are sliced out, keeping the per-tensor scale of the full embedding. For inputs that use only kept tokens, the output is identical to the unpruned 8-bit model.
    • Quantizing after pruning instead re-derives the embedding scale from the kept rows. That version moved one borderline check sentence from P(T) 0.429 to 0.501 and agreed on 258/259, so it was not used.
  • ONNX vs pruned PyTorch (onnxruntime 1.30 CPU, batch 1, 259 items): same decision on 259/259, largest probability difference 0.064.
  • Browser (onnxruntime-web 1.30 WASM, transformers.js 4.3 tokenizer, headless Chromium): token ids identical to Python on 259/259, largest difference from Python onnxruntime 2.8e-6, no flipped decision. Single-threaded inference took a median 1.0 s per item; session start from a local server took 1.1 s.

Files

File Note
onnx/model_q8.onnx 184 MB, single graph
tokenizer.json, tokenizer_config.json pruned 39,001-token tokenizer, 1.7 MB (truncation reset to none)
rl_agent_config.json max_len, head_max_len, temperature
questions.json the exact question and option text used in training

Inference

  • Inputs: input_ids, attention_mask, marker_pos, marker_mask, qtype (choice = 0). Output logits is before temperature.
  • Sequence: [CLS] <choice> question: Q [SEP] [MASK] option0 [MASK] option1 [SEP] state [SEP], following the official laya-ts client. State format used in training: ์ƒํ™ฉ: <situation>\n๋Œ€๋‹ต: <reply>.
  • Replies were 5 to 50 characters in training and testing; longer input is outside what was evaluated.
  • Keep the option order T first, as evaluated. P(T) = 1 / (1 + exp(-(logits[0] - logits[1]) / 2.879)).
  • Run one item per call (batch 1).

License

Apache-2.0, inherited from convaiinnovations/laya-multilingual. See LICENSE.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Dunde/laya-tf-judge

Quantized
(36)
this model