tnh0527 commited on
Commit
4289dff
·
verified ·
1 Parent(s): 09ee555

Publish gliclass-std-base-v3-daecore-5facet-qint8-v2 (qualified ONNX export and complete attribution)

Browse files
MODIFICATIONS.md CHANGED
@@ -21,6 +21,11 @@ What changed:
21
  - **Calibration.** Per-facet temperature scaling and two frozen threshold
22
  tables (`recall_leaning`, `contract`) travel in `classifier-metadata.json`
23
  and are part of the artifact's identity.
 
 
 
 
 
24
 
25
  Unchanged: the tokenizer vocabulary and the `<<LABEL>>` / `<<SEP>>` prompt
26
  convention of the upstream model.
 
21
  - **Calibration.** Per-facet temperature scaling and two frozen threshold
22
  tables (`recall_leaning`, `contract`) travel in `classifier-metadata.json`
23
  and are part of the artifact's identity.
24
+ - **Vulkan graph preparation.** Bounded integer/boolean mask calculations
25
+ use exact FP32 equivalents before their original output types are restored.
26
+ Learned parameters, tokenizers, label prompts, calibration and thresholds
27
+ are unchanged. `vulkan-derivation.json` binds the source and derived graph;
28
+ `classifier-metadata.json` records the new graph's byte identity.
29
 
30
  Unchanged: the tokenizer vocabulary and the `<<LABEL>>` / `<<SEP>>` prompt
31
  convention of the upstream model.
README.md CHANGED
@@ -1,511 +1,126 @@
1
- ---
2
- license: apache-2.0
3
- base_model: knowledgator/gliclass-base-v3.0
4
- pipeline_tag: text-classification
5
- tags:
6
- - gliclass
7
- - deberta-v3
8
- - multi-label-classification
9
- - fine-tuned
10
- - onnx
11
- - int8
12
- ---
13
-
14
- # GLiClass Base v3, five-facet memory classifier
15
-
16
- ## Summary
17
-
18
- This model tags passages of working notes and documentation with five
19
- independent facets: `trap`, `decision`, `constraint`, `mechanism`, and
20
- `procedure`. It is a supervised fine-tune of
21
- [`knowledgator/gliclass-base-v3.0`](https://huggingface.co/knowledgator/gliclass-base-v3.0)
22
- trained on 59,886 labeled passages and shipped as a 452.8 MB ONNX graph inside a
23
- 461.5 MB bundle, qualified on CPU, CUDA, and DirectML.
24
-
25
- On a frozen 1,300-row held-out panel drawn from three unseen document families,
26
- it reached macro average precision 0.9676 against 0.8526 for the previous
27
- classifier, a paired gain of 0.1150 with every facet improving. At 90% recall it
28
- is 0.902 precise on average against 0.719 before. The gain concentrates on the
29
- two hard facets: decision rose from 0.678 to 0.914 average precision and
30
- procedure from 0.817 to 0.958.
31
-
32
- Two caveats travel with every number below. All labels, for training and for
33
- evaluation, are frontier-model judgments under a frozen protocol, not human
34
- annotations. All evaluation panels are generated documents, chosen because they
35
- span far more domains and document shapes than the operator's own corpus; how
36
- the model behaves on unrelated real documents has not been measured.
37
-
38
- ## Model details
39
-
40
- | Item | Value |
41
- |---|---|
42
- | Architecture | GLiClass single-pass label-conditioned classifier on a DeBERTa-v3 base backbone |
43
- | Parameters | 186.5M |
44
- | Input contract | 768 tokens per passage, five label prompts |
45
- | Output | five independent posteriors in [0, 1], temperature-scaled per facet |
46
- | Serialized artifact | `model.onnx`, 452,812,018 bytes, opset 17 |
47
- | Quantization | signed INT8 on the token-embedding matrix only; the transformer body stays FP32 |
48
- | Threshold tables | two frozen operating views, `contract` and `recall_leaning` |
49
- | Deployment state | the installed default since 2026-09-11; published 2026-09-06 |
50
-
51
- The artifact size follows from the quantization choice: the 128k-entry
52
- embedding table, about 98M parameters, is stored at one byte per weight while
53
- the 88M-parameter body stays at four. Broader quantization was measured, not
54
- skipped: on the predecessor fit of the same recipe, six arms were scored
55
- against the FP32 reference on a 5,298-row panel,
56
- and every one beyond the embedding table cost quality — embedding plus attention
57
- −0.0013, full per-channel reduced-range −0.0531, and full per-channel −0.0599
58
- macro average precision.
59
-
60
- ## Intended use
61
-
62
- The model stamps passages inside a local-first memory system so that retrieval
63
- can weight evidence by kind. The full posterior map is the product; thresholds
64
- are advisory views for consumers that need a hard label. It is not a safety
65
- classifier, a factuality judge, or an authority detector, and it should not be
66
- used to filter what a retrieval system is allowed to return.
67
-
68
- ## Task and labels
69
-
70
- Each facet answers one question about a passage:
71
-
72
- - `trap`: it names a specific mistake, hazard, pitfall, or failure mode.
73
- - `decision`: it makes or records a choice among alternatives.
74
- - `constraint`: it states a rule, policy, invariant, or restricting condition.
75
- - `mechanism`: it explains how something works.
76
- - `procedure`: it gives sequenced steps for carrying out a task.
77
-
78
- Facets overlap freely. The hardest boundary is between a passage that makes a
79
- decision and one that merely discusses or mentions one. Labels distinguish
80
- "performs" from "mentions only", and mentions-only rows count as negatives for
81
- both training and evaluation.
82
-
83
- **Label provenance.** Every label was produced by a language model applying the
84
- same one-sentence facet definitions the classifier sees. Training rows received
85
- two independent judgments each with a higher-effort tiebreak on semantic
86
- disagreement, except for 3,750 inherited rows that received one primary judgment
87
- from GPT-5.6 (Sol, medium reasoning) with a frozen 400-row cross-model audit by
88
- GPT-5.6 (Terra, high) and a Sol-high tiebreak where the two disagreed.
89
- Calibration and promotion panels used two
90
- independent Sol-high sessions per row plus a third Sol-xhigh session that saw
91
- only the disputed fields. Unresolved fields stay unresolved and are masked per
92
- facet; nothing unresolved becomes a negative.
93
-
94
- Before adopting model judges, the operator labeled about fifty passages by
95
- hand and compared. The hand labels were not reliable enough to serve as an
96
- answer key: attention drifted toward one or two facets per passage, and rows
97
- came out either under- or over-labeled relative to the definitions. Frontier
98
- judgments applied the definitions more evenly, so they became the label
99
- instrument. The consequence is stated plainly in every result: scores measure
100
- agreement with that instrument, not with human ground truth.
101
-
102
- ## Training data
103
-
104
- | Lane | Rows | Origin |
105
- |---|---:|---|
106
- | Final training set | 59,886 | 54,803 generated; 4,583 real rows from the operator's self-owned corpora and 500 real rows of public-domain US Federal Register text |
107
-
108
- The 59,886 passages come from 9,376 documents, about 6.4 passages each: 8,212
109
- generated documents spanning 143 fictional organizations, and 1,164 real
110
- documents — 1,042 from the operator's two workspaces and 122 from the Federal
111
- Register set.
112
-
113
- Generated documents come from fictional organizations written as complete
114
- documents (procedures, incident reports, decision records, design notes) and
115
- then passed through the production parser. They were preferred over the
116
- real lane for both training and evaluation because they cover many more domains,
117
- document shapes, and organizational settings. The real lane is narrower than its
118
- family count suggests: 315 of its 437 source families are directory groupings inside
119
- two workspaces belonging to one person, and the other 122 are single Federal
120
- Register dockets, not 437 organizational settings; the Federal Register rows are
121
- also regulatory prose from a single genre. Whole-family
122
- splitting is therefore a strong isolation guarantee in the generated lane, where
123
- a family is a whole organization, and a weak one in the real lane.
124
-
125
- Every real row is redistributable — the operator's corpora by ruling,
126
- the Federal Register text as public-domain US government work. The Federal
127
- Register rows are training supervision only: no panel that calibrates, gates, or
128
- reports on this model contains them, or any other real row.
129
-
130
- Splits are by whole source family, never by row. The seven families reserved
131
- for calibration and promotion were removed from training along with 3,875
132
- predecessor training rows that shared them. Calibration and promotion families
133
- are disjoint from each other and from training by family, document lineage, and
134
- exact passage identity. Assignment was label-blind and checked for support only
135
- afterward.
136
-
137
- The final fit used weighted binary cross-entropy with a 2.0 multiplier on
138
- mentions-only hard negatives, learning rate 2e-5, batch 4 with 8 accumulation
139
- steps, four epochs, and an equal-weight average of the checkpoints from epochs
140
- 2 through 4. Generated and real rows were weighted equally.
141
-
142
- ## Evaluation data
143
-
144
- | Panel | Rows | Families | Origin | Role |
145
- |---|---:|---:|---|---|
146
- | Calibration | 1,700 | 4 | generated | fit temperatures and threshold tables; no evaluation authority |
147
- | Promotion | 1,300 | 3 | generated | one sealed read; the numbers reported below |
148
-
149
- Promotion-panel support per facet, with mentions-only rows counted as
150
- negatives. Prevalence equals the average precision a random ranker would score.
151
-
152
- | Facet | Positive | Mentions only | Negative | Unresolved | Prevalence |
153
- |---|---:|---:|---:|---:|---:|
154
- | trap | 933 | 62 | 305 | 0 | 0.718 |
155
- | decision | 598 | 398 | 303 | 1 | 0.460 |
156
- | constraint | 1,187 | 32 | 81 | 0 | 0.913 |
157
- | mechanism | 892 | 62 | 343 | 3 | 0.688 |
158
- | procedure | 485 | 277 | 531 | 7 | 0.375 |
159
-
160
- The macro random floor is 0.631. Constraint is near saturation at 0.913
161
- prevalence, so its average precision carries little information and the macro
162
- average leans on the other four facets. The panel was deliberately enriched for
163
- hard decision and procedure rows; it is not a natural-traffic sample.
164
-
165
- ## Evaluation protocol
166
-
167
- The promotion plan was frozen on 2026-08-12 before the candidate was scored:
168
- one read of the panel, bootstrap seed 20260812, 10,000 draws resampling whole
169
- source families, 95% one-sided bounds, and zero permitted stamp flips between
170
- the reference scorer and the production ONNX adapter. Pass criteria were
171
- pre-registered on the decision facet under its frozen contract threshold:
172
- precision at least 0.70 with a one-sided 95% lower bound at least 0.60, and
173
- recall at least 0.50. Threshold tables were fitted on the calibration panel and
174
- bound into the prediction receipt; nothing was tuned on the promotion panel.
175
-
176
- Average precision and precision at fixed recall are the primary metrics because
177
- they do not depend on a threshold. Brier score and 10-bin expected calibration
178
- error report probability quality separately from ranking quality.
179
-
180
- ## Results
181
-
182
- Promotion panel, 1,300 rows, three families, single sealed read. "Previous"
183
- is the classifier this model replaces.
184
-
185
- | Facet | Prevalence | AP, this model | AP, previous | P@R90, this model | P@R90, previous | P@R95, this model | P@R95, previous |
186
- |---|---:|---:|---:|---:|---:|---:|---:|
187
- | trap | 0.718 | 0.9864 | 0.8931 | 0.960 | 0.832 | 0.935 | 0.784 |
188
- | decision | 0.460 | 0.9135 | 0.6783 | 0.741 | 0.504 | 0.658 | 0.484 |
189
- | constraint | 0.913 | 0.9986 | 0.9884 | 0.996 | 0.956 | 0.995 | 0.942 |
190
- | mechanism | 0.688 | 0.9817 | 0.8865 | 0.926 | 0.777 | 0.891 | 0.736 |
191
- | procedure | 0.375 | 0.9578 | 0.8167 | 0.886 | 0.528 | 0.822 | 0.489 |
192
- | macro | 0.631 | 0.9676 | 0.8526 | 0.902 | 0.719 | 0.860 | 0.687 |
193
-
194
- The macro gain is 0.1150. The one-sided 95%
195
- lower bound on the macro precision gain under the frozen contract thresholds is
196
- 0.009; the decision-facet precision gain has a negative lower bound, because the
197
- previous classifier's decision threshold was so strict that it predicted only 26
198
- positives at 92% precision and 4% recall.
199
-
200
- ![Per-facet precision-recall curves on the promotion panel for the promoted model, the previous classifier, and the linear baseline](figures/promotion_precision_recall.svg)
201
-
202
- **Frozen-threshold operating points on the promotion panel.** The `contract`
203
- table is the stricter view; the pre-registered gates apply to decision.
204
-
205
- | Facet | Threshold | Precision | Recall | Predicted positive |
206
- |---|---:|---:|---:|---:|
207
- | trap | 0.976 | 0.995 | 0.636 | 596 |
208
- | decision | 0.485 | 0.890 | 0.703 | 473 |
209
- | constraint | 0.993 | 1.000 | 0.802 | 952 |
210
- | mechanism | 0.900 | 0.992 | 0.655 | 591 |
211
- | procedure | 0.578 | 0.946 | 0.715 | 371 |
212
-
213
- The recall denominators in this table do not all match the support table above.
214
- Where a facet has unresolved rows, the promoted model's recall is computed over a
215
- slightly larger set: decision 599 against 598 positives, mechanism 895 against
216
- 892, procedure 491 against 485. The previous classifier is scored over the
217
- support counts, except procedure at 487. The gaps track the per-facet unresolved
218
- rows, which the sealed read did not mask identically for the two models. The
219
- effect is conservative for the promoted model — a larger denominator lowers its
220
- recall — and no gate outcome changes.
221
-
222
- Decision passed all three gates: precision 0.890 against a 0.70 floor, a
223
- one-sided 95% lower bound of 0.872 against 0.60, and recall 0.703 against 0.50.
224
- The operating point encodes a product judgment: surfacing a non-decision as a
225
- decision is treated as worse than missing one. The `recall_leaning` table
226
- (decision threshold 0.193) exists for consumers with the opposite preference.
227
-
228
- **Calibration.** Per-facet temperatures fitted on the calibration panel were
229
- 1.89 (trap), 2.40 (decision), 1.69 (constraint), 1.50 (mechanism), and 2.05
230
- (procedure). On the promotion panel the calibrated posteriors score macro Brier
231
- 0.072 and macro expected calibration error 0.039, against 0.349 and 0.374 for
232
- the previous classifier.
233
-
234
- ![Per-facet reliability diagrams of the promoted model's calibrated probabilities on the promotion panel](figures/promotion_reliability.svg)
235
-
236
- Held-family slices are diagnostic only: macro average precision was 0.962,
237
- 0.967, and 0.974 across the three families for this model and 0.851, 0.850, and
238
- 0.859 for the previous one.
239
-
240
- ## Baselines and comparisons
241
-
242
- Ordered by how directly each comparison can be checked.
243
-
244
- **Previous classifier.** A ModernBERT-base sequence-classification head used as
245
- a zero-shot entailment judge: one pass per facet, five passes per passage. Its
246
- staged bundle declares `ModernBertForSequenceClassification` over the two
247
- labels `entailment` and `not_entailment`, 22 layers at hidden size 768, and a
248
- 599,027,211-byte FP32 `model.onnx` — about 1.3 times the promoted artifact for
249
- five times the work per passage. Its frozen thresholds were fitted on an older
250
- population and artifact, which is why its contract recall collapsed to 0.001 on
251
- trap and 0.040 on decision on this panel; the threshold-free columns above are
252
- the fair comparison.
253
-
254
- **Lexical baseline.** A word-level TF-IDF (1-2 grams, 200k features) with one
255
- balanced logistic-regression head per facet, trained on the same 59,886 rows
256
- and scored on the same 1,300-row panel under the same label rules.
257
-
258
- | Facet | Floor | TF-IDF + LR | This model | Previous |
259
- |---|---:|---:|---:|---:|
260
- | trap | 0.718 | 0.970 | 0.986 | 0.893 |
261
- | decision | 0.460 | 0.847 | 0.914 | 0.678 |
262
- | constraint | 0.913 | 0.997 | 0.999 | 0.988 |
263
- | mechanism | 0.688 | 0.963 | 0.982 | 0.886 |
264
- | procedure | 0.375 | 0.914 | 0.958 | 0.817 |
265
- | macro AP | 0.631 | 0.938 | 0.968 | 0.853 |
266
- | macro P@R90 | | 0.834 | 0.902 | 0.719 |
267
- | macro P@R95 | | 0.786 | 0.860 | 0.687 |
268
-
269
- The linear model lands 0.029 macro average precision below this model and well
270
- above the previous classifier. Much of the task is lexical on generated text.
271
- The fine-tune earns its size at high recall on the hard facets: at 95% recall it
272
- is 9 points more precise on decision and 17 points more precise on procedure
273
- than the linear model, and those are the operating regions the memory system
274
- runs in.
275
-
276
- ![Average precision by facet on the promotion panel against the random floor](figures/promotion_average_precision.svg)
277
-
278
- **Zero-shot frontier models.** A score-blind 300-row subsample of the promotion
279
- panel, 100 rows per family, scored under one shared prompt, output schema, and
280
- label set; each external model returned one confidence per facet. This sample
281
- is not independent of the promotion read, and each external configuration was
282
- run once, so run-to-run variance is not estimated. Subsample support: trap
283
- 213/87, decision 137/163, constraint 271/29, mechanism 194/105 with one
284
- unresolved, procedure 106/194 (positive/negative).
285
-
286
- | Model | Setup | Macro AP | Decision AP | Procedure AP | Macro P@R90 |
287
- |---|---|---:|---:|---:|---:|
288
- | This model | supervised fine-tune | 0.9671 | 0.9140 | 0.9500 | 0.8872 |
289
- | GPT-5.4 | zero-shot, high reasoning | 0.9242 | 0.7885 | 0.9177 | 0.8016 |
290
- | GPT-5.5 | zero-shot, high reasoning | 0.9162 | 0.7691 | 0.9210 | 0.8064 |
291
- | GPT-5.5 | zero-shot, low reasoning | 0.9082 | 0.7127 | 0.9209 | 0.7924 |
292
- | GPT-5.4 | zero-shot, low reasoning | 0.8734 | 0.6985 | 0.8337 | 0.6705 |
293
- | Previous classifier | prior local model | 0.8368 | 0.5992 | 0.8387 | 0.7007 |
294
- | GPT-5.3-Codex-Spark | zero-shot, low reasoning | 0.8086 | 0.5387 | 0.7684 | 0.7156 |
295
- | Nemotron 3 Ultra 550B-A55B | zero-shot, reasoning off | 0.7356 | 0.5474 | 0.7163 | 0.6144 |
296
- | Gemma 4 26B-A4B IT | zero-shot, reasoning off | 0.7103 | 0.4970 | 0.6707 | 0.6144 |
297
-
298
- Paired macro average-precision differences from this model, with family-grouped
299
- 95% intervals: GPT-5.4 high −0.0428 [−0.0523, −0.0339]; GPT-5.5 high −0.0508
300
- [−0.0548, −0.0366]; GPT-5.5 low −0.0588 [−0.0666, −0.0442]; GPT-5.4 low
301
- −0.0937 [−0.1015, −0.0827]. Every interval excludes parity. Raising reasoning
302
- effort improved GPT-5.4 by 0.0509 [0.0456, 0.0589] and GPT-5.5 by 0.0080
303
- [0.0014, 0.0204]; it narrows the gap without closing it. Three accepted runs
304
- of the Spark configuration spanned 0.8086 to 0.8496 — the reported 0.8086 plus
305
- repeats at 0.8180 and 0.8496 — a spread of 0.041 that nearly
306
- matches the smallest GPT gap and exceeds every one of the four interval widths,
307
- which is the reason single runs are flagged above.
308
-
309
- The open-weight rows are kept for completeness but carry a scoring caveat:
310
- their confidences were much coarser than the classifier's posteriors, and once
311
- tied zero scores had to be admitted at high recall their P@R90 fell to the
312
- sample's macro prevalence. Their deltas (−0.23 and −0.26) are partly an artifact
313
- of that coarseness.
314
-
315
- This comparison shows task specialization, not general capability. The
316
- classifier was trained on roughly 60,000 examples of how the label instrument
317
- applies the definitions; the external models saw the definitions once. Where a
318
- frontier model and this classifier disagree on a row, the panel cannot say which
319
- is right.
320
-
321
- ## How the recipe was chosen
322
-
323
- Selection ran on generated development panels held out by whole document
324
- family; none of them carried promotion authority. Numbers are macro average
325
- precision unless stated.
326
-
327
- 1. **Data first.** On a 1,000-row eight-family development panel (858 fully
328
- resolved rows, random floor 0.654) with the larger GLiClass backbone, training
329
- on the predecessor's data alone scored 0.929 and adding more of the operator's
330
- real documents scored 0.929; adding twelve generated document families scored
331
- 0.968. Two loss-side branches on top of that baseline were rejected: a data
332
- curriculum tied it (+0.0005, interval spanning zero) with worse calibration,
333
- and a robust-loss variant lost (−0.0018, interval entirely below zero).
334
- 2. **Targeted reserve.** Appending the independently reviewed generated reserve
335
- added +0.0025 [+0.0009, +0.0039].
336
- 3. **Context window.** Widening the input from 416 to 768 tokens added +0.0048
337
- [+0.0028, +0.0068] on development and +0.0033 [+0.0011, +0.0065] on a
338
- separate held-out set. A counterfactual-projector branch on minimal pairs
339
- changed nothing against its own replay control (+0.0001) and was dropped; the
340
- run's window also did not match its draft plan, so it could not have
341
- authorized a recipe change either way.
342
- 4. **Backbone size.** On the ModernBERT GLiClass pair, with 44,049 training
343
- rows and a matched recipe on a 5,298-row eight-family panel, the large arm beat
344
- the base by 0.0042 average precision and 0.0152 at 90% recall [+0.0015,
345
- +0.0451] — at 95% recall the interval spans zero — but cost 2.6 times
346
- the artifact bytes, 2.4 times the training time, and 1.7 times peak GPU
347
- memory, while the base scored 1.8 times as many rows per second. The base was
348
- selected as the product architecture.
349
- 5. **Backbone variant and window.** On the standard GLiClass Base v3
350
- (DeBERTa-v3) backbone, 768 tokens scored 0.928 against 0.914 at 512, with
351
- procedure moving from 0.765 to 0.821. The backbone declares a 512-position
352
- limit; the 768-token contract is a measured improvement, retained because
353
- every trial confirmed it.
354
- 6. **Final fit.** The frozen recipe was trained once on all 59,886 rows,
355
- calibrated, compressed, and read once on the promotion panel.
356
-
357
- What did not help: more real documents without generated diversity, the data
358
- curriculum, the robust loss, the counterfactual projector, and the larger
359
- backbone at its cost. A distribution-balanced loss, R-Drop, SMART smoothness, a
360
- three-state auxiliary head, and a 512-token window were also measured and
361
- rejected, and the experiment log carries each with its numbers. Two further
362
- branches closed without a scored comparison: a frozen large-NLI recipe, stopped
363
- after 23.2 hours with no completed epoch, and a decision-only rationale
364
- auxiliary, which trained four epochs and reduced all three training losses but
365
- failed artifact assembly on incomplete checkpoint-average candidates, so no
366
- panel was mounted and no per-row scores were emitted. It is retired as
367
- inconclusive. Better and more diverse examples, boundary cases, and a wider
368
- context did the work.
369
-
370
- ## Runtime
371
-
372
- The shipped artifact is the INT8 bundle. Quantization was qualified as a compression step, not
373
- a retrain: the FP32 export and the INT8 export were scored on the same 5,298-row eight-family
374
- panel and compared directly.
375
-
376
- | Item | Value |
377
- |---|---|
378
- | FP32 export | 747,319,179 bytes |
379
- | INT8 export | 452,812,018 bytes, 39.4% smaller |
380
- | Quantization | signed INT8 weights, per-channel off, reduce-range off, one operator type; 24 constant-identity nodes folded |
381
- | Toolchain | onnx 1.21.0, onnxruntime 1.24.4, opset 17 |
382
- | Macro average-precision change, INT8 minus FP32 | +0.000119, family-grouped 95% interval [-0.000038, +0.000187] |
383
- | Posterior absolute error, INT8 against FP32 | mean 0.00088, p95 0.00483, maximum 0.0516 |
384
- | Production adapter against the reference scorer | maximum posterior delta 9.06e-6 against an allowed 1e-4; zero stamp flips under both threshold tables across 1,300 rows |
385
- | Local CPU scoring rate | 3.74 passages per second at batch 16 over 5,298 rows, peak resident set 4.18 GB |
386
-
387
- The compression interval spans zero, so INT8 is not measurably worse than FP32 on that panel.
388
- That read is a diagnostic: it carries no promotion authority and set no thresholds.
389
-
390
- The scoring rate above comes from the run that produced the compression score receipt
391
- `5505e117…`, on this artifact, through the production CPU path, as an aggregate over a whole
392
- panel.
393
-
394
- Provider qualification (2026-09-06, ONNX Runtime 1.24.4, one NVIDIA RTX 3060 Ti with 8 GB)
395
- ran the same graph through the production adapter on three execution providers, with CUDA's
396
- TensorFloat-32 disabled (enabled, posteriors drifted 7.5e-4 from the CPU reference). Every
397
- matrix product stays on the accelerator; the only CPU-resident work under CUDA or DirectML
398
- is shape bookkeeping and the rank-0 scalars of the attention scale.
399
-
400
- | Provider | Posterior delta vs CPU (418 passages) | Stamp flips | Single passage p50 | Batched rate |
401
- |---|---|---|---|---|
402
- | CPU | reference | 0 | 331 ms | 1.5 passages/s |
403
- | CUDA | 2e-6 | 0 | 19 ms | 56 passages/s |
404
- | DirectML | 3e-6 | 0 | 74 ms | 33 passages/s |
405
-
406
- Batched rates use eight-row batches planned against the device's free memory: throughput is
407
- flat beyond four rows on both accelerators, and larger batches at the 768-token ceiling exceed
408
- 8 GB and page silently on Windows. Device memory per row at 768 tokens is about 208 MiB on
409
- CUDA and 277 MiB on DirectML above a resident session of 657 MiB and 481 MiB respectively.
410
-
411
- ## Evidence status
412
-
413
- | Evidence | Status |
414
- |---|---|
415
- | Promotion results, gates, and calibration | as recorded from the sealed 2026-08-12 read |
416
- | Artifact identity, size, and quantization settings | verified 2026-09-04 against the quantization receipt |
417
- | Compression equivalence | as recorded; diagnostic authority only |
418
- | Production-adapter parity | as recorded from the promotion read |
419
- | Linear baseline | re-run 2026-09-04 from the recorded inputs; byte-identical |
420
- | Figures | regenerated 2026-09-04 from the recorded inputs; byte-identical |
421
- | Previous-classifier identity | verified 2026-09-04 against the staged bundle's own config |
422
- | Frontier-model comparisons | as recorded; single runs on a subsample that is not independent of the promotion read |
423
- | Serving throughput | 3.74 passages per second, as recorded from the scoring run bound to the compression score receipt |
424
- | Per-passage latency percentiles | single-passage p50 per provider only; no p95 or p99 |
425
-
426
- ## Limitations, ranked
427
-
428
- 1. **Labels are model judgments.** All scores measure agreement with a frozen
429
- frontier-model instrument. There is no human-adjudicated answer key, by the
430
- operator's explicit decision after the hand-labeling trial described above.
431
- Resolving evidence would be a human-adjudicated sample; none is planned.
432
- 2. **Panels are generated text.** Absolute performance on unrelated real
433
- documents is unmeasured for this artifact. Resolving evidence would be a
434
- multi-domain real-document panel labeled under the same protocol.
435
- 3. **Three families in the promotion panel.** The reported intervals capture
436
- sampling noise within three similar generated families, not variation across
437
- domains. More families would widen and honest-size the intervals.
438
- 4. **No seed replicate.** Every winner was trained once. Hardware and compute
439
- budget, not design, set that limit; the smallest selection margins above
440
- (0.0042 between backbones, 0.0005 for the curriculum) are within a plausible
441
- seed effect.
442
- 5. **Lexical headroom is thin.** A linear model sits 0.029 macro average
443
- precision below this model. The advantage is real but lives at high recall
444
- on decision and procedure, and should be read that way.
445
- 6. **Frontier comparisons are single runs** on a non-independent subsample with
446
- coarse external confidences.
447
- 7. **Decision stays the hardest facet.** 398 of the 1,300 promotion rows mention
448
- a decision without making one; that boundary is where the remaining error
449
- concentrates.
450
- 8. **Serving cost is measured on one machine.** The provider table in the runtime
451
- section comes from a single 8 GB NVIDIA card and a four-core CPU budget; other
452
- hosts will land elsewhere, and cold-start cost is not separated from warm batches.
453
- 9. **Second sealed campaign in its family.** An earlier fit was read on a
454
- sealed panel, returned NO-GO on a per-origin recall gate, and the contract was
455
- then amended so origin slices are diagnostics rather than independent vetoes.
456
- That fit's calibration and promotion rows were folded into this model's own
457
- training set, and this campaign therefore grants its panel promotion authority
458
- while explicitly declining a claim to a pristine sealed alpha-spending draw.
459
- The earlier panel carried 333 operator-workspace rows; this one is entirely
460
- generated, so the per-origin recall that failed on the earlier read cannot be
461
- measured on this model at all. Resolving evidence would be a promotion panel
462
- drawn from unspent reserve that includes real rows.
463
-
464
- ## Reproducibility
465
-
466
- - Base model: `knowledgator/gliclass-base-v3.0`, revision pinned in
467
- `gliclass_final_fit_campaign.json`.
468
- - Training rows: 59,886; manifest and lane digests in
469
- `gliclass_final_fit_campaign.json` and `history/gliclass_std_base_v3_5facet_promotion_results.json`.
470
- - Promotion protocol: `gliclass_std_base_v3_5facet_promotion.json` (frozen
471
- 2026-08-12, seed 20260812, 10,000 family-cluster draws). Three digests in the
472
- results record do not reproduce from anything published or retained.
473
- `promotion_plan_sha256` describes the private sealed plan, not this tracked
474
- public-safe copy. `panel_sha256` and `predictions_sha256` were transcribed
475
- faithfully from the sealed promotion report, but the private panel and
476
- predictions files that survive today hash to different values, so the exact
477
- bytes the sealed read scored are gone. The results themselves are the record.
478
- - Calibration and thresholds: `gliclass_std_base_v3_5facet_calibration.json`,
479
- `gliclass_std_base_v3_5facet_thresholds.json`; fitted values in the promotion
480
- results record.
481
- - Artifact: `model.onnx` SHA-256
482
- `690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c`,
483
- 452,812,018 bytes; candidate manifest
484
- `ed8c1f8febe961624a698c97a3035e8ca5caa582718fd5a8f93f962c8106b942`.
485
- - Baselines: `history/linear_baseline_results.json`, reproduced byte-identical
486
- from the promoted training rows and the frozen evaluation lane with
487
- `python -m classifier.scripts.qualification.linear_baseline --train
488
- <private>/classifier/campaigns/gliclass_standard_base_v3_final_fit_v2_training/train.jsonl
489
- --panel <private>/source_allocations/classifier_final_repartition_v1/lanes/classifier-evaluation.jsonl`.
490
- The external-model protocols and results sit under `history/`. Those
491
- comparison plans pin the SHA-256 of the four scripts that produced them, and
492
- two of the four — `eval/scripts/decision_grade_runner.py` and
493
- `eval/scripts/isolated_codex.py` — have changed since. The plans therefore no
494
- longer load and are sealed records of what was run, not re-runnable commands.
495
- - Figures: all three regenerate byte-identical with
496
- `python eval/classifier/scripts/qualification/figures.py` against the promotion
497
- run's private `predictions.json`, `panel.json`, and `linear-baseline-scores.json`.
498
- - Entry point: `python -m classifier.scripts.promotion.evaluation`, which
499
- reproduces the plan's live bindings and verifies each against its recorded
500
- digest.
501
-
502
- ## License and attribution
503
-
504
- This derivative is published under the Apache License, Version 2.0, following the
505
- upstream model. It is a modified derivative of
506
- [`knowledgator/gliclass-base-v3.0`](https://huggingface.co/knowledgator/gliclass-base-v3.0)
507
- (Apache-2.0, revision `77a70e6cd52e602ed18184ef37d18bdd3741e3d5`), which builds on
508
- [`microsoft/deberta-v3-base`](https://huggingface.co/microsoft/deberta-v3-base) (MIT).
509
- Neither upstream author endorses this derivative. The changes — a supervised five-facet
510
- fine-tune, an ONNX export, INT8 quantization of the token-embedding table, and per-facet
511
- calibration — are listed in the `NOTICE` file that accompanies this model.
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: knowledgator/gliclass-base-v3.0
4
+ pipeline_tag: text-classification
5
+ language:
6
+ - en
7
+ tags:
8
+ - gliclass
9
+ - deberta-v3
10
+ - multi-label-classification
11
+ - fine-tuned
12
+ - onnx
13
+ - int8
14
+ ---
15
+
16
+ # Five-facet passage classifier
17
+
18
+ A GLiClass Base v3 fine-tune that describes what an English passage contains: a **trap, decision, constraint, mechanism or procedure**. A passage can have several facets. The model produces all five scores in one pass and runs locally through ONNX Runtime.
19
+
20
+ This revision's ONNX graph adds Vulkan compatibility. Learned weights, prompts, calibration and thresholds are unchanged.
21
+
22
+ On 1,300 generated passages from three document families excluded from training, macro average precision is **0.9676** for the fine-tune, 0.9384 for a TF-IDF baseline trained on the same rows, and 0.7709 for the zero-shot upstream checkpoint.
23
+
24
+ ## Quick start
25
+
26
+ Install `onnxruntime==1.24.4`, `tokenizers` and `numpy`. This CPU example assumes the package is in `downloaded-model` and uses its published prompts and calibration rather than new label wording:
27
+
28
+ ```python
29
+ import json
30
+ from pathlib import Path
31
+ import numpy as np
32
+ import onnxruntime as ort
33
+ from tokenizers import Tokenizer
34
+
35
+ root = Path("downloaded-model")
36
+ metadata = json.load(open(f"{root}/classifier-metadata.json", encoding="utf-8"))
37
+ facets = metadata["facet_order"]
38
+ prep = metadata["preprocessing"]
39
+ prefix = "".join(
40
+ prep["label_token"] + prep["label_prompts"][facet] for facet in facets
41
+ ) + prep["separator_token"]
42
+ tokenizer = Tokenizer.from_file(f"{root}/tokenizer.json")
43
+ tokenizer.enable_truncation(max_length=prep["max_length"])
44
+ encoded = tokenizer.encode(prefix + "We chose weekly releases to reduce rollout risk.")
45
+ session = ort.InferenceSession(
46
+ f"{root}/model.onnx", providers=["CPUExecutionProvider"]
47
+ )
48
+ logits = session.run(["logits"], {
49
+ "input_ids": np.array([encoded.ids], dtype=np.int64),
50
+ "attention_mask": np.array([encoded.attention_mask], dtype=np.int64),
51
+ })[0][0].astype(np.float64)
52
+ calibration = metadata["probability_calibration"]
53
+ clip = calibration["probability_clip"]
54
+ raw = np.clip(1 / (1 + np.exp(-np.clip(logits, -80, 80))), clip, 1 - clip)
55
+ temperatures = np.array([calibration["temperatures"][f] for f in facets])
56
+ probabilities = 1 / (1 + np.exp(-np.log(raw / (1 - raw)) / temperatures))
57
+ print(dict(zip(facets, probabilities.tolist())))
58
+ ```
59
+
60
+ The output is five independent calibrated scores, not a distribution that sums to one. For applications that need labels, the metadata supplies two threshold tables: `contract` favors precision and `recall_leaning` retains more candidates. Choose between them by the relative cost of missed and incorrect labels.
61
+
62
+ ## What the labels mean
63
+
64
+ | Facet | The passage… |
65
+ |---|---|
66
+ | `trap` | Warns about a specific mistake, hazard or failure mode |
67
+ | `decision` | Records a choice or commitment among alternatives |
68
+ | `constraint` | States a rule or condition that materially restricts action |
69
+ | `mechanism` | Explains how something is organized, connected or works |
70
+ | `procedure` | Gives sequenced actions for carrying out a task |
71
+
72
+ Mentioning a decision is different from making one; training labels treat “mentions only” as negative for that facet. Facets overlap: a procedure may also contain a constraint and warn about a trap. The scores help organize passages or inspect what a search returns. They do not establish relevance, truth or authority and should not serve as a hard retrieval filter.
73
+
74
+ ## Before and after fine-tuning
75
+
76
+ “Upstream” is the exact `knowledgator/gliclass-base-v3.0` checkpoint used to start training, applied zero-shot with the same label definitions and **no Daecore fine-tuning**. The word-feature baseline is TF-IDF plus logistic regression, trained on the same labeled rows as the fine-tune. All three use the same 1,300-passage panel; unresolved labels are excluded per facet for every model.
77
+
78
+ | Metric, macro average over five facets | Upstream GLiClass | Word-feature baseline | Daecore fine-tune |
79
+ |---|---:|---:|---:|
80
+ | Average precision | 0.7709 | 0.9384 | **0.9676** |
81
+ | Precision at 90% recall | 66.11% | 83.35% | **90.21%** |
82
+ | Precision at 95% recall | 64.51% | 78.62% | **86.02%** |
83
+
84
+ Average precision summarizes how well scores rank positives across thresholds. Precision at 90% recall is the cleanest measured prediction set that still keeps at least 90% of positive labels; it comes from this evaluation's curve, not from a threshold chosen in advance.
85
+
86
+ | Facet | Resolved passages | Positive prevalence | Upstream AP | Fine-tuned AP |
87
+ |---|---:|---:|---:|---:|
88
+ | Trap | 1,300 | 71.77% | 0.9028 | **0.9864** |
89
+ | Decision | 1,299 | 46.04% | 0.5219 | **0.9135** |
90
+ | Constraint | 1,300 | 91.31% | 0.9544 | **0.9986** |
91
+ | Mechanism | 1,297 | 68.77% | 0.7692 | **0.9817** |
92
+ | Procedure | 1,293 | 37.51% | 0.7064 | **0.9578** |
93
+
94
+ ![Upstream, word-feature and fine-tuned classifier results](figures/classifier-comparison.svg)
95
+
96
+ As the baseline shows, much of this task can be learned from word cues; the fine-tune's clearest benefit is at high recall. Constraint is common in this panel, so its near-perfect AP says less than the decision and procedure results. The panel was deliberately enriched for difficult decision and procedure cases and is not a sample of natural traffic. The [evaluation companion](evaluation/README.md) contains the text-free labels and scores, model identities and calculation code.
97
+
98
+ ## Data and training choices
99
+
100
+ The final fit used **59,886 passages from 9,376 documents**: 54,803 generated and 5,083 real. The real portion is 4,583 passages from self-owned workspaces and 500 from public-domain US Federal Register documents; all real rows were used for training, none for evaluation. Workspace documents may also be AI-written; real describes their source, not human authorship. Generated material was written as complete organizational documents and then parsed into passages, covering more settings and document types than the available workspaces.
101
+
102
+ Models applied fixed facet definitions to produce the labels. The main label campaigns used independent judgments with conflict resolution; 3,750 inherited rows used a single primary judge with a separate audit. Splits separate whole source families: the three evaluation families and four calibration families are disjoint from training and from each other. Calibration uses 1,700 passages to set per-facet temperatures and the two threshold tables.
103
+
104
+ | Training choice | Value and reason |
105
+ |---|---|
106
+ | Starting model | GLiClass Base v3 on DeBERTa-v3, 186.5M parameters |
107
+ | Input length | 768 tokens including label prompts; context-length experiments favored it over 512 |
108
+ | Objective | Weighted binary cross-entropy; mentions-only negatives receive 2× weight |
109
+ | Optimizer schedule | Learning rate 2e-5; batch 4, accumulated over 8 steps |
110
+ | Duration | Four fixed epochs; equal-weight average of epochs 2, 3 and 4, chosen before the final fit |
111
+ | Serving compression | INT8 token embeddings with the transformer body in FP32 |
112
+
113
+ Quantizing only the token embeddings shrinks the large vocabulary table while keeping transformer calculations in FP32; broader quantization reduced quality. The ONNX graph is **452.8 MB**, and the package is about **461.5 MB** excluding runtime libraries.
114
+
115
+ ## Runtime and limits
116
+
117
+ On the 1,300 evaluation passages, the derived graph produced the same labels under both threshold tables on CPU, CUDA and Vulkan as the previous graph on CPU, with maximum score differences below 5.7e-6. Vulkan uses `onnxruntime==1.24.4` with the native WebGPU plugin [`onnxruntime-ep-webgpu==0.4.0`](https://pypi.org/project/onnxruntime-ep-webgpu/), registered explicitly, with `dawnBackendType=Vulkan` and no providers list at session creation. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
118
+
119
+ - The model is trained and evaluated only on English text.
120
+ - Labels are model judgments without human validation.
121
+ - The evaluation covers generated text from three held-out families and was reused during development; performance on unrelated real documents is unmeasured.
122
+ - Training was run once; variation across seeds is unmeasured.
123
+
124
+ ## License
125
+
126
+ Apache-2.0. The package includes upstream attribution and a modification notice.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
classifier-metadata.json CHANGED
@@ -1,78 +1,78 @@
1
- {
2
- "architecture": "gliclass-single-pass",
3
- "classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
4
- "facet_order": [
5
- "trap",
6
- "decision",
7
- "constraint",
8
- "mechanism",
9
- "procedure"
10
- ],
11
- "files": {
12
- "config.json": {
13
- "sha256": "9bad60c87ac3e047bc838ca7e6897a6fe265180eeffce4e62aabc1e9c828d606",
14
- "size": 2118
15
- },
16
- "model.onnx": {
17
- "sha256": "690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c",
18
- "size": 452812018
19
- },
20
- "tokenizer.json": {
21
- "sha256": "519648948c4c59da1af88f2cf2c8b4f84417b5c673981bc9809abf84cda1b7cc",
22
- "size": 8649234
23
- },
24
- "tokenizer_config.json": {
25
- "sha256": "9d5220b355d2cb9a7df69deccca45509ce1c9a857990fb1047c48403b0f6ddfb",
26
- "size": 1692
27
- }
28
- },
29
- "preprocessing": {
30
- "label_prompts": {
31
- "constraint": "This text states a rule, policy, convention, invariant, or factual condition that materially restricts action.",
32
- "decision": "This text records a decision or commitment that chooses one option over alternatives.",
33
- "mechanism": "This text explains how a system, process, component, or structure is organized, connected, or works.",
34
- "procedure": "This text gives sequenced actions or steps for carrying out a task or operation.",
35
- "trap": "This text warns about a specific mistake, hazard, pitfall, or failure mode."
36
- },
37
- "label_token": "<<LABEL>>",
38
- "max_length": 768,
39
- "prompt_first": true,
40
- "schema": "daecore.gliclass-preprocessing.v1",
41
- "separator_token": "<<SEP>>"
42
- },
43
- "probability_calibration": {
44
- "method": "per-facet-temperature-scaling",
45
- "probability_clip": 1e-06,
46
- "temperatures": {
47
- "constraint": 1.6869790276494951,
48
- "decision": 2.4036524709979896,
49
- "mechanism": 1.501092485229896,
50
- "procedure": 2.0450110042733725,
51
- "trap": 1.8898163065802225
52
- }
53
- },
54
- "schema_version": 3,
55
- "source_candidate_sha256": "ed8c1f8febe961624a698c97a3035e8ca5caa582718fd5a8f93f962c8106b942",
56
- "threshold_tables": {
57
- "contract": {
58
- "values": {
59
- "constraint": 0.9928175,
60
- "decision": 0.48485,
61
- "mechanism": 0.8998575,
62
- "procedure": 0.5776285,
63
- "trap": 0.9759885
64
- },
65
- "version": "gliclass-std-base-v3-daecore-5facet-qint8-v2-contract"
66
- },
67
- "recall_leaning": {
68
- "values": {
69
- "constraint": 0.9059865,
70
- "decision": 0.192514,
71
- "mechanism": 0.5053545,
72
- "procedure": 0.201766,
73
- "trap": 0.715311
74
- },
75
- "version": "gliclass-std-base-v3-daecore-5facet-qint8-v2-recall"
76
- }
77
- }
78
- }
 
1
+ {
2
+ "architecture": "gliclass-single-pass",
3
+ "classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
4
+ "facet_order": [
5
+ "trap",
6
+ "decision",
7
+ "constraint",
8
+ "mechanism",
9
+ "procedure"
10
+ ],
11
+ "files": {
12
+ "config.json": {
13
+ "sha256": "9bad60c87ac3e047bc838ca7e6897a6fe265180eeffce4e62aabc1e9c828d606",
14
+ "size": 2118
15
+ },
16
+ "model.onnx": {
17
+ "sha256": "7e9b7e37afc06a23a89ba064e1255363cdc3140ff4b2823f5dae5865945621dc",
18
+ "size": 452831012
19
+ },
20
+ "tokenizer.json": {
21
+ "sha256": "519648948c4c59da1af88f2cf2c8b4f84417b5c673981bc9809abf84cda1b7cc",
22
+ "size": 8649234
23
+ },
24
+ "tokenizer_config.json": {
25
+ "sha256": "9d5220b355d2cb9a7df69deccca45509ce1c9a857990fb1047c48403b0f6ddfb",
26
+ "size": 1692
27
+ }
28
+ },
29
+ "preprocessing": {
30
+ "label_prompts": {
31
+ "constraint": "This text states a rule, policy, convention, invariant, or factual condition that materially restricts action.",
32
+ "decision": "This text records a decision or commitment that chooses one option over alternatives.",
33
+ "mechanism": "This text explains how a system, process, component, or structure is organized, connected, or works.",
34
+ "procedure": "This text gives sequenced actions or steps for carrying out a task or operation.",
35
+ "trap": "This text warns about a specific mistake, hazard, pitfall, or failure mode."
36
+ },
37
+ "label_token": "<<LABEL>>",
38
+ "max_length": 768,
39
+ "prompt_first": true,
40
+ "schema": "daecore.gliclass-preprocessing.v1",
41
+ "separator_token": "<<SEP>>"
42
+ },
43
+ "probability_calibration": {
44
+ "method": "per-facet-temperature-scaling",
45
+ "probability_clip": 1e-06,
46
+ "temperatures": {
47
+ "constraint": 1.6869790276494951,
48
+ "decision": 2.4036524709979896,
49
+ "mechanism": 1.501092485229896,
50
+ "procedure": 2.0450110042733725,
51
+ "trap": 1.8898163065802225
52
+ }
53
+ },
54
+ "schema_version": 3,
55
+ "source_candidate_sha256": "ed8c1f8febe961624a698c97a3035e8ca5caa582718fd5a8f93f962c8106b942",
56
+ "threshold_tables": {
57
+ "contract": {
58
+ "values": {
59
+ "constraint": 0.9928175,
60
+ "decision": 0.48485,
61
+ "mechanism": 0.8998575,
62
+ "procedure": 0.5776285,
63
+ "trap": 0.9759885
64
+ },
65
+ "version": "gliclass-std-base-v3-daecore-5facet-qint8-v2-contract"
66
+ },
67
+ "recall_leaning": {
68
+ "values": {
69
+ "constraint": 0.9059865,
70
+ "decision": 0.192514,
71
+ "mechanism": 0.5053545,
72
+ "procedure": 0.201766,
73
+ "trap": 0.715311
74
+ },
75
+ "version": "gliclass-std-base-v3-daecore-5facet-qint8-v2-recall"
76
+ }
77
+ }
78
+ }
evaluation/README.md ADDED
@@ -0,0 +1,83 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Model evaluation records
2
+
3
+ These files let readers recompute the comparisons in the Daecore model cards without model weights. Task-specific records contain anonymized row and group IDs, labels, and scores or ranked grades; none contains query or document text, private source identifiers or workspace paths. Model identities are recorded inside each file.
4
+
5
+ ## Recompute
6
+
7
+ Python 3.11 or later is enough; no packages or network access are needed. From the directory that holds the records:
8
+
9
+ ```sh
10
+ python metrics.py --directory .
11
+ ```
12
+
13
+ The command summarizes every record present and stops with an error if there is none. Each entry in its output equals the matching `*-summary.json`. The classifier and Ettin packages also include `figures.py`; `python figures.py --output <directory>` redraws their card chart. Gemma uses a table. Each model package contains only its own records and summaries; the source repository holds all four record types. Fractions are stored at full precision; cards round percentages to two decimals and ranking metrics to four.
14
+
15
+ The records support metric recomputation. Re-running inference would need the private text and corpus, which are not included. They come from task-specific development evaluations, not a new untouched test.
16
+
17
+ ## Records
18
+
19
+ | Record | Inputs and denominator | Comparison |
20
+ |---|---|---|
21
+ | `classifier.json` | 1,300 generated passages from three held-out source families; resolved labels only, per facet | Upstream GLiClass, a trained word-feature baseline and the fine-tune |
22
+ | `reranker.json` | 970 fixed pools of 50 passages, all 48,500 pairs judged; 873 pools have useful evidence | Upstream Ettin and the fine-tune, plus the expected result of random ordering |
23
+ | `promotion.json` | 970 hybrid-search queries over 82,719 passage texts, with reviewed labels and declared exclusions; five public dense-retrieval panels | Previous and updated Gemma with unchanged Ettin; the public panels also include upstream Gemma |
24
+ | `serving.json` | The same 970 queries with updated-Gemma candidate pools; reference extended by 32 grades | Ettin through CUDA FP16 and Vulkan FP32, with identical weights and score mapping |
25
+
26
+ `serving-qualification.json` summarizes provider and recovery checks by receipt hash and keeps the aggregate FiQA and SciFact results for upstream and fine-tuned Ettin.
27
+
28
+ “Upstream” means no Daecore fine-tuning, not an untrained network; upstream Ettin is already a trained reranker. Interim training checkpoints are not included.
29
+
30
+ ## How each comparison was run
31
+
32
+ **Classifier.** Evaluation families were excluded from training. Upstream uses the same five label definitions and 768-token input limit as the fine-tune. Its raw logits and the fine-tune's calibrated probabilities are used only to rank within each facet; their scales are not compared. All 1,300 saved predictions matched fresh CPU inference of the released graph within 5.62e-6. The baseline uses word unigram and bigram TF-IDF with one balanced logistic-regression model per facet, fitted on the same 59,886 training passages; it sees full passage text, while the neural models apply their 768-token limit.
33
+
34
+ **Ettin, fixed pools.** Candidates and grades are held fixed. Both models run in PyTorch FP16 with the same 1,153-token pair construction, Transformers 5.2.0 and Sentence Transformers 5.5.1; upstream's metadata names a newer library version, but both use the same reference runtime here. The fine-tune's saved scores were checked against fresh inference on three complete pools. This evaluates the trained models, not ONNX serving or latency. The public FiQA and SciFact results rerank fixed 50-candidate pools from upstream Gemma.
35
+
36
+ **Gemma update.** Only Gemma changes; BM25, fusion and Ettin's weights stay fixed. All 970 queries are retained, and seven reviewed grade corrections apply to both models. Within each original top-k or selected prefix, 65 unresolved query–passage abstentions are excluded without backfilling: `ranked_grades` keeps their positions as `null`, with the required `excluded` flag. Precision is pooled over retained positions; nDCG reindexes them and takes its ideal ordering at the retained depth from `reference_grade_counts`. Selected depths refer to the original score-selected prefixes.
37
+
38
+ | Metric, 970 queries | Previous Gemma | Updated Gemma |
39
+ |---|---:|---:|
40
+ | Hit@3 | 83.92% | 85.36% |
41
+ | Hit@5 | 85.88% | 88.45% |
42
+ | Hit@10 | 87.84% | 89.90% |
43
+ | Hit@20 | 90.10% | 91.03% |
44
+ | nDCG@10 | 0.6764 | 0.6611 |
45
+ | Selected-prefix precision | 80.41% | 80.34% |
46
+
47
+ The same record holds per-query dense nDCG@10 for the five public panels. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); missing baselines stay absent. These sets supplied no training examples but informed development, and recomputing their means is different from rerunning retrieval on the public corpora.
48
+
49
+ **Ettin providers.** Updated-Gemma candidate pools for the same 970 queries were scored through CUDA FP16 and Vulkan FP32 with unchanged weights and score mapping. The two paths' top-20 results included 32 query–passage pairs without a grade. Of these, 21 reuse grades from the hard-contrast Ettin evaluation. The other 11 received two independent GPT-6 Sol judgments plus a resolution step and were then reviewed against the full passage text by GPT-6 Astra; no person reviewed them. No existing grade changed, and the 65 abstentions remain excluded. The added grades slightly change nDCG's ideal ordering, so the promotion record keeps its original reference.
50
+
51
+ | Metric, 970 queries | CUDA FP16 | Vulkan FP32 |
52
+ |---|---:|---:|
53
+ | Hit@3 | 85.36% | 85.26% |
54
+ | Hit@5 | 88.45% | 88.45% |
55
+ | Hit@10 | 89.90% | 89.90% |
56
+ | Hit@20 | 91.03% | 91.03% |
57
+ | nDCG@10 | 0.6611 | 0.6615 |
58
+ | Selected-prefix precision | 80.34% | 80.28% |
59
+
60
+ Small numerical differences between the paths can reorder close scores. On 20 matched pools replayed twice, second-pass reranking took 1.56 s median and 2.11 s at the 95th percentile with Vulkan, against 1.92 s and 2.70 s with the previous package's DirectML graph. All provider measurements come from one Windows x64 machine with an NVIDIA RTX 3060 Ti (8 GB) and do not transfer to other GPUs or platforms.
61
+
62
+ ## Metric definitions
63
+
64
+ Relevance grades are 0–3, and grades **2 and 3** count as useful. An “answerable” query has at least one useful labeled passage in the specified pool; the absence of a useful judgment does not prove that no answer exists.
65
+
66
+ - **Hit@k:** fraction of queries with at least one useful passage in the first k positions.
67
+ - **Precision@k:** useful passages among the first k. Fixed-pool tables average it over queries; the promotion and serving records pool it over retained positions. Selected-prefix precision applies the same pooling to all returned passages.
68
+ - **Recall@k:** useful passages retrieved divided by the query's known useful passages, averaged over queries. It counts labeled passages, not every fact an answer needs.
69
+ - **nDCG@k:** gain `2**grade - 1`, discount `1/log2(rank + 1)`, divided by the ideal ordering at the same cutoff. Grade 1 adds a small gain although it is not useful for Hit, precision or recall. Each cutoff has its own ideal, so values need not change monotonically with k.
70
+ - **Average precision (AP):** area under the stepwise precision–recall curve, with tied scores grouped at one threshold; macro AP weights the five facets equally.
71
+ - **Precision at 90% or 95% recall:** the best measured precision at any threshold reaching that recall, taken from the evaluation curve rather than a threshold chosen in advance.
72
+
73
+ Ties keep the original candidate order. Classifier fields marked `null` are excluded identically for every model. Conditional reranker nDCG and recall use the 873 answerable pools; all-query Hit and precision use all 970. The random-order reference is exact within each pool: with N candidates and R useful passages, expected precision is R/N and Hit@k is `1 - C(N-R, k)/C(N, k)`. Random ordering already reaches 97.30% Hit@20 on the answerable reranker pools.
74
+
75
+ ## Limits
76
+
77
+ Task-specific labels are language-model judgments without a human-adjudicated reference. Much of the source material is generated. Project-document sources are narrow and often AI-written; they do not establish generalization across users. The panels were reused during development, and each released model was trained once. Several relevant passages can repeat one fact, so passage-level scores do not measure unique-fact coverage, answer completeness or downstream agent success.
78
+
79
+ ## Cards
80
+
81
+ - [Five-facet classifier](https://huggingface.co/Daecore/gliclass-std-base-v3-5facet-qint8-v2)
82
+ - [EmbeddingGemma retriever](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v1)
83
+ - [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1)
evaluation/classifier-summary.json ADDED
@@ -0,0 +1,138 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "rows": 1300,
3
+ "families": 3,
4
+ "models": {
5
+ "upstream": {
6
+ "facets": {
7
+ "trap": {
8
+ "rows": 1300,
9
+ "average_precision": 0.9027959606064969,
10
+ "precision_at_recall_90": 0.7787037037037037,
11
+ "precision_at_recall_95": 0.7485232067510549,
12
+ "prevalence": 0.7176923076923077
13
+ },
14
+ "decision": {
15
+ "rows": 1299,
16
+ "average_precision": 0.5219192438473815,
17
+ "precision_at_recall_90": 0.46072874493927124,
18
+ "precision_at_recall_95": 0.46072874493927124,
19
+ "prevalence": 0.4603541185527329
20
+ },
21
+ "constraint": {
22
+ "rows": 1300,
23
+ "average_precision": 0.954374886008778,
24
+ "precision_at_recall_90": 0.9130769230769231,
25
+ "precision_at_recall_95": 0.9130769230769231,
26
+ "prevalence": 0.9130769230769231
27
+ },
28
+ "mechanism": {
29
+ "rows": 1297,
30
+ "average_precision": 0.7691508918347779,
31
+ "precision_at_recall_90": 0.6880804953560371,
32
+ "precision_at_recall_95": 0.6880804953560371,
33
+ "prevalence": 0.6877409406322282
34
+ },
35
+ "procedure": {
36
+ "rows": 1293,
37
+ "average_precision": 0.7063681153993947,
38
+ "precision_at_recall_90": 0.4648936170212766,
39
+ "precision_at_recall_95": 0.4153153153153153,
40
+ "prevalence": 0.3750966744006187
41
+ }
42
+ },
43
+ "macro": {
44
+ "average_precision": 0.7709218195393658,
45
+ "precision_at_recall_90": 0.6610966968194424,
46
+ "precision_at_recall_95": 0.6451449370877204
47
+ }
48
+ },
49
+ "linear": {
50
+ "facets": {
51
+ "trap": {
52
+ "rows": 1300,
53
+ "average_precision": 0.9704508234286988,
54
+ "precision_at_recall_90": 0.9231613611416026,
55
+ "precision_at_recall_95": 0.895959595959596,
56
+ "prevalence": 0.7176923076923077
57
+ },
58
+ "decision": {
59
+ "rows": 1299,
60
+ "average_precision": 0.8471593175324644,
61
+ "precision_at_recall_90": 0.6304093567251462,
62
+ "precision_at_recall_95": 0.571,
63
+ "prevalence": 0.4603541185527329
64
+ },
65
+ "constraint": {
66
+ "rows": 1300,
67
+ "average_precision": 0.9969713734939275,
68
+ "precision_at_recall_90": 0.9916743755781684,
69
+ "precision_at_recall_95": 0.9791666666666666,
70
+ "prevalence": 0.9130769230769231
71
+ },
72
+ "mechanism": {
73
+ "rows": 1297,
74
+ "average_precision": 0.9633161957308005,
75
+ "precision_at_recall_90": 0.8865638766519823,
76
+ "precision_at_recall_95": 0.8330058939096268,
77
+ "prevalence": 0.6877409406322282
78
+ },
79
+ "procedure": {
80
+ "rows": 1293,
81
+ "average_precision": 0.914259121155341,
82
+ "precision_at_recall_90": 0.7356902356902357,
83
+ "precision_at_recall_95": 0.652050919377652,
84
+ "prevalence": 0.3750966744006187
85
+ }
86
+ },
87
+ "macro": {
88
+ "average_precision": 0.9384313662682464,
89
+ "precision_at_recall_90": 0.833499841157427,
90
+ "precision_at_recall_95": 0.7862366151827083
91
+ }
92
+ },
93
+ "finetuned": {
94
+ "facets": {
95
+ "trap": {
96
+ "rows": 1300,
97
+ "average_precision": 0.9863618253424686,
98
+ "precision_at_recall_90": 0.96,
99
+ "precision_at_recall_95": 0.934668071654373,
100
+ "prevalence": 0.7176923076923077
101
+ },
102
+ "decision": {
103
+ "rows": 1299,
104
+ "average_precision": 0.9135015910515543,
105
+ "precision_at_recall_90": 0.7414030261348006,
106
+ "precision_at_recall_95": 0.6581986143187067,
107
+ "prevalence": 0.4603541185527329
108
+ },
109
+ "constraint": {
110
+ "rows": 1300,
111
+ "average_precision": 0.9985879374608636,
112
+ "precision_at_recall_90": 0.9962825278810409,
113
+ "precision_at_recall_95": 0.9947183098591549,
114
+ "prevalence": 0.9130769230769231
115
+ },
116
+ "mechanism": {
117
+ "rows": 1297,
118
+ "average_precision": 0.9816644729408911,
119
+ "precision_at_recall_90": 0.9261822376009228,
120
+ "precision_at_recall_95": 0.8908709338929696,
121
+ "prevalence": 0.6877409406322282
122
+ },
123
+ "procedure": {
124
+ "rows": 1293,
125
+ "average_precision": 0.9578217870935172,
126
+ "precision_at_recall_90": 0.8868686868686869,
127
+ "precision_at_recall_95": 0.822380106571936,
128
+ "prevalence": 0.3750966744006187
129
+ }
130
+ },
131
+ "macro": {
132
+ "average_precision": 0.967587522777859,
133
+ "precision_at_recall_90": 0.9021472956970902,
134
+ "precision_at_recall_95": 0.8601672072594281
135
+ }
136
+ }
137
+ }
138
+ }
evaluation/classifier.json ADDED
The diff for this file is too large to render. See raw diff
 
evaluation/figures.py ADDED
@@ -0,0 +1,89 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Regenerate the classifier and reranker SVGs from the public labels and scores.
2
+
3
+ Run with Python 3.11+: python figures.py --output ../figures
4
+ The repository shares its renderer from eval/lib; HF packages include that
5
+ same renderer next to this script.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import argparse
11
+ import importlib.util
12
+ import json
13
+ import sys
14
+ from pathlib import Path
15
+
16
+ import metrics
17
+
18
+ HERE = Path(__file__).resolve().parent
19
+ renderer = HERE / 'svg_figures.py'
20
+ if not renderer.is_file():
21
+ renderer = HERE.parents[1] / 'lib/svg_figures.py'
22
+ spec = importlib.util.spec_from_file_location('model_card_svg', renderer)
23
+ svg = importlib.util.module_from_spec(spec)
24
+ sys.modules[spec.name] = svg
25
+ spec.loader.exec_module(svg)
26
+
27
+ UPSTREAM, FIT, REFERENCE = '#0072B2', '#D55E00', '#777777'
28
+
29
+
30
+ def classifier(data: dict) -> str:
31
+ summary = metrics.summarize_classifier(data)
32
+ facets = [*metrics.FACETS, 'macro']
33
+ def values(model):
34
+ result = summary['models'][model]
35
+ return [result['facets'][f]['average_precision'] for f in metrics.FACETS] + [result['macro']['average_precision']]
36
+ return svg.render_bars(
37
+ facets, [('Upstream', values('upstream'), UPSTREAM),
38
+ ('TF-IDF + logistic regression', values('linear'), REFERENCE),
39
+ ('Daecore fine-tune', values('finetuned'), FIT)],
40
+ title='Classifier: matched before and after fine-tuning',
41
+ subtitle='1,300 generated passages · 3 held-out families · resolved labels only',
42
+ ylabel='average precision', ylim=(0, 1.05), separator_before=5,
43
+ width=840, height=360,
44
+ notes=['Model-generated labels; these results do not establish accuracy on unrelated real documents.'],
45
+ )
46
+
47
+
48
+ def reranker(data: dict) -> str:
49
+ summary = metrics.summarize_reranker(data)
50
+ definitions = [('hit', 'At least one useful passage'), ('precision', 'Useful passages / returned passages'), ('ndcg', 'Graded ranking quality')]
51
+ panels = []
52
+ legend = [('Upstream', UPSTREAM, None), ('Daecore fine-tune', FIT, None), ('Random pool order', REFERENCE, '5 3')]
53
+ for field, title in definitions:
54
+ series = []
55
+ for model, (label, color, dash) in zip(('upstream', 'finetuned', 'random'), legend, strict=True):
56
+ values = summary['models'][model]['answerable']
57
+ series.append(svg.Series(label, list(range(5)), [values[str(k)][field] for k in metrics.CUTOFFS], color, dash=dash))
58
+ panels.append(svg.Panel(title, series, xlabel='rank cutoff',
59
+ ylabel={'hit': 'Hit@k', 'precision': 'Precision@k', 'ndcg': 'nDCG@k'}[field],
60
+ xlim=(0, 4), ylim=(0, 1.02), xticks=list(enumerate(map(str, metrics.CUTOFFS)))))
61
+ return svg.render_grid(panels, columns=3, title='Ettin: identical candidates, different ordering',
62
+ subtitle='873 answerable pools of 50 · 970 queries in total · useful = grade 2 or 3',
63
+ legend=legend, panel_width=300, panel_height=270)
64
+
65
+
66
+ def main() -> None:
67
+ parser = argparse.ArgumentParser(description=__doc__)
68
+ parser.add_argument('--directory', type=Path, default=HERE)
69
+ parser.add_argument('--output', type=Path, required=True)
70
+ parser.add_argument('--model', choices=('classifier', 'reranker'))
71
+ args = parser.parse_args()
72
+ args.output.mkdir(parents=True, exist_ok=True)
73
+ definitions = {'classifier': ('classifier', 'classifier-comparison.svg'),
74
+ 'reranker': ('reranker', 'ettin-comparison.svg')}
75
+ names = [args.model] if args.model else [
76
+ name for name, (record, _) in definitions.items()
77
+ if (args.directory / f'{record}.json').is_file()
78
+ ]
79
+ if not names:
80
+ parser.error('No model comparison records found in the selected directory')
81
+ for name in names:
82
+ record, filename = definitions[name]
83
+ data = json.loads((args.directory / f'{record}.json').read_text(encoding='utf-8'))
84
+ (args.output / filename).write_text(globals()[name](data), encoding='utf-8', newline='\n')
85
+
86
+
87
+
88
+ if __name__ == '__main__':
89
+ main()
evaluation/metrics.py ADDED
@@ -0,0 +1,219 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Recompute the published model comparisons from text-free evaluation records.
2
+
3
+ Python 3.11+, standard library only. Run: python metrics.py --directory .
4
+ """
5
+
6
+ from __future__ import annotations
7
+
8
+ import argparse
9
+ import json
10
+ import math
11
+ from pathlib import Path
12
+ from statistics import mean
13
+
14
+ CUTOFFS = (1, 3, 5, 10, 20)
15
+ FACETS = ("trap", "decision", "constraint", "mechanism", "procedure")
16
+
17
+
18
+ def classification_metrics(labels: list[int], scores: list[float]) -> dict[str, float]:
19
+ """Threshold-grouped AP and best precision at or above each recall target."""
20
+ if len(labels) != len(scores) or not labels or set(labels) - {0, 1}:
21
+ raise ValueError("Classification labels and scores must be aligned binary rows")
22
+ if not all(math.isfinite(s) for s in scores) or sum(labels) == 0:
23
+ raise ValueError("Classification scores must be finite with positive support")
24
+ order = sorted(range(len(labels)), key=lambda i: -scores[i])
25
+ positives = sum(labels)
26
+ true_positives = 0
27
+ previous_recall = 0.0
28
+ ap = 0.0
29
+ points = []
30
+ for rank, index in enumerate(order, 1):
31
+ true_positives += labels[index]
32
+ if rank < len(order) and scores[order[rank]] == scores[index]:
33
+ continue
34
+ precision = true_positives / rank
35
+ recall = true_positives / positives
36
+ ap += (recall - previous_recall) * precision
37
+ previous_recall = recall
38
+ points.append((precision, recall))
39
+ return {
40
+ "average_precision": ap,
41
+ "precision_at_recall_90": max(p for p, r in points if r >= 0.9),
42
+ "precision_at_recall_95": max(p for p, r in points if r >= 0.95),
43
+ "prevalence": positives / len(labels),
44
+ }
45
+
46
+
47
+ def dcg(grades: list[int], k: int) -> float:
48
+ return sum((2**g - 1) / math.log2(i + 2) for i, g in enumerate(grades[:k]))
49
+
50
+
51
+ def ranking_metrics(grades: list[int], scores: list[float], k: int) -> dict[str, float]:
52
+ if len(grades) != len(scores) or len(grades) < k or k < 1:
53
+ raise ValueError("Ranking inputs must be aligned and cover the cutoff")
54
+ if any(type(g) is not int or g not in range(4) for g in grades):
55
+ raise ValueError("Ranking requires fully judged integer grades 0 through 3")
56
+ if not all(math.isfinite(s) for s in scores):
57
+ raise ValueError("Ranking scores must be finite")
58
+ order = sorted(range(len(scores)), key=lambda i: (-scores[i], i))
59
+ ranked = [grades[i] for i in order]
60
+ useful = sum(g >= 2 for g in ranked[:k])
61
+ relevant = sum(g >= 2 for g in grades)
62
+ ideal = dcg(sorted(grades, reverse=True), k)
63
+ result = {"hit": float(useful > 0), "precision": useful / k, "useful": float(useful)}
64
+ if relevant:
65
+ result["recall"] = useful / relevant
66
+ result["ndcg"] = dcg(ranked, k) / ideal
67
+ return result
68
+
69
+
70
+ def random_ranking_metrics(grades: list[int], k: int) -> dict[str, float]:
71
+ """Exact expectation under a uniform permutation of one fixed judged pool."""
72
+ n = len(grades)
73
+ if k < 1 or k > n or any(type(g) is not int or g not in range(4) for g in grades):
74
+ raise ValueError("Random reference requires judged grades and a valid cutoff")
75
+ relevant = sum(g >= 2 for g in grades)
76
+ misses = math.comb(n - relevant, k) if n - relevant >= k else 0
77
+ result = {"hit": 1 - misses / math.comb(n, k), "precision": relevant / n, "useful": k * relevant / n}
78
+ if relevant:
79
+ expected_dcg = mean(2**g - 1 for g in grades) * sum(1 / math.log2(i + 2) for i in range(k))
80
+ result["recall"] = k / n
81
+ result["ndcg"] = expected_dcg / dcg(sorted(grades, reverse=True), k)
82
+ return result
83
+
84
+
85
+ def summarize_classifier(data: dict) -> dict:
86
+ rows = data["rows"]
87
+ result = {"rows": len(rows), "families": len({r["group"] for r in rows}), "models": {}}
88
+ for model in data["model_order"]:
89
+ by_facet = {}
90
+ for facet in FACETS:
91
+ resolved = [r for r in rows if r["labels"][facet] is not None]
92
+ if any(r["labels"][facet] not in (0, 1) for r in resolved):
93
+ raise ValueError("Invalid resolved classifier label")
94
+ by_facet[facet] = {"rows": len(resolved), **classification_metrics(
95
+ [r["labels"][facet] for r in resolved], [r["scores"][model][facet] for r in resolved]
96
+ )}
97
+ macro = {k: mean(v[k] for v in by_facet.values()) for k in ("average_precision", "precision_at_recall_90", "precision_at_recall_95")}
98
+ result["models"][model] = {"facets": by_facet, "macro": macro}
99
+ return result
100
+
101
+
102
+ def summarize_reranker(data: dict) -> dict:
103
+ rows = data["rows"]
104
+ answerable = [r for r in rows if any(g >= 2 for g in r["grades"])]
105
+ output = {"queries": len(rows), "answerable": len(answerable), "models": {}}
106
+ for model in ["random", *data["model_order"]]:
107
+ scopes = {}
108
+ for scope, subset in [("answerable", answerable), ("all", rows)]:
109
+ cuts = {}
110
+ for k in CUTOFFS:
111
+ values = [random_ranking_metrics(r["grades"], k) if model == "random" else ranking_metrics(r["grades"], r["scores"][model], k) for r in subset]
112
+ fields = ("hit", "precision", "useful", "recall", "ndcg") if scope == "answerable" else ("hit", "precision", "useful")
113
+ cuts[str(k)] = {field: mean(v[field] for v in values) for field in fields}
114
+ scopes[scope] = cuts
115
+ output["models"][model] = scopes
116
+ return output
117
+
118
+
119
+ def summarize_promotion(data: dict) -> dict:
120
+ """Hybrid search with reviewed exclusions inside each original prefix.
121
+
122
+ A null is permitted only for an explicitly excluded judging abstention.
123
+ Cut first, remove exclusions second, and never backfill from a deeper rank.
124
+ """
125
+ rows = data['rows']
126
+ if not rows or len({row['id'] for row in rows}) != len(rows):
127
+ raise ValueError('Promotion rows require unique nonempty query identities')
128
+ output = {'queries': len(rows), 'models': {}, 'public': {}}
129
+ for model in data['model_order']:
130
+ cutoffs = {}
131
+ for cutoff in (3, 5, 10, 20, 'selected'):
132
+ per_query = []
133
+ for row in rows:
134
+ counts = row['reference_grade_counts']
135
+ if set(counts) != {'0', '1', '2', '3'} or any(type(n) is not int or n < 0 for n in counts.values()):
136
+ raise ValueError('Reference grade counts must cover grades zero through three')
137
+ ranked = row['ranked_grades'][model]
138
+ excluded = row['excluded'][model]
139
+ depth = row['selected_depth'][model] if cutoff == 'selected' else cutoff
140
+ if (len(ranked) != 20 or len(excluded) != 20 or type(depth) is not int
141
+ or not 3 <= depth <= 20 or any(type(x) is not bool for x in excluded)):
142
+ raise ValueError('Promotion rows require a bounded original top twenty')
143
+ if any((grade is not None if drop else type(grade) is not int or grade not in range(4))
144
+ for grade, drop in zip(ranked, excluded, strict=True)):
145
+ raise ValueError('Only declared abstentions may lack grades')
146
+ kept = [grade for grade, drop in zip(ranked[:depth], excluded[:depth], strict=True) if not drop]
147
+ if not kept:
148
+ raise ValueError('Every scored prefix must retain judged passages')
149
+ ideal_grades = [g for g in (3, 2, 1, 0) for _ in range(min(counts[str(g)], len(kept)))][:len(kept)]
150
+ useful = sum(g >= 2 for g in kept)
151
+ positives = counts['2'] + counts['3']
152
+ ideal = dcg(ideal_grades, len(kept))
153
+ per_query.append({
154
+ 'hit': float(useful > 0), 'precision': useful / len(kept),
155
+ 'useful': useful, 'retained': len(kept), 'excluded': depth - len(kept),
156
+ 'ndcg': dcg(kept, len(kept)) / ideal if ideal else 0.0,
157
+ 'known_recall': useful / positives if positives else None,
158
+ })
159
+ useful = sum(row['useful'] for row in per_query)
160
+ retained = sum(row['retained'] for row in per_query)
161
+ recalls = [row['known_recall'] for row in per_query if row['known_recall'] is not None]
162
+ cutoffs[str(cutoff)] = {
163
+ 'hit': mean(row['hit'] for row in per_query), 'precision': useful / retained,
164
+ 'macro_precision': mean(row['precision'] for row in per_query),
165
+ 'ndcg': mean(row['ndcg'] for row in per_query),
166
+ 'known_positive_recall': mean(recalls) if recalls else None,
167
+ 'recall_queries': len(recalls), 'useful': useful, 'retained': retained,
168
+ 'excluded_positions': sum(row['excluded'] for row in per_query),
169
+ 'mean_useful': useful / len(rows), 'mean_retained': retained / len(rows),
170
+ }
171
+ output['models'][model] = cutoffs
172
+ for dataset, panel in data['public'].items():
173
+ if not panel['rows'] or len({row['id'] for row in panel['rows']}) != len(panel['rows']):
174
+ raise ValueError('Public panel requires distinct query identities')
175
+ measured = panel['model_order']
176
+ for row in panel['rows']:
177
+ if set(row['ndcg@10']) != set(measured) or any(
178
+ isinstance(value, bool) or not isinstance(value, int | float)
179
+ or not math.isfinite(value) or not 0 <= value <= 1
180
+ for value in row['ndcg@10'].values()
181
+ ):
182
+ raise ValueError('Public nDCG values must be finite measured scores')
183
+ output['public'][dataset] = {
184
+ 'queries': len(panel['rows']),
185
+ 'ndcg@10': {model: mean(row['ndcg@10'][model] for row in panel['rows']) for model in measured},
186
+ }
187
+ return output
188
+
189
+
190
+ def main() -> None:
191
+ parser = argparse.ArgumentParser(description=__doc__)
192
+ parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
193
+ parser.add_argument("--output", type=Path)
194
+ parser.add_argument("--model", choices=("classifier", "reranker", "promotion", "serving"), help="Recompute one comparison; default: every record in this package")
195
+ args = parser.parse_args()
196
+ summarizers = {"classifier": summarize_classifier, "reranker": summarize_reranker,
197
+ "promotion": summarize_promotion, "serving": summarize_promotion}
198
+ names = [args.model] if args.model else [
199
+ name for name in summarizers if (args.directory / f"{name}.json").is_file()
200
+ ]
201
+ if not names:
202
+ parser.error("No evaluation records found in the selected directory")
203
+ result = {}
204
+ for name in names:
205
+ summarize = summarizers[name]
206
+ data = json.loads((args.directory / f"{name}.json").read_text(encoding="utf-8"))
207
+ ids = [r["id"] for r in data["rows"]]
208
+ if len(set(ids)) != len(ids):
209
+ raise ValueError(f"Repeated query/passage identity in {name}")
210
+ result[name] = summarize(data)
211
+ text = json.dumps(result, indent=2, allow_nan=False) + "\n"
212
+ if args.output:
213
+ args.output.write_text(text, encoding="utf-8")
214
+ else:
215
+ print(text, end="")
216
+
217
+
218
+ if __name__ == "__main__":
219
+ main()
evaluation/serving-qualification.json ADDED
@@ -0,0 +1,134 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "schema": "daecore.serving-preparation-summary.v1",
3
+ "date": "2026-09-28",
4
+ "status": "local-model-serving-checks-complete-release-transition-pending",
5
+ "hardware": "Windows x64, NVIDIA RTX 3060 Ti 8 GiB",
6
+ "unmeasured_targets": [
7
+ "AMD",
8
+ "Intel",
9
+ "Linux x64"
10
+ ],
11
+ "runtimes": {
12
+ "onnxruntime": "1.24.4",
13
+ "vulkan_plugin": "0.4.0",
14
+ "backend": "Vulkan",
15
+ "storage_buffer_cache": "lazyRelease"
16
+ },
17
+ "models": {
18
+ "gemma": {
19
+ "vectors": 74,
20
+ "providers": {
21
+ "cpu": {
22
+ "receipt_sha256": "31851791104fa95d6b7b25d35f950b25e239a3d6f8fa458addb2a44917c25cb4",
23
+ "maximum_vector_delta": 3.0174851417541504e-07,
24
+ "concurrent_delta": 0.0
25
+ },
26
+ "cuda": {
27
+ "receipt_sha256": "6977d0aa0ebf9c39ea0fd24078c47367d1d50e81184653225fcd10758961903e",
28
+ "maximum_vector_delta": 2.644956111907959e-07,
29
+ "concurrent_delta": 2.2351741790771484e-08
30
+ },
31
+ "vulkan": {
32
+ "receipt_sha256": "567fb89ea195e8b188203d2334f3eda82994ae8c7fc97f2126fac3e6f4766079",
33
+ "maximum_vector_delta": 5.438923835754395e-07,
34
+ "concurrent_delta": 0.0
35
+ }
36
+ },
37
+ "admission": {
38
+ "cases": 3841,
39
+ "semantic_fallbacks": 679,
40
+ "new_useful_exclusions": 0,
41
+ "additional_noise_retained": 2,
42
+ "receipt_sha256": "71550f606c95c1e1192dd5939488b280d5f2f2f091f99413f6b63f4408a5bc2c"
43
+ }
44
+ },
45
+ "classifier": {
46
+ "cpu": {
47
+ "rows": 1300,
48
+ "maximum_posterior_delta": 5.612167303103988e-06,
49
+ "changed_threshold_labels": {
50
+ "contract": 0,
51
+ "recall_leaning": 0
52
+ },
53
+ "receipt_sha256": "ea545fb0a30ff25bf3da178b60593e934e6025e789c2db67290ff5360e8669a4"
54
+ },
55
+ "cuda": {
56
+ "rows": 1300,
57
+ "maximum_posterior_delta": 5.612167303103988e-06,
58
+ "changed_threshold_labels": {
59
+ "contract": 0,
60
+ "recall_leaning": 0
61
+ },
62
+ "receipt_sha256": "5616891e9dfdc3b6615951901cb24aa31189d6a81ed84b0080f461b85c045065"
63
+ },
64
+ "vulkan": {
65
+ "rows": 1300,
66
+ "maximum_posterior_delta": 5.612167303103988e-06,
67
+ "changed_threshold_labels": {
68
+ "contract": 0,
69
+ "recall_leaning": 0
70
+ },
71
+ "receipt_sha256": "c7ec971e020d9771adf0756316f468629cc05e300d7edf9abb52612901adf6b0"
72
+ }
73
+ },
74
+ "ettin": {
75
+ "queries": 970,
76
+ "unjudged_top20": 0,
77
+ "quality_evidence": "serving.json",
78
+ "quality_receipt_sha256": "805a1b0acd7baaf248e182908bc23cd52171e40aa8f1281799a20d15b58c6905",
79
+ "full_replay_seconds_p50_p95": [
80
+ 1.7189999999827705,
81
+ 2.610000000044238
82
+ ],
83
+ "matched_20_pool_seconds_p50_p95": {
84
+ "vulkan": [
85
+ 1.5565476999909151,
86
+ 2.1063475799834124
87
+ ],
88
+ "predecessor_directml": [
89
+ 1.9197023500164505,
90
+ 2.701977299965802
91
+ ]
92
+ },
93
+ "pressure": "One preflight refusal after 322 queries with lazy release; resumed all remaining rows with zero further retries. An earlier default-cache run stopped after 424. No native OOM observed.",
94
+ "precision": "FP32 Vulkan versus FP16 CUDA; rankings not bit-exact",
95
+ "selector": "Unchanged CUDA mapping transferred for measurement; new package/provider binding pending"
96
+ }
97
+ },
98
+ "operator_runtime_changed": false,
99
+ "published": false,
100
+ "remaining": [
101
+ "immutable-publication-revisions",
102
+ "profile-and-calibration-release-bindings",
103
+ "isolated-package-update-and-index-rebuild-recovery",
104
+ "DirectML-current-route-retirement"
105
+ ],
106
+ "worker_recovery": {
107
+ "receipt_sha256": "46205d39521bde40d1228c837e57f88e1a941d42f6846f4a244c48b513347cc6",
108
+ "all_three_consumer_outputs_identical_after_owned_worker_crash": true,
109
+ "ettin_maximum_envelope": {
110
+ "tokens": 1153,
111
+ "max_abs_logit_delta_vs_cpu": 5.91278076171875e-05
112
+ }
113
+ },
114
+ "retained_reranker_public": {
115
+ "scope": "Retained matched public pools retrieved by upstream Gemma; upstream and production Ettin PyTorch FP16; not a new Vulkan public replay",
116
+ "source_sha256": "0814db560081af8376e4b1f3e3d2bcd5c4d9b8b208f7fb8b7172f113e8123ef3",
117
+ "panels": {
118
+ "fiqa": {
119
+ "queries": 648,
120
+ "ndcg@10": {
121
+ "upstream": 0.48603492061219766,
122
+ "finetuned": 0.452680922415071
123
+ }
124
+ },
125
+ "scifact": {
126
+ "queries": 300,
127
+ "ndcg@10": {
128
+ "upstream": 0.7487436478294831,
129
+ "finetuned": 0.7543833016238701
130
+ }
131
+ }
132
+ }
133
+ }
134
+ }
evaluation/svg_figures.py ADDED
@@ -0,0 +1,576 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Dependency-free SVG charts for tracked evaluation figures.
2
+
3
+ The evaluation figures are generated from private per-row evidence into tracked SVG files.
4
+ Keeping the renderer inside the repository, with no plotting dependency, means a figure can be
5
+ regenerated by any contributor with the private inputs and compared byte for byte.
6
+
7
+ Everything is emitted as presentation attributes rather than CSS. Markdown hosts sanitize
8
+ embedded stylesheets and scripts out of SVG, so a figure that carries its styling in attributes
9
+ renders the same in the repository, on a model-card host, and in a local viewer.
10
+
11
+ Layout is measured rather than assumed: legend entries, panel titles, and reference-line labels
12
+ are placed from an estimated text width, so a longer label reflows instead of overlapping its
13
+ neighbour. ``text_width`` approximates a sans-serif advance table, which is enough to keep
14
+ elements apart but is not a substitute for a real font metric.
15
+ """
16
+
17
+ from __future__ import annotations
18
+
19
+ import math
20
+ from dataclasses import dataclass, field
21
+ from xml.sax.saxutils import escape
22
+
23
+ # Okabe-Ito, chosen because it stays distinguishable under the common colour-vision
24
+ # deficiencies and prints legibly in greyscale.
25
+ PALETTE = ("#0072B2", "#D55E00", "#009E73", "#CC79A7", "#E69F00", "#56B4E9", "#000000")
26
+ FONT_STACK = "system-ui, Segoe UI, Roboto, Helvetica, Arial, sans-serif"
27
+ FONT = f'font-family="{FONT_STACK}"'
28
+
29
+ INK = "#1a1a1a"
30
+ MUTED = "#5c5c5c"
31
+ AXIS = "#8a8a8a"
32
+ GRID = "#e8e8e8"
33
+ GRID_STRONG = "#d0d0d0"
34
+
35
+ # Per-character advance as a fraction of font size, for a humanist sans at normal weight.
36
+ _NARROW = set("iljft.,;:|!()[]{}I '")
37
+ _WIDE = set("mwMW@%")
38
+ _DIGIT = set("0123456789")
39
+
40
+
41
+ def text_width(text: str, size: float, *, weight: str = "normal") -> float:
42
+ """Estimate rendered width in user units."""
43
+
44
+ total = 0.0
45
+ for character in text:
46
+ if character in _NARROW:
47
+ total += 0.30
48
+ elif character in _WIDE:
49
+ total += 0.86
50
+ elif character in _DIGIT:
51
+ total += 0.56
52
+ elif character.isupper():
53
+ total += 0.66
54
+ else:
55
+ total += 0.52
56
+ if weight in {"600", "700", "bold"}:
57
+ total *= 1.05
58
+ return total * size
59
+
60
+
61
+ def _fit_lines(text: str, size: float, limit: float, *, weight: str = "normal") -> list[str]:
62
+ """Wrap to at most two lines, breaking on whitespace."""
63
+
64
+ if text_width(text, size, weight=weight) <= limit:
65
+ return [text]
66
+ words = text.split(" ")
67
+ line: list[str] = []
68
+ for index, word in enumerate(words):
69
+ candidate = " ".join([*line, word])
70
+ if line and text_width(candidate, size, weight=weight) > limit:
71
+ return [" ".join(line), " ".join(words[index:])]
72
+ line.append(word)
73
+ return [" ".join(line)]
74
+
75
+
76
+ @dataclass
77
+ class Series:
78
+ label: str
79
+ x: list[float]
80
+ y: list[float]
81
+ color: str = PALETTE[0]
82
+ dash: str | None = None
83
+ marker: bool = True
84
+ width: float = 1.9
85
+ marker_size: float = 2.6
86
+
87
+
88
+ @dataclass
89
+ class Band:
90
+ """A shaded interval drawn behind its series."""
91
+
92
+ x: list[float]
93
+ low: list[float]
94
+ high: list[float]
95
+ color: str = PALETTE[0]
96
+ opacity: float = 0.16
97
+
98
+
99
+ @dataclass
100
+ class Counts:
101
+ """A support strip under the plot: how many rows sit behind each x position."""
102
+
103
+ x: list[float]
104
+ values: list[float]
105
+ color: str = MUTED
106
+ label: str = "rows per bin"
107
+
108
+
109
+ @dataclass
110
+ class Panel:
111
+ title: str
112
+ series: list[Series] = field(default_factory=list)
113
+ bands: list[Band] = field(default_factory=list)
114
+ xlabel: str = ""
115
+ ylabel: str = ""
116
+ xlim: tuple[float, float] | None = None
117
+ ylim: tuple[float, float] | None = None
118
+ xticks: list[tuple[float, str]] | None = None
119
+ yticks: list[tuple[float, str]] | None = None
120
+ xscale: str = "linear"
121
+ hlines: list[tuple[float, str, str]] = field(default_factory=list)
122
+ diagonal: bool = False
123
+ notes: list[str] = field(default_factory=list)
124
+ counts: Counts | None = None
125
+
126
+
127
+ def _fmt(value: float) -> str:
128
+ if abs(value) >= 1e6:
129
+ return f"{value:.3g}"
130
+ text = f"{value:.3f}".rstrip("0").rstrip(".")
131
+ return text or "0"
132
+
133
+
134
+ def _auto_ticks(low: float, high: float, count: int = 5) -> list[tuple[float, str]]:
135
+ if high <= low:
136
+ high = low + 1.0
137
+ step = (high - low) / count
138
+ magnitude = 10 ** math.floor(math.log10(step)) if step > 0 else 1.0
139
+ for factor in (1, 2, 2.5, 5, 10):
140
+ if step <= factor * magnitude:
141
+ step = factor * magnitude
142
+ break
143
+ start = math.ceil(low / step) * step
144
+ ticks = []
145
+ value = start
146
+ while value <= high + 1e-9:
147
+ ticks.append((value, _fmt(0.0 if abs(value) < step * 1e-6 else value)))
148
+ value += step
149
+ return ticks
150
+
151
+
152
+ def _text(
153
+ x: float,
154
+ y: float,
155
+ body: str,
156
+ *,
157
+ size: float,
158
+ fill: str = INK,
159
+ anchor: str = "start",
160
+ weight: str | None = None,
161
+ ) -> str:
162
+ weight_attr = f' font-weight="{weight}"' if weight else ""
163
+ return (
164
+ f'<text x="{x:.1f}" y="{y:.1f}" text-anchor="{anchor}" font-size="{size}" '
165
+ f'fill="{fill}"{weight_attr} {FONT}>{escape(body)}</text>'
166
+ )
167
+
168
+
169
+ _TITLE_SIZE = 12.5
170
+
171
+ # Everything below the plot box is stacked in fixed bands rather than placed at absolute
172
+ # offsets, so an x-axis label, a support strip, and a note can coexist without overlapping.
173
+ _TICK_BAND = 18.0
174
+ _XLABEL_BAND = 16.0
175
+ _COUNTS_BAND = 30.0
176
+ _NOTE_BAND = 12.0
177
+ _NOTE_LEAD = 6.0
178
+ _FLOOR_SLACK = 8.0
179
+
180
+
181
+ def _panel_bottom(panel: Panel) -> float:
182
+ bottom = _TICK_BAND + _FLOOR_SLACK
183
+ if panel.xlabel:
184
+ bottom += _XLABEL_BAND
185
+ if panel.counts:
186
+ bottom += _COUNTS_BAND
187
+ if panel.notes:
188
+ bottom += _NOTE_LEAD + _NOTE_BAND * len(panel.notes)
189
+ return bottom
190
+
191
+
192
+ def _panel_svg(panel: Panel, width: float, height: float, *, title_rows: int | None = None) -> str:
193
+ left, right = 54.0, 14.0
194
+ plot_w = width - left - right
195
+ title_lines = _fit_lines(panel.title, _TITLE_SIZE, plot_w, weight="600")
196
+ top = 18.0 + 14.0 * (title_rows or len(title_lines))
197
+ plot_h = height - top - _panel_bottom(panel)
198
+
199
+ xs = [x for s in panel.series for x in s.x] or [0.0, 1.0]
200
+ ys = [y for s in panel.series for y in s.y] or [0.0, 1.0]
201
+ ys += [value for band in panel.bands for value in (*band.low, *band.high)]
202
+ ys += [level for level, _, _ in panel.hlines]
203
+ xlim = panel.xlim or (min(xs), max(xs))
204
+ ylim = panel.ylim or (min(ys), max(ys))
205
+ if ylim[0] == ylim[1]:
206
+ ylim = (ylim[0] - 0.5, ylim[1] + 0.5)
207
+
208
+ def tx(value: float) -> float:
209
+ if panel.xscale == "log":
210
+ lo, hi = math.log(xlim[0]), math.log(xlim[1])
211
+ return left + (math.log(value) - lo) / (hi - lo) * plot_w
212
+ return left + (value - xlim[0]) / (xlim[1] - xlim[0]) * plot_w
213
+
214
+ def ty(value: float) -> float:
215
+ return top + plot_h - (value - ylim[0]) / (ylim[1] - ylim[0]) * plot_h
216
+
217
+ parts: list[str] = []
218
+ for index, line in enumerate(title_lines):
219
+ parts.append(
220
+ _text(
221
+ left + plot_w / 2,
222
+ 16.0 + 14.0 * index,
223
+ line,
224
+ size=_TITLE_SIZE,
225
+ anchor="middle",
226
+ weight="600",
227
+ )
228
+ )
229
+
230
+ xticks = panel.xticks or _auto_ticks(*xlim)
231
+ yticks = panel.yticks or _auto_ticks(*ylim)
232
+ for value, label in yticks:
233
+ if not ylim[0] - 1e-9 <= value <= ylim[1] + 1e-9:
234
+ continue
235
+ y = ty(value)
236
+ parts.append(
237
+ f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
238
+ f'stroke="{GRID}" stroke-width="1"/>'
239
+ )
240
+ parts.append(_text(left - 6, y + 3.4, label, size=10, fill=MUTED, anchor="end"))
241
+ for value, label in xticks:
242
+ if not xlim[0] - 1e-9 <= value <= xlim[1] + 1e-9:
243
+ continue
244
+ x = tx(value)
245
+ parts.append(
246
+ f'<line x1="{x:.1f}" y1="{top}" x2="{x:.1f}" y2="{top + plot_h:.1f}" '
247
+ f'stroke="{GRID}" stroke-width="1"/>'
248
+ )
249
+ parts.append(_text(x, top + plot_h + 13, label, size=10, fill=MUTED, anchor="middle"))
250
+
251
+ if panel.diagonal:
252
+ parts.append(
253
+ f'<line x1="{tx(xlim[0]):.1f}" y1="{ty(ylim[0]):.1f}" '
254
+ f'x2="{tx(xlim[1]):.1f}" y2="{ty(ylim[1]):.1f}" stroke="{AXIS}" '
255
+ f'stroke-width="1.1" stroke-dasharray="4 3"/>'
256
+ )
257
+
258
+ for band in panel.bands:
259
+ forward = " ".join(
260
+ f"{tx(x):.1f},{ty(y):.1f}" for x, y in zip(band.x, band.high, strict=True)
261
+ )
262
+ backward = " ".join(
263
+ f"{tx(x):.1f},{ty(y):.1f}"
264
+ for x, y in zip(reversed(band.x), reversed(band.low), strict=True)
265
+ )
266
+ parts.append(
267
+ f'<polygon points="{forward} {backward}" fill="{band.color}" '
268
+ f'fill-opacity="{band.opacity}" stroke="none"/>'
269
+ )
270
+
271
+ # A reference line carries its label at the left margin over a solid backing box, so the
272
+ # label never lands on the data it is a reference for.
273
+ for value, label, color in panel.hlines:
274
+ y = ty(value)
275
+ parts.append(
276
+ f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
277
+ f'stroke="{color}" stroke-width="1.2" stroke-dasharray="5 3"/>'
278
+ )
279
+ if not label:
280
+ continue
281
+ label_w = text_width(label, 9.5) + 8.0
282
+ parts.append(
283
+ f'<rect x="{left + 3:.1f}" y="{y - 12:.1f}" width="{label_w:.1f}" height="11.5" '
284
+ f'fill="white" fill-opacity="0.9" stroke="none"/>'
285
+ )
286
+ parts.append(_text(left + 7, y - 3.5, label, size=9.5, fill=color))
287
+
288
+ for series in panel.series:
289
+ points = " ".join(
290
+ f"{tx(x):.1f},{ty(y):.1f}" for x, y in zip(series.x, series.y, strict=True)
291
+ )
292
+ dash = f' stroke-dasharray="{series.dash}"' if series.dash else ""
293
+ parts.append(
294
+ f'<polyline points="{points}" fill="none" stroke="{series.color}" '
295
+ f'stroke-width="{series.width}" stroke-linejoin="round" '
296
+ f'stroke-linecap="round"{dash}/>'
297
+ )
298
+ if series.marker:
299
+ for x, y in zip(series.x, series.y, strict=True):
300
+ parts.append(
301
+ f'<circle cx="{tx(x):.1f}" cy="{ty(y):.1f}" r="{series.marker_size}" '
302
+ f'fill="{series.color}"/>'
303
+ )
304
+
305
+ parts.append(
306
+ f'<rect x="{left}" y="{top}" width="{plot_w:.1f}" height="{plot_h:.1f}" fill="none" '
307
+ f'stroke="{AXIS}" stroke-width="1"/>'
308
+ )
309
+
310
+ cursor = top + plot_h + _TICK_BAND
311
+ if panel.xlabel:
312
+ parts.append(
313
+ _text(
314
+ left + plot_w / 2,
315
+ cursor + 11,
316
+ panel.xlabel,
317
+ size=10.5,
318
+ fill=MUTED,
319
+ anchor="middle",
320
+ )
321
+ )
322
+ cursor += _XLABEL_BAND
323
+ if panel.counts:
324
+ counts = panel.counts
325
+ if not counts.values:
326
+ raise ValueError("a support strip needs at least one count")
327
+ strip_top = cursor + 2.0
328
+ strip_h = 15.0
329
+ peak = max(counts.values) or 1.0
330
+ slot = plot_w / max(len(counts.x), 1) * 0.7
331
+ for x, value in zip(counts.x, counts.values, strict=True):
332
+ bar_h = (value / peak) * strip_h
333
+ parts.append(
334
+ f'<rect x="{tx(x) - slot / 2:.1f}" y="{strip_top + strip_h - bar_h:.1f}" '
335
+ f'width="{slot:.1f}" height="{bar_h:.1f}" fill="{counts.color}" '
336
+ f'fill-opacity="0.5"/>'
337
+ )
338
+ parts.append(
339
+ f'<line x1="{left}" y1="{strip_top + strip_h:.1f}" x2="{left + plot_w:.1f}" '
340
+ f'y2="{strip_top + strip_h:.1f}" stroke="{GRID_STRONG}" stroke-width="1"/>'
341
+ )
342
+ parts.append(_text(left, strip_top + strip_h + 9, counts.label, size=9, fill=MUTED))
343
+ parts.append(
344
+ _text(
345
+ left + plot_w,
346
+ strip_top + strip_h + 9,
347
+ f"tallest {int(peak):,}",
348
+ size=9,
349
+ fill=MUTED,
350
+ anchor="end",
351
+ )
352
+ )
353
+ cursor += _COUNTS_BAND
354
+ for index, note in enumerate(panel.notes):
355
+ parts.append(
356
+ _text(left, cursor + _NOTE_LEAD + 9 + _NOTE_BAND * index, note, size=9.5, fill=MUTED)
357
+ )
358
+ if panel.ylabel:
359
+ parts.append(
360
+ f'<text transform="translate(13,{top + plot_h / 2:.1f}) rotate(-90)" '
361
+ f'text-anchor="middle" font-size="10.5" fill="{MUTED}" {FONT}>'
362
+ f"{escape(panel.ylabel)}</text>"
363
+ )
364
+ return "\n".join(parts)
365
+
366
+
367
+ def _legend_svg(
368
+ entries: list[tuple[str, str, str | None]],
369
+ *,
370
+ x: float,
371
+ y: float,
372
+ max_width: float,
373
+ ) -> tuple[str, float]:
374
+ """Flow legend entries across as many rows as their measured widths need."""
375
+
376
+ swatch, gap, pad = 22.0, 7.0, 22.0
377
+ parts: list[str] = []
378
+ cursor_x, cursor_y, rows = x, y, 1
379
+ for label, color, dash in entries:
380
+ entry_w = swatch + gap + text_width(label, 11) + pad
381
+ if cursor_x > x and cursor_x + entry_w - pad > x + max_width:
382
+ cursor_x, cursor_y, rows = x, cursor_y + 16.0, rows + 1
383
+ dash_attr = f' stroke-dasharray="{dash}"' if dash else ""
384
+ parts.append(
385
+ f'<line x1="{cursor_x:.1f}" y1="{cursor_y:.1f}" x2="{cursor_x + swatch:.1f}" '
386
+ f'y2="{cursor_y:.1f}" stroke="{color}" stroke-width="2.4" '
387
+ f'stroke-linecap="round"{dash_attr}/>'
388
+ )
389
+ parts.append(_text(cursor_x + swatch + gap, cursor_y + 3.8, label, size=11))
390
+ cursor_x += entry_w
391
+ return "\n".join(parts), 16.0 * rows
392
+
393
+
394
+ def _dedupe(entries: list[tuple[str, str, str | None]]) -> list[tuple[str, str, str | None]]:
395
+ seen: set[tuple[str, str, str | None]] = set()
396
+ unique = []
397
+ for entry in entries:
398
+ if entry not in seen:
399
+ seen.add(entry)
400
+ unique.append(entry)
401
+ return unique
402
+
403
+
404
+ def _open_svg(width: float, height: float, title: str, description: str) -> list[str]:
405
+ return [
406
+ f'<svg xmlns="http://www.w3.org/2000/svg" width="{width:.0f}" height="{height:.0f}" '
407
+ f'viewBox="0 0 {width:.0f} {height:.0f}" role="img" aria-label="{escape(title)}">',
408
+ f"<title>{escape(title)}</title>",
409
+ f"<desc>{escape(description)}</desc>",
410
+ '<rect width="100%" height="100%" fill="white"/>',
411
+ ]
412
+
413
+
414
+ def render_grid(
415
+ panels: list[Panel],
416
+ *,
417
+ columns: int,
418
+ title: str,
419
+ subtitle: str = "",
420
+ panel_width: float = 352.0,
421
+ panel_height: float = 256.0,
422
+ legend: list[tuple[str, str, str | None]] | None = None,
423
+ description: str = "",
424
+ ) -> str:
425
+ """Render panels on a grid with one shared title and a flowed legend.
426
+
427
+ A final short row is centred, so a five-panel figure on three columns has no empty cell.
428
+ """
429
+
430
+ if not panels:
431
+ raise ValueError("render_grid needs at least one panel")
432
+ if columns < 1:
433
+ raise ValueError("render_grid needs at least one column")
434
+ legend = _dedupe(legend or [])
435
+ rows = math.ceil(len(panels) / columns)
436
+ width = columns * panel_width
437
+ header = 26.0 + (16.0 if subtitle else 0.0)
438
+ legend_svg, legend_h = "", 0.0
439
+ if legend:
440
+ legend_svg, legend_h = _legend_svg(legend, x=18.0, y=header + 10.0, max_width=width - 36.0)
441
+ legend_h += 8.0
442
+ height = header + legend_h + rows * panel_height
443
+ parts = _open_svg(width, height, title, description or title)
444
+ parts.append(_text(width / 2, 19, title, size=14.5, anchor="middle", weight="700"))
445
+ if subtitle:
446
+ parts.append(_text(width / 2, 34, subtitle, size=11, fill=MUTED, anchor="middle"))
447
+ if legend_svg:
448
+ parts.append(legend_svg)
449
+ # One title row count for the whole grid, so a panel whose title wraps does not push its
450
+ # plot box below its neighbours'.
451
+ title_rows = max(
452
+ len(_fit_lines(panel.title, _TITLE_SIZE, panel_width - 68.0, weight="600"))
453
+ for panel in panels
454
+ )
455
+ for index, panel in enumerate(panels):
456
+ row, column = divmod(index, columns)
457
+ in_row = min(columns, len(panels) - row * columns)
458
+ offset = (columns - in_row) * panel_width / 2.0
459
+ px = offset + column * panel_width
460
+ py = header + legend_h + row * panel_height
461
+ parts.append(f'<g transform="translate({px:.1f},{py:.1f})">')
462
+ parts.append(_panel_svg(panel, panel_width, panel_height, title_rows=title_rows))
463
+ parts.append("</g>")
464
+ parts.append("</svg>")
465
+ return "\n".join(parts) + "\n"
466
+
467
+
468
+ def render_bars(
469
+ groups: list[str],
470
+ series: list[tuple[str, list[float], str]],
471
+ *,
472
+ title: str,
473
+ ylabel: str,
474
+ subtitle: str = "",
475
+ ylim: tuple[float, float] = (0.0, 1.0),
476
+ reference: list[tuple[str, list[float], str]] | None = None,
477
+ notes: list[str] | None = None,
478
+ separator_before: int | None = None,
479
+ width: float = 780.0,
480
+ height: float = 350.0,
481
+ description: str = "",
482
+ ) -> str:
483
+ """Render grouped bars with per-group dashed reference levels.
484
+
485
+ Bars keep a zero baseline. Value labels are drawn only where a bar is wide enough to hold
486
+ one, because a crowded label is worse than none.
487
+ """
488
+
489
+ if not groups or not series:
490
+ raise ValueError("render_bars needs at least one group and one series")
491
+ if any(len(values) != len(groups) for _, values, _ in series):
492
+ raise ValueError("every bar series must carry one value per group")
493
+ notes = notes or []
494
+ left, right, top = 54.0, 16.0, 26.0 + (15.0 if subtitle else 0.0)
495
+ legend_entries = [(label, color, None) for label, _, color in series]
496
+ legend_entries += [(label, color, "5 3") for label, _, color in reference or []]
497
+ legend_svg, legend_h = _legend_svg(
498
+ _dedupe(legend_entries), x=18.0, y=top + 12.0, max_width=width - 36.0
499
+ )
500
+ top += legend_h + 10.0
501
+ bottom = 46.0 + 12.0 * len(notes)
502
+ plot_w, plot_h = width - left - right, height - top - bottom
503
+ group_w = plot_w / len(groups)
504
+ bar_w = group_w * 0.74 / len(series)
505
+
506
+ def ty(value: float) -> float:
507
+ return top + plot_h - (value - ylim[0]) / (ylim[1] - ylim[0]) * plot_h
508
+
509
+ parts = _open_svg(width, height, title, description or title)
510
+ parts.append(_text(width / 2, 19, title, size=14.5, anchor="middle", weight="700"))
511
+ if subtitle:
512
+ parts.append(_text(width / 2, 34, subtitle, size=11, fill=MUTED, anchor="middle"))
513
+ parts.append(legend_svg)
514
+ for value, label in _auto_ticks(*ylim):
515
+ y = ty(value)
516
+ parts.append(
517
+ f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
518
+ f'stroke="{GRID}" stroke-width="1"/>'
519
+ )
520
+ parts.append(_text(left - 6, y + 3.4, label, size=10, fill=MUTED, anchor="end"))
521
+
522
+ label_fits = bar_w - 2 >= text_width("0.000", 8.5) + 2
523
+ for g_index, group in enumerate(groups):
524
+ gx = left + g_index * group_w + group_w * 0.13
525
+ for s_index, (_, values, color) in enumerate(series):
526
+ value = values[g_index]
527
+ x = gx + s_index * bar_w
528
+ parts.append(
529
+ f'<rect x="{x:.1f}" y="{ty(value):.1f}" width="{bar_w - 2:.1f}" '
530
+ f'height="{ty(ylim[0]) - ty(value):.1f}" fill="{color}"/>'
531
+ )
532
+ if label_fits:
533
+ parts.append(
534
+ _text(
535
+ x + (bar_w - 2) / 2,
536
+ ty(value) - 4,
537
+ f"{value:.3f}",
538
+ size=8.5,
539
+ anchor="middle",
540
+ )
541
+ )
542
+ for _, values, color in reference or []:
543
+ y = ty(values[g_index])
544
+ parts.append(
545
+ f'<line x1="{gx - 3:.1f}" y1="{y:.1f}" '
546
+ f'x2="{gx + group_w * 0.74 + 3:.1f}" y2="{y:.1f}" stroke="{color}" '
547
+ f'stroke-width="1.6" stroke-dasharray="5 3"/>'
548
+ )
549
+ parts.append(
550
+ _text(
551
+ left + g_index * group_w + group_w / 2,
552
+ top + plot_h + 16,
553
+ group,
554
+ size=11,
555
+ anchor="middle",
556
+ )
557
+ )
558
+ if separator_before is not None and 0 < separator_before < len(groups):
559
+ x = left + separator_before * group_w
560
+ parts.append(
561
+ f'<line x1="{x:.1f}" y1="{top}" x2="{x:.1f}" y2="{top + plot_h + 6:.1f}" '
562
+ f'stroke="{GRID_STRONG}" stroke-width="1.4"/>'
563
+ )
564
+ parts.append(
565
+ f'<rect x="{left}" y="{top}" width="{plot_w:.1f}" height="{plot_h:.1f}" fill="none" '
566
+ f'stroke="{AXIS}" stroke-width="1"/>'
567
+ )
568
+ parts.append(
569
+ f'<text transform="translate(13,{top + plot_h / 2:.1f}) rotate(-90)" '
570
+ f'text-anchor="middle" font-size="10.5" fill="{MUTED}" {FONT}>'
571
+ f"{escape(ylabel)}</text>"
572
+ )
573
+ for index, note in enumerate(notes):
574
+ parts.append(_text(left, top + plot_h + 34 + 12 * index, note, size=9.5, fill=MUTED))
575
+ parts.append("</svg>")
576
+ return "\n".join(parts) + "\n"
figures/classifier-comparison.svg ADDED
model.onnx CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c
3
- size 452812018
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7e9b7e37afc06a23a89ba064e1255363cdc3140ff4b2823f5dae5865945621dc
3
+ size 452831012
publication-manifest.json CHANGED
@@ -1,6 +1,6 @@
1
  {
2
  "schema": "daecore.classifier-publication-manifest",
3
- "tool_sha256": "ec0f21b2c945a07b0a1f7be9895c91f30a985fb981a5fb83f409604b33879af4",
4
  "public_repo": "Daecore/gliclass-std-base-v3-5facet-qint8-v2",
5
  "classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
6
  "upstream": {
@@ -8,38 +8,39 @@
8
  "revision": "77a70e6cd52e602ed18184ef37d18bdd3741e3d5"
9
  },
10
  "nominee_receipt_sha256": "d463fc167b2a26f8f5b56e03c18e4c604b231eb25b5093064dd8a9bd7ffa6373",
11
- "staged_at": "2026-09-18T18:58:21+00:00",
 
 
12
  "files": {
 
 
 
 
 
 
 
 
 
 
13
  "model.onnx": {
14
- "sha256": "690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c",
15
- "size": 452812018,
16
- "binding": "candidate manifest (file hash)"
17
  },
18
  "tokenizer.json": {
19
  "sha256": "519648948c4c59da1af88f2cf2c8b4f84417b5c673981bc9809abf84cda1b7cc",
20
  "size": 8649234,
21
- "binding": "candidate manifest (file hash)"
22
- },
23
- "config.json": {
24
- "sha256": "9bad60c87ac3e047bc838ca7e6897a6fe265180eeffce4e62aabc1e9c828d606",
25
- "size": 2118,
26
- "binding": "candidate manifest (file hash)"
27
  },
28
  "tokenizer_config.json": {
29
  "sha256": "9d5220b355d2cb9a7df69deccca45509ce1c9a857990fb1047c48403b0f6ddfb",
30
  "size": 1692,
31
- "binding": "candidate manifest (file hash)"
32
- },
33
- "classifier-metadata.json": {
34
- "sha256": "00d8c3b36be47f96ddd113bd6ddbf525c3599de27064c68ef7a1aacf4c0408e5",
35
- "size": 2596,
36
- "canonical_sha256": "1918aa0c392e6582a7028a4a2482166714080958602f6234f145dd294fb01996",
37
- "binding": "nominee receipt runtime_metadata_sha256 (canonical JSON hash)"
38
  },
39
- "README.md": {
40
- "sha256": "48cc3f3e7659463c7bd6031d923a96f3822e7ebd62a2d5f660d779d95612f76c",
41
- "size": 29930,
42
- "binding": "packaging record"
43
  },
44
  "LICENSE": {
45
  "sha256": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
@@ -52,9 +53,55 @@
52
  "binding": "packaging record"
53
  },
54
  "MODIFICATIONS.md": {
55
- "sha256": "fca339612182863205137d75b60cb23dc921090c0e91a4541a382c3889cd7e95",
56
- "size": 1399,
 
 
 
 
 
57
  "binding": "packaging record"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
58
  }
59
  }
60
  }
 
1
  {
2
  "schema": "daecore.classifier-publication-manifest",
3
+ "tool_sha256": "5ac2ef9465fe9a3454bfd2e61c7aa651b08fbdfef9473fb5b3455585ba560a7b",
4
  "public_repo": "Daecore/gliclass-std-base-v3-5facet-qint8-v2",
5
  "classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
6
  "upstream": {
 
8
  "revision": "77a70e6cd52e602ed18184ef37d18bdd3741e3d5"
9
  },
10
  "nominee_receipt_sha256": "d463fc167b2a26f8f5b56e03c18e4c604b231eb25b5093064dd8a9bd7ffa6373",
11
+ "derivation_preparation_sha256": "bfd902e06b7a984348283dc279f5e46d76398ff59633a9b5b426f33d872a67d3",
12
+ "source_nominee_binding": "unchanged learned model and calibration; graph bytes separately derived",
13
+ "staged_at": "2026-09-29T03:04:21+00:00",
14
  "files": {
15
+ "classifier-metadata.json": {
16
+ "sha256": "736d8b615fb8c4bd7cc36d1c2e0e86bbe096d63dc373d047ca75af9cc6119a11",
17
+ "size": 2674,
18
+ "binding": "pinned graph derivation after original nominee verification"
19
+ },
20
+ "config.json": {
21
+ "sha256": "9bad60c87ac3e047bc838ca7e6897a6fe265180eeffce4e62aabc1e9c828d606",
22
+ "size": 2118,
23
+ "binding": "pinned graph derivation after original nominee verification"
24
+ },
25
  "model.onnx": {
26
+ "sha256": "7e9b7e37afc06a23a89ba064e1255363cdc3140ff4b2823f5dae5865945621dc",
27
+ "size": 452831012,
28
+ "binding": "pinned graph derivation after original nominee verification"
29
  },
30
  "tokenizer.json": {
31
  "sha256": "519648948c4c59da1af88f2cf2c8b4f84417b5c673981bc9809abf84cda1b7cc",
32
  "size": 8649234,
33
+ "binding": "pinned graph derivation after original nominee verification"
 
 
 
 
 
34
  },
35
  "tokenizer_config.json": {
36
  "sha256": "9d5220b355d2cb9a7df69deccca45509ce1c9a857990fb1047c48403b0f6ddfb",
37
  "size": 1692,
38
+ "binding": "pinned graph derivation after original nominee verification"
 
 
 
 
 
 
39
  },
40
+ "vulkan-derivation.json": {
41
+ "sha256": "472be46b2c127c4bf38dcf1db3d5dfe65672a6ea66164677d91a2cbb2d05ee70",
42
+ "size": 5782,
43
+ "binding": "pinned graph derivation after original nominee verification"
44
  },
45
  "LICENSE": {
46
  "sha256": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
 
53
  "binding": "packaging record"
54
  },
55
  "MODIFICATIONS.md": {
56
+ "sha256": "429cfef847f3816fe9cc4715d2431f344d5b34376e03b214eae52c1626021f39",
57
+ "size": 1773,
58
+ "binding": "packaging record"
59
+ },
60
+ "evaluation/README.md": {
61
+ "sha256": "3215e6b0e302fdecd86ef1457e46e7d033fa18fb1ebcbd0dcb47c2b8abe7c174",
62
+ "size": 8825,
63
  "binding": "packaging record"
64
+ },
65
+ "evaluation/metrics.py": {
66
+ "sha256": "7c25f29a03e0f61d5f0e59f7781269489f80512ae6862920f91a2d333c058bc4",
67
+ "size": 11135,
68
+ "binding": "packaging record"
69
+ },
70
+ "evaluation/classifier.json": {
71
+ "sha256": "5498bb3f5f15f587bc93fc77c9a56308e6ccbc85976bd8e38530b643614628bc",
72
+ "size": 721988,
73
+ "binding": "packaging record"
74
+ },
75
+ "evaluation/classifier-summary.json": {
76
+ "sha256": "e89f48dff2e29f9d8f4c27b221b88c2d380f725a682d6792116bb0d3dff7bd8e",
77
+ "size": 4799,
78
+ "binding": "packaging record"
79
+ },
80
+ "evaluation/serving-qualification.json": {
81
+ "sha256": "bf7deda3f25306d15454798ad4c8e785c25fba970fc774efbe557225fc5d42e0",
82
+ "size": 4667,
83
+ "binding": "packaging record"
84
+ },
85
+ "evaluation/figures.py": {
86
+ "sha256": "9944765ee42148f2bb135f428ee1dacb822af0ca5c77be9f82b507eb8da236d4",
87
+ "size": 4097,
88
+ "binding": "packaging record"
89
+ },
90
+ "evaluation/svg_figures.py": {
91
+ "sha256": "c68b7e26e072e3875f111b592f8031322261272fcf285ae2a3edbe378e70740d",
92
+ "size": 21536,
93
+ "binding": "packaging record"
94
+ },
95
+ "figures/classifier-comparison.svg": {
96
+ "sha256": "a465b6ab22d60b50543c29dd7f16fe440b948eb0d12c38df7facc34fecfee509",
97
+ "size": 8642,
98
+ "binding": "packaging record"
99
+ },
100
+ "README.md": {
101
+ "sha256": "eec5fe5f0d257201d4d517f6a3edb97c5a6cfbe352cc8835c2be42c3baf0ad14",
102
+ "size": 8321,
103
+ "binding": "model card with upload-relative links",
104
+ "source_sha256": "f62d093e04a97c57080285d08d64c66906b500ee2124d86013b2dde0c120b592"
105
  }
106
  }
107
  }
vulkan-derivation.json ADDED
@@ -0,0 +1,227 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "artifacts": [
3
+ {
4
+ "bytes": 452831012,
5
+ "path": "model.onnx",
6
+ "sha256": "7e9b7e37afc06a23a89ba064e1255363cdc3140ff4b2823f5dae5865945621dc"
7
+ }
8
+ ],
9
+ "changes": [
10
+ {
11
+ "dtype": 7,
12
+ "name": "/inner/model/encoder_model/encoder/Squeeze",
13
+ "operation": "Squeeze"
14
+ },
15
+ {
16
+ "dtype": 7,
17
+ "name": "/inner/model/CumSum",
18
+ "operation": "CumSum"
19
+ },
20
+ {
21
+ "dtype": 7,
22
+ "name": "Sign_1892",
23
+ "operation": "Sign"
24
+ },
25
+ {
26
+ "dtype": 7,
27
+ "name": "Abs_1898",
28
+ "operation": "Abs"
29
+ },
30
+ {
31
+ "dtype": 7,
32
+ "name": "/inner/model/encoder_model/encoder/Slice_1",
33
+ "operation": "Slice"
34
+ },
35
+ {
36
+ "dtype": 7,
37
+ "name": "/inner/model/encoder_model/encoder/layer.0/attention/self/Squeeze_3",
38
+ "operation": "Squeeze"
39
+ },
40
+ {
41
+ "dtype": 7,
42
+ "name": "Identity_2252",
43
+ "operation": "Identity"
44
+ },
45
+ {
46
+ "dtype": 7,
47
+ "name": "/inner/model/encoder_model/encoder/layer.0/attention/self/Neg",
48
+ "operation": "Neg"
49
+ },
50
+ {
51
+ "dtype": 7,
52
+ "name": "/inner/model/encoder_model/encoder/layer.0/attention/self/Squeeze_5",
53
+ "operation": "Squeeze"
54
+ },
55
+ {
56
+ "dtype": 7,
57
+ "name": "Identity_2696",
58
+ "operation": "Identity"
59
+ },
60
+ {
61
+ "dtype": 7,
62
+ "name": "/inner/model/encoder_model/encoder/layer.1/attention/self/Neg",
63
+ "operation": "Neg"
64
+ },
65
+ {
66
+ "dtype": 7,
67
+ "name": "/inner/model/encoder_model/encoder/layer.1/attention/self/Squeeze_3",
68
+ "operation": "Squeeze"
69
+ },
70
+ {
71
+ "dtype": 7,
72
+ "name": "Identity_3138",
73
+ "operation": "Identity"
74
+ },
75
+ {
76
+ "dtype": 7,
77
+ "name": "/inner/model/encoder_model/encoder/layer.2/attention/self/Neg",
78
+ "operation": "Neg"
79
+ },
80
+ {
81
+ "dtype": 7,
82
+ "name": "/inner/model/encoder_model/encoder/layer.2/attention/self/Squeeze_3",
83
+ "operation": "Squeeze"
84
+ },
85
+ {
86
+ "dtype": 7,
87
+ "name": "Identity_3580",
88
+ "operation": "Identity"
89
+ },
90
+ {
91
+ "dtype": 7,
92
+ "name": "/inner/model/encoder_model/encoder/layer.3/attention/self/Neg",
93
+ "operation": "Neg"
94
+ },
95
+ {
96
+ "dtype": 7,
97
+ "name": "/inner/model/encoder_model/encoder/layer.3/attention/self/Squeeze_3",
98
+ "operation": "Squeeze"
99
+ },
100
+ {
101
+ "dtype": 7,
102
+ "name": "Identity_4022",
103
+ "operation": "Identity"
104
+ },
105
+ {
106
+ "dtype": 7,
107
+ "name": "/inner/model/encoder_model/encoder/layer.4/attention/self/Neg",
108
+ "operation": "Neg"
109
+ },
110
+ {
111
+ "dtype": 7,
112
+ "name": "/inner/model/encoder_model/encoder/layer.4/attention/self/Squeeze_3",
113
+ "operation": "Squeeze"
114
+ },
115
+ {
116
+ "dtype": 7,
117
+ "name": "Identity_4464",
118
+ "operation": "Identity"
119
+ },
120
+ {
121
+ "dtype": 7,
122
+ "name": "/inner/model/encoder_model/encoder/layer.5/attention/self/Neg",
123
+ "operation": "Neg"
124
+ },
125
+ {
126
+ "dtype": 7,
127
+ "name": "/inner/model/encoder_model/encoder/layer.5/attention/self/Squeeze_3",
128
+ "operation": "Squeeze"
129
+ },
130
+ {
131
+ "dtype": 7,
132
+ "name": "Identity_4906",
133
+ "operation": "Identity"
134
+ },
135
+ {
136
+ "dtype": 7,
137
+ "name": "/inner/model/encoder_model/encoder/layer.6/attention/self/Neg",
138
+ "operation": "Neg"
139
+ },
140
+ {
141
+ "dtype": 7,
142
+ "name": "/inner/model/encoder_model/encoder/layer.6/attention/self/Squeeze_3",
143
+ "operation": "Squeeze"
144
+ },
145
+ {
146
+ "dtype": 7,
147
+ "name": "Identity_5348",
148
+ "operation": "Identity"
149
+ },
150
+ {
151
+ "dtype": 7,
152
+ "name": "/inner/model/encoder_model/encoder/layer.7/attention/self/Neg",
153
+ "operation": "Neg"
154
+ },
155
+ {
156
+ "dtype": 7,
157
+ "name": "/inner/model/encoder_model/encoder/layer.7/attention/self/Squeeze_3",
158
+ "operation": "Squeeze"
159
+ },
160
+ {
161
+ "dtype": 7,
162
+ "name": "Identity_5790",
163
+ "operation": "Identity"
164
+ },
165
+ {
166
+ "dtype": 7,
167
+ "name": "/inner/model/encoder_model/encoder/layer.8/attention/self/Neg",
168
+ "operation": "Neg"
169
+ },
170
+ {
171
+ "dtype": 7,
172
+ "name": "/inner/model/encoder_model/encoder/layer.8/attention/self/Squeeze_3",
173
+ "operation": "Squeeze"
174
+ },
175
+ {
176
+ "dtype": 7,
177
+ "name": "Identity_6232",
178
+ "operation": "Identity"
179
+ },
180
+ {
181
+ "dtype": 7,
182
+ "name": "/inner/model/encoder_model/encoder/layer.9/attention/self/Neg",
183
+ "operation": "Neg"
184
+ },
185
+ {
186
+ "dtype": 7,
187
+ "name": "/inner/model/encoder_model/encoder/layer.9/attention/self/Squeeze_3",
188
+ "operation": "Squeeze"
189
+ },
190
+ {
191
+ "dtype": 7,
192
+ "name": "Identity_6674",
193
+ "operation": "Identity"
194
+ },
195
+ {
196
+ "dtype": 7,
197
+ "name": "/inner/model/encoder_model/encoder/layer.10/attention/self/Neg",
198
+ "operation": "Neg"
199
+ },
200
+ {
201
+ "dtype": 7,
202
+ "name": "/inner/model/encoder_model/encoder/layer.10/attention/self/Squeeze_3",
203
+ "operation": "Squeeze"
204
+ },
205
+ {
206
+ "dtype": 7,
207
+ "name": "Identity_7116",
208
+ "operation": "Identity"
209
+ },
210
+ {
211
+ "dtype": 7,
212
+ "name": "/inner/model/encoder_model/encoder/layer.11/attention/self/Neg",
213
+ "operation": "Neg"
214
+ },
215
+ {
216
+ "dtype": 7,
217
+ "name": "/inner/model/encoder_model/encoder/layer.11/attention/self/Squeeze_3",
218
+ "operation": "Squeeze"
219
+ }
220
+ ],
221
+ "max_sequence": 768,
222
+ "role": "gliclass",
223
+ "schema": "daecore.vulkan-mask-derivation.v1",
224
+ "source_graph_sha256": "690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c",
225
+ "source_weights_sha256": null,
226
+ "weights_unchanged": true
227
+ }