Publish gliclass-std-base-v3-daecore-5facet-qint8-v2 (qualified ONNX export and complete attribution)
Browse files- MODIFICATIONS.md +5 -0
- README.md +126 -511
- classifier-metadata.json +78 -78
- evaluation/README.md +83 -0
- evaluation/classifier-summary.json +138 -0
- evaluation/classifier.json +0 -0
- evaluation/figures.py +89 -0
- evaluation/metrics.py +219 -0
- evaluation/serving-qualification.json +134 -0
- evaluation/svg_figures.py +576 -0
- figures/classifier-comparison.svg +69 -0
- model.onnx +2 -2
- publication-manifest.json +71 -24
- vulkan-derivation.json +227 -0
MODIFICATIONS.md
CHANGED
|
@@ -21,6 +21,11 @@ What changed:
|
|
| 21 |
- **Calibration.** Per-facet temperature scaling and two frozen threshold
|
| 22 |
tables (`recall_leaning`, `contract`) travel in `classifier-metadata.json`
|
| 23 |
and are part of the artifact's identity.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
|
| 25 |
Unchanged: the tokenizer vocabulary and the `<<LABEL>>` / `<<SEP>>` prompt
|
| 26 |
convention of the upstream model.
|
|
|
|
| 21 |
- **Calibration.** Per-facet temperature scaling and two frozen threshold
|
| 22 |
tables (`recall_leaning`, `contract`) travel in `classifier-metadata.json`
|
| 23 |
and are part of the artifact's identity.
|
| 24 |
+
- **Vulkan graph preparation.** Bounded integer/boolean mask calculations
|
| 25 |
+
use exact FP32 equivalents before their original output types are restored.
|
| 26 |
+
Learned parameters, tokenizers, label prompts, calibration and thresholds
|
| 27 |
+
are unchanged. `vulkan-derivation.json` binds the source and derived graph;
|
| 28 |
+
`classifier-metadata.json` records the new graph's byte identity.
|
| 29 |
|
| 30 |
Unchanged: the tokenizer vocabulary and the `<<LABEL>>` / `<<SEP>>` prompt
|
| 31 |
convention of the upstream model.
|
README.md
CHANGED
|
@@ -1,511 +1,126 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: apache-2.0
|
| 3 |
-
base_model: knowledgator/gliclass-base-v3.0
|
| 4 |
-
pipeline_tag: text-classification
|
| 5 |
-
|
| 6 |
-
-
|
| 7 |
-
|
| 8 |
-
-
|
| 9 |
-
-
|
| 10 |
-
-
|
| 11 |
-
-
|
| 12 |
-
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
#
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
|
| 105 |
-
|---|---
|
| 106 |
-
|
|
| 107 |
-
|
| 108 |
-
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
|
| 115 |
-
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
Register rows are training supervision only: no panel that calibrates, gates, or
|
| 128 |
-
reports on this model contains them, or any other real row.
|
| 129 |
-
|
| 130 |
-
Splits are by whole source family, never by row. The seven families reserved
|
| 131 |
-
for calibration and promotion were removed from training along with 3,875
|
| 132 |
-
predecessor training rows that shared them. Calibration and promotion families
|
| 133 |
-
are disjoint from each other and from training by family, document lineage, and
|
| 134 |
-
exact passage identity. Assignment was label-blind and checked for support only
|
| 135 |
-
afterward.
|
| 136 |
-
|
| 137 |
-
The final fit used weighted binary cross-entropy with a 2.0 multiplier on
|
| 138 |
-
mentions-only hard negatives, learning rate 2e-5, batch 4 with 8 accumulation
|
| 139 |
-
steps, four epochs, and an equal-weight average of the checkpoints from epochs
|
| 140 |
-
2 through 4. Generated and real rows were weighted equally.
|
| 141 |
-
|
| 142 |
-
## Evaluation data
|
| 143 |
-
|
| 144 |
-
| Panel | Rows | Families | Origin | Role |
|
| 145 |
-
|---|---:|---:|---|---|
|
| 146 |
-
| Calibration | 1,700 | 4 | generated | fit temperatures and threshold tables; no evaluation authority |
|
| 147 |
-
| Promotion | 1,300 | 3 | generated | one sealed read; the numbers reported below |
|
| 148 |
-
|
| 149 |
-
Promotion-panel support per facet, with mentions-only rows counted as
|
| 150 |
-
negatives. Prevalence equals the average precision a random ranker would score.
|
| 151 |
-
|
| 152 |
-
| Facet | Positive | Mentions only | Negative | Unresolved | Prevalence |
|
| 153 |
-
|---|---:|---:|---:|---:|---:|
|
| 154 |
-
| trap | 933 | 62 | 305 | 0 | 0.718 |
|
| 155 |
-
| decision | 598 | 398 | 303 | 1 | 0.460 |
|
| 156 |
-
| constraint | 1,187 | 32 | 81 | 0 | 0.913 |
|
| 157 |
-
| mechanism | 892 | 62 | 343 | 3 | 0.688 |
|
| 158 |
-
| procedure | 485 | 277 | 531 | 7 | 0.375 |
|
| 159 |
-
|
| 160 |
-
The macro random floor is 0.631. Constraint is near saturation at 0.913
|
| 161 |
-
prevalence, so its average precision carries little information and the macro
|
| 162 |
-
average leans on the other four facets. The panel was deliberately enriched for
|
| 163 |
-
hard decision and procedure rows; it is not a natural-traffic sample.
|
| 164 |
-
|
| 165 |
-
## Evaluation protocol
|
| 166 |
-
|
| 167 |
-
The promotion plan was frozen on 2026-08-12 before the candidate was scored:
|
| 168 |
-
one read of the panel, bootstrap seed 20260812, 10,000 draws resampling whole
|
| 169 |
-
source families, 95% one-sided bounds, and zero permitted stamp flips between
|
| 170 |
-
the reference scorer and the production ONNX adapter. Pass criteria were
|
| 171 |
-
pre-registered on the decision facet under its frozen contract threshold:
|
| 172 |
-
precision at least 0.70 with a one-sided 95% lower bound at least 0.60, and
|
| 173 |
-
recall at least 0.50. Threshold tables were fitted on the calibration panel and
|
| 174 |
-
bound into the prediction receipt; nothing was tuned on the promotion panel.
|
| 175 |
-
|
| 176 |
-
Average precision and precision at fixed recall are the primary metrics because
|
| 177 |
-
they do not depend on a threshold. Brier score and 10-bin expected calibration
|
| 178 |
-
error report probability quality separately from ranking quality.
|
| 179 |
-
|
| 180 |
-
## Results
|
| 181 |
-
|
| 182 |
-
Promotion panel, 1,300 rows, three families, single sealed read. "Previous"
|
| 183 |
-
is the classifier this model replaces.
|
| 184 |
-
|
| 185 |
-
| Facet | Prevalence | AP, this model | AP, previous | P@R90, this model | P@R90, previous | P@R95, this model | P@R95, previous |
|
| 186 |
-
|---|---:|---:|---:|---:|---:|---:|---:|
|
| 187 |
-
| trap | 0.718 | 0.9864 | 0.8931 | 0.960 | 0.832 | 0.935 | 0.784 |
|
| 188 |
-
| decision | 0.460 | 0.9135 | 0.6783 | 0.741 | 0.504 | 0.658 | 0.484 |
|
| 189 |
-
| constraint | 0.913 | 0.9986 | 0.9884 | 0.996 | 0.956 | 0.995 | 0.942 |
|
| 190 |
-
| mechanism | 0.688 | 0.9817 | 0.8865 | 0.926 | 0.777 | 0.891 | 0.736 |
|
| 191 |
-
| procedure | 0.375 | 0.9578 | 0.8167 | 0.886 | 0.528 | 0.822 | 0.489 |
|
| 192 |
-
| macro | 0.631 | 0.9676 | 0.8526 | 0.902 | 0.719 | 0.860 | 0.687 |
|
| 193 |
-
|
| 194 |
-
The macro gain is 0.1150. The one-sided 95%
|
| 195 |
-
lower bound on the macro precision gain under the frozen contract thresholds is
|
| 196 |
-
0.009; the decision-facet precision gain has a negative lower bound, because the
|
| 197 |
-
previous classifier's decision threshold was so strict that it predicted only 26
|
| 198 |
-
positives at 92% precision and 4% recall.
|
| 199 |
-
|
| 200 |
-

|
| 201 |
-
|
| 202 |
-
**Frozen-threshold operating points on the promotion panel.** The `contract`
|
| 203 |
-
table is the stricter view; the pre-registered gates apply to decision.
|
| 204 |
-
|
| 205 |
-
| Facet | Threshold | Precision | Recall | Predicted positive |
|
| 206 |
-
|---|---:|---:|---:|---:|
|
| 207 |
-
| trap | 0.976 | 0.995 | 0.636 | 596 |
|
| 208 |
-
| decision | 0.485 | 0.890 | 0.703 | 473 |
|
| 209 |
-
| constraint | 0.993 | 1.000 | 0.802 | 952 |
|
| 210 |
-
| mechanism | 0.900 | 0.992 | 0.655 | 591 |
|
| 211 |
-
| procedure | 0.578 | 0.946 | 0.715 | 371 |
|
| 212 |
-
|
| 213 |
-
The recall denominators in this table do not all match the support table above.
|
| 214 |
-
Where a facet has unresolved rows, the promoted model's recall is computed over a
|
| 215 |
-
slightly larger set: decision 599 against 598 positives, mechanism 895 against
|
| 216 |
-
892, procedure 491 against 485. The previous classifier is scored over the
|
| 217 |
-
support counts, except procedure at 487. The gaps track the per-facet unresolved
|
| 218 |
-
rows, which the sealed read did not mask identically for the two models. The
|
| 219 |
-
effect is conservative for the promoted model — a larger denominator lowers its
|
| 220 |
-
recall — and no gate outcome changes.
|
| 221 |
-
|
| 222 |
-
Decision passed all three gates: precision 0.890 against a 0.70 floor, a
|
| 223 |
-
one-sided 95% lower bound of 0.872 against 0.60, and recall 0.703 against 0.50.
|
| 224 |
-
The operating point encodes a product judgment: surfacing a non-decision as a
|
| 225 |
-
decision is treated as worse than missing one. The `recall_leaning` table
|
| 226 |
-
(decision threshold 0.193) exists for consumers with the opposite preference.
|
| 227 |
-
|
| 228 |
-
**Calibration.** Per-facet temperatures fitted on the calibration panel were
|
| 229 |
-
1.89 (trap), 2.40 (decision), 1.69 (constraint), 1.50 (mechanism), and 2.05
|
| 230 |
-
(procedure). On the promotion panel the calibrated posteriors score macro Brier
|
| 231 |
-
0.072 and macro expected calibration error 0.039, against 0.349 and 0.374 for
|
| 232 |
-
the previous classifier.
|
| 233 |
-
|
| 234 |
-

|
| 235 |
-
|
| 236 |
-
Held-family slices are diagnostic only: macro average precision was 0.962,
|
| 237 |
-
0.967, and 0.974 across the three families for this model and 0.851, 0.850, and
|
| 238 |
-
0.859 for the previous one.
|
| 239 |
-
|
| 240 |
-
## Baselines and comparisons
|
| 241 |
-
|
| 242 |
-
Ordered by how directly each comparison can be checked.
|
| 243 |
-
|
| 244 |
-
**Previous classifier.** A ModernBERT-base sequence-classification head used as
|
| 245 |
-
a zero-shot entailment judge: one pass per facet, five passes per passage. Its
|
| 246 |
-
staged bundle declares `ModernBertForSequenceClassification` over the two
|
| 247 |
-
labels `entailment` and `not_entailment`, 22 layers at hidden size 768, and a
|
| 248 |
-
599,027,211-byte FP32 `model.onnx` — about 1.3 times the promoted artifact for
|
| 249 |
-
five times the work per passage. Its frozen thresholds were fitted on an older
|
| 250 |
-
population and artifact, which is why its contract recall collapsed to 0.001 on
|
| 251 |
-
trap and 0.040 on decision on this panel; the threshold-free columns above are
|
| 252 |
-
the fair comparison.
|
| 253 |
-
|
| 254 |
-
**Lexical baseline.** A word-level TF-IDF (1-2 grams, 200k features) with one
|
| 255 |
-
balanced logistic-regression head per facet, trained on the same 59,886 rows
|
| 256 |
-
and scored on the same 1,300-row panel under the same label rules.
|
| 257 |
-
|
| 258 |
-
| Facet | Floor | TF-IDF + LR | This model | Previous |
|
| 259 |
-
|---|---:|---:|---:|---:|
|
| 260 |
-
| trap | 0.718 | 0.970 | 0.986 | 0.893 |
|
| 261 |
-
| decision | 0.460 | 0.847 | 0.914 | 0.678 |
|
| 262 |
-
| constraint | 0.913 | 0.997 | 0.999 | 0.988 |
|
| 263 |
-
| mechanism | 0.688 | 0.963 | 0.982 | 0.886 |
|
| 264 |
-
| procedure | 0.375 | 0.914 | 0.958 | 0.817 |
|
| 265 |
-
| macro AP | 0.631 | 0.938 | 0.968 | 0.853 |
|
| 266 |
-
| macro P@R90 | | 0.834 | 0.902 | 0.719 |
|
| 267 |
-
| macro P@R95 | | 0.786 | 0.860 | 0.687 |
|
| 268 |
-
|
| 269 |
-
The linear model lands 0.029 macro average precision below this model and well
|
| 270 |
-
above the previous classifier. Much of the task is lexical on generated text.
|
| 271 |
-
The fine-tune earns its size at high recall on the hard facets: at 95% recall it
|
| 272 |
-
is 9 points more precise on decision and 17 points more precise on procedure
|
| 273 |
-
than the linear model, and those are the operating regions the memory system
|
| 274 |
-
runs in.
|
| 275 |
-
|
| 276 |
-

|
| 277 |
-
|
| 278 |
-
**Zero-shot frontier models.** A score-blind 300-row subsample of the promotion
|
| 279 |
-
panel, 100 rows per family, scored under one shared prompt, output schema, and
|
| 280 |
-
label set; each external model returned one confidence per facet. This sample
|
| 281 |
-
is not independent of the promotion read, and each external configuration was
|
| 282 |
-
run once, so run-to-run variance is not estimated. Subsample support: trap
|
| 283 |
-
213/87, decision 137/163, constraint 271/29, mechanism 194/105 with one
|
| 284 |
-
unresolved, procedure 106/194 (positive/negative).
|
| 285 |
-
|
| 286 |
-
| Model | Setup | Macro AP | Decision AP | Procedure AP | Macro P@R90 |
|
| 287 |
-
|---|---|---:|---:|---:|---:|
|
| 288 |
-
| This model | supervised fine-tune | 0.9671 | 0.9140 | 0.9500 | 0.8872 |
|
| 289 |
-
| GPT-5.4 | zero-shot, high reasoning | 0.9242 | 0.7885 | 0.9177 | 0.8016 |
|
| 290 |
-
| GPT-5.5 | zero-shot, high reasoning | 0.9162 | 0.7691 | 0.9210 | 0.8064 |
|
| 291 |
-
| GPT-5.5 | zero-shot, low reasoning | 0.9082 | 0.7127 | 0.9209 | 0.7924 |
|
| 292 |
-
| GPT-5.4 | zero-shot, low reasoning | 0.8734 | 0.6985 | 0.8337 | 0.6705 |
|
| 293 |
-
| Previous classifier | prior local model | 0.8368 | 0.5992 | 0.8387 | 0.7007 |
|
| 294 |
-
| GPT-5.3-Codex-Spark | zero-shot, low reasoning | 0.8086 | 0.5387 | 0.7684 | 0.7156 |
|
| 295 |
-
| Nemotron 3 Ultra 550B-A55B | zero-shot, reasoning off | 0.7356 | 0.5474 | 0.7163 | 0.6144 |
|
| 296 |
-
| Gemma 4 26B-A4B IT | zero-shot, reasoning off | 0.7103 | 0.4970 | 0.6707 | 0.6144 |
|
| 297 |
-
|
| 298 |
-
Paired macro average-precision differences from this model, with family-grouped
|
| 299 |
-
95% intervals: GPT-5.4 high −0.0428 [−0.0523, −0.0339]; GPT-5.5 high −0.0508
|
| 300 |
-
[−0.0548, −0.0366]; GPT-5.5 low −0.0588 [−0.0666, −0.0442]; GPT-5.4 low
|
| 301 |
-
−0.0937 [−0.1015, −0.0827]. Every interval excludes parity. Raising reasoning
|
| 302 |
-
effort improved GPT-5.4 by 0.0509 [0.0456, 0.0589] and GPT-5.5 by 0.0080
|
| 303 |
-
[0.0014, 0.0204]; it narrows the gap without closing it. Three accepted runs
|
| 304 |
-
of the Spark configuration spanned 0.8086 to 0.8496 — the reported 0.8086 plus
|
| 305 |
-
repeats at 0.8180 and 0.8496 — a spread of 0.041 that nearly
|
| 306 |
-
matches the smallest GPT gap and exceeds every one of the four interval widths,
|
| 307 |
-
which is the reason single runs are flagged above.
|
| 308 |
-
|
| 309 |
-
The open-weight rows are kept for completeness but carry a scoring caveat:
|
| 310 |
-
their confidences were much coarser than the classifier's posteriors, and once
|
| 311 |
-
tied zero scores had to be admitted at high recall their P@R90 fell to the
|
| 312 |
-
sample's macro prevalence. Their deltas (−0.23 and −0.26) are partly an artifact
|
| 313 |
-
of that coarseness.
|
| 314 |
-
|
| 315 |
-
This comparison shows task specialization, not general capability. The
|
| 316 |
-
classifier was trained on roughly 60,000 examples of how the label instrument
|
| 317 |
-
applies the definitions; the external models saw the definitions once. Where a
|
| 318 |
-
frontier model and this classifier disagree on a row, the panel cannot say which
|
| 319 |
-
is right.
|
| 320 |
-
|
| 321 |
-
## How the recipe was chosen
|
| 322 |
-
|
| 323 |
-
Selection ran on generated development panels held out by whole document
|
| 324 |
-
family; none of them carried promotion authority. Numbers are macro average
|
| 325 |
-
precision unless stated.
|
| 326 |
-
|
| 327 |
-
1. **Data first.** On a 1,000-row eight-family development panel (858 fully
|
| 328 |
-
resolved rows, random floor 0.654) with the larger GLiClass backbone, training
|
| 329 |
-
on the predecessor's data alone scored 0.929 and adding more of the operator's
|
| 330 |
-
real documents scored 0.929; adding twelve generated document families scored
|
| 331 |
-
0.968. Two loss-side branches on top of that baseline were rejected: a data
|
| 332 |
-
curriculum tied it (+0.0005, interval spanning zero) with worse calibration,
|
| 333 |
-
and a robust-loss variant lost (−0.0018, interval entirely below zero).
|
| 334 |
-
2. **Targeted reserve.** Appending the independently reviewed generated reserve
|
| 335 |
-
added +0.0025 [+0.0009, +0.0039].
|
| 336 |
-
3. **Context window.** Widening the input from 416 to 768 tokens added +0.0048
|
| 337 |
-
[+0.0028, +0.0068] on development and +0.0033 [+0.0011, +0.0065] on a
|
| 338 |
-
separate held-out set. A counterfactual-projector branch on minimal pairs
|
| 339 |
-
changed nothing against its own replay control (+0.0001) and was dropped; the
|
| 340 |
-
run's window also did not match its draft plan, so it could not have
|
| 341 |
-
authorized a recipe change either way.
|
| 342 |
-
4. **Backbone size.** On the ModernBERT GLiClass pair, with 44,049 training
|
| 343 |
-
rows and a matched recipe on a 5,298-row eight-family panel, the large arm beat
|
| 344 |
-
the base by 0.0042 average precision and 0.0152 at 90% recall [+0.0015,
|
| 345 |
-
+0.0451] — at 95% recall the interval spans zero — but cost 2.6 times
|
| 346 |
-
the artifact bytes, 2.4 times the training time, and 1.7 times peak GPU
|
| 347 |
-
memory, while the base scored 1.8 times as many rows per second. The base was
|
| 348 |
-
selected as the product architecture.
|
| 349 |
-
5. **Backbone variant and window.** On the standard GLiClass Base v3
|
| 350 |
-
(DeBERTa-v3) backbone, 768 tokens scored 0.928 against 0.914 at 512, with
|
| 351 |
-
procedure moving from 0.765 to 0.821. The backbone declares a 512-position
|
| 352 |
-
limit; the 768-token contract is a measured improvement, retained because
|
| 353 |
-
every trial confirmed it.
|
| 354 |
-
6. **Final fit.** The frozen recipe was trained once on all 59,886 rows,
|
| 355 |
-
calibrated, compressed, and read once on the promotion panel.
|
| 356 |
-
|
| 357 |
-
What did not help: more real documents without generated diversity, the data
|
| 358 |
-
curriculum, the robust loss, the counterfactual projector, and the larger
|
| 359 |
-
backbone at its cost. A distribution-balanced loss, R-Drop, SMART smoothness, a
|
| 360 |
-
three-state auxiliary head, and a 512-token window were also measured and
|
| 361 |
-
rejected, and the experiment log carries each with its numbers. Two further
|
| 362 |
-
branches closed without a scored comparison: a frozen large-NLI recipe, stopped
|
| 363 |
-
after 23.2 hours with no completed epoch, and a decision-only rationale
|
| 364 |
-
auxiliary, which trained four epochs and reduced all three training losses but
|
| 365 |
-
failed artifact assembly on incomplete checkpoint-average candidates, so no
|
| 366 |
-
panel was mounted and no per-row scores were emitted. It is retired as
|
| 367 |
-
inconclusive. Better and more diverse examples, boundary cases, and a wider
|
| 368 |
-
context did the work.
|
| 369 |
-
|
| 370 |
-
## Runtime
|
| 371 |
-
|
| 372 |
-
The shipped artifact is the INT8 bundle. Quantization was qualified as a compression step, not
|
| 373 |
-
a retrain: the FP32 export and the INT8 export were scored on the same 5,298-row eight-family
|
| 374 |
-
panel and compared directly.
|
| 375 |
-
|
| 376 |
-
| Item | Value |
|
| 377 |
-
|---|---|
|
| 378 |
-
| FP32 export | 747,319,179 bytes |
|
| 379 |
-
| INT8 export | 452,812,018 bytes, 39.4% smaller |
|
| 380 |
-
| Quantization | signed INT8 weights, per-channel off, reduce-range off, one operator type; 24 constant-identity nodes folded |
|
| 381 |
-
| Toolchain | onnx 1.21.0, onnxruntime 1.24.4, opset 17 |
|
| 382 |
-
| Macro average-precision change, INT8 minus FP32 | +0.000119, family-grouped 95% interval [-0.000038, +0.000187] |
|
| 383 |
-
| Posterior absolute error, INT8 against FP32 | mean 0.00088, p95 0.00483, maximum 0.0516 |
|
| 384 |
-
| Production adapter against the reference scorer | maximum posterior delta 9.06e-6 against an allowed 1e-4; zero stamp flips under both threshold tables across 1,300 rows |
|
| 385 |
-
| Local CPU scoring rate | 3.74 passages per second at batch 16 over 5,298 rows, peak resident set 4.18 GB |
|
| 386 |
-
|
| 387 |
-
The compression interval spans zero, so INT8 is not measurably worse than FP32 on that panel.
|
| 388 |
-
That read is a diagnostic: it carries no promotion authority and set no thresholds.
|
| 389 |
-
|
| 390 |
-
The scoring rate above comes from the run that produced the compression score receipt
|
| 391 |
-
`5505e117…`, on this artifact, through the production CPU path, as an aggregate over a whole
|
| 392 |
-
panel.
|
| 393 |
-
|
| 394 |
-
Provider qualification (2026-09-06, ONNX Runtime 1.24.4, one NVIDIA RTX 3060 Ti with 8 GB)
|
| 395 |
-
ran the same graph through the production adapter on three execution providers, with CUDA's
|
| 396 |
-
TensorFloat-32 disabled (enabled, posteriors drifted 7.5e-4 from the CPU reference). Every
|
| 397 |
-
matrix product stays on the accelerator; the only CPU-resident work under CUDA or DirectML
|
| 398 |
-
is shape bookkeeping and the rank-0 scalars of the attention scale.
|
| 399 |
-
|
| 400 |
-
| Provider | Posterior delta vs CPU (418 passages) | Stamp flips | Single passage p50 | Batched rate |
|
| 401 |
-
|---|---|---|---|---|
|
| 402 |
-
| CPU | reference | 0 | 331 ms | 1.5 passages/s |
|
| 403 |
-
| CUDA | 2e-6 | 0 | 19 ms | 56 passages/s |
|
| 404 |
-
| DirectML | 3e-6 | 0 | 74 ms | 33 passages/s |
|
| 405 |
-
|
| 406 |
-
Batched rates use eight-row batches planned against the device's free memory: throughput is
|
| 407 |
-
flat beyond four rows on both accelerators, and larger batches at the 768-token ceiling exceed
|
| 408 |
-
8 GB and page silently on Windows. Device memory per row at 768 tokens is about 208 MiB on
|
| 409 |
-
CUDA and 277 MiB on DirectML above a resident session of 657 MiB and 481 MiB respectively.
|
| 410 |
-
|
| 411 |
-
## Evidence status
|
| 412 |
-
|
| 413 |
-
| Evidence | Status |
|
| 414 |
-
|---|---|
|
| 415 |
-
| Promotion results, gates, and calibration | as recorded from the sealed 2026-08-12 read |
|
| 416 |
-
| Artifact identity, size, and quantization settings | verified 2026-09-04 against the quantization receipt |
|
| 417 |
-
| Compression equivalence | as recorded; diagnostic authority only |
|
| 418 |
-
| Production-adapter parity | as recorded from the promotion read |
|
| 419 |
-
| Linear baseline | re-run 2026-09-04 from the recorded inputs; byte-identical |
|
| 420 |
-
| Figures | regenerated 2026-09-04 from the recorded inputs; byte-identical |
|
| 421 |
-
| Previous-classifier identity | verified 2026-09-04 against the staged bundle's own config |
|
| 422 |
-
| Frontier-model comparisons | as recorded; single runs on a subsample that is not independent of the promotion read |
|
| 423 |
-
| Serving throughput | 3.74 passages per second, as recorded from the scoring run bound to the compression score receipt |
|
| 424 |
-
| Per-passage latency percentiles | single-passage p50 per provider only; no p95 or p99 |
|
| 425 |
-
|
| 426 |
-
## Limitations, ranked
|
| 427 |
-
|
| 428 |
-
1. **Labels are model judgments.** All scores measure agreement with a frozen
|
| 429 |
-
frontier-model instrument. There is no human-adjudicated answer key, by the
|
| 430 |
-
operator's explicit decision after the hand-labeling trial described above.
|
| 431 |
-
Resolving evidence would be a human-adjudicated sample; none is planned.
|
| 432 |
-
2. **Panels are generated text.** Absolute performance on unrelated real
|
| 433 |
-
documents is unmeasured for this artifact. Resolving evidence would be a
|
| 434 |
-
multi-domain real-document panel labeled under the same protocol.
|
| 435 |
-
3. **Three families in the promotion panel.** The reported intervals capture
|
| 436 |
-
sampling noise within three similar generated families, not variation across
|
| 437 |
-
domains. More families would widen and honest-size the intervals.
|
| 438 |
-
4. **No seed replicate.** Every winner was trained once. Hardware and compute
|
| 439 |
-
budget, not design, set that limit; the smallest selection margins above
|
| 440 |
-
(0.0042 between backbones, 0.0005 for the curriculum) are within a plausible
|
| 441 |
-
seed effect.
|
| 442 |
-
5. **Lexical headroom is thin.** A linear model sits 0.029 macro average
|
| 443 |
-
precision below this model. The advantage is real but lives at high recall
|
| 444 |
-
on decision and procedure, and should be read that way.
|
| 445 |
-
6. **Frontier comparisons are single runs** on a non-independent subsample with
|
| 446 |
-
coarse external confidences.
|
| 447 |
-
7. **Decision stays the hardest facet.** 398 of the 1,300 promotion rows mention
|
| 448 |
-
a decision without making one; that boundary is where the remaining error
|
| 449 |
-
concentrates.
|
| 450 |
-
8. **Serving cost is measured on one machine.** The provider table in the runtime
|
| 451 |
-
section comes from a single 8 GB NVIDIA card and a four-core CPU budget; other
|
| 452 |
-
hosts will land elsewhere, and cold-start cost is not separated from warm batches.
|
| 453 |
-
9. **Second sealed campaign in its family.** An earlier fit was read on a
|
| 454 |
-
sealed panel, returned NO-GO on a per-origin recall gate, and the contract was
|
| 455 |
-
then amended so origin slices are diagnostics rather than independent vetoes.
|
| 456 |
-
That fit's calibration and promotion rows were folded into this model's own
|
| 457 |
-
training set, and this campaign therefore grants its panel promotion authority
|
| 458 |
-
while explicitly declining a claim to a pristine sealed alpha-spending draw.
|
| 459 |
-
The earlier panel carried 333 operator-workspace rows; this one is entirely
|
| 460 |
-
generated, so the per-origin recall that failed on the earlier read cannot be
|
| 461 |
-
measured on this model at all. Resolving evidence would be a promotion panel
|
| 462 |
-
drawn from unspent reserve that includes real rows.
|
| 463 |
-
|
| 464 |
-
## Reproducibility
|
| 465 |
-
|
| 466 |
-
- Base model: `knowledgator/gliclass-base-v3.0`, revision pinned in
|
| 467 |
-
`gliclass_final_fit_campaign.json`.
|
| 468 |
-
- Training rows: 59,886; manifest and lane digests in
|
| 469 |
-
`gliclass_final_fit_campaign.json` and `history/gliclass_std_base_v3_5facet_promotion_results.json`.
|
| 470 |
-
- Promotion protocol: `gliclass_std_base_v3_5facet_promotion.json` (frozen
|
| 471 |
-
2026-08-12, seed 20260812, 10,000 family-cluster draws). Three digests in the
|
| 472 |
-
results record do not reproduce from anything published or retained.
|
| 473 |
-
`promotion_plan_sha256` describes the private sealed plan, not this tracked
|
| 474 |
-
public-safe copy. `panel_sha256` and `predictions_sha256` were transcribed
|
| 475 |
-
faithfully from the sealed promotion report, but the private panel and
|
| 476 |
-
predictions files that survive today hash to different values, so the exact
|
| 477 |
-
bytes the sealed read scored are gone. The results themselves are the record.
|
| 478 |
-
- Calibration and thresholds: `gliclass_std_base_v3_5facet_calibration.json`,
|
| 479 |
-
`gliclass_std_base_v3_5facet_thresholds.json`; fitted values in the promotion
|
| 480 |
-
results record.
|
| 481 |
-
- Artifact: `model.onnx` SHA-256
|
| 482 |
-
`690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c`,
|
| 483 |
-
452,812,018 bytes; candidate manifest
|
| 484 |
-
`ed8c1f8febe961624a698c97a3035e8ca5caa582718fd5a8f93f962c8106b942`.
|
| 485 |
-
- Baselines: `history/linear_baseline_results.json`, reproduced byte-identical
|
| 486 |
-
from the promoted training rows and the frozen evaluation lane with
|
| 487 |
-
`python -m classifier.scripts.qualification.linear_baseline --train
|
| 488 |
-
<private>/classifier/campaigns/gliclass_standard_base_v3_final_fit_v2_training/train.jsonl
|
| 489 |
-
--panel <private>/source_allocations/classifier_final_repartition_v1/lanes/classifier-evaluation.jsonl`.
|
| 490 |
-
The external-model protocols and results sit under `history/`. Those
|
| 491 |
-
comparison plans pin the SHA-256 of the four scripts that produced them, and
|
| 492 |
-
two of the four — `eval/scripts/decision_grade_runner.py` and
|
| 493 |
-
`eval/scripts/isolated_codex.py` — have changed since. The plans therefore no
|
| 494 |
-
longer load and are sealed records of what was run, not re-runnable commands.
|
| 495 |
-
- Figures: all three regenerate byte-identical with
|
| 496 |
-
`python eval/classifier/scripts/qualification/figures.py` against the promotion
|
| 497 |
-
run's private `predictions.json`, `panel.json`, and `linear-baseline-scores.json`.
|
| 498 |
-
- Entry point: `python -m classifier.scripts.promotion.evaluation`, which
|
| 499 |
-
reproduces the plan's live bindings and verifies each against its recorded
|
| 500 |
-
digest.
|
| 501 |
-
|
| 502 |
-
## License and attribution
|
| 503 |
-
|
| 504 |
-
This derivative is published under the Apache License, Version 2.0, following the
|
| 505 |
-
upstream model. It is a modified derivative of
|
| 506 |
-
[`knowledgator/gliclass-base-v3.0`](https://huggingface.co/knowledgator/gliclass-base-v3.0)
|
| 507 |
-
(Apache-2.0, revision `77a70e6cd52e602ed18184ef37d18bdd3741e3d5`), which builds on
|
| 508 |
-
[`microsoft/deberta-v3-base`](https://huggingface.co/microsoft/deberta-v3-base) (MIT).
|
| 509 |
-
Neither upstream author endorses this derivative. The changes — a supervised five-facet
|
| 510 |
-
fine-tune, an ONNX export, INT8 quantization of the token-embedding table, and per-facet
|
| 511 |
-
calibration — are listed in the `NOTICE` file that accompanies this model.
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: knowledgator/gliclass-base-v3.0
|
| 4 |
+
pipeline_tag: text-classification
|
| 5 |
+
language:
|
| 6 |
+
- en
|
| 7 |
+
tags:
|
| 8 |
+
- gliclass
|
| 9 |
+
- deberta-v3
|
| 10 |
+
- multi-label-classification
|
| 11 |
+
- fine-tuned
|
| 12 |
+
- onnx
|
| 13 |
+
- int8
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# Five-facet passage classifier
|
| 17 |
+
|
| 18 |
+
A GLiClass Base v3 fine-tune that describes what an English passage contains: a **trap, decision, constraint, mechanism or procedure**. A passage can have several facets. The model produces all five scores in one pass and runs locally through ONNX Runtime.
|
| 19 |
+
|
| 20 |
+
This revision's ONNX graph adds Vulkan compatibility. Learned weights, prompts, calibration and thresholds are unchanged.
|
| 21 |
+
|
| 22 |
+
On 1,300 generated passages from three document families excluded from training, macro average precision is **0.9676** for the fine-tune, 0.9384 for a TF-IDF baseline trained on the same rows, and 0.7709 for the zero-shot upstream checkpoint.
|
| 23 |
+
|
| 24 |
+
## Quick start
|
| 25 |
+
|
| 26 |
+
Install `onnxruntime==1.24.4`, `tokenizers` and `numpy`. This CPU example assumes the package is in `downloaded-model` and uses its published prompts and calibration rather than new label wording:
|
| 27 |
+
|
| 28 |
+
```python
|
| 29 |
+
import json
|
| 30 |
+
from pathlib import Path
|
| 31 |
+
import numpy as np
|
| 32 |
+
import onnxruntime as ort
|
| 33 |
+
from tokenizers import Tokenizer
|
| 34 |
+
|
| 35 |
+
root = Path("downloaded-model")
|
| 36 |
+
metadata = json.load(open(f"{root}/classifier-metadata.json", encoding="utf-8"))
|
| 37 |
+
facets = metadata["facet_order"]
|
| 38 |
+
prep = metadata["preprocessing"]
|
| 39 |
+
prefix = "".join(
|
| 40 |
+
prep["label_token"] + prep["label_prompts"][facet] for facet in facets
|
| 41 |
+
) + prep["separator_token"]
|
| 42 |
+
tokenizer = Tokenizer.from_file(f"{root}/tokenizer.json")
|
| 43 |
+
tokenizer.enable_truncation(max_length=prep["max_length"])
|
| 44 |
+
encoded = tokenizer.encode(prefix + "We chose weekly releases to reduce rollout risk.")
|
| 45 |
+
session = ort.InferenceSession(
|
| 46 |
+
f"{root}/model.onnx", providers=["CPUExecutionProvider"]
|
| 47 |
+
)
|
| 48 |
+
logits = session.run(["logits"], {
|
| 49 |
+
"input_ids": np.array([encoded.ids], dtype=np.int64),
|
| 50 |
+
"attention_mask": np.array([encoded.attention_mask], dtype=np.int64),
|
| 51 |
+
})[0][0].astype(np.float64)
|
| 52 |
+
calibration = metadata["probability_calibration"]
|
| 53 |
+
clip = calibration["probability_clip"]
|
| 54 |
+
raw = np.clip(1 / (1 + np.exp(-np.clip(logits, -80, 80))), clip, 1 - clip)
|
| 55 |
+
temperatures = np.array([calibration["temperatures"][f] for f in facets])
|
| 56 |
+
probabilities = 1 / (1 + np.exp(-np.log(raw / (1 - raw)) / temperatures))
|
| 57 |
+
print(dict(zip(facets, probabilities.tolist())))
|
| 58 |
+
```
|
| 59 |
+
|
| 60 |
+
The output is five independent calibrated scores, not a distribution that sums to one. For applications that need labels, the metadata supplies two threshold tables: `contract` favors precision and `recall_leaning` retains more candidates. Choose between them by the relative cost of missed and incorrect labels.
|
| 61 |
+
|
| 62 |
+
## What the labels mean
|
| 63 |
+
|
| 64 |
+
| Facet | The passage… |
|
| 65 |
+
|---|---|
|
| 66 |
+
| `trap` | Warns about a specific mistake, hazard or failure mode |
|
| 67 |
+
| `decision` | Records a choice or commitment among alternatives |
|
| 68 |
+
| `constraint` | States a rule or condition that materially restricts action |
|
| 69 |
+
| `mechanism` | Explains how something is organized, connected or works |
|
| 70 |
+
| `procedure` | Gives sequenced actions for carrying out a task |
|
| 71 |
+
|
| 72 |
+
Mentioning a decision is different from making one; training labels treat “mentions only” as negative for that facet. Facets overlap: a procedure may also contain a constraint and warn about a trap. The scores help organize passages or inspect what a search returns. They do not establish relevance, truth or authority and should not serve as a hard retrieval filter.
|
| 73 |
+
|
| 74 |
+
## Before and after fine-tuning
|
| 75 |
+
|
| 76 |
+
“Upstream” is the exact `knowledgator/gliclass-base-v3.0` checkpoint used to start training, applied zero-shot with the same label definitions and **no Daecore fine-tuning**. The word-feature baseline is TF-IDF plus logistic regression, trained on the same labeled rows as the fine-tune. All three use the same 1,300-passage panel; unresolved labels are excluded per facet for every model.
|
| 77 |
+
|
| 78 |
+
| Metric, macro average over five facets | Upstream GLiClass | Word-feature baseline | Daecore fine-tune |
|
| 79 |
+
|---|---:|---:|---:|
|
| 80 |
+
| Average precision | 0.7709 | 0.9384 | **0.9676** |
|
| 81 |
+
| Precision at 90% recall | 66.11% | 83.35% | **90.21%** |
|
| 82 |
+
| Precision at 95% recall | 64.51% | 78.62% | **86.02%** |
|
| 83 |
+
|
| 84 |
+
Average precision summarizes how well scores rank positives across thresholds. Precision at 90% recall is the cleanest measured prediction set that still keeps at least 90% of positive labels; it comes from this evaluation's curve, not from a threshold chosen in advance.
|
| 85 |
+
|
| 86 |
+
| Facet | Resolved passages | Positive prevalence | Upstream AP | Fine-tuned AP |
|
| 87 |
+
|---|---:|---:|---:|---:|
|
| 88 |
+
| Trap | 1,300 | 71.77% | 0.9028 | **0.9864** |
|
| 89 |
+
| Decision | 1,299 | 46.04% | 0.5219 | **0.9135** |
|
| 90 |
+
| Constraint | 1,300 | 91.31% | 0.9544 | **0.9986** |
|
| 91 |
+
| Mechanism | 1,297 | 68.77% | 0.7692 | **0.9817** |
|
| 92 |
+
| Procedure | 1,293 | 37.51% | 0.7064 | **0.9578** |
|
| 93 |
+
|
| 94 |
+

|
| 95 |
+
|
| 96 |
+
As the baseline shows, much of this task can be learned from word cues; the fine-tune's clearest benefit is at high recall. Constraint is common in this panel, so its near-perfect AP says less than the decision and procedure results. The panel was deliberately enriched for difficult decision and procedure cases and is not a sample of natural traffic. The [evaluation companion](evaluation/README.md) contains the text-free labels and scores, model identities and calculation code.
|
| 97 |
+
|
| 98 |
+
## Data and training choices
|
| 99 |
+
|
| 100 |
+
The final fit used **59,886 passages from 9,376 documents**: 54,803 generated and 5,083 real. The real portion is 4,583 passages from self-owned workspaces and 500 from public-domain US Federal Register documents; all real rows were used for training, none for evaluation. Workspace documents may also be AI-written; real describes their source, not human authorship. Generated material was written as complete organizational documents and then parsed into passages, covering more settings and document types than the available workspaces.
|
| 101 |
+
|
| 102 |
+
Models applied fixed facet definitions to produce the labels. The main label campaigns used independent judgments with conflict resolution; 3,750 inherited rows used a single primary judge with a separate audit. Splits separate whole source families: the three evaluation families and four calibration families are disjoint from training and from each other. Calibration uses 1,700 passages to set per-facet temperatures and the two threshold tables.
|
| 103 |
+
|
| 104 |
+
| Training choice | Value and reason |
|
| 105 |
+
|---|---|
|
| 106 |
+
| Starting model | GLiClass Base v3 on DeBERTa-v3, 186.5M parameters |
|
| 107 |
+
| Input length | 768 tokens including label prompts; context-length experiments favored it over 512 |
|
| 108 |
+
| Objective | Weighted binary cross-entropy; mentions-only negatives receive 2× weight |
|
| 109 |
+
| Optimizer schedule | Learning rate 2e-5; batch 4, accumulated over 8 steps |
|
| 110 |
+
| Duration | Four fixed epochs; equal-weight average of epochs 2, 3 and 4, chosen before the final fit |
|
| 111 |
+
| Serving compression | INT8 token embeddings with the transformer body in FP32 |
|
| 112 |
+
|
| 113 |
+
Quantizing only the token embeddings shrinks the large vocabulary table while keeping transformer calculations in FP32; broader quantization reduced quality. The ONNX graph is **452.8 MB**, and the package is about **461.5 MB** excluding runtime libraries.
|
| 114 |
+
|
| 115 |
+
## Runtime and limits
|
| 116 |
+
|
| 117 |
+
On the 1,300 evaluation passages, the derived graph produced the same labels under both threshold tables on CPU, CUDA and Vulkan as the previous graph on CPU, with maximum score differences below 5.7e-6. Vulkan uses `onnxruntime==1.24.4` with the native WebGPU plugin [`onnxruntime-ep-webgpu==0.4.0`](https://pypi.org/project/onnxruntime-ep-webgpu/), registered explicitly, with `dawnBackendType=Vulkan` and no providers list at session creation. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
|
| 118 |
+
|
| 119 |
+
- The model is trained and evaluated only on English text.
|
| 120 |
+
- Labels are model judgments without human validation.
|
| 121 |
+
- The evaluation covers generated text from three held-out families and was reused during development; performance on unrelated real documents is unmeasured.
|
| 122 |
+
- Training was run once; variation across seeds is unmeasured.
|
| 123 |
+
|
| 124 |
+
## License
|
| 125 |
+
|
| 126 |
+
Apache-2.0. The package includes upstream attribution and a modification notice.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
classifier-metadata.json
CHANGED
|
@@ -1,78 +1,78 @@
|
|
| 1 |
-
{
|
| 2 |
-
"architecture": "gliclass-single-pass",
|
| 3 |
-
"classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
|
| 4 |
-
"facet_order": [
|
| 5 |
-
"trap",
|
| 6 |
-
"decision",
|
| 7 |
-
"constraint",
|
| 8 |
-
"mechanism",
|
| 9 |
-
"procedure"
|
| 10 |
-
],
|
| 11 |
-
"files": {
|
| 12 |
-
"config.json": {
|
| 13 |
-
"sha256": "9bad60c87ac3e047bc838ca7e6897a6fe265180eeffce4e62aabc1e9c828d606",
|
| 14 |
-
"size": 2118
|
| 15 |
-
},
|
| 16 |
-
"model.onnx": {
|
| 17 |
-
"sha256": "
|
| 18 |
-
"size":
|
| 19 |
-
},
|
| 20 |
-
"tokenizer.json": {
|
| 21 |
-
"sha256": "519648948c4c59da1af88f2cf2c8b4f84417b5c673981bc9809abf84cda1b7cc",
|
| 22 |
-
"size": 8649234
|
| 23 |
-
},
|
| 24 |
-
"tokenizer_config.json": {
|
| 25 |
-
"sha256": "9d5220b355d2cb9a7df69deccca45509ce1c9a857990fb1047c48403b0f6ddfb",
|
| 26 |
-
"size": 1692
|
| 27 |
-
}
|
| 28 |
-
},
|
| 29 |
-
"preprocessing": {
|
| 30 |
-
"label_prompts": {
|
| 31 |
-
"constraint": "This text states a rule, policy, convention, invariant, or factual condition that materially restricts action.",
|
| 32 |
-
"decision": "This text records a decision or commitment that chooses one option over alternatives.",
|
| 33 |
-
"mechanism": "This text explains how a system, process, component, or structure is organized, connected, or works.",
|
| 34 |
-
"procedure": "This text gives sequenced actions or steps for carrying out a task or operation.",
|
| 35 |
-
"trap": "This text warns about a specific mistake, hazard, pitfall, or failure mode."
|
| 36 |
-
},
|
| 37 |
-
"label_token": "<<LABEL>>",
|
| 38 |
-
"max_length": 768,
|
| 39 |
-
"prompt_first": true,
|
| 40 |
-
"schema": "daecore.gliclass-preprocessing.v1",
|
| 41 |
-
"separator_token": "<<SEP>>"
|
| 42 |
-
},
|
| 43 |
-
"probability_calibration": {
|
| 44 |
-
"method": "per-facet-temperature-scaling",
|
| 45 |
-
"probability_clip": 1e-06,
|
| 46 |
-
"temperatures": {
|
| 47 |
-
"constraint": 1.6869790276494951,
|
| 48 |
-
"decision": 2.4036524709979896,
|
| 49 |
-
"mechanism": 1.501092485229896,
|
| 50 |
-
"procedure": 2.0450110042733725,
|
| 51 |
-
"trap": 1.8898163065802225
|
| 52 |
-
}
|
| 53 |
-
},
|
| 54 |
-
"schema_version": 3,
|
| 55 |
-
"source_candidate_sha256": "ed8c1f8febe961624a698c97a3035e8ca5caa582718fd5a8f93f962c8106b942",
|
| 56 |
-
"threshold_tables": {
|
| 57 |
-
"contract": {
|
| 58 |
-
"values": {
|
| 59 |
-
"constraint": 0.9928175,
|
| 60 |
-
"decision": 0.48485,
|
| 61 |
-
"mechanism": 0.8998575,
|
| 62 |
-
"procedure": 0.5776285,
|
| 63 |
-
"trap": 0.9759885
|
| 64 |
-
},
|
| 65 |
-
"version": "gliclass-std-base-v3-daecore-5facet-qint8-v2-contract"
|
| 66 |
-
},
|
| 67 |
-
"recall_leaning": {
|
| 68 |
-
"values": {
|
| 69 |
-
"constraint": 0.9059865,
|
| 70 |
-
"decision": 0.192514,
|
| 71 |
-
"mechanism": 0.5053545,
|
| 72 |
-
"procedure": 0.201766,
|
| 73 |
-
"trap": 0.715311
|
| 74 |
-
},
|
| 75 |
-
"version": "gliclass-std-base-v3-daecore-5facet-qint8-v2-recall"
|
| 76 |
-
}
|
| 77 |
-
}
|
| 78 |
-
}
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architecture": "gliclass-single-pass",
|
| 3 |
+
"classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
|
| 4 |
+
"facet_order": [
|
| 5 |
+
"trap",
|
| 6 |
+
"decision",
|
| 7 |
+
"constraint",
|
| 8 |
+
"mechanism",
|
| 9 |
+
"procedure"
|
| 10 |
+
],
|
| 11 |
+
"files": {
|
| 12 |
+
"config.json": {
|
| 13 |
+
"sha256": "9bad60c87ac3e047bc838ca7e6897a6fe265180eeffce4e62aabc1e9c828d606",
|
| 14 |
+
"size": 2118
|
| 15 |
+
},
|
| 16 |
+
"model.onnx": {
|
| 17 |
+
"sha256": "7e9b7e37afc06a23a89ba064e1255363cdc3140ff4b2823f5dae5865945621dc",
|
| 18 |
+
"size": 452831012
|
| 19 |
+
},
|
| 20 |
+
"tokenizer.json": {
|
| 21 |
+
"sha256": "519648948c4c59da1af88f2cf2c8b4f84417b5c673981bc9809abf84cda1b7cc",
|
| 22 |
+
"size": 8649234
|
| 23 |
+
},
|
| 24 |
+
"tokenizer_config.json": {
|
| 25 |
+
"sha256": "9d5220b355d2cb9a7df69deccca45509ce1c9a857990fb1047c48403b0f6ddfb",
|
| 26 |
+
"size": 1692
|
| 27 |
+
}
|
| 28 |
+
},
|
| 29 |
+
"preprocessing": {
|
| 30 |
+
"label_prompts": {
|
| 31 |
+
"constraint": "This text states a rule, policy, convention, invariant, or factual condition that materially restricts action.",
|
| 32 |
+
"decision": "This text records a decision or commitment that chooses one option over alternatives.",
|
| 33 |
+
"mechanism": "This text explains how a system, process, component, or structure is organized, connected, or works.",
|
| 34 |
+
"procedure": "This text gives sequenced actions or steps for carrying out a task or operation.",
|
| 35 |
+
"trap": "This text warns about a specific mistake, hazard, pitfall, or failure mode."
|
| 36 |
+
},
|
| 37 |
+
"label_token": "<<LABEL>>",
|
| 38 |
+
"max_length": 768,
|
| 39 |
+
"prompt_first": true,
|
| 40 |
+
"schema": "daecore.gliclass-preprocessing.v1",
|
| 41 |
+
"separator_token": "<<SEP>>"
|
| 42 |
+
},
|
| 43 |
+
"probability_calibration": {
|
| 44 |
+
"method": "per-facet-temperature-scaling",
|
| 45 |
+
"probability_clip": 1e-06,
|
| 46 |
+
"temperatures": {
|
| 47 |
+
"constraint": 1.6869790276494951,
|
| 48 |
+
"decision": 2.4036524709979896,
|
| 49 |
+
"mechanism": 1.501092485229896,
|
| 50 |
+
"procedure": 2.0450110042733725,
|
| 51 |
+
"trap": 1.8898163065802225
|
| 52 |
+
}
|
| 53 |
+
},
|
| 54 |
+
"schema_version": 3,
|
| 55 |
+
"source_candidate_sha256": "ed8c1f8febe961624a698c97a3035e8ca5caa582718fd5a8f93f962c8106b942",
|
| 56 |
+
"threshold_tables": {
|
| 57 |
+
"contract": {
|
| 58 |
+
"values": {
|
| 59 |
+
"constraint": 0.9928175,
|
| 60 |
+
"decision": 0.48485,
|
| 61 |
+
"mechanism": 0.8998575,
|
| 62 |
+
"procedure": 0.5776285,
|
| 63 |
+
"trap": 0.9759885
|
| 64 |
+
},
|
| 65 |
+
"version": "gliclass-std-base-v3-daecore-5facet-qint8-v2-contract"
|
| 66 |
+
},
|
| 67 |
+
"recall_leaning": {
|
| 68 |
+
"values": {
|
| 69 |
+
"constraint": 0.9059865,
|
| 70 |
+
"decision": 0.192514,
|
| 71 |
+
"mechanism": 0.5053545,
|
| 72 |
+
"procedure": 0.201766,
|
| 73 |
+
"trap": 0.715311
|
| 74 |
+
},
|
| 75 |
+
"version": "gliclass-std-base-v3-daecore-5facet-qint8-v2-recall"
|
| 76 |
+
}
|
| 77 |
+
}
|
| 78 |
+
}
|
evaluation/README.md
ADDED
|
@@ -0,0 +1,83 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Model evaluation records
|
| 2 |
+
|
| 3 |
+
These files let readers recompute the comparisons in the Daecore model cards without model weights. Task-specific records contain anonymized row and group IDs, labels, and scores or ranked grades; none contains query or document text, private source identifiers or workspace paths. Model identities are recorded inside each file.
|
| 4 |
+
|
| 5 |
+
## Recompute
|
| 6 |
+
|
| 7 |
+
Python 3.11 or later is enough; no packages or network access are needed. From the directory that holds the records:
|
| 8 |
+
|
| 9 |
+
```sh
|
| 10 |
+
python metrics.py --directory .
|
| 11 |
+
```
|
| 12 |
+
|
| 13 |
+
The command summarizes every record present and stops with an error if there is none. Each entry in its output equals the matching `*-summary.json`. The classifier and Ettin packages also include `figures.py`; `python figures.py --output <directory>` redraws their card chart. Gemma uses a table. Each model package contains only its own records and summaries; the source repository holds all four record types. Fractions are stored at full precision; cards round percentages to two decimals and ranking metrics to four.
|
| 14 |
+
|
| 15 |
+
The records support metric recomputation. Re-running inference would need the private text and corpus, which are not included. They come from task-specific development evaluations, not a new untouched test.
|
| 16 |
+
|
| 17 |
+
## Records
|
| 18 |
+
|
| 19 |
+
| Record | Inputs and denominator | Comparison |
|
| 20 |
+
|---|---|---|
|
| 21 |
+
| `classifier.json` | 1,300 generated passages from three held-out source families; resolved labels only, per facet | Upstream GLiClass, a trained word-feature baseline and the fine-tune |
|
| 22 |
+
| `reranker.json` | 970 fixed pools of 50 passages, all 48,500 pairs judged; 873 pools have useful evidence | Upstream Ettin and the fine-tune, plus the expected result of random ordering |
|
| 23 |
+
| `promotion.json` | 970 hybrid-search queries over 82,719 passage texts, with reviewed labels and declared exclusions; five public dense-retrieval panels | Previous and updated Gemma with unchanged Ettin; the public panels also include upstream Gemma |
|
| 24 |
+
| `serving.json` | The same 970 queries with updated-Gemma candidate pools; reference extended by 32 grades | Ettin through CUDA FP16 and Vulkan FP32, with identical weights and score mapping |
|
| 25 |
+
|
| 26 |
+
`serving-qualification.json` summarizes provider and recovery checks by receipt hash and keeps the aggregate FiQA and SciFact results for upstream and fine-tuned Ettin.
|
| 27 |
+
|
| 28 |
+
“Upstream” means no Daecore fine-tuning, not an untrained network; upstream Ettin is already a trained reranker. Interim training checkpoints are not included.
|
| 29 |
+
|
| 30 |
+
## How each comparison was run
|
| 31 |
+
|
| 32 |
+
**Classifier.** Evaluation families were excluded from training. Upstream uses the same five label definitions and 768-token input limit as the fine-tune. Its raw logits and the fine-tune's calibrated probabilities are used only to rank within each facet; their scales are not compared. All 1,300 saved predictions matched fresh CPU inference of the released graph within 5.62e-6. The baseline uses word unigram and bigram TF-IDF with one balanced logistic-regression model per facet, fitted on the same 59,886 training passages; it sees full passage text, while the neural models apply their 768-token limit.
|
| 33 |
+
|
| 34 |
+
**Ettin, fixed pools.** Candidates and grades are held fixed. Both models run in PyTorch FP16 with the same 1,153-token pair construction, Transformers 5.2.0 and Sentence Transformers 5.5.1; upstream's metadata names a newer library version, but both use the same reference runtime here. The fine-tune's saved scores were checked against fresh inference on three complete pools. This evaluates the trained models, not ONNX serving or latency. The public FiQA and SciFact results rerank fixed 50-candidate pools from upstream Gemma.
|
| 35 |
+
|
| 36 |
+
**Gemma update.** Only Gemma changes; BM25, fusion and Ettin's weights stay fixed. All 970 queries are retained, and seven reviewed grade corrections apply to both models. Within each original top-k or selected prefix, 65 unresolved query–passage abstentions are excluded without backfilling: `ranked_grades` keeps their positions as `null`, with the required `excluded` flag. Precision is pooled over retained positions; nDCG reindexes them and takes its ideal ordering at the retained depth from `reference_grade_counts`. Selected depths refer to the original score-selected prefixes.
|
| 37 |
+
|
| 38 |
+
| Metric, 970 queries | Previous Gemma | Updated Gemma |
|
| 39 |
+
|---|---:|---:|
|
| 40 |
+
| Hit@3 | 83.92% | 85.36% |
|
| 41 |
+
| Hit@5 | 85.88% | 88.45% |
|
| 42 |
+
| Hit@10 | 87.84% | 89.90% |
|
| 43 |
+
| Hit@20 | 90.10% | 91.03% |
|
| 44 |
+
| nDCG@10 | 0.6764 | 0.6611 |
|
| 45 |
+
| Selected-prefix precision | 80.41% | 80.34% |
|
| 46 |
+
|
| 47 |
+
The same record holds per-query dense nDCG@10 for the five public panels. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); missing baselines stay absent. These sets supplied no training examples but informed development, and recomputing their means is different from rerunning retrieval on the public corpora.
|
| 48 |
+
|
| 49 |
+
**Ettin providers.** Updated-Gemma candidate pools for the same 970 queries were scored through CUDA FP16 and Vulkan FP32 with unchanged weights and score mapping. The two paths' top-20 results included 32 query–passage pairs without a grade. Of these, 21 reuse grades from the hard-contrast Ettin evaluation. The other 11 received two independent GPT-6 Sol judgments plus a resolution step and were then reviewed against the full passage text by GPT-6 Astra; no person reviewed them. No existing grade changed, and the 65 abstentions remain excluded. The added grades slightly change nDCG's ideal ordering, so the promotion record keeps its original reference.
|
| 50 |
+
|
| 51 |
+
| Metric, 970 queries | CUDA FP16 | Vulkan FP32 |
|
| 52 |
+
|---|---:|---:|
|
| 53 |
+
| Hit@3 | 85.36% | 85.26% |
|
| 54 |
+
| Hit@5 | 88.45% | 88.45% |
|
| 55 |
+
| Hit@10 | 89.90% | 89.90% |
|
| 56 |
+
| Hit@20 | 91.03% | 91.03% |
|
| 57 |
+
| nDCG@10 | 0.6611 | 0.6615 |
|
| 58 |
+
| Selected-prefix precision | 80.34% | 80.28% |
|
| 59 |
+
|
| 60 |
+
Small numerical differences between the paths can reorder close scores. On 20 matched pools replayed twice, second-pass reranking took 1.56 s median and 2.11 s at the 95th percentile with Vulkan, against 1.92 s and 2.70 s with the previous package's DirectML graph. All provider measurements come from one Windows x64 machine with an NVIDIA RTX 3060 Ti (8 GB) and do not transfer to other GPUs or platforms.
|
| 61 |
+
|
| 62 |
+
## Metric definitions
|
| 63 |
+
|
| 64 |
+
Relevance grades are 0–3, and grades **2 and 3** count as useful. An “answerable” query has at least one useful labeled passage in the specified pool; the absence of a useful judgment does not prove that no answer exists.
|
| 65 |
+
|
| 66 |
+
- **Hit@k:** fraction of queries with at least one useful passage in the first k positions.
|
| 67 |
+
- **Precision@k:** useful passages among the first k. Fixed-pool tables average it over queries; the promotion and serving records pool it over retained positions. Selected-prefix precision applies the same pooling to all returned passages.
|
| 68 |
+
- **Recall@k:** useful passages retrieved divided by the query's known useful passages, averaged over queries. It counts labeled passages, not every fact an answer needs.
|
| 69 |
+
- **nDCG@k:** gain `2**grade - 1`, discount `1/log2(rank + 1)`, divided by the ideal ordering at the same cutoff. Grade 1 adds a small gain although it is not useful for Hit, precision or recall. Each cutoff has its own ideal, so values need not change monotonically with k.
|
| 70 |
+
- **Average precision (AP):** area under the stepwise precision–recall curve, with tied scores grouped at one threshold; macro AP weights the five facets equally.
|
| 71 |
+
- **Precision at 90% or 95% recall:** the best measured precision at any threshold reaching that recall, taken from the evaluation curve rather than a threshold chosen in advance.
|
| 72 |
+
|
| 73 |
+
Ties keep the original candidate order. Classifier fields marked `null` are excluded identically for every model. Conditional reranker nDCG and recall use the 873 answerable pools; all-query Hit and precision use all 970. The random-order reference is exact within each pool: with N candidates and R useful passages, expected precision is R/N and Hit@k is `1 - C(N-R, k)/C(N, k)`. Random ordering already reaches 97.30% Hit@20 on the answerable reranker pools.
|
| 74 |
+
|
| 75 |
+
## Limits
|
| 76 |
+
|
| 77 |
+
Task-specific labels are language-model judgments without a human-adjudicated reference. Much of the source material is generated. Project-document sources are narrow and often AI-written; they do not establish generalization across users. The panels were reused during development, and each released model was trained once. Several relevant passages can repeat one fact, so passage-level scores do not measure unique-fact coverage, answer completeness or downstream agent success.
|
| 78 |
+
|
| 79 |
+
## Cards
|
| 80 |
+
|
| 81 |
+
- [Five-facet classifier](https://huggingface.co/Daecore/gliclass-std-base-v3-5facet-qint8-v2)
|
| 82 |
+
- [EmbeddingGemma retriever](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v1)
|
| 83 |
+
- [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1)
|
evaluation/classifier-summary.json
ADDED
|
@@ -0,0 +1,138 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"rows": 1300,
|
| 3 |
+
"families": 3,
|
| 4 |
+
"models": {
|
| 5 |
+
"upstream": {
|
| 6 |
+
"facets": {
|
| 7 |
+
"trap": {
|
| 8 |
+
"rows": 1300,
|
| 9 |
+
"average_precision": 0.9027959606064969,
|
| 10 |
+
"precision_at_recall_90": 0.7787037037037037,
|
| 11 |
+
"precision_at_recall_95": 0.7485232067510549,
|
| 12 |
+
"prevalence": 0.7176923076923077
|
| 13 |
+
},
|
| 14 |
+
"decision": {
|
| 15 |
+
"rows": 1299,
|
| 16 |
+
"average_precision": 0.5219192438473815,
|
| 17 |
+
"precision_at_recall_90": 0.46072874493927124,
|
| 18 |
+
"precision_at_recall_95": 0.46072874493927124,
|
| 19 |
+
"prevalence": 0.4603541185527329
|
| 20 |
+
},
|
| 21 |
+
"constraint": {
|
| 22 |
+
"rows": 1300,
|
| 23 |
+
"average_precision": 0.954374886008778,
|
| 24 |
+
"precision_at_recall_90": 0.9130769230769231,
|
| 25 |
+
"precision_at_recall_95": 0.9130769230769231,
|
| 26 |
+
"prevalence": 0.9130769230769231
|
| 27 |
+
},
|
| 28 |
+
"mechanism": {
|
| 29 |
+
"rows": 1297,
|
| 30 |
+
"average_precision": 0.7691508918347779,
|
| 31 |
+
"precision_at_recall_90": 0.6880804953560371,
|
| 32 |
+
"precision_at_recall_95": 0.6880804953560371,
|
| 33 |
+
"prevalence": 0.6877409406322282
|
| 34 |
+
},
|
| 35 |
+
"procedure": {
|
| 36 |
+
"rows": 1293,
|
| 37 |
+
"average_precision": 0.7063681153993947,
|
| 38 |
+
"precision_at_recall_90": 0.4648936170212766,
|
| 39 |
+
"precision_at_recall_95": 0.4153153153153153,
|
| 40 |
+
"prevalence": 0.3750966744006187
|
| 41 |
+
}
|
| 42 |
+
},
|
| 43 |
+
"macro": {
|
| 44 |
+
"average_precision": 0.7709218195393658,
|
| 45 |
+
"precision_at_recall_90": 0.6610966968194424,
|
| 46 |
+
"precision_at_recall_95": 0.6451449370877204
|
| 47 |
+
}
|
| 48 |
+
},
|
| 49 |
+
"linear": {
|
| 50 |
+
"facets": {
|
| 51 |
+
"trap": {
|
| 52 |
+
"rows": 1300,
|
| 53 |
+
"average_precision": 0.9704508234286988,
|
| 54 |
+
"precision_at_recall_90": 0.9231613611416026,
|
| 55 |
+
"precision_at_recall_95": 0.895959595959596,
|
| 56 |
+
"prevalence": 0.7176923076923077
|
| 57 |
+
},
|
| 58 |
+
"decision": {
|
| 59 |
+
"rows": 1299,
|
| 60 |
+
"average_precision": 0.8471593175324644,
|
| 61 |
+
"precision_at_recall_90": 0.6304093567251462,
|
| 62 |
+
"precision_at_recall_95": 0.571,
|
| 63 |
+
"prevalence": 0.4603541185527329
|
| 64 |
+
},
|
| 65 |
+
"constraint": {
|
| 66 |
+
"rows": 1300,
|
| 67 |
+
"average_precision": 0.9969713734939275,
|
| 68 |
+
"precision_at_recall_90": 0.9916743755781684,
|
| 69 |
+
"precision_at_recall_95": 0.9791666666666666,
|
| 70 |
+
"prevalence": 0.9130769230769231
|
| 71 |
+
},
|
| 72 |
+
"mechanism": {
|
| 73 |
+
"rows": 1297,
|
| 74 |
+
"average_precision": 0.9633161957308005,
|
| 75 |
+
"precision_at_recall_90": 0.8865638766519823,
|
| 76 |
+
"precision_at_recall_95": 0.8330058939096268,
|
| 77 |
+
"prevalence": 0.6877409406322282
|
| 78 |
+
},
|
| 79 |
+
"procedure": {
|
| 80 |
+
"rows": 1293,
|
| 81 |
+
"average_precision": 0.914259121155341,
|
| 82 |
+
"precision_at_recall_90": 0.7356902356902357,
|
| 83 |
+
"precision_at_recall_95": 0.652050919377652,
|
| 84 |
+
"prevalence": 0.3750966744006187
|
| 85 |
+
}
|
| 86 |
+
},
|
| 87 |
+
"macro": {
|
| 88 |
+
"average_precision": 0.9384313662682464,
|
| 89 |
+
"precision_at_recall_90": 0.833499841157427,
|
| 90 |
+
"precision_at_recall_95": 0.7862366151827083
|
| 91 |
+
}
|
| 92 |
+
},
|
| 93 |
+
"finetuned": {
|
| 94 |
+
"facets": {
|
| 95 |
+
"trap": {
|
| 96 |
+
"rows": 1300,
|
| 97 |
+
"average_precision": 0.9863618253424686,
|
| 98 |
+
"precision_at_recall_90": 0.96,
|
| 99 |
+
"precision_at_recall_95": 0.934668071654373,
|
| 100 |
+
"prevalence": 0.7176923076923077
|
| 101 |
+
},
|
| 102 |
+
"decision": {
|
| 103 |
+
"rows": 1299,
|
| 104 |
+
"average_precision": 0.9135015910515543,
|
| 105 |
+
"precision_at_recall_90": 0.7414030261348006,
|
| 106 |
+
"precision_at_recall_95": 0.6581986143187067,
|
| 107 |
+
"prevalence": 0.4603541185527329
|
| 108 |
+
},
|
| 109 |
+
"constraint": {
|
| 110 |
+
"rows": 1300,
|
| 111 |
+
"average_precision": 0.9985879374608636,
|
| 112 |
+
"precision_at_recall_90": 0.9962825278810409,
|
| 113 |
+
"precision_at_recall_95": 0.9947183098591549,
|
| 114 |
+
"prevalence": 0.9130769230769231
|
| 115 |
+
},
|
| 116 |
+
"mechanism": {
|
| 117 |
+
"rows": 1297,
|
| 118 |
+
"average_precision": 0.9816644729408911,
|
| 119 |
+
"precision_at_recall_90": 0.9261822376009228,
|
| 120 |
+
"precision_at_recall_95": 0.8908709338929696,
|
| 121 |
+
"prevalence": 0.6877409406322282
|
| 122 |
+
},
|
| 123 |
+
"procedure": {
|
| 124 |
+
"rows": 1293,
|
| 125 |
+
"average_precision": 0.9578217870935172,
|
| 126 |
+
"precision_at_recall_90": 0.8868686868686869,
|
| 127 |
+
"precision_at_recall_95": 0.822380106571936,
|
| 128 |
+
"prevalence": 0.3750966744006187
|
| 129 |
+
}
|
| 130 |
+
},
|
| 131 |
+
"macro": {
|
| 132 |
+
"average_precision": 0.967587522777859,
|
| 133 |
+
"precision_at_recall_90": 0.9021472956970902,
|
| 134 |
+
"precision_at_recall_95": 0.8601672072594281
|
| 135 |
+
}
|
| 136 |
+
}
|
| 137 |
+
}
|
| 138 |
+
}
|
evaluation/classifier.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evaluation/figures.py
ADDED
|
@@ -0,0 +1,89 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Regenerate the classifier and reranker SVGs from the public labels and scores.
|
| 2 |
+
|
| 3 |
+
Run with Python 3.11+: python figures.py --output ../figures
|
| 4 |
+
The repository shares its renderer from eval/lib; HF packages include that
|
| 5 |
+
same renderer next to this script.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from __future__ import annotations
|
| 9 |
+
|
| 10 |
+
import argparse
|
| 11 |
+
import importlib.util
|
| 12 |
+
import json
|
| 13 |
+
import sys
|
| 14 |
+
from pathlib import Path
|
| 15 |
+
|
| 16 |
+
import metrics
|
| 17 |
+
|
| 18 |
+
HERE = Path(__file__).resolve().parent
|
| 19 |
+
renderer = HERE / 'svg_figures.py'
|
| 20 |
+
if not renderer.is_file():
|
| 21 |
+
renderer = HERE.parents[1] / 'lib/svg_figures.py'
|
| 22 |
+
spec = importlib.util.spec_from_file_location('model_card_svg', renderer)
|
| 23 |
+
svg = importlib.util.module_from_spec(spec)
|
| 24 |
+
sys.modules[spec.name] = svg
|
| 25 |
+
spec.loader.exec_module(svg)
|
| 26 |
+
|
| 27 |
+
UPSTREAM, FIT, REFERENCE = '#0072B2', '#D55E00', '#777777'
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
def classifier(data: dict) -> str:
|
| 31 |
+
summary = metrics.summarize_classifier(data)
|
| 32 |
+
facets = [*metrics.FACETS, 'macro']
|
| 33 |
+
def values(model):
|
| 34 |
+
result = summary['models'][model]
|
| 35 |
+
return [result['facets'][f]['average_precision'] for f in metrics.FACETS] + [result['macro']['average_precision']]
|
| 36 |
+
return svg.render_bars(
|
| 37 |
+
facets, [('Upstream', values('upstream'), UPSTREAM),
|
| 38 |
+
('TF-IDF + logistic regression', values('linear'), REFERENCE),
|
| 39 |
+
('Daecore fine-tune', values('finetuned'), FIT)],
|
| 40 |
+
title='Classifier: matched before and after fine-tuning',
|
| 41 |
+
subtitle='1,300 generated passages · 3 held-out families · resolved labels only',
|
| 42 |
+
ylabel='average precision', ylim=(0, 1.05), separator_before=5,
|
| 43 |
+
width=840, height=360,
|
| 44 |
+
notes=['Model-generated labels; these results do not establish accuracy on unrelated real documents.'],
|
| 45 |
+
)
|
| 46 |
+
|
| 47 |
+
|
| 48 |
+
def reranker(data: dict) -> str:
|
| 49 |
+
summary = metrics.summarize_reranker(data)
|
| 50 |
+
definitions = [('hit', 'At least one useful passage'), ('precision', 'Useful passages / returned passages'), ('ndcg', 'Graded ranking quality')]
|
| 51 |
+
panels = []
|
| 52 |
+
legend = [('Upstream', UPSTREAM, None), ('Daecore fine-tune', FIT, None), ('Random pool order', REFERENCE, '5 3')]
|
| 53 |
+
for field, title in definitions:
|
| 54 |
+
series = []
|
| 55 |
+
for model, (label, color, dash) in zip(('upstream', 'finetuned', 'random'), legend, strict=True):
|
| 56 |
+
values = summary['models'][model]['answerable']
|
| 57 |
+
series.append(svg.Series(label, list(range(5)), [values[str(k)][field] for k in metrics.CUTOFFS], color, dash=dash))
|
| 58 |
+
panels.append(svg.Panel(title, series, xlabel='rank cutoff',
|
| 59 |
+
ylabel={'hit': 'Hit@k', 'precision': 'Precision@k', 'ndcg': 'nDCG@k'}[field],
|
| 60 |
+
xlim=(0, 4), ylim=(0, 1.02), xticks=list(enumerate(map(str, metrics.CUTOFFS)))))
|
| 61 |
+
return svg.render_grid(panels, columns=3, title='Ettin: identical candidates, different ordering',
|
| 62 |
+
subtitle='873 answerable pools of 50 · 970 queries in total · useful = grade 2 or 3',
|
| 63 |
+
legend=legend, panel_width=300, panel_height=270)
|
| 64 |
+
|
| 65 |
+
|
| 66 |
+
def main() -> None:
|
| 67 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 68 |
+
parser.add_argument('--directory', type=Path, default=HERE)
|
| 69 |
+
parser.add_argument('--output', type=Path, required=True)
|
| 70 |
+
parser.add_argument('--model', choices=('classifier', 'reranker'))
|
| 71 |
+
args = parser.parse_args()
|
| 72 |
+
args.output.mkdir(parents=True, exist_ok=True)
|
| 73 |
+
definitions = {'classifier': ('classifier', 'classifier-comparison.svg'),
|
| 74 |
+
'reranker': ('reranker', 'ettin-comparison.svg')}
|
| 75 |
+
names = [args.model] if args.model else [
|
| 76 |
+
name for name, (record, _) in definitions.items()
|
| 77 |
+
if (args.directory / f'{record}.json').is_file()
|
| 78 |
+
]
|
| 79 |
+
if not names:
|
| 80 |
+
parser.error('No model comparison records found in the selected directory')
|
| 81 |
+
for name in names:
|
| 82 |
+
record, filename = definitions[name]
|
| 83 |
+
data = json.loads((args.directory / f'{record}.json').read_text(encoding='utf-8'))
|
| 84 |
+
(args.output / filename).write_text(globals()[name](data), encoding='utf-8', newline='\n')
|
| 85 |
+
|
| 86 |
+
|
| 87 |
+
|
| 88 |
+
if __name__ == '__main__':
|
| 89 |
+
main()
|
evaluation/metrics.py
ADDED
|
@@ -0,0 +1,219 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Recompute the published model comparisons from text-free evaluation records.
|
| 2 |
+
|
| 3 |
+
Python 3.11+, standard library only. Run: python metrics.py --directory .
|
| 4 |
+
"""
|
| 5 |
+
|
| 6 |
+
from __future__ import annotations
|
| 7 |
+
|
| 8 |
+
import argparse
|
| 9 |
+
import json
|
| 10 |
+
import math
|
| 11 |
+
from pathlib import Path
|
| 12 |
+
from statistics import mean
|
| 13 |
+
|
| 14 |
+
CUTOFFS = (1, 3, 5, 10, 20)
|
| 15 |
+
FACETS = ("trap", "decision", "constraint", "mechanism", "procedure")
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
def classification_metrics(labels: list[int], scores: list[float]) -> dict[str, float]:
|
| 19 |
+
"""Threshold-grouped AP and best precision at or above each recall target."""
|
| 20 |
+
if len(labels) != len(scores) or not labels or set(labels) - {0, 1}:
|
| 21 |
+
raise ValueError("Classification labels and scores must be aligned binary rows")
|
| 22 |
+
if not all(math.isfinite(s) for s in scores) or sum(labels) == 0:
|
| 23 |
+
raise ValueError("Classification scores must be finite with positive support")
|
| 24 |
+
order = sorted(range(len(labels)), key=lambda i: -scores[i])
|
| 25 |
+
positives = sum(labels)
|
| 26 |
+
true_positives = 0
|
| 27 |
+
previous_recall = 0.0
|
| 28 |
+
ap = 0.0
|
| 29 |
+
points = []
|
| 30 |
+
for rank, index in enumerate(order, 1):
|
| 31 |
+
true_positives += labels[index]
|
| 32 |
+
if rank < len(order) and scores[order[rank]] == scores[index]:
|
| 33 |
+
continue
|
| 34 |
+
precision = true_positives / rank
|
| 35 |
+
recall = true_positives / positives
|
| 36 |
+
ap += (recall - previous_recall) * precision
|
| 37 |
+
previous_recall = recall
|
| 38 |
+
points.append((precision, recall))
|
| 39 |
+
return {
|
| 40 |
+
"average_precision": ap,
|
| 41 |
+
"precision_at_recall_90": max(p for p, r in points if r >= 0.9),
|
| 42 |
+
"precision_at_recall_95": max(p for p, r in points if r >= 0.95),
|
| 43 |
+
"prevalence": positives / len(labels),
|
| 44 |
+
}
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
def dcg(grades: list[int], k: int) -> float:
|
| 48 |
+
return sum((2**g - 1) / math.log2(i + 2) for i, g in enumerate(grades[:k]))
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
def ranking_metrics(grades: list[int], scores: list[float], k: int) -> dict[str, float]:
|
| 52 |
+
if len(grades) != len(scores) or len(grades) < k or k < 1:
|
| 53 |
+
raise ValueError("Ranking inputs must be aligned and cover the cutoff")
|
| 54 |
+
if any(type(g) is not int or g not in range(4) for g in grades):
|
| 55 |
+
raise ValueError("Ranking requires fully judged integer grades 0 through 3")
|
| 56 |
+
if not all(math.isfinite(s) for s in scores):
|
| 57 |
+
raise ValueError("Ranking scores must be finite")
|
| 58 |
+
order = sorted(range(len(scores)), key=lambda i: (-scores[i], i))
|
| 59 |
+
ranked = [grades[i] for i in order]
|
| 60 |
+
useful = sum(g >= 2 for g in ranked[:k])
|
| 61 |
+
relevant = sum(g >= 2 for g in grades)
|
| 62 |
+
ideal = dcg(sorted(grades, reverse=True), k)
|
| 63 |
+
result = {"hit": float(useful > 0), "precision": useful / k, "useful": float(useful)}
|
| 64 |
+
if relevant:
|
| 65 |
+
result["recall"] = useful / relevant
|
| 66 |
+
result["ndcg"] = dcg(ranked, k) / ideal
|
| 67 |
+
return result
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
def random_ranking_metrics(grades: list[int], k: int) -> dict[str, float]:
|
| 71 |
+
"""Exact expectation under a uniform permutation of one fixed judged pool."""
|
| 72 |
+
n = len(grades)
|
| 73 |
+
if k < 1 or k > n or any(type(g) is not int or g not in range(4) for g in grades):
|
| 74 |
+
raise ValueError("Random reference requires judged grades and a valid cutoff")
|
| 75 |
+
relevant = sum(g >= 2 for g in grades)
|
| 76 |
+
misses = math.comb(n - relevant, k) if n - relevant >= k else 0
|
| 77 |
+
result = {"hit": 1 - misses / math.comb(n, k), "precision": relevant / n, "useful": k * relevant / n}
|
| 78 |
+
if relevant:
|
| 79 |
+
expected_dcg = mean(2**g - 1 for g in grades) * sum(1 / math.log2(i + 2) for i in range(k))
|
| 80 |
+
result["recall"] = k / n
|
| 81 |
+
result["ndcg"] = expected_dcg / dcg(sorted(grades, reverse=True), k)
|
| 82 |
+
return result
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
def summarize_classifier(data: dict) -> dict:
|
| 86 |
+
rows = data["rows"]
|
| 87 |
+
result = {"rows": len(rows), "families": len({r["group"] for r in rows}), "models": {}}
|
| 88 |
+
for model in data["model_order"]:
|
| 89 |
+
by_facet = {}
|
| 90 |
+
for facet in FACETS:
|
| 91 |
+
resolved = [r for r in rows if r["labels"][facet] is not None]
|
| 92 |
+
if any(r["labels"][facet] not in (0, 1) for r in resolved):
|
| 93 |
+
raise ValueError("Invalid resolved classifier label")
|
| 94 |
+
by_facet[facet] = {"rows": len(resolved), **classification_metrics(
|
| 95 |
+
[r["labels"][facet] for r in resolved], [r["scores"][model][facet] for r in resolved]
|
| 96 |
+
)}
|
| 97 |
+
macro = {k: mean(v[k] for v in by_facet.values()) for k in ("average_precision", "precision_at_recall_90", "precision_at_recall_95")}
|
| 98 |
+
result["models"][model] = {"facets": by_facet, "macro": macro}
|
| 99 |
+
return result
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
def summarize_reranker(data: dict) -> dict:
|
| 103 |
+
rows = data["rows"]
|
| 104 |
+
answerable = [r for r in rows if any(g >= 2 for g in r["grades"])]
|
| 105 |
+
output = {"queries": len(rows), "answerable": len(answerable), "models": {}}
|
| 106 |
+
for model in ["random", *data["model_order"]]:
|
| 107 |
+
scopes = {}
|
| 108 |
+
for scope, subset in [("answerable", answerable), ("all", rows)]:
|
| 109 |
+
cuts = {}
|
| 110 |
+
for k in CUTOFFS:
|
| 111 |
+
values = [random_ranking_metrics(r["grades"], k) if model == "random" else ranking_metrics(r["grades"], r["scores"][model], k) for r in subset]
|
| 112 |
+
fields = ("hit", "precision", "useful", "recall", "ndcg") if scope == "answerable" else ("hit", "precision", "useful")
|
| 113 |
+
cuts[str(k)] = {field: mean(v[field] for v in values) for field in fields}
|
| 114 |
+
scopes[scope] = cuts
|
| 115 |
+
output["models"][model] = scopes
|
| 116 |
+
return output
|
| 117 |
+
|
| 118 |
+
|
| 119 |
+
def summarize_promotion(data: dict) -> dict:
|
| 120 |
+
"""Hybrid search with reviewed exclusions inside each original prefix.
|
| 121 |
+
|
| 122 |
+
A null is permitted only for an explicitly excluded judging abstention.
|
| 123 |
+
Cut first, remove exclusions second, and never backfill from a deeper rank.
|
| 124 |
+
"""
|
| 125 |
+
rows = data['rows']
|
| 126 |
+
if not rows or len({row['id'] for row in rows}) != len(rows):
|
| 127 |
+
raise ValueError('Promotion rows require unique nonempty query identities')
|
| 128 |
+
output = {'queries': len(rows), 'models': {}, 'public': {}}
|
| 129 |
+
for model in data['model_order']:
|
| 130 |
+
cutoffs = {}
|
| 131 |
+
for cutoff in (3, 5, 10, 20, 'selected'):
|
| 132 |
+
per_query = []
|
| 133 |
+
for row in rows:
|
| 134 |
+
counts = row['reference_grade_counts']
|
| 135 |
+
if set(counts) != {'0', '1', '2', '3'} or any(type(n) is not int or n < 0 for n in counts.values()):
|
| 136 |
+
raise ValueError('Reference grade counts must cover grades zero through three')
|
| 137 |
+
ranked = row['ranked_grades'][model]
|
| 138 |
+
excluded = row['excluded'][model]
|
| 139 |
+
depth = row['selected_depth'][model] if cutoff == 'selected' else cutoff
|
| 140 |
+
if (len(ranked) != 20 or len(excluded) != 20 or type(depth) is not int
|
| 141 |
+
or not 3 <= depth <= 20 or any(type(x) is not bool for x in excluded)):
|
| 142 |
+
raise ValueError('Promotion rows require a bounded original top twenty')
|
| 143 |
+
if any((grade is not None if drop else type(grade) is not int or grade not in range(4))
|
| 144 |
+
for grade, drop in zip(ranked, excluded, strict=True)):
|
| 145 |
+
raise ValueError('Only declared abstentions may lack grades')
|
| 146 |
+
kept = [grade for grade, drop in zip(ranked[:depth], excluded[:depth], strict=True) if not drop]
|
| 147 |
+
if not kept:
|
| 148 |
+
raise ValueError('Every scored prefix must retain judged passages')
|
| 149 |
+
ideal_grades = [g for g in (3, 2, 1, 0) for _ in range(min(counts[str(g)], len(kept)))][:len(kept)]
|
| 150 |
+
useful = sum(g >= 2 for g in kept)
|
| 151 |
+
positives = counts['2'] + counts['3']
|
| 152 |
+
ideal = dcg(ideal_grades, len(kept))
|
| 153 |
+
per_query.append({
|
| 154 |
+
'hit': float(useful > 0), 'precision': useful / len(kept),
|
| 155 |
+
'useful': useful, 'retained': len(kept), 'excluded': depth - len(kept),
|
| 156 |
+
'ndcg': dcg(kept, len(kept)) / ideal if ideal else 0.0,
|
| 157 |
+
'known_recall': useful / positives if positives else None,
|
| 158 |
+
})
|
| 159 |
+
useful = sum(row['useful'] for row in per_query)
|
| 160 |
+
retained = sum(row['retained'] for row in per_query)
|
| 161 |
+
recalls = [row['known_recall'] for row in per_query if row['known_recall'] is not None]
|
| 162 |
+
cutoffs[str(cutoff)] = {
|
| 163 |
+
'hit': mean(row['hit'] for row in per_query), 'precision': useful / retained,
|
| 164 |
+
'macro_precision': mean(row['precision'] for row in per_query),
|
| 165 |
+
'ndcg': mean(row['ndcg'] for row in per_query),
|
| 166 |
+
'known_positive_recall': mean(recalls) if recalls else None,
|
| 167 |
+
'recall_queries': len(recalls), 'useful': useful, 'retained': retained,
|
| 168 |
+
'excluded_positions': sum(row['excluded'] for row in per_query),
|
| 169 |
+
'mean_useful': useful / len(rows), 'mean_retained': retained / len(rows),
|
| 170 |
+
}
|
| 171 |
+
output['models'][model] = cutoffs
|
| 172 |
+
for dataset, panel in data['public'].items():
|
| 173 |
+
if not panel['rows'] or len({row['id'] for row in panel['rows']}) != len(panel['rows']):
|
| 174 |
+
raise ValueError('Public panel requires distinct query identities')
|
| 175 |
+
measured = panel['model_order']
|
| 176 |
+
for row in panel['rows']:
|
| 177 |
+
if set(row['ndcg@10']) != set(measured) or any(
|
| 178 |
+
isinstance(value, bool) or not isinstance(value, int | float)
|
| 179 |
+
or not math.isfinite(value) or not 0 <= value <= 1
|
| 180 |
+
for value in row['ndcg@10'].values()
|
| 181 |
+
):
|
| 182 |
+
raise ValueError('Public nDCG values must be finite measured scores')
|
| 183 |
+
output['public'][dataset] = {
|
| 184 |
+
'queries': len(panel['rows']),
|
| 185 |
+
'ndcg@10': {model: mean(row['ndcg@10'][model] for row in panel['rows']) for model in measured},
|
| 186 |
+
}
|
| 187 |
+
return output
|
| 188 |
+
|
| 189 |
+
|
| 190 |
+
def main() -> None:
|
| 191 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 192 |
+
parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
|
| 193 |
+
parser.add_argument("--output", type=Path)
|
| 194 |
+
parser.add_argument("--model", choices=("classifier", "reranker", "promotion", "serving"), help="Recompute one comparison; default: every record in this package")
|
| 195 |
+
args = parser.parse_args()
|
| 196 |
+
summarizers = {"classifier": summarize_classifier, "reranker": summarize_reranker,
|
| 197 |
+
"promotion": summarize_promotion, "serving": summarize_promotion}
|
| 198 |
+
names = [args.model] if args.model else [
|
| 199 |
+
name for name in summarizers if (args.directory / f"{name}.json").is_file()
|
| 200 |
+
]
|
| 201 |
+
if not names:
|
| 202 |
+
parser.error("No evaluation records found in the selected directory")
|
| 203 |
+
result = {}
|
| 204 |
+
for name in names:
|
| 205 |
+
summarize = summarizers[name]
|
| 206 |
+
data = json.loads((args.directory / f"{name}.json").read_text(encoding="utf-8"))
|
| 207 |
+
ids = [r["id"] for r in data["rows"]]
|
| 208 |
+
if len(set(ids)) != len(ids):
|
| 209 |
+
raise ValueError(f"Repeated query/passage identity in {name}")
|
| 210 |
+
result[name] = summarize(data)
|
| 211 |
+
text = json.dumps(result, indent=2, allow_nan=False) + "\n"
|
| 212 |
+
if args.output:
|
| 213 |
+
args.output.write_text(text, encoding="utf-8")
|
| 214 |
+
else:
|
| 215 |
+
print(text, end="")
|
| 216 |
+
|
| 217 |
+
|
| 218 |
+
if __name__ == "__main__":
|
| 219 |
+
main()
|
evaluation/serving-qualification.json
ADDED
|
@@ -0,0 +1,134 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schema": "daecore.serving-preparation-summary.v1",
|
| 3 |
+
"date": "2026-09-28",
|
| 4 |
+
"status": "local-model-serving-checks-complete-release-transition-pending",
|
| 5 |
+
"hardware": "Windows x64, NVIDIA RTX 3060 Ti 8 GiB",
|
| 6 |
+
"unmeasured_targets": [
|
| 7 |
+
"AMD",
|
| 8 |
+
"Intel",
|
| 9 |
+
"Linux x64"
|
| 10 |
+
],
|
| 11 |
+
"runtimes": {
|
| 12 |
+
"onnxruntime": "1.24.4",
|
| 13 |
+
"vulkan_plugin": "0.4.0",
|
| 14 |
+
"backend": "Vulkan",
|
| 15 |
+
"storage_buffer_cache": "lazyRelease"
|
| 16 |
+
},
|
| 17 |
+
"models": {
|
| 18 |
+
"gemma": {
|
| 19 |
+
"vectors": 74,
|
| 20 |
+
"providers": {
|
| 21 |
+
"cpu": {
|
| 22 |
+
"receipt_sha256": "31851791104fa95d6b7b25d35f950b25e239a3d6f8fa458addb2a44917c25cb4",
|
| 23 |
+
"maximum_vector_delta": 3.0174851417541504e-07,
|
| 24 |
+
"concurrent_delta": 0.0
|
| 25 |
+
},
|
| 26 |
+
"cuda": {
|
| 27 |
+
"receipt_sha256": "6977d0aa0ebf9c39ea0fd24078c47367d1d50e81184653225fcd10758961903e",
|
| 28 |
+
"maximum_vector_delta": 2.644956111907959e-07,
|
| 29 |
+
"concurrent_delta": 2.2351741790771484e-08
|
| 30 |
+
},
|
| 31 |
+
"vulkan": {
|
| 32 |
+
"receipt_sha256": "567fb89ea195e8b188203d2334f3eda82994ae8c7fc97f2126fac3e6f4766079",
|
| 33 |
+
"maximum_vector_delta": 5.438923835754395e-07,
|
| 34 |
+
"concurrent_delta": 0.0
|
| 35 |
+
}
|
| 36 |
+
},
|
| 37 |
+
"admission": {
|
| 38 |
+
"cases": 3841,
|
| 39 |
+
"semantic_fallbacks": 679,
|
| 40 |
+
"new_useful_exclusions": 0,
|
| 41 |
+
"additional_noise_retained": 2,
|
| 42 |
+
"receipt_sha256": "71550f606c95c1e1192dd5939488b280d5f2f2f091f99413f6b63f4408a5bc2c"
|
| 43 |
+
}
|
| 44 |
+
},
|
| 45 |
+
"classifier": {
|
| 46 |
+
"cpu": {
|
| 47 |
+
"rows": 1300,
|
| 48 |
+
"maximum_posterior_delta": 5.612167303103988e-06,
|
| 49 |
+
"changed_threshold_labels": {
|
| 50 |
+
"contract": 0,
|
| 51 |
+
"recall_leaning": 0
|
| 52 |
+
},
|
| 53 |
+
"receipt_sha256": "ea545fb0a30ff25bf3da178b60593e934e6025e789c2db67290ff5360e8669a4"
|
| 54 |
+
},
|
| 55 |
+
"cuda": {
|
| 56 |
+
"rows": 1300,
|
| 57 |
+
"maximum_posterior_delta": 5.612167303103988e-06,
|
| 58 |
+
"changed_threshold_labels": {
|
| 59 |
+
"contract": 0,
|
| 60 |
+
"recall_leaning": 0
|
| 61 |
+
},
|
| 62 |
+
"receipt_sha256": "5616891e9dfdc3b6615951901cb24aa31189d6a81ed84b0080f461b85c045065"
|
| 63 |
+
},
|
| 64 |
+
"vulkan": {
|
| 65 |
+
"rows": 1300,
|
| 66 |
+
"maximum_posterior_delta": 5.612167303103988e-06,
|
| 67 |
+
"changed_threshold_labels": {
|
| 68 |
+
"contract": 0,
|
| 69 |
+
"recall_leaning": 0
|
| 70 |
+
},
|
| 71 |
+
"receipt_sha256": "c7ec971e020d9771adf0756316f468629cc05e300d7edf9abb52612901adf6b0"
|
| 72 |
+
}
|
| 73 |
+
},
|
| 74 |
+
"ettin": {
|
| 75 |
+
"queries": 970,
|
| 76 |
+
"unjudged_top20": 0,
|
| 77 |
+
"quality_evidence": "serving.json",
|
| 78 |
+
"quality_receipt_sha256": "805a1b0acd7baaf248e182908bc23cd52171e40aa8f1281799a20d15b58c6905",
|
| 79 |
+
"full_replay_seconds_p50_p95": [
|
| 80 |
+
1.7189999999827705,
|
| 81 |
+
2.610000000044238
|
| 82 |
+
],
|
| 83 |
+
"matched_20_pool_seconds_p50_p95": {
|
| 84 |
+
"vulkan": [
|
| 85 |
+
1.5565476999909151,
|
| 86 |
+
2.1063475799834124
|
| 87 |
+
],
|
| 88 |
+
"predecessor_directml": [
|
| 89 |
+
1.9197023500164505,
|
| 90 |
+
2.701977299965802
|
| 91 |
+
]
|
| 92 |
+
},
|
| 93 |
+
"pressure": "One preflight refusal after 322 queries with lazy release; resumed all remaining rows with zero further retries. An earlier default-cache run stopped after 424. No native OOM observed.",
|
| 94 |
+
"precision": "FP32 Vulkan versus FP16 CUDA; rankings not bit-exact",
|
| 95 |
+
"selector": "Unchanged CUDA mapping transferred for measurement; new package/provider binding pending"
|
| 96 |
+
}
|
| 97 |
+
},
|
| 98 |
+
"operator_runtime_changed": false,
|
| 99 |
+
"published": false,
|
| 100 |
+
"remaining": [
|
| 101 |
+
"immutable-publication-revisions",
|
| 102 |
+
"profile-and-calibration-release-bindings",
|
| 103 |
+
"isolated-package-update-and-index-rebuild-recovery",
|
| 104 |
+
"DirectML-current-route-retirement"
|
| 105 |
+
],
|
| 106 |
+
"worker_recovery": {
|
| 107 |
+
"receipt_sha256": "46205d39521bde40d1228c837e57f88e1a941d42f6846f4a244c48b513347cc6",
|
| 108 |
+
"all_three_consumer_outputs_identical_after_owned_worker_crash": true,
|
| 109 |
+
"ettin_maximum_envelope": {
|
| 110 |
+
"tokens": 1153,
|
| 111 |
+
"max_abs_logit_delta_vs_cpu": 5.91278076171875e-05
|
| 112 |
+
}
|
| 113 |
+
},
|
| 114 |
+
"retained_reranker_public": {
|
| 115 |
+
"scope": "Retained matched public pools retrieved by upstream Gemma; upstream and production Ettin PyTorch FP16; not a new Vulkan public replay",
|
| 116 |
+
"source_sha256": "0814db560081af8376e4b1f3e3d2bcd5c4d9b8b208f7fb8b7172f113e8123ef3",
|
| 117 |
+
"panels": {
|
| 118 |
+
"fiqa": {
|
| 119 |
+
"queries": 648,
|
| 120 |
+
"ndcg@10": {
|
| 121 |
+
"upstream": 0.48603492061219766,
|
| 122 |
+
"finetuned": 0.452680922415071
|
| 123 |
+
}
|
| 124 |
+
},
|
| 125 |
+
"scifact": {
|
| 126 |
+
"queries": 300,
|
| 127 |
+
"ndcg@10": {
|
| 128 |
+
"upstream": 0.7487436478294831,
|
| 129 |
+
"finetuned": 0.7543833016238701
|
| 130 |
+
}
|
| 131 |
+
}
|
| 132 |
+
}
|
| 133 |
+
}
|
| 134 |
+
}
|
evaluation/svg_figures.py
ADDED
|
@@ -0,0 +1,576 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Dependency-free SVG charts for tracked evaluation figures.
|
| 2 |
+
|
| 3 |
+
The evaluation figures are generated from private per-row evidence into tracked SVG files.
|
| 4 |
+
Keeping the renderer inside the repository, with no plotting dependency, means a figure can be
|
| 5 |
+
regenerated by any contributor with the private inputs and compared byte for byte.
|
| 6 |
+
|
| 7 |
+
Everything is emitted as presentation attributes rather than CSS. Markdown hosts sanitize
|
| 8 |
+
embedded stylesheets and scripts out of SVG, so a figure that carries its styling in attributes
|
| 9 |
+
renders the same in the repository, on a model-card host, and in a local viewer.
|
| 10 |
+
|
| 11 |
+
Layout is measured rather than assumed: legend entries, panel titles, and reference-line labels
|
| 12 |
+
are placed from an estimated text width, so a longer label reflows instead of overlapping its
|
| 13 |
+
neighbour. ``text_width`` approximates a sans-serif advance table, which is enough to keep
|
| 14 |
+
elements apart but is not a substitute for a real font metric.
|
| 15 |
+
"""
|
| 16 |
+
|
| 17 |
+
from __future__ import annotations
|
| 18 |
+
|
| 19 |
+
import math
|
| 20 |
+
from dataclasses import dataclass, field
|
| 21 |
+
from xml.sax.saxutils import escape
|
| 22 |
+
|
| 23 |
+
# Okabe-Ito, chosen because it stays distinguishable under the common colour-vision
|
| 24 |
+
# deficiencies and prints legibly in greyscale.
|
| 25 |
+
PALETTE = ("#0072B2", "#D55E00", "#009E73", "#CC79A7", "#E69F00", "#56B4E9", "#000000")
|
| 26 |
+
FONT_STACK = "system-ui, Segoe UI, Roboto, Helvetica, Arial, sans-serif"
|
| 27 |
+
FONT = f'font-family="{FONT_STACK}"'
|
| 28 |
+
|
| 29 |
+
INK = "#1a1a1a"
|
| 30 |
+
MUTED = "#5c5c5c"
|
| 31 |
+
AXIS = "#8a8a8a"
|
| 32 |
+
GRID = "#e8e8e8"
|
| 33 |
+
GRID_STRONG = "#d0d0d0"
|
| 34 |
+
|
| 35 |
+
# Per-character advance as a fraction of font size, for a humanist sans at normal weight.
|
| 36 |
+
_NARROW = set("iljft.,;:|!()[]{}I '")
|
| 37 |
+
_WIDE = set("mwMW@%")
|
| 38 |
+
_DIGIT = set("0123456789")
|
| 39 |
+
|
| 40 |
+
|
| 41 |
+
def text_width(text: str, size: float, *, weight: str = "normal") -> float:
|
| 42 |
+
"""Estimate rendered width in user units."""
|
| 43 |
+
|
| 44 |
+
total = 0.0
|
| 45 |
+
for character in text:
|
| 46 |
+
if character in _NARROW:
|
| 47 |
+
total += 0.30
|
| 48 |
+
elif character in _WIDE:
|
| 49 |
+
total += 0.86
|
| 50 |
+
elif character in _DIGIT:
|
| 51 |
+
total += 0.56
|
| 52 |
+
elif character.isupper():
|
| 53 |
+
total += 0.66
|
| 54 |
+
else:
|
| 55 |
+
total += 0.52
|
| 56 |
+
if weight in {"600", "700", "bold"}:
|
| 57 |
+
total *= 1.05
|
| 58 |
+
return total * size
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
def _fit_lines(text: str, size: float, limit: float, *, weight: str = "normal") -> list[str]:
|
| 62 |
+
"""Wrap to at most two lines, breaking on whitespace."""
|
| 63 |
+
|
| 64 |
+
if text_width(text, size, weight=weight) <= limit:
|
| 65 |
+
return [text]
|
| 66 |
+
words = text.split(" ")
|
| 67 |
+
line: list[str] = []
|
| 68 |
+
for index, word in enumerate(words):
|
| 69 |
+
candidate = " ".join([*line, word])
|
| 70 |
+
if line and text_width(candidate, size, weight=weight) > limit:
|
| 71 |
+
return [" ".join(line), " ".join(words[index:])]
|
| 72 |
+
line.append(word)
|
| 73 |
+
return [" ".join(line)]
|
| 74 |
+
|
| 75 |
+
|
| 76 |
+
@dataclass
|
| 77 |
+
class Series:
|
| 78 |
+
label: str
|
| 79 |
+
x: list[float]
|
| 80 |
+
y: list[float]
|
| 81 |
+
color: str = PALETTE[0]
|
| 82 |
+
dash: str | None = None
|
| 83 |
+
marker: bool = True
|
| 84 |
+
width: float = 1.9
|
| 85 |
+
marker_size: float = 2.6
|
| 86 |
+
|
| 87 |
+
|
| 88 |
+
@dataclass
|
| 89 |
+
class Band:
|
| 90 |
+
"""A shaded interval drawn behind its series."""
|
| 91 |
+
|
| 92 |
+
x: list[float]
|
| 93 |
+
low: list[float]
|
| 94 |
+
high: list[float]
|
| 95 |
+
color: str = PALETTE[0]
|
| 96 |
+
opacity: float = 0.16
|
| 97 |
+
|
| 98 |
+
|
| 99 |
+
@dataclass
|
| 100 |
+
class Counts:
|
| 101 |
+
"""A support strip under the plot: how many rows sit behind each x position."""
|
| 102 |
+
|
| 103 |
+
x: list[float]
|
| 104 |
+
values: list[float]
|
| 105 |
+
color: str = MUTED
|
| 106 |
+
label: str = "rows per bin"
|
| 107 |
+
|
| 108 |
+
|
| 109 |
+
@dataclass
|
| 110 |
+
class Panel:
|
| 111 |
+
title: str
|
| 112 |
+
series: list[Series] = field(default_factory=list)
|
| 113 |
+
bands: list[Band] = field(default_factory=list)
|
| 114 |
+
xlabel: str = ""
|
| 115 |
+
ylabel: str = ""
|
| 116 |
+
xlim: tuple[float, float] | None = None
|
| 117 |
+
ylim: tuple[float, float] | None = None
|
| 118 |
+
xticks: list[tuple[float, str]] | None = None
|
| 119 |
+
yticks: list[tuple[float, str]] | None = None
|
| 120 |
+
xscale: str = "linear"
|
| 121 |
+
hlines: list[tuple[float, str, str]] = field(default_factory=list)
|
| 122 |
+
diagonal: bool = False
|
| 123 |
+
notes: list[str] = field(default_factory=list)
|
| 124 |
+
counts: Counts | None = None
|
| 125 |
+
|
| 126 |
+
|
| 127 |
+
def _fmt(value: float) -> str:
|
| 128 |
+
if abs(value) >= 1e6:
|
| 129 |
+
return f"{value:.3g}"
|
| 130 |
+
text = f"{value:.3f}".rstrip("0").rstrip(".")
|
| 131 |
+
return text or "0"
|
| 132 |
+
|
| 133 |
+
|
| 134 |
+
def _auto_ticks(low: float, high: float, count: int = 5) -> list[tuple[float, str]]:
|
| 135 |
+
if high <= low:
|
| 136 |
+
high = low + 1.0
|
| 137 |
+
step = (high - low) / count
|
| 138 |
+
magnitude = 10 ** math.floor(math.log10(step)) if step > 0 else 1.0
|
| 139 |
+
for factor in (1, 2, 2.5, 5, 10):
|
| 140 |
+
if step <= factor * magnitude:
|
| 141 |
+
step = factor * magnitude
|
| 142 |
+
break
|
| 143 |
+
start = math.ceil(low / step) * step
|
| 144 |
+
ticks = []
|
| 145 |
+
value = start
|
| 146 |
+
while value <= high + 1e-9:
|
| 147 |
+
ticks.append((value, _fmt(0.0 if abs(value) < step * 1e-6 else value)))
|
| 148 |
+
value += step
|
| 149 |
+
return ticks
|
| 150 |
+
|
| 151 |
+
|
| 152 |
+
def _text(
|
| 153 |
+
x: float,
|
| 154 |
+
y: float,
|
| 155 |
+
body: str,
|
| 156 |
+
*,
|
| 157 |
+
size: float,
|
| 158 |
+
fill: str = INK,
|
| 159 |
+
anchor: str = "start",
|
| 160 |
+
weight: str | None = None,
|
| 161 |
+
) -> str:
|
| 162 |
+
weight_attr = f' font-weight="{weight}"' if weight else ""
|
| 163 |
+
return (
|
| 164 |
+
f'<text x="{x:.1f}" y="{y:.1f}" text-anchor="{anchor}" font-size="{size}" '
|
| 165 |
+
f'fill="{fill}"{weight_attr} {FONT}>{escape(body)}</text>'
|
| 166 |
+
)
|
| 167 |
+
|
| 168 |
+
|
| 169 |
+
_TITLE_SIZE = 12.5
|
| 170 |
+
|
| 171 |
+
# Everything below the plot box is stacked in fixed bands rather than placed at absolute
|
| 172 |
+
# offsets, so an x-axis label, a support strip, and a note can coexist without overlapping.
|
| 173 |
+
_TICK_BAND = 18.0
|
| 174 |
+
_XLABEL_BAND = 16.0
|
| 175 |
+
_COUNTS_BAND = 30.0
|
| 176 |
+
_NOTE_BAND = 12.0
|
| 177 |
+
_NOTE_LEAD = 6.0
|
| 178 |
+
_FLOOR_SLACK = 8.0
|
| 179 |
+
|
| 180 |
+
|
| 181 |
+
def _panel_bottom(panel: Panel) -> float:
|
| 182 |
+
bottom = _TICK_BAND + _FLOOR_SLACK
|
| 183 |
+
if panel.xlabel:
|
| 184 |
+
bottom += _XLABEL_BAND
|
| 185 |
+
if panel.counts:
|
| 186 |
+
bottom += _COUNTS_BAND
|
| 187 |
+
if panel.notes:
|
| 188 |
+
bottom += _NOTE_LEAD + _NOTE_BAND * len(panel.notes)
|
| 189 |
+
return bottom
|
| 190 |
+
|
| 191 |
+
|
| 192 |
+
def _panel_svg(panel: Panel, width: float, height: float, *, title_rows: int | None = None) -> str:
|
| 193 |
+
left, right = 54.0, 14.0
|
| 194 |
+
plot_w = width - left - right
|
| 195 |
+
title_lines = _fit_lines(panel.title, _TITLE_SIZE, plot_w, weight="600")
|
| 196 |
+
top = 18.0 + 14.0 * (title_rows or len(title_lines))
|
| 197 |
+
plot_h = height - top - _panel_bottom(panel)
|
| 198 |
+
|
| 199 |
+
xs = [x for s in panel.series for x in s.x] or [0.0, 1.0]
|
| 200 |
+
ys = [y for s in panel.series for y in s.y] or [0.0, 1.0]
|
| 201 |
+
ys += [value for band in panel.bands for value in (*band.low, *band.high)]
|
| 202 |
+
ys += [level for level, _, _ in panel.hlines]
|
| 203 |
+
xlim = panel.xlim or (min(xs), max(xs))
|
| 204 |
+
ylim = panel.ylim or (min(ys), max(ys))
|
| 205 |
+
if ylim[0] == ylim[1]:
|
| 206 |
+
ylim = (ylim[0] - 0.5, ylim[1] + 0.5)
|
| 207 |
+
|
| 208 |
+
def tx(value: float) -> float:
|
| 209 |
+
if panel.xscale == "log":
|
| 210 |
+
lo, hi = math.log(xlim[0]), math.log(xlim[1])
|
| 211 |
+
return left + (math.log(value) - lo) / (hi - lo) * plot_w
|
| 212 |
+
return left + (value - xlim[0]) / (xlim[1] - xlim[0]) * plot_w
|
| 213 |
+
|
| 214 |
+
def ty(value: float) -> float:
|
| 215 |
+
return top + plot_h - (value - ylim[0]) / (ylim[1] - ylim[0]) * plot_h
|
| 216 |
+
|
| 217 |
+
parts: list[str] = []
|
| 218 |
+
for index, line in enumerate(title_lines):
|
| 219 |
+
parts.append(
|
| 220 |
+
_text(
|
| 221 |
+
left + plot_w / 2,
|
| 222 |
+
16.0 + 14.0 * index,
|
| 223 |
+
line,
|
| 224 |
+
size=_TITLE_SIZE,
|
| 225 |
+
anchor="middle",
|
| 226 |
+
weight="600",
|
| 227 |
+
)
|
| 228 |
+
)
|
| 229 |
+
|
| 230 |
+
xticks = panel.xticks or _auto_ticks(*xlim)
|
| 231 |
+
yticks = panel.yticks or _auto_ticks(*ylim)
|
| 232 |
+
for value, label in yticks:
|
| 233 |
+
if not ylim[0] - 1e-9 <= value <= ylim[1] + 1e-9:
|
| 234 |
+
continue
|
| 235 |
+
y = ty(value)
|
| 236 |
+
parts.append(
|
| 237 |
+
f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
|
| 238 |
+
f'stroke="{GRID}" stroke-width="1"/>'
|
| 239 |
+
)
|
| 240 |
+
parts.append(_text(left - 6, y + 3.4, label, size=10, fill=MUTED, anchor="end"))
|
| 241 |
+
for value, label in xticks:
|
| 242 |
+
if not xlim[0] - 1e-9 <= value <= xlim[1] + 1e-9:
|
| 243 |
+
continue
|
| 244 |
+
x = tx(value)
|
| 245 |
+
parts.append(
|
| 246 |
+
f'<line x1="{x:.1f}" y1="{top}" x2="{x:.1f}" y2="{top + plot_h:.1f}" '
|
| 247 |
+
f'stroke="{GRID}" stroke-width="1"/>'
|
| 248 |
+
)
|
| 249 |
+
parts.append(_text(x, top + plot_h + 13, label, size=10, fill=MUTED, anchor="middle"))
|
| 250 |
+
|
| 251 |
+
if panel.diagonal:
|
| 252 |
+
parts.append(
|
| 253 |
+
f'<line x1="{tx(xlim[0]):.1f}" y1="{ty(ylim[0]):.1f}" '
|
| 254 |
+
f'x2="{tx(xlim[1]):.1f}" y2="{ty(ylim[1]):.1f}" stroke="{AXIS}" '
|
| 255 |
+
f'stroke-width="1.1" stroke-dasharray="4 3"/>'
|
| 256 |
+
)
|
| 257 |
+
|
| 258 |
+
for band in panel.bands:
|
| 259 |
+
forward = " ".join(
|
| 260 |
+
f"{tx(x):.1f},{ty(y):.1f}" for x, y in zip(band.x, band.high, strict=True)
|
| 261 |
+
)
|
| 262 |
+
backward = " ".join(
|
| 263 |
+
f"{tx(x):.1f},{ty(y):.1f}"
|
| 264 |
+
for x, y in zip(reversed(band.x), reversed(band.low), strict=True)
|
| 265 |
+
)
|
| 266 |
+
parts.append(
|
| 267 |
+
f'<polygon points="{forward} {backward}" fill="{band.color}" '
|
| 268 |
+
f'fill-opacity="{band.opacity}" stroke="none"/>'
|
| 269 |
+
)
|
| 270 |
+
|
| 271 |
+
# A reference line carries its label at the left margin over a solid backing box, so the
|
| 272 |
+
# label never lands on the data it is a reference for.
|
| 273 |
+
for value, label, color in panel.hlines:
|
| 274 |
+
y = ty(value)
|
| 275 |
+
parts.append(
|
| 276 |
+
f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
|
| 277 |
+
f'stroke="{color}" stroke-width="1.2" stroke-dasharray="5 3"/>'
|
| 278 |
+
)
|
| 279 |
+
if not label:
|
| 280 |
+
continue
|
| 281 |
+
label_w = text_width(label, 9.5) + 8.0
|
| 282 |
+
parts.append(
|
| 283 |
+
f'<rect x="{left + 3:.1f}" y="{y - 12:.1f}" width="{label_w:.1f}" height="11.5" '
|
| 284 |
+
f'fill="white" fill-opacity="0.9" stroke="none"/>'
|
| 285 |
+
)
|
| 286 |
+
parts.append(_text(left + 7, y - 3.5, label, size=9.5, fill=color))
|
| 287 |
+
|
| 288 |
+
for series in panel.series:
|
| 289 |
+
points = " ".join(
|
| 290 |
+
f"{tx(x):.1f},{ty(y):.1f}" for x, y in zip(series.x, series.y, strict=True)
|
| 291 |
+
)
|
| 292 |
+
dash = f' stroke-dasharray="{series.dash}"' if series.dash else ""
|
| 293 |
+
parts.append(
|
| 294 |
+
f'<polyline points="{points}" fill="none" stroke="{series.color}" '
|
| 295 |
+
f'stroke-width="{series.width}" stroke-linejoin="round" '
|
| 296 |
+
f'stroke-linecap="round"{dash}/>'
|
| 297 |
+
)
|
| 298 |
+
if series.marker:
|
| 299 |
+
for x, y in zip(series.x, series.y, strict=True):
|
| 300 |
+
parts.append(
|
| 301 |
+
f'<circle cx="{tx(x):.1f}" cy="{ty(y):.1f}" r="{series.marker_size}" '
|
| 302 |
+
f'fill="{series.color}"/>'
|
| 303 |
+
)
|
| 304 |
+
|
| 305 |
+
parts.append(
|
| 306 |
+
f'<rect x="{left}" y="{top}" width="{plot_w:.1f}" height="{plot_h:.1f}" fill="none" '
|
| 307 |
+
f'stroke="{AXIS}" stroke-width="1"/>'
|
| 308 |
+
)
|
| 309 |
+
|
| 310 |
+
cursor = top + plot_h + _TICK_BAND
|
| 311 |
+
if panel.xlabel:
|
| 312 |
+
parts.append(
|
| 313 |
+
_text(
|
| 314 |
+
left + plot_w / 2,
|
| 315 |
+
cursor + 11,
|
| 316 |
+
panel.xlabel,
|
| 317 |
+
size=10.5,
|
| 318 |
+
fill=MUTED,
|
| 319 |
+
anchor="middle",
|
| 320 |
+
)
|
| 321 |
+
)
|
| 322 |
+
cursor += _XLABEL_BAND
|
| 323 |
+
if panel.counts:
|
| 324 |
+
counts = panel.counts
|
| 325 |
+
if not counts.values:
|
| 326 |
+
raise ValueError("a support strip needs at least one count")
|
| 327 |
+
strip_top = cursor + 2.0
|
| 328 |
+
strip_h = 15.0
|
| 329 |
+
peak = max(counts.values) or 1.0
|
| 330 |
+
slot = plot_w / max(len(counts.x), 1) * 0.7
|
| 331 |
+
for x, value in zip(counts.x, counts.values, strict=True):
|
| 332 |
+
bar_h = (value / peak) * strip_h
|
| 333 |
+
parts.append(
|
| 334 |
+
f'<rect x="{tx(x) - slot / 2:.1f}" y="{strip_top + strip_h - bar_h:.1f}" '
|
| 335 |
+
f'width="{slot:.1f}" height="{bar_h:.1f}" fill="{counts.color}" '
|
| 336 |
+
f'fill-opacity="0.5"/>'
|
| 337 |
+
)
|
| 338 |
+
parts.append(
|
| 339 |
+
f'<line x1="{left}" y1="{strip_top + strip_h:.1f}" x2="{left + plot_w:.1f}" '
|
| 340 |
+
f'y2="{strip_top + strip_h:.1f}" stroke="{GRID_STRONG}" stroke-width="1"/>'
|
| 341 |
+
)
|
| 342 |
+
parts.append(_text(left, strip_top + strip_h + 9, counts.label, size=9, fill=MUTED))
|
| 343 |
+
parts.append(
|
| 344 |
+
_text(
|
| 345 |
+
left + plot_w,
|
| 346 |
+
strip_top + strip_h + 9,
|
| 347 |
+
f"tallest {int(peak):,}",
|
| 348 |
+
size=9,
|
| 349 |
+
fill=MUTED,
|
| 350 |
+
anchor="end",
|
| 351 |
+
)
|
| 352 |
+
)
|
| 353 |
+
cursor += _COUNTS_BAND
|
| 354 |
+
for index, note in enumerate(panel.notes):
|
| 355 |
+
parts.append(
|
| 356 |
+
_text(left, cursor + _NOTE_LEAD + 9 + _NOTE_BAND * index, note, size=9.5, fill=MUTED)
|
| 357 |
+
)
|
| 358 |
+
if panel.ylabel:
|
| 359 |
+
parts.append(
|
| 360 |
+
f'<text transform="translate(13,{top + plot_h / 2:.1f}) rotate(-90)" '
|
| 361 |
+
f'text-anchor="middle" font-size="10.5" fill="{MUTED}" {FONT}>'
|
| 362 |
+
f"{escape(panel.ylabel)}</text>"
|
| 363 |
+
)
|
| 364 |
+
return "\n".join(parts)
|
| 365 |
+
|
| 366 |
+
|
| 367 |
+
def _legend_svg(
|
| 368 |
+
entries: list[tuple[str, str, str | None]],
|
| 369 |
+
*,
|
| 370 |
+
x: float,
|
| 371 |
+
y: float,
|
| 372 |
+
max_width: float,
|
| 373 |
+
) -> tuple[str, float]:
|
| 374 |
+
"""Flow legend entries across as many rows as their measured widths need."""
|
| 375 |
+
|
| 376 |
+
swatch, gap, pad = 22.0, 7.0, 22.0
|
| 377 |
+
parts: list[str] = []
|
| 378 |
+
cursor_x, cursor_y, rows = x, y, 1
|
| 379 |
+
for label, color, dash in entries:
|
| 380 |
+
entry_w = swatch + gap + text_width(label, 11) + pad
|
| 381 |
+
if cursor_x > x and cursor_x + entry_w - pad > x + max_width:
|
| 382 |
+
cursor_x, cursor_y, rows = x, cursor_y + 16.0, rows + 1
|
| 383 |
+
dash_attr = f' stroke-dasharray="{dash}"' if dash else ""
|
| 384 |
+
parts.append(
|
| 385 |
+
f'<line x1="{cursor_x:.1f}" y1="{cursor_y:.1f}" x2="{cursor_x + swatch:.1f}" '
|
| 386 |
+
f'y2="{cursor_y:.1f}" stroke="{color}" stroke-width="2.4" '
|
| 387 |
+
f'stroke-linecap="round"{dash_attr}/>'
|
| 388 |
+
)
|
| 389 |
+
parts.append(_text(cursor_x + swatch + gap, cursor_y + 3.8, label, size=11))
|
| 390 |
+
cursor_x += entry_w
|
| 391 |
+
return "\n".join(parts), 16.0 * rows
|
| 392 |
+
|
| 393 |
+
|
| 394 |
+
def _dedupe(entries: list[tuple[str, str, str | None]]) -> list[tuple[str, str, str | None]]:
|
| 395 |
+
seen: set[tuple[str, str, str | None]] = set()
|
| 396 |
+
unique = []
|
| 397 |
+
for entry in entries:
|
| 398 |
+
if entry not in seen:
|
| 399 |
+
seen.add(entry)
|
| 400 |
+
unique.append(entry)
|
| 401 |
+
return unique
|
| 402 |
+
|
| 403 |
+
|
| 404 |
+
def _open_svg(width: float, height: float, title: str, description: str) -> list[str]:
|
| 405 |
+
return [
|
| 406 |
+
f'<svg xmlns="http://www.w3.org/2000/svg" width="{width:.0f}" height="{height:.0f}" '
|
| 407 |
+
f'viewBox="0 0 {width:.0f} {height:.0f}" role="img" aria-label="{escape(title)}">',
|
| 408 |
+
f"<title>{escape(title)}</title>",
|
| 409 |
+
f"<desc>{escape(description)}</desc>",
|
| 410 |
+
'<rect width="100%" height="100%" fill="white"/>',
|
| 411 |
+
]
|
| 412 |
+
|
| 413 |
+
|
| 414 |
+
def render_grid(
|
| 415 |
+
panels: list[Panel],
|
| 416 |
+
*,
|
| 417 |
+
columns: int,
|
| 418 |
+
title: str,
|
| 419 |
+
subtitle: str = "",
|
| 420 |
+
panel_width: float = 352.0,
|
| 421 |
+
panel_height: float = 256.0,
|
| 422 |
+
legend: list[tuple[str, str, str | None]] | None = None,
|
| 423 |
+
description: str = "",
|
| 424 |
+
) -> str:
|
| 425 |
+
"""Render panels on a grid with one shared title and a flowed legend.
|
| 426 |
+
|
| 427 |
+
A final short row is centred, so a five-panel figure on three columns has no empty cell.
|
| 428 |
+
"""
|
| 429 |
+
|
| 430 |
+
if not panels:
|
| 431 |
+
raise ValueError("render_grid needs at least one panel")
|
| 432 |
+
if columns < 1:
|
| 433 |
+
raise ValueError("render_grid needs at least one column")
|
| 434 |
+
legend = _dedupe(legend or [])
|
| 435 |
+
rows = math.ceil(len(panels) / columns)
|
| 436 |
+
width = columns * panel_width
|
| 437 |
+
header = 26.0 + (16.0 if subtitle else 0.0)
|
| 438 |
+
legend_svg, legend_h = "", 0.0
|
| 439 |
+
if legend:
|
| 440 |
+
legend_svg, legend_h = _legend_svg(legend, x=18.0, y=header + 10.0, max_width=width - 36.0)
|
| 441 |
+
legend_h += 8.0
|
| 442 |
+
height = header + legend_h + rows * panel_height
|
| 443 |
+
parts = _open_svg(width, height, title, description or title)
|
| 444 |
+
parts.append(_text(width / 2, 19, title, size=14.5, anchor="middle", weight="700"))
|
| 445 |
+
if subtitle:
|
| 446 |
+
parts.append(_text(width / 2, 34, subtitle, size=11, fill=MUTED, anchor="middle"))
|
| 447 |
+
if legend_svg:
|
| 448 |
+
parts.append(legend_svg)
|
| 449 |
+
# One title row count for the whole grid, so a panel whose title wraps does not push its
|
| 450 |
+
# plot box below its neighbours'.
|
| 451 |
+
title_rows = max(
|
| 452 |
+
len(_fit_lines(panel.title, _TITLE_SIZE, panel_width - 68.0, weight="600"))
|
| 453 |
+
for panel in panels
|
| 454 |
+
)
|
| 455 |
+
for index, panel in enumerate(panels):
|
| 456 |
+
row, column = divmod(index, columns)
|
| 457 |
+
in_row = min(columns, len(panels) - row * columns)
|
| 458 |
+
offset = (columns - in_row) * panel_width / 2.0
|
| 459 |
+
px = offset + column * panel_width
|
| 460 |
+
py = header + legend_h + row * panel_height
|
| 461 |
+
parts.append(f'<g transform="translate({px:.1f},{py:.1f})">')
|
| 462 |
+
parts.append(_panel_svg(panel, panel_width, panel_height, title_rows=title_rows))
|
| 463 |
+
parts.append("</g>")
|
| 464 |
+
parts.append("</svg>")
|
| 465 |
+
return "\n".join(parts) + "\n"
|
| 466 |
+
|
| 467 |
+
|
| 468 |
+
def render_bars(
|
| 469 |
+
groups: list[str],
|
| 470 |
+
series: list[tuple[str, list[float], str]],
|
| 471 |
+
*,
|
| 472 |
+
title: str,
|
| 473 |
+
ylabel: str,
|
| 474 |
+
subtitle: str = "",
|
| 475 |
+
ylim: tuple[float, float] = (0.0, 1.0),
|
| 476 |
+
reference: list[tuple[str, list[float], str]] | None = None,
|
| 477 |
+
notes: list[str] | None = None,
|
| 478 |
+
separator_before: int | None = None,
|
| 479 |
+
width: float = 780.0,
|
| 480 |
+
height: float = 350.0,
|
| 481 |
+
description: str = "",
|
| 482 |
+
) -> str:
|
| 483 |
+
"""Render grouped bars with per-group dashed reference levels.
|
| 484 |
+
|
| 485 |
+
Bars keep a zero baseline. Value labels are drawn only where a bar is wide enough to hold
|
| 486 |
+
one, because a crowded label is worse than none.
|
| 487 |
+
"""
|
| 488 |
+
|
| 489 |
+
if not groups or not series:
|
| 490 |
+
raise ValueError("render_bars needs at least one group and one series")
|
| 491 |
+
if any(len(values) != len(groups) for _, values, _ in series):
|
| 492 |
+
raise ValueError("every bar series must carry one value per group")
|
| 493 |
+
notes = notes or []
|
| 494 |
+
left, right, top = 54.0, 16.0, 26.0 + (15.0 if subtitle else 0.0)
|
| 495 |
+
legend_entries = [(label, color, None) for label, _, color in series]
|
| 496 |
+
legend_entries += [(label, color, "5 3") for label, _, color in reference or []]
|
| 497 |
+
legend_svg, legend_h = _legend_svg(
|
| 498 |
+
_dedupe(legend_entries), x=18.0, y=top + 12.0, max_width=width - 36.0
|
| 499 |
+
)
|
| 500 |
+
top += legend_h + 10.0
|
| 501 |
+
bottom = 46.0 + 12.0 * len(notes)
|
| 502 |
+
plot_w, plot_h = width - left - right, height - top - bottom
|
| 503 |
+
group_w = plot_w / len(groups)
|
| 504 |
+
bar_w = group_w * 0.74 / len(series)
|
| 505 |
+
|
| 506 |
+
def ty(value: float) -> float:
|
| 507 |
+
return top + plot_h - (value - ylim[0]) / (ylim[1] - ylim[0]) * plot_h
|
| 508 |
+
|
| 509 |
+
parts = _open_svg(width, height, title, description or title)
|
| 510 |
+
parts.append(_text(width / 2, 19, title, size=14.5, anchor="middle", weight="700"))
|
| 511 |
+
if subtitle:
|
| 512 |
+
parts.append(_text(width / 2, 34, subtitle, size=11, fill=MUTED, anchor="middle"))
|
| 513 |
+
parts.append(legend_svg)
|
| 514 |
+
for value, label in _auto_ticks(*ylim):
|
| 515 |
+
y = ty(value)
|
| 516 |
+
parts.append(
|
| 517 |
+
f'<line x1="{left}" y1="{y:.1f}" x2="{left + plot_w:.1f}" y2="{y:.1f}" '
|
| 518 |
+
f'stroke="{GRID}" stroke-width="1"/>'
|
| 519 |
+
)
|
| 520 |
+
parts.append(_text(left - 6, y + 3.4, label, size=10, fill=MUTED, anchor="end"))
|
| 521 |
+
|
| 522 |
+
label_fits = bar_w - 2 >= text_width("0.000", 8.5) + 2
|
| 523 |
+
for g_index, group in enumerate(groups):
|
| 524 |
+
gx = left + g_index * group_w + group_w * 0.13
|
| 525 |
+
for s_index, (_, values, color) in enumerate(series):
|
| 526 |
+
value = values[g_index]
|
| 527 |
+
x = gx + s_index * bar_w
|
| 528 |
+
parts.append(
|
| 529 |
+
f'<rect x="{x:.1f}" y="{ty(value):.1f}" width="{bar_w - 2:.1f}" '
|
| 530 |
+
f'height="{ty(ylim[0]) - ty(value):.1f}" fill="{color}"/>'
|
| 531 |
+
)
|
| 532 |
+
if label_fits:
|
| 533 |
+
parts.append(
|
| 534 |
+
_text(
|
| 535 |
+
x + (bar_w - 2) / 2,
|
| 536 |
+
ty(value) - 4,
|
| 537 |
+
f"{value:.3f}",
|
| 538 |
+
size=8.5,
|
| 539 |
+
anchor="middle",
|
| 540 |
+
)
|
| 541 |
+
)
|
| 542 |
+
for _, values, color in reference or []:
|
| 543 |
+
y = ty(values[g_index])
|
| 544 |
+
parts.append(
|
| 545 |
+
f'<line x1="{gx - 3:.1f}" y1="{y:.1f}" '
|
| 546 |
+
f'x2="{gx + group_w * 0.74 + 3:.1f}" y2="{y:.1f}" stroke="{color}" '
|
| 547 |
+
f'stroke-width="1.6" stroke-dasharray="5 3"/>'
|
| 548 |
+
)
|
| 549 |
+
parts.append(
|
| 550 |
+
_text(
|
| 551 |
+
left + g_index * group_w + group_w / 2,
|
| 552 |
+
top + plot_h + 16,
|
| 553 |
+
group,
|
| 554 |
+
size=11,
|
| 555 |
+
anchor="middle",
|
| 556 |
+
)
|
| 557 |
+
)
|
| 558 |
+
if separator_before is not None and 0 < separator_before < len(groups):
|
| 559 |
+
x = left + separator_before * group_w
|
| 560 |
+
parts.append(
|
| 561 |
+
f'<line x1="{x:.1f}" y1="{top}" x2="{x:.1f}" y2="{top + plot_h + 6:.1f}" '
|
| 562 |
+
f'stroke="{GRID_STRONG}" stroke-width="1.4"/>'
|
| 563 |
+
)
|
| 564 |
+
parts.append(
|
| 565 |
+
f'<rect x="{left}" y="{top}" width="{plot_w:.1f}" height="{plot_h:.1f}" fill="none" '
|
| 566 |
+
f'stroke="{AXIS}" stroke-width="1"/>'
|
| 567 |
+
)
|
| 568 |
+
parts.append(
|
| 569 |
+
f'<text transform="translate(13,{top + plot_h / 2:.1f}) rotate(-90)" '
|
| 570 |
+
f'text-anchor="middle" font-size="10.5" fill="{MUTED}" {FONT}>'
|
| 571 |
+
f"{escape(ylabel)}</text>"
|
| 572 |
+
)
|
| 573 |
+
for index, note in enumerate(notes):
|
| 574 |
+
parts.append(_text(left, top + plot_h + 34 + 12 * index, note, size=9.5, fill=MUTED))
|
| 575 |
+
parts.append("</svg>")
|
| 576 |
+
return "\n".join(parts) + "\n"
|
figures/classifier-comparison.svg
ADDED
|
|
model.onnx
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7e9b7e37afc06a23a89ba064e1255363cdc3140ff4b2823f5dae5865945621dc
|
| 3 |
+
size 452831012
|
publication-manifest.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.classifier-publication-manifest",
|
| 3 |
-
"tool_sha256": "
|
| 4 |
"public_repo": "Daecore/gliclass-std-base-v3-5facet-qint8-v2",
|
| 5 |
"classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
|
| 6 |
"upstream": {
|
|
@@ -8,38 +8,39 @@
|
|
| 8 |
"revision": "77a70e6cd52e602ed18184ef37d18bdd3741e3d5"
|
| 9 |
},
|
| 10 |
"nominee_receipt_sha256": "d463fc167b2a26f8f5b56e03c18e4c604b231eb25b5093064dd8a9bd7ffa6373",
|
| 11 |
-
"
|
|
|
|
|
|
|
| 12 |
"files": {
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 13 |
"model.onnx": {
|
| 14 |
-
"sha256": "
|
| 15 |
-
"size":
|
| 16 |
-
"binding": "
|
| 17 |
},
|
| 18 |
"tokenizer.json": {
|
| 19 |
"sha256": "519648948c4c59da1af88f2cf2c8b4f84417b5c673981bc9809abf84cda1b7cc",
|
| 20 |
"size": 8649234,
|
| 21 |
-
"binding": "
|
| 22 |
-
},
|
| 23 |
-
"config.json": {
|
| 24 |
-
"sha256": "9bad60c87ac3e047bc838ca7e6897a6fe265180eeffce4e62aabc1e9c828d606",
|
| 25 |
-
"size": 2118,
|
| 26 |
-
"binding": "candidate manifest (file hash)"
|
| 27 |
},
|
| 28 |
"tokenizer_config.json": {
|
| 29 |
"sha256": "9d5220b355d2cb9a7df69deccca45509ce1c9a857990fb1047c48403b0f6ddfb",
|
| 30 |
"size": 1692,
|
| 31 |
-
"binding": "
|
| 32 |
-
},
|
| 33 |
-
"classifier-metadata.json": {
|
| 34 |
-
"sha256": "00d8c3b36be47f96ddd113bd6ddbf525c3599de27064c68ef7a1aacf4c0408e5",
|
| 35 |
-
"size": 2596,
|
| 36 |
-
"canonical_sha256": "1918aa0c392e6582a7028a4a2482166714080958602f6234f145dd294fb01996",
|
| 37 |
-
"binding": "nominee receipt runtime_metadata_sha256 (canonical JSON hash)"
|
| 38 |
},
|
| 39 |
-
"
|
| 40 |
-
"sha256": "
|
| 41 |
-
"size":
|
| 42 |
-
"binding": "
|
| 43 |
},
|
| 44 |
"LICENSE": {
|
| 45 |
"sha256": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
|
|
@@ -52,9 +53,55 @@
|
|
| 52 |
"binding": "packaging record"
|
| 53 |
},
|
| 54 |
"MODIFICATIONS.md": {
|
| 55 |
-
"sha256": "
|
| 56 |
-
"size":
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
"binding": "packaging record"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
}
|
| 59 |
}
|
| 60 |
}
|
|
|
|
| 1 |
{
|
| 2 |
"schema": "daecore.classifier-publication-manifest",
|
| 3 |
+
"tool_sha256": "5ac2ef9465fe9a3454bfd2e61c7aa651b08fbdfef9473fb5b3455585ba560a7b",
|
| 4 |
"public_repo": "Daecore/gliclass-std-base-v3-5facet-qint8-v2",
|
| 5 |
"classifier_version": "gliclass-std-base-v3-daecore-5facet-qint8-v2",
|
| 6 |
"upstream": {
|
|
|
|
| 8 |
"revision": "77a70e6cd52e602ed18184ef37d18bdd3741e3d5"
|
| 9 |
},
|
| 10 |
"nominee_receipt_sha256": "d463fc167b2a26f8f5b56e03c18e4c604b231eb25b5093064dd8a9bd7ffa6373",
|
| 11 |
+
"derivation_preparation_sha256": "bfd902e06b7a984348283dc279f5e46d76398ff59633a9b5b426f33d872a67d3",
|
| 12 |
+
"source_nominee_binding": "unchanged learned model and calibration; graph bytes separately derived",
|
| 13 |
+
"staged_at": "2026-09-29T03:04:21+00:00",
|
| 14 |
"files": {
|
| 15 |
+
"classifier-metadata.json": {
|
| 16 |
+
"sha256": "736d8b615fb8c4bd7cc36d1c2e0e86bbe096d63dc373d047ca75af9cc6119a11",
|
| 17 |
+
"size": 2674,
|
| 18 |
+
"binding": "pinned graph derivation after original nominee verification"
|
| 19 |
+
},
|
| 20 |
+
"config.json": {
|
| 21 |
+
"sha256": "9bad60c87ac3e047bc838ca7e6897a6fe265180eeffce4e62aabc1e9c828d606",
|
| 22 |
+
"size": 2118,
|
| 23 |
+
"binding": "pinned graph derivation after original nominee verification"
|
| 24 |
+
},
|
| 25 |
"model.onnx": {
|
| 26 |
+
"sha256": "7e9b7e37afc06a23a89ba064e1255363cdc3140ff4b2823f5dae5865945621dc",
|
| 27 |
+
"size": 452831012,
|
| 28 |
+
"binding": "pinned graph derivation after original nominee verification"
|
| 29 |
},
|
| 30 |
"tokenizer.json": {
|
| 31 |
"sha256": "519648948c4c59da1af88f2cf2c8b4f84417b5c673981bc9809abf84cda1b7cc",
|
| 32 |
"size": 8649234,
|
| 33 |
+
"binding": "pinned graph derivation after original nominee verification"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
},
|
| 35 |
"tokenizer_config.json": {
|
| 36 |
"sha256": "9d5220b355d2cb9a7df69deccca45509ce1c9a857990fb1047c48403b0f6ddfb",
|
| 37 |
"size": 1692,
|
| 38 |
+
"binding": "pinned graph derivation after original nominee verification"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
},
|
| 40 |
+
"vulkan-derivation.json": {
|
| 41 |
+
"sha256": "472be46b2c127c4bf38dcf1db3d5dfe65672a6ea66164677d91a2cbb2d05ee70",
|
| 42 |
+
"size": 5782,
|
| 43 |
+
"binding": "pinned graph derivation after original nominee verification"
|
| 44 |
},
|
| 45 |
"LICENSE": {
|
| 46 |
"sha256": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4",
|
|
|
|
| 53 |
"binding": "packaging record"
|
| 54 |
},
|
| 55 |
"MODIFICATIONS.md": {
|
| 56 |
+
"sha256": "429cfef847f3816fe9cc4715d2431f344d5b34376e03b214eae52c1626021f39",
|
| 57 |
+
"size": 1773,
|
| 58 |
+
"binding": "packaging record"
|
| 59 |
+
},
|
| 60 |
+
"evaluation/README.md": {
|
| 61 |
+
"sha256": "3215e6b0e302fdecd86ef1457e46e7d033fa18fb1ebcbd0dcb47c2b8abe7c174",
|
| 62 |
+
"size": 8825,
|
| 63 |
"binding": "packaging record"
|
| 64 |
+
},
|
| 65 |
+
"evaluation/metrics.py": {
|
| 66 |
+
"sha256": "7c25f29a03e0f61d5f0e59f7781269489f80512ae6862920f91a2d333c058bc4",
|
| 67 |
+
"size": 11135,
|
| 68 |
+
"binding": "packaging record"
|
| 69 |
+
},
|
| 70 |
+
"evaluation/classifier.json": {
|
| 71 |
+
"sha256": "5498bb3f5f15f587bc93fc77c9a56308e6ccbc85976bd8e38530b643614628bc",
|
| 72 |
+
"size": 721988,
|
| 73 |
+
"binding": "packaging record"
|
| 74 |
+
},
|
| 75 |
+
"evaluation/classifier-summary.json": {
|
| 76 |
+
"sha256": "e89f48dff2e29f9d8f4c27b221b88c2d380f725a682d6792116bb0d3dff7bd8e",
|
| 77 |
+
"size": 4799,
|
| 78 |
+
"binding": "packaging record"
|
| 79 |
+
},
|
| 80 |
+
"evaluation/serving-qualification.json": {
|
| 81 |
+
"sha256": "bf7deda3f25306d15454798ad4c8e785c25fba970fc774efbe557225fc5d42e0",
|
| 82 |
+
"size": 4667,
|
| 83 |
+
"binding": "packaging record"
|
| 84 |
+
},
|
| 85 |
+
"evaluation/figures.py": {
|
| 86 |
+
"sha256": "9944765ee42148f2bb135f428ee1dacb822af0ca5c77be9f82b507eb8da236d4",
|
| 87 |
+
"size": 4097,
|
| 88 |
+
"binding": "packaging record"
|
| 89 |
+
},
|
| 90 |
+
"evaluation/svg_figures.py": {
|
| 91 |
+
"sha256": "c68b7e26e072e3875f111b592f8031322261272fcf285ae2a3edbe378e70740d",
|
| 92 |
+
"size": 21536,
|
| 93 |
+
"binding": "packaging record"
|
| 94 |
+
},
|
| 95 |
+
"figures/classifier-comparison.svg": {
|
| 96 |
+
"sha256": "a465b6ab22d60b50543c29dd7f16fe440b948eb0d12c38df7facc34fecfee509",
|
| 97 |
+
"size": 8642,
|
| 98 |
+
"binding": "packaging record"
|
| 99 |
+
},
|
| 100 |
+
"README.md": {
|
| 101 |
+
"sha256": "eec5fe5f0d257201d4d517f6a3edb97c5a6cfbe352cc8835c2be42c3baf0ad14",
|
| 102 |
+
"size": 8321,
|
| 103 |
+
"binding": "model card with upload-relative links",
|
| 104 |
+
"source_sha256": "f62d093e04a97c57080285d08d64c66906b500ee2124d86013b2dde0c120b592"
|
| 105 |
}
|
| 106 |
}
|
| 107 |
}
|
vulkan-derivation.json
ADDED
|
@@ -0,0 +1,227 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"artifacts": [
|
| 3 |
+
{
|
| 4 |
+
"bytes": 452831012,
|
| 5 |
+
"path": "model.onnx",
|
| 6 |
+
"sha256": "7e9b7e37afc06a23a89ba064e1255363cdc3140ff4b2823f5dae5865945621dc"
|
| 7 |
+
}
|
| 8 |
+
],
|
| 9 |
+
"changes": [
|
| 10 |
+
{
|
| 11 |
+
"dtype": 7,
|
| 12 |
+
"name": "/inner/model/encoder_model/encoder/Squeeze",
|
| 13 |
+
"operation": "Squeeze"
|
| 14 |
+
},
|
| 15 |
+
{
|
| 16 |
+
"dtype": 7,
|
| 17 |
+
"name": "/inner/model/CumSum",
|
| 18 |
+
"operation": "CumSum"
|
| 19 |
+
},
|
| 20 |
+
{
|
| 21 |
+
"dtype": 7,
|
| 22 |
+
"name": "Sign_1892",
|
| 23 |
+
"operation": "Sign"
|
| 24 |
+
},
|
| 25 |
+
{
|
| 26 |
+
"dtype": 7,
|
| 27 |
+
"name": "Abs_1898",
|
| 28 |
+
"operation": "Abs"
|
| 29 |
+
},
|
| 30 |
+
{
|
| 31 |
+
"dtype": 7,
|
| 32 |
+
"name": "/inner/model/encoder_model/encoder/Slice_1",
|
| 33 |
+
"operation": "Slice"
|
| 34 |
+
},
|
| 35 |
+
{
|
| 36 |
+
"dtype": 7,
|
| 37 |
+
"name": "/inner/model/encoder_model/encoder/layer.0/attention/self/Squeeze_3",
|
| 38 |
+
"operation": "Squeeze"
|
| 39 |
+
},
|
| 40 |
+
{
|
| 41 |
+
"dtype": 7,
|
| 42 |
+
"name": "Identity_2252",
|
| 43 |
+
"operation": "Identity"
|
| 44 |
+
},
|
| 45 |
+
{
|
| 46 |
+
"dtype": 7,
|
| 47 |
+
"name": "/inner/model/encoder_model/encoder/layer.0/attention/self/Neg",
|
| 48 |
+
"operation": "Neg"
|
| 49 |
+
},
|
| 50 |
+
{
|
| 51 |
+
"dtype": 7,
|
| 52 |
+
"name": "/inner/model/encoder_model/encoder/layer.0/attention/self/Squeeze_5",
|
| 53 |
+
"operation": "Squeeze"
|
| 54 |
+
},
|
| 55 |
+
{
|
| 56 |
+
"dtype": 7,
|
| 57 |
+
"name": "Identity_2696",
|
| 58 |
+
"operation": "Identity"
|
| 59 |
+
},
|
| 60 |
+
{
|
| 61 |
+
"dtype": 7,
|
| 62 |
+
"name": "/inner/model/encoder_model/encoder/layer.1/attention/self/Neg",
|
| 63 |
+
"operation": "Neg"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"dtype": 7,
|
| 67 |
+
"name": "/inner/model/encoder_model/encoder/layer.1/attention/self/Squeeze_3",
|
| 68 |
+
"operation": "Squeeze"
|
| 69 |
+
},
|
| 70 |
+
{
|
| 71 |
+
"dtype": 7,
|
| 72 |
+
"name": "Identity_3138",
|
| 73 |
+
"operation": "Identity"
|
| 74 |
+
},
|
| 75 |
+
{
|
| 76 |
+
"dtype": 7,
|
| 77 |
+
"name": "/inner/model/encoder_model/encoder/layer.2/attention/self/Neg",
|
| 78 |
+
"operation": "Neg"
|
| 79 |
+
},
|
| 80 |
+
{
|
| 81 |
+
"dtype": 7,
|
| 82 |
+
"name": "/inner/model/encoder_model/encoder/layer.2/attention/self/Squeeze_3",
|
| 83 |
+
"operation": "Squeeze"
|
| 84 |
+
},
|
| 85 |
+
{
|
| 86 |
+
"dtype": 7,
|
| 87 |
+
"name": "Identity_3580",
|
| 88 |
+
"operation": "Identity"
|
| 89 |
+
},
|
| 90 |
+
{
|
| 91 |
+
"dtype": 7,
|
| 92 |
+
"name": "/inner/model/encoder_model/encoder/layer.3/attention/self/Neg",
|
| 93 |
+
"operation": "Neg"
|
| 94 |
+
},
|
| 95 |
+
{
|
| 96 |
+
"dtype": 7,
|
| 97 |
+
"name": "/inner/model/encoder_model/encoder/layer.3/attention/self/Squeeze_3",
|
| 98 |
+
"operation": "Squeeze"
|
| 99 |
+
},
|
| 100 |
+
{
|
| 101 |
+
"dtype": 7,
|
| 102 |
+
"name": "Identity_4022",
|
| 103 |
+
"operation": "Identity"
|
| 104 |
+
},
|
| 105 |
+
{
|
| 106 |
+
"dtype": 7,
|
| 107 |
+
"name": "/inner/model/encoder_model/encoder/layer.4/attention/self/Neg",
|
| 108 |
+
"operation": "Neg"
|
| 109 |
+
},
|
| 110 |
+
{
|
| 111 |
+
"dtype": 7,
|
| 112 |
+
"name": "/inner/model/encoder_model/encoder/layer.4/attention/self/Squeeze_3",
|
| 113 |
+
"operation": "Squeeze"
|
| 114 |
+
},
|
| 115 |
+
{
|
| 116 |
+
"dtype": 7,
|
| 117 |
+
"name": "Identity_4464",
|
| 118 |
+
"operation": "Identity"
|
| 119 |
+
},
|
| 120 |
+
{
|
| 121 |
+
"dtype": 7,
|
| 122 |
+
"name": "/inner/model/encoder_model/encoder/layer.5/attention/self/Neg",
|
| 123 |
+
"operation": "Neg"
|
| 124 |
+
},
|
| 125 |
+
{
|
| 126 |
+
"dtype": 7,
|
| 127 |
+
"name": "/inner/model/encoder_model/encoder/layer.5/attention/self/Squeeze_3",
|
| 128 |
+
"operation": "Squeeze"
|
| 129 |
+
},
|
| 130 |
+
{
|
| 131 |
+
"dtype": 7,
|
| 132 |
+
"name": "Identity_4906",
|
| 133 |
+
"operation": "Identity"
|
| 134 |
+
},
|
| 135 |
+
{
|
| 136 |
+
"dtype": 7,
|
| 137 |
+
"name": "/inner/model/encoder_model/encoder/layer.6/attention/self/Neg",
|
| 138 |
+
"operation": "Neg"
|
| 139 |
+
},
|
| 140 |
+
{
|
| 141 |
+
"dtype": 7,
|
| 142 |
+
"name": "/inner/model/encoder_model/encoder/layer.6/attention/self/Squeeze_3",
|
| 143 |
+
"operation": "Squeeze"
|
| 144 |
+
},
|
| 145 |
+
{
|
| 146 |
+
"dtype": 7,
|
| 147 |
+
"name": "Identity_5348",
|
| 148 |
+
"operation": "Identity"
|
| 149 |
+
},
|
| 150 |
+
{
|
| 151 |
+
"dtype": 7,
|
| 152 |
+
"name": "/inner/model/encoder_model/encoder/layer.7/attention/self/Neg",
|
| 153 |
+
"operation": "Neg"
|
| 154 |
+
},
|
| 155 |
+
{
|
| 156 |
+
"dtype": 7,
|
| 157 |
+
"name": "/inner/model/encoder_model/encoder/layer.7/attention/self/Squeeze_3",
|
| 158 |
+
"operation": "Squeeze"
|
| 159 |
+
},
|
| 160 |
+
{
|
| 161 |
+
"dtype": 7,
|
| 162 |
+
"name": "Identity_5790",
|
| 163 |
+
"operation": "Identity"
|
| 164 |
+
},
|
| 165 |
+
{
|
| 166 |
+
"dtype": 7,
|
| 167 |
+
"name": "/inner/model/encoder_model/encoder/layer.8/attention/self/Neg",
|
| 168 |
+
"operation": "Neg"
|
| 169 |
+
},
|
| 170 |
+
{
|
| 171 |
+
"dtype": 7,
|
| 172 |
+
"name": "/inner/model/encoder_model/encoder/layer.8/attention/self/Squeeze_3",
|
| 173 |
+
"operation": "Squeeze"
|
| 174 |
+
},
|
| 175 |
+
{
|
| 176 |
+
"dtype": 7,
|
| 177 |
+
"name": "Identity_6232",
|
| 178 |
+
"operation": "Identity"
|
| 179 |
+
},
|
| 180 |
+
{
|
| 181 |
+
"dtype": 7,
|
| 182 |
+
"name": "/inner/model/encoder_model/encoder/layer.9/attention/self/Neg",
|
| 183 |
+
"operation": "Neg"
|
| 184 |
+
},
|
| 185 |
+
{
|
| 186 |
+
"dtype": 7,
|
| 187 |
+
"name": "/inner/model/encoder_model/encoder/layer.9/attention/self/Squeeze_3",
|
| 188 |
+
"operation": "Squeeze"
|
| 189 |
+
},
|
| 190 |
+
{
|
| 191 |
+
"dtype": 7,
|
| 192 |
+
"name": "Identity_6674",
|
| 193 |
+
"operation": "Identity"
|
| 194 |
+
},
|
| 195 |
+
{
|
| 196 |
+
"dtype": 7,
|
| 197 |
+
"name": "/inner/model/encoder_model/encoder/layer.10/attention/self/Neg",
|
| 198 |
+
"operation": "Neg"
|
| 199 |
+
},
|
| 200 |
+
{
|
| 201 |
+
"dtype": 7,
|
| 202 |
+
"name": "/inner/model/encoder_model/encoder/layer.10/attention/self/Squeeze_3",
|
| 203 |
+
"operation": "Squeeze"
|
| 204 |
+
},
|
| 205 |
+
{
|
| 206 |
+
"dtype": 7,
|
| 207 |
+
"name": "Identity_7116",
|
| 208 |
+
"operation": "Identity"
|
| 209 |
+
},
|
| 210 |
+
{
|
| 211 |
+
"dtype": 7,
|
| 212 |
+
"name": "/inner/model/encoder_model/encoder/layer.11/attention/self/Neg",
|
| 213 |
+
"operation": "Neg"
|
| 214 |
+
},
|
| 215 |
+
{
|
| 216 |
+
"dtype": 7,
|
| 217 |
+
"name": "/inner/model/encoder_model/encoder/layer.11/attention/self/Squeeze_3",
|
| 218 |
+
"operation": "Squeeze"
|
| 219 |
+
}
|
| 220 |
+
],
|
| 221 |
+
"max_sequence": 768,
|
| 222 |
+
"role": "gliclass",
|
| 223 |
+
"schema": "daecore.vulkan-mask-derivation.v1",
|
| 224 |
+
"source_graph_sha256": "690e50920e7780db5ecd14e4b209f0d3c86199214063e8bb59d399c93861324c",
|
| 225 |
+
"source_weights_sha256": null,
|
| 226 |
+
"weights_unchanged": true
|
| 227 |
+
}
|