AI & ML interests
Canada Quant Labs — Canada's open-weight model lab. We train, quantize, and ship sovereign reference models for regulated industries (legal, medical, defence, finance) on a DGX B300 at Equinix Vancouver. Upstream contributors to vLLM and llm-compressor. Recipes: W4A16, NVFP4, MXFP4. Built in Victoria, BC. partnerships@cql.ca · cql.ca
Recent Activity
Canada Quant Labs
Canada's open-weight model lab.
We train, quantize, and deploy sovereign AI models on Canadian Blackwell silicon — for the regulated industries that can't run on someone else's API.
What we do
- Post-training on open base models (SFT, DPO, GRPO, RLAIF)
- Production quantization recipes (W4A16, NVFP4, MXFP4)
- Self-trained speculative-decoding drafters (DFlash2) for the models we serve — three generations for GLM-5.3-Flash
- Audited, air-gapped deployment with eval evidence and MRM docs
Where we work
- Legal · Medical · Defence · Finance
- Headquarters: Victoria, BC
- Compute: NVIDIA DGX B300 at Equinix Vancouver · 2× DGX Spark (GB10, SM121) serving cluster
Upstream
- Contributors to vLLM, llm-compressor, compressed-tensors
Open artifacts — 10 checkpoints on this org, every one shipped with its recipe, eval methodology, and engineering log
Partnerships · partnerships@cql.ca Press · press@cql.ca Web · cql.ca
GLM-5.3-Flash DFlash2-G — our own speculative drafter · Sep 2026
Our self-trained drafter, not a quant: a DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash, trained against our own W4A16-MTP quant on 736,675 self-generated samples. Every completion was regenerated by the target model itself, from public-instruction-set prompts including real user chat, tool calling and hard math. No third-party drafter weights or traces anywhere in the training path. Apache-2.0.
→ canada-quant/GLM-5.3-Flash-DFlash2-G · serving stack: canada-quant/vllm-glm53-flash-sm121 (Apache-2.0 — one-command serving of GLM-5.3-Flash W4A16 + DFlash2 drafter on 2× DGX Spark: TP=2 + expert-parallel, 262K context, fp8 KV, K=7; its drafter/ folder holds the eval kit and every raw measurement)
Measured — ~1.84B parameters, block size 8 → K=7 speculative tokens, 9 hidden-state taps on the target. Same hardware, same protocol (500 held-out prompts, thinking ON, 8× B300): 3.676 mean acceptance at K=7 vs 3.632 for the incoai reference, and 3.697 vs 3.667 single-stream. Both leads are outside the ±0.015 run-to-run noise; G is the first of our drafters to beat the reference on acceptance. Where it is still level or short: K=4 3.088 vs 3.136; 4,096-token traces 3.478 vs 3.473 (parity); raw c16 throughput 1,410 vs 1,425 tok/s. Lineage: -E 3.561 (15 Sep) → -F 3.626 (24 Sep, parity) → -G 3.676 (25 Sep). All three stay up with their records. G has been the drafter-of-record on our 2× DGX Spark stack since 25 Sep, and it is the latency drafter in the 8× H100 recipe below (181.2 tok/s single-stream, +34% vs TP=8 spec-off).
Memory and long context (corrected 4 Oct 2026) — E, F and G have eight full-attention layers, so the drafter's cache grows with the whole context: about 3× the KV cache per token (366,749 tokens in an 8 GiB pool and 888,729 in 16 GiB, against 1,360,420 in 9 GiB with the reference drafter). We run G at 262K on 2× DGX Spark, up to 800K with a 16 GiB pool; 1M on Spark is validated with the reference drafter only. For long context or non-English text, use the model's built-in MTP head (k=3). A community user also found G 20–30% slower on non-English text and faster on English code. Our own same-pair Spark comparison has been running since 4 Oct.
Writeup: cql.ca/news/glm-5-3-flash-h100-nvfp4-vision.html.
GLM-5.3-Flash W4A16-MTP · Aug 2026
INT4 weight-only quantization of GLM-5.3-Flash with the BF16 MTP draft head preserved for speculative decoding — 177.7 GiB (−70%). It serves the full 1M-token context on 2× DGX Spark (validated with the reference drafter) or 2× H200, and 262K–512K on 4- and 8-GPU H100, H200 and RTX PRO 6000 configurations. MIT.
→ canada-quant/GLM-5.3-Flash-W4A16-MTP
Measured — only the 36,288 routed-expert GEMMs are INT4 (group-128, GPTQ); attention, router, shared experts, embeddings, the vision tower, and the MTP head stay BF16.
- 8× H100, vs NVIDIA's own NVFP4 checkpoint on the same rig: 2,418 vs 2,299.84 at c256 (+5.1%), both without speculative decoding; c128 is a dead heat. Single-stream 181.2 vs 122.56 tok/s (+47.8%) is our drafter recipe against NVFP4 run without speculation. The checkpoint does ship an MTP layer, which we missed (corrected 4 Oct 2026); a with-MTP NVFP4 arm on H100 has not been run. On Hopper, NVFP4 runs through FP4 emulation.
- First formal vision evals (same rig, lmms-eval 0.7.3 specs): MMMU val 0.747 vs 0.68 NVFP4 · OCRBench 888 vs 882.
- Quality: AIME 2026 85.0% (102/120, +5.0 pt over the EXL3 arm, matched protocol); AIME 2025 0.883 vs NVFP4 0.900 on H100 (within noise); GPQA-Diamond within noise; GSM8K 0.97.
- 2× DGX Spark, honestly: vs EXL3 we are +10.4% single-stream and +31–39% on long prefill, while EXL3 leads mid-concurrency. In our 27 Sep matched head-to-head, NVIDIA's NVFP4 checkpoint on its own serving stack (with the incoai drafter) led all 28 cells, widest at deep context. Our one measured win there was stability at depth.
- Published 4× H200 recipe: 195 → 1,954 output tok/s from c1 to c256; 2× H200 admits a 929K-token prompt on two GPUs. RTX PRO 6000: parity with NVFP4 (±0.7%).
Every grid is in the model card's BENCHMARKS.md. Writeups: launch · a month in.
Hy3 W4A16-MTP · Jul 2026
A 4-bit weight quantization of Tencent Hy3 (295B-parameter MoE, 21B active) that keeps the multi-token-prediction (MTP) draft layer in BF16. It matches the FP8 release on quality at 57% of the footprint (≈598 GB BF16 → 172 GB), runs on 8×H100-80GB with headroom (4×H100 at TP=4), and its preserved MTP layer is worth +40% single-stream throughput over the same-scheme quant that shipped without it. Apache-2.0.
→ canada-quant/hy3-w4a16-mtp · calibration data: canada-quant/hy3-w4a16-mtp-calibration
Quality is statistically indistinguishable from tencent/Hy3-FP8 on the same harness (8×H100): AIME24/25 70.0/83.3, MATH-500 94.0, GPQA-diamond 88.4. Speed: 144 output tok/s at concurrency 1 with MTP k=2; k=1 tops the whole 12-config matrix at c=8 (822 tok/s) and c=32 (2,146 tok/s, +21% over FP8). Serving: stock vLLM ≥ 0.25, no patches; native 256K context validated.
Writeup: cql.ca/news/hy3-w4a16-mtp.html.
GLM-5.2 W4A16-MTP · Jun 2026
GLM-5.2 (744B MoE, ~40B active, MLA + DeepSeek Sparse Attention, native 1M context, MIT). We shipped one artifact: W4A16 INT4 on the routed experts with the BF16 MTP draft head preserved for speculative decoding (attention, dense layers 0–2, shared experts, router, embeddings, lm_head, and MTP layer 78 stay BF16).
| Repo | Routed experts | MTP | On-disk | Hardware | Notes |
|---|---|---|---|---|---|
| GLM-5.2-W4A16-MTP | W4A16 INT4 g=128 (GPTQ) | yes (BF16, layer 78) | ~405 GB | 4×H200 (≤128K) · 8×H200 (1M) | matches FP8 quality; fastest 4-bit GLM-5.2 quant at interactive concurrency |
Validated against the official FP8 release (same harness, 8×H200): matches on GSM8K (0.960), IFEval (0.909/0.911), MATH-500 (0.954), RULER@32K/64K, and SWE-bench Verified (82.0% vs 82.2%) — serves at 1M context on 8×H200, and on 4×H200 up to ~128K. Throughput vs the most-downloaded community 4-bit quants: leads the interactive regime — +69–79% at concurrency 1 vs AWQ-INT4 / NVFP4 — from the MTP draft head (neither competitor ships one); the no-MTP quants edge ahead only at full saturation.
Writeup: cql.ca/news/glm-5-2-w4a16-mtp.html.
DeepSeek-V4 quantization family · May 2026
Four artifacts in the same lineage. One base model in two sizes (V4-Flash, V4-Pro); two routed-expert formats (W4A16, NVFP4); MTP draft head retained on three of four. Attention is FP8 block 128×128 across all four. Upstream reference recipes: RedHatAI/DeepSeek-V4-Flash-NVFP4-FP8 (Flash NVFP4) and nvidia/DeepSeek-V3.2-NVFP4 (Pro NVFP4, MTP-exclusion topology).
| Model | Base | Routed experts | MTP | On-disk | Min hardware (TP=2) | When to pick |
|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash-W4A16-FP8 | V4-Flash | W4A16 INT4 g=128 | no | ≈143 GB | H200 / DGX Spark / RTX PRO 6000 | maximum compatibility, no MTP needed |
| DeepSeek-V4-Flash-W4A16-FP8-MTP | V4-Flash | W4A16 INT4 g=128 | yes (BF16) | 159 GB | H200 / RTX PRO 6000 | best $/token interactive on V4-Flash |
| DeepSeek-V4-Flash-NVFP4-FP8-MTP | V4-Flash | NVFP4 g=16 | yes (BF16) | 172 GB | RTX PRO 6000 / B300 | best Blackwell-native interactive on V4-Flash |
| DeepSeek-V4-Pro-NVFP4-FP8-MTP | V4-Pro | NVFP4 g=16 | yes (byte-identical) | 913 GiB | 8× B300 (TP=8 + EP) | only choice for V4-Pro deployment; +25–37% throughput vs upstream MXFP4 |
Hardware shorthand
- H200 — 8× NVIDIA H200 SXM5 (Hopper SM 9.0a, 141 GB HBM3e/GPU)
- DGX Spark — 2× NVIDIA DGX Spark (GB10, Blackwell SM 12.1a)
- RTX PRO 6000 — NVIDIA RTX PRO 6000 Blackwell Server Edition (SM 12.0, sm_120, 96 GB HBM)
- B300 — NVIDIA B300 SXM6 AC (Blackwell SM 10.3, sm_103a, 288 GB HBM3e/GPU)
Reproduction repos
Every artifact has a public reproduction repo with calibration scripts, vLLM patches, bench harnesses, and findings docs:
canada-quant/dsv4-flash-w4a16-fp8canada-quant/dsv4-flash-w4a16-fp8-mtpcanada-quant/dsv4-flash-nvfp4-fp8-mtpcanada-quant/dsv4-pro-nvfp4-fp8-mtp(recipe repo not yet public)
Upstream contributions filed during this work
- vLLM: PRs #42209 (merged — NVFP4 MoE for DSV4), #43248, #43288, #43290, #43319, #43467, #41511, #41700 (landed via
jasl/vllm@1d6f5c4) - vLLM issues filed from the GLM-5.3-Flash work: #55800 (DFlash2 sliding-window drafter KV admission wedge, fix validated to 256K prompts), #56064 (Marlin no-split-K illegal memory access on SM121)
- llm-compressor: #2745
- compressed-tensors: #711