AI & ML interests

Canada Quant Labs — Canada's open-weight model lab. We train, quantize, and ship sovereign reference models for regulated industries (legal, medical, defence, finance) on a DGX B300 at Equinix Vancouver. Upstream contributors to vLLM and llm-compressor. Recipes: W4A16, NVFP4, MXFP4. Built in Victoria, BC. partnerships@cql.ca · cql.ca

Recent Activity

Organization Card

Canada Quant Labs

Canada's open-weight model lab.

We train, quantize, and deploy sovereign AI models on Canadian Blackwell silicon — for the regulated industries that can't run on someone else's API.

What we do

  • Post-training on open base models (SFT, DPO, GRPO, RLAIF)
  • Production quantization recipes (W4A16, NVFP4, MXFP4)
  • Self-trained speculative-decoding drafters (DFlash2) for the models we serve — three generations for GLM-5.3-Flash
  • Audited, air-gapped deployment with eval evidence and MRM docs

Where we work

  • Legal · Medical · Defence · Finance
  • Headquarters: Victoria, BC
  • Compute: NVIDIA DGX B300 at Equinix Vancouver · 2× DGX Spark (GB10, SM121) serving cluster

Upstream

  • Contributors to vLLM, llm-compressor, compressed-tensors

Open artifacts — 10 checkpoints on this org, every one shipped with its recipe, eval methodology, and engineering log

Partnerships · partnerships@cql.ca Press · press@cql.ca Web · cql.ca


GLM-5.3-Flash DFlash2-G — our own speculative drafter · Sep 2026

Our self-trained drafter, not a quant: a DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash, trained against our own W4A16-MTP quant on 736,675 self-generated samples. Every completion was regenerated by the target model itself, from public-instruction-set prompts including real user chat, tool calling and hard math. No third-party drafter weights or traces anywhere in the training path. Apache-2.0.

→ canada-quant/GLM-5.3-Flash-DFlash2-G · serving stack: canada-quant/vllm-glm53-flash-sm121 (Apache-2.0 — one-command serving of GLM-5.3-Flash W4A16 + DFlash2 drafter on 2× DGX Spark: TP=2 + expert-parallel, 262K context, fp8 KV, K=7; its drafter/ folder holds the eval kit and every raw measurement)

Measured — ~1.84B parameters, block size 8 → K=7 speculative tokens, 9 hidden-state taps on the target. Same hardware, same protocol (500 held-out prompts, thinking ON, 8× B300): 3.676 mean acceptance at K=7 vs 3.632 for the incoai reference, and 3.697 vs 3.667 single-stream. Both leads are outside the ±0.015 run-to-run noise; G is the first of our drafters to beat the reference on acceptance. Where it is still level or short: K=4 3.088 vs 3.136; 4,096-token traces 3.478 vs 3.473 (parity); raw c16 throughput 1,410 vs 1,425 tok/s. Lineage: -E 3.561 (15 Sep) → -F 3.626 (24 Sep, parity) → -G 3.676 (25 Sep). All three stay up with their records. G has been the drafter-of-record on our 2× DGX Spark stack since 25 Sep, and it is the latency drafter in the 8× H100 recipe below (181.2 tok/s single-stream, +34% vs TP=8 spec-off).

Memory and long context (corrected 4 Oct 2026) — E, F and G have eight full-attention layers, so the drafter's cache grows with the whole context: about 3× the KV cache per token (366,749 tokens in an 8 GiB pool and 888,729 in 16 GiB, against 1,360,420 in 9 GiB with the reference drafter). We run G at 262K on 2× DGX Spark, up to 800K with a 16 GiB pool; 1M on Spark is validated with the reference drafter only. For long context or non-English text, use the model's built-in MTP head (k=3). A community user also found G 20–30% slower on non-English text and faster on English code. Our own same-pair Spark comparison has been running since 4 Oct.

Writeup: cql.ca/news/glm-5-3-flash-h100-nvfp4-vision.html.


GLM-5.3-Flash W4A16-MTP · Aug 2026

INT4 weight-only quantization of GLM-5.3-Flash with the BF16 MTP draft head preserved for speculative decoding — 177.7 GiB (−70%). It serves the full 1M-token context on 2× DGX Spark (validated with the reference drafter) or 2× H200, and 262K–512K on 4- and 8-GPU H100, H200 and RTX PRO 6000 configurations. MIT.

→ canada-quant/GLM-5.3-Flash-W4A16-MTP

Measured — only the 36,288 routed-expert GEMMs are INT4 (group-128, GPTQ); attention, router, shared experts, embeddings, the vision tower, and the MTP head stay BF16.

  • 8× H100, vs NVIDIA's own NVFP4 checkpoint on the same rig: 2,418 vs 2,299.84 at c256 (+5.1%), both without speculative decoding; c128 is a dead heat. Single-stream 181.2 vs 122.56 tok/s (+47.8%) is our drafter recipe against NVFP4 run without speculation. The checkpoint does ship an MTP layer, which we missed (corrected 4 Oct 2026); a with-MTP NVFP4 arm on H100 has not been run. On Hopper, NVFP4 runs through FP4 emulation.
  • First formal vision evals (same rig, lmms-eval 0.7.3 specs): MMMU val 0.747 vs 0.68 NVFP4 · OCRBench 888 vs 882.
  • Quality: AIME 2026 85.0% (102/120, +5.0 pt over the EXL3 arm, matched protocol); AIME 2025 0.883 vs NVFP4 0.900 on H100 (within noise); GPQA-Diamond within noise; GSM8K 0.97.
  • 2× DGX Spark, honestly: vs EXL3 we are +10.4% single-stream and +31–39% on long prefill, while EXL3 leads mid-concurrency. In our 27 Sep matched head-to-head, NVIDIA's NVFP4 checkpoint on its own serving stack (with the incoai drafter) led all 28 cells, widest at deep context. Our one measured win there was stability at depth.
  • Published 4× H200 recipe: 195 → 1,954 output tok/s from c1 to c256; 2× H200 admits a 929K-token prompt on two GPUs. RTX PRO 6000: parity with NVFP4 (±0.7%).

Every grid is in the model card's BENCHMARKS.md. Writeups: launch · a month in.


Hy3 W4A16-MTP · Jul 2026

A 4-bit weight quantization of Tencent Hy3 (295B-parameter MoE, 21B active) that keeps the multi-token-prediction (MTP) draft layer in BF16. It matches the FP8 release on quality at 57% of the footprint (≈598 GB BF16 → 172 GB), runs on 8×H100-80GB with headroom (4×H100 at TP=4), and its preserved MTP layer is worth +40% single-stream throughput over the same-scheme quant that shipped without it. Apache-2.0.

→ canada-quant/hy3-w4a16-mtp · calibration data: canada-quant/hy3-w4a16-mtp-calibration

Quality is statistically indistinguishable from tencent/Hy3-FP8 on the same harness (8×H100): AIME24/25 70.0/83.3, MATH-500 94.0, GPQA-diamond 88.4. Speed: 144 output tok/s at concurrency 1 with MTP k=2; k=1 tops the whole 12-config matrix at c=8 (822 tok/s) and c=32 (2,146 tok/s, +21% over FP8). Serving: stock vLLM ≥ 0.25, no patches; native 256K context validated.

Writeup: cql.ca/news/hy3-w4a16-mtp.html.


GLM-5.2 W4A16-MTP · Jun 2026

GLM-5.2 (744B MoE, ~40B active, MLA + DeepSeek Sparse Attention, native 1M context, MIT). We shipped one artifact: W4A16 INT4 on the routed experts with the BF16 MTP draft head preserved for speculative decoding (attention, dense layers 0–2, shared experts, router, embeddings, lm_head, and MTP layer 78 stay BF16).

Repo Routed experts MTP On-disk Hardware Notes
GLM-5.2-W4A16-MTP W4A16 INT4 g=128 (GPTQ) yes (BF16, layer 78) ~405 GB 4×H200 (≤128K) · 8×H200 (1M) matches FP8 quality; fastest 4-bit GLM-5.2 quant at interactive concurrency

Validated against the official FP8 release (same harness, 8×H200): matches on GSM8K (0.960), IFEval (0.909/0.911), MATH-500 (0.954), RULER@32K/64K, and SWE-bench Verified (82.0% vs 82.2%) — serves at 1M context on 8×H200, and on 4×H200 up to ~128K. Throughput vs the most-downloaded community 4-bit quants: leads the interactive regime — +69–79% at concurrency 1 vs AWQ-INT4 / NVFP4 — from the MTP draft head (neither competitor ships one); the no-MTP quants edge ahead only at full saturation.

Writeup: cql.ca/news/glm-5-2-w4a16-mtp.html.


DeepSeek-V4 quantization family · May 2026

Four artifacts in the same lineage. One base model in two sizes (V4-Flash, V4-Pro); two routed-expert formats (W4A16, NVFP4); MTP draft head retained on three of four. Attention is FP8 block 128×128 across all four. Upstream reference recipes: RedHatAI/DeepSeek-V4-Flash-NVFP4-FP8 (Flash NVFP4) and nvidia/DeepSeek-V3.2-NVFP4 (Pro NVFP4, MTP-exclusion topology).

Model Base Routed experts MTP On-disk Min hardware (TP=2) When to pick
DeepSeek-V4-Flash-W4A16-FP8 V4-Flash W4A16 INT4 g=128 no ≈143 GB H200 / DGX Spark / RTX PRO 6000 maximum compatibility, no MTP needed
DeepSeek-V4-Flash-W4A16-FP8-MTP V4-Flash W4A16 INT4 g=128 yes (BF16) 159 GB H200 / RTX PRO 6000 best $/token interactive on V4-Flash
DeepSeek-V4-Flash-NVFP4-FP8-MTP V4-Flash NVFP4 g=16 yes (BF16) 172 GB RTX PRO 6000 / B300 best Blackwell-native interactive on V4-Flash
DeepSeek-V4-Pro-NVFP4-FP8-MTP V4-Pro NVFP4 g=16 yes (byte-identical) 913 GiB 8× B300 (TP=8 + EP) only choice for V4-Pro deployment; +25–37% throughput vs upstream MXFP4

Hardware shorthand

  • H200 — 8× NVIDIA H200 SXM5 (Hopper SM 9.0a, 141 GB HBM3e/GPU)
  • DGX Spark — 2× NVIDIA DGX Spark (GB10, Blackwell SM 12.1a)
  • RTX PRO 6000 — NVIDIA RTX PRO 6000 Blackwell Server Edition (SM 12.0, sm_120, 96 GB HBM)
  • B300 — NVIDIA B300 SXM6 AC (Blackwell SM 10.3, sm_103a, 288 GB HBM3e/GPU)

Reproduction repos

Every artifact has a public reproduction repo with calibration scripts, vLLM patches, bench harnesses, and findings docs:

Upstream contributions filed during this work

  • vLLM: PRs #42209 (merged — NVFP4 MoE for DSV4), #43248, #43288, #43290, #43319, #43467, #41511, #41700 (landed via jasl/vllm@1d6f5c4)
  • vLLM issues filed from the GLM-5.3-Flash work: #55800 (DFlash2 sliding-window drafter KV admission wedge, fix validated to 256K prompts), #56064 (Marlin no-split-K illegal memory access on SM121)
  • llm-compressor: #2745
  • compressed-tensors: #711