Qwen3.6-35B-A3B — APEX GGUF (torch imatrix)

MoE-aware, mixed-precision APEX quantization of Qwen/Qwen3.6-35B-A3B — architecturally identical to Qwen3.5-35B-A3B (same qwen3_5_moe: 40 layers, 256 routed + 1 shared expert, hybrid GatedDeltaNet + periodic full attention, NextN/MTP head).

The importance matrix here is generated with a PyTorch band-serialized generator rather than llama-imatrix, for two concrete reasons on this architecture:

  1. llama.cpp's imatrix tool is impractical for qwen35moe. The GatedDeltaNet linear-attention is a serial state-space recurrence; the imatrix collection callback breaks the GPU path that makes normal inference fast, so it falls back to a single CPU thread that no thread count can parallelize.
  2. The torch generator covers the MTP/NextN head (blk.40.*) that llama-imatrix does not. On Qwen3.6 the MTP head is a full MoE decoder layer with fused experts; both the GGUF converter and the imatrix generator handle that layout.

The -torch imatrix — what it is

Qwen3.6-35B-A3B-torch.imatrix is a standard GGUF-format importance matrix (in_sum2 + counts per tensor), bit-compatible with llama-quantize --imatrix. Generated from the HF safetensors with a band-serialized PyTorch forward on a general text corpus, covering all 40 transformer layers plus the NextN/MTP head.

Validation vs a reference llama.cpp imatrix

No public llama-imatrix exists for Qwen3.6 (and it is impractical to compute locally, see above). Since the imatrix's per-channel importance is largely architecture-driven, we validate per-tensor against bartowski's canonical Qwen3.5 llama.cpp imatrix (same arch):

Metric Value
Tensors covered (torch) 523 (incl. 13 MTP-head tensors)
Per-tensor correlation (median) 0.956
Per-tensor correlation (mean) 0.861
Lowest-correlation tensors blk.*.ssm_out.weight only

Median 0.956 against an independently-computed imatrix of the same architecture (different fine-tune and different calibration data) confirms the mapping/values; the only weakly-correlated tensors are ssm_out (post-nonlinearity, does not affect quantization quality). Torch coverage is a strict superset (only-real = []).

PPL parity (measured on Qwen3.5, same architecture)

The torch-vs-llama.cpp PPL parity was measured on the sibling Qwen3.5 (identical qwen3_5_moe architecture): i-compact quantized with each imatrix, perplexity over 200×512-token wikitext-2 windows —

i-compact quantized with PPL Δ vs bf16
bf16 (reference) 6.620
bartowski llama.cpp imatrix 6.756 +2.05%
torch imatrix 6.775 +2.34%

The torch and llama.cpp quants differ by 0.02 PPL — inside the ±0.073 error bars (statistically indistinguishable). Since Qwen3.6 shares the architecture and the torch imatrix correlates at 0.956 here, the same parity holds.

Confirmed on a diverse code-heavy corpus (HumanEval + MBPP + GSM8K + prose, 342 windows), on the sibling Qwen3.5 i-compact: bf16 2.247 · llama.cpp-imatrix 2.310 · torch-imatrix 2.318 — parity holds (Δ0.008, inside ±0.014). The imatrix records per-channel activation magnitudes, so the method is domain-agnostic.

Calibration used the diverse calibration_datav3 (prose + code + multilingual), so these quants are not domain-handicapped. The imatrix's calibration corpus is a knob: a code-weighted calibration would favor coding channels further, at a small cost elsewhere — useful if you're specializing for a single domain.

Files

  • Qwen3.6-35B-A3B-torch.imatrix — the PyTorch-generated importance matrix (this repo).
  • Qwen3.6-35B-A3B-APEX-i-quality-torch.gguf (23.5 GB) — largest/highest-fidelity tier.
  • Qwen3.6-35B-A3B-APEX-i-compact-torch.gguf (17.4 GB) — smaller, more aggressive tier.
  • llama.cpp reference imatrix — not re-hosted; see bartowski/Qwen_Qwen3.5-35B-A3B-GGUF.

Attribution

Unofficial community quantization; not affiliated with or endorsed by Qwen.

Downloads last month
2,228
GGUF
Model size
36B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Myric/Qwen3.6-35B-A3B-APEX-GGUF

Quantized
(759)
this model