Atlas3D/JEV-27B-VL-NVFP4
Accuracy warning: this 4-bit build is less accurate than the 8-bit releases. On our decision test its top choice matched the full-precision model on 90.5% of decisions (266 of 294). For accuracy-sensitive use, prefer Atlas3D/JEV-27B-VL-FP8 (97.1%) or the GGUF Q8_0 in Atlas3D/JEV-27B-VL-GGUF (99.0%).
An NVFP4 (W4A4, compressed-tensors) checkpoint of autotrust/JEV-27B-VL, made with llm-compressor using 384 calibration samples of 2,048 tokens. It needs an NVIDIA Blackwell GPU with FP4 tensor cores.
What is quantized: the language model's Linear layers in its full-attention blocks and MLPs.
What stays bf16: the vision tower, the linear-attention blocks, the embeddings, lm_head, MTP, and
the JEV System 1 decision adapter (adapter_vllm/, unchanged). So the weights are about 28 GB, not a
quarter of the original.
Evaluation (System 1 /v1/decide)
Same harness and states as the bf16 model, compared one decision at a time:
| NVFP4 (this) | FP8 | bf16 | |
|---|---|---|---|
| Top choice agrees with bf16 | 90.5% (266/294) | 97.1% | reference |
| Largest probability change on any option | 0.164 | 0.061 | none |
| Sanity scenarios | 6/6 | 6/6 | 6/6 |
| VRAM in use, RTX PRO 6000 | 37.6 GB | 42.6 GB | about 78 GB |
| 24-decision tick, p50 | 1130 ms | 1152 ms | 1253 ms |
This is a narrow test of the decision head. System 2 (free-text generation) quality was not separately benchmarked.
Serving
Serve it like the original (vLLM with --quantization compressed-tensors). serve_decide.py is
included. FlashInfer compiles its FP4 kernels on first use, so set CUDA_HOME to your CUDA toolkit.
License
Apache-2.0, as with the original. This is a modified (quantized) version of autotrust/JEV-27B-VL,
which is itself based on Qwen/Qwen3.8-27B. LICENSE is included unchanged.
- Downloads last month
- 71