Atlas3D/JEV-27B-VL-NVFP4

Accuracy warning: this 4-bit build is less accurate than the 8-bit releases. On our decision test its top choice matched the full-precision model on 90.5% of decisions (266 of 294). For accuracy-sensitive use, prefer Atlas3D/JEV-27B-VL-FP8 (97.1%) or the GGUF Q8_0 in Atlas3D/JEV-27B-VL-GGUF (99.0%).

An NVFP4 (W4A4, compressed-tensors) checkpoint of autotrust/JEV-27B-VL, made with llm-compressor using 384 calibration samples of 2,048 tokens. It needs an NVIDIA Blackwell GPU with FP4 tensor cores.

What is quantized: the language model's Linear layers in its full-attention blocks and MLPs. What stays bf16: the vision tower, the linear-attention blocks, the embeddings, lm_head, MTP, and the JEV System 1 decision adapter (adapter_vllm/, unchanged). So the weights are about 28 GB, not a quarter of the original.

Evaluation (System 1 /v1/decide)

Same harness and states as the bf16 model, compared one decision at a time:

NVFP4 (this) FP8 bf16
Top choice agrees with bf16 90.5% (266/294) 97.1% reference
Largest probability change on any option 0.164 0.061 none
Sanity scenarios 6/6 6/6 6/6
VRAM in use, RTX PRO 6000 37.6 GB 42.6 GB about 78 GB
24-decision tick, p50 1130 ms 1152 ms 1253 ms

This is a narrow test of the decision head. System 2 (free-text generation) quality was not separately benchmarked.

Serving

Serve it like the original (vLLM with --quantization compressed-tensors). serve_decide.py is included. FlashInfer compiles its FP4 kernels on first use, so set CUDA_HOME to your CUDA toolkit.

License

Apache-2.0, as with the original. This is a modified (quantized) version of autotrust/JEV-27B-VL, which is itself based on Qwen/Qwen3.8-27B. LICENSE is included unchanged.

Downloads last month
71
Safetensors
Model size
28B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Atlas3D/JEV-27B-VL-NVFP4

Base model

Qwen/Qwen3.8-27B
Quantized
(9)
this model