nota-ai/Qwen3.8-2.4T-A95B-Nota-NVFP4-Global-Pruned-40 Text Generation • 1.5T • Updated 7 days ago • 431 • 23
nota-ai/Nemotron-3.5-Lightning-30B-A3B-NVFP4-Global-Pruned-15 Text Generation • 16B • Updated 9 days ago • 875 • 20
Efficient NVIDIA Models Collection Compressed NVIDIA Models (Nemotron and Cosmos) • 1 item • Updated 7 days ago • 1
Efficient Solar Open Family Collection Solar Open models officially optimized by Nota AI for Korea’s government-led Sovereign AI initiative as part of the Upstage consortium. • 7 items • Updated Jul 23 • 26
nota-ai/Solar-Open2-250B-Nota-NVFP4-GlobalPruned Text Generation • 117B • Updated 19 days ago • 545 • 34
nota-ai/Solar-Open2-250B-Nota-INT4-GlobalPruned Text Generation • 35B • Updated 26 days ago • 657 • 43
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B Paper • 2607.04244 • Published Jul 5 • 1
Efficient MoE-based LLM Collection Mixture-of-Experts Large Language Models with Advanced Quantization • 5 items • Updated Mar 11 • 25
Efficient Large Vision-Language Model Collection ERGO: LVLM trained with RL on efficiency objectives; https://github.com/nota-github/ERGO • 3 items • Updated Feb 22 • 27