-
empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF
Image-Text-to-Text • 9B • Updated • 554k • 2.79k -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520 -
moonshotai/Kimi-K3
Image-Text-to-Text • 2.8T • Updated • 1.2M • • 11.6k -
google/timesfm-3.0-pytorch
Time Series Forecasting • 0.3B • Updated • 1.43M • 914
Collections
Discover the best community collections!
Collections including paper arxiv:2510.04871
-
Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
Paper • 2509.15591 • Published • 46 -
A Survey on Latent Reasoning
Paper • 2507.06203 • Published • 95 -
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
Paper • 2602.03120 • Published • 1 -
TADA! Tuning Audio Diffusion Models through Activation Steering
Paper • 2602.11910 • Published • 2
-
Attention Is All You Need
Paper • 1706.03762 • Published • 144 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 12 -
Training Compute-Optimal Large Language Models
Paper • 2203.15556 • Published • 13 -
Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT
Paper • 2210.04186 • Published
-
HRM-Text: Efficient Pretraining Beyond Scaling
Paper • 2605.20613 • Published • 322 -
LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models
Paper • 2605.11011 • Published • 10 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520 -
Characterizing, Evaluating, and Optimizing Complex Reasoning
Paper • 2602.08498 • Published • 1
-
General Intelligence Requires Rethinking Exploration
Paper • 2211.07819 • Published • 1 -
Chain of Mindset: Reasoning with Adaptive Cognitive Modes
Paper • 2602.10063 • Published • 38 -
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
Paper • 2601.10160 • Published • 1 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520
-
A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Paper • 2309.06497 • Published • 6 -
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Paper • 2402.17764 • Published • 630 -
Llama 2: Open Foundation and Fine-Tuned Chat Models
Paper • 2307.09288 • Published • 252
-
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
Paper • 2511.18538 • Published • 307 -
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
Paper • 2511.14993 • Published • 236 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520
-
empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF
Image-Text-to-Text • 9B • Updated • 554k • 2.79k -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520 -
moonshotai/Kimi-K3
Image-Text-to-Text • 2.8T • Updated • 1.2M • • 11.6k -
google/timesfm-3.0-pytorch
Time Series Forecasting • 0.3B • Updated • 1.43M • 914
-
HRM-Text: Efficient Pretraining Beyond Scaling
Paper • 2605.20613 • Published • 322 -
LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models
Paper • 2605.11011 • Published • 10 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520 -
Characterizing, Evaluating, and Optimizing Complex Reasoning
Paper • 2602.08498 • Published • 1
-
Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
Paper • 2509.15591 • Published • 46 -
A Survey on Latent Reasoning
Paper • 2507.06203 • Published • 95 -
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
Paper • 2602.03120 • Published • 1 -
TADA! Tuning Audio Diffusion Models through Activation Steering
Paper • 2602.11910 • Published • 2
-
General Intelligence Requires Rethinking Exploration
Paper • 2211.07819 • Published • 1 -
Chain of Mindset: Reasoning with Adaptive Cognitive Modes
Paper • 2602.10063 • Published • 38 -
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
Paper • 2601.10160 • Published • 1 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520
-
A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Paper • 2309.06497 • Published • 6 -
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Paper • 2402.17764 • Published • 630 -
Llama 2: Open Foundation and Fine-Tuned Chat Models
Paper • 2307.09288 • Published • 252
-
Attention Is All You Need
Paper • 1706.03762 • Published • 144 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 12 -
Training Compute-Optimal Large Language Models
Paper • 2203.15556 • Published • 13 -
Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT
Paper • 2210.04186 • Published
-
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
Paper • 2511.18538 • Published • 307 -
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
Paper • 2511.14993 • Published • 236 -
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520