Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution Paper • 2610.07641 • Published 6 days ago • 15
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training Paper • 2610.07510 • Published 7 days ago • 14
ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context Paper • 2609.36684 • Published 13 days ago • 20
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts Paper • 2610.01153 • Published 11 days ago • 24
ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization Paper • 2610.00906 • Published 11 days ago • 85
The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends Paper • 2609.39661 • Published 12 days ago • 13
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 16 days ago • 326
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 16 days ago • 35
AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs Paper • 2609.31590 • Published 17 days ago • 13
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 18 days ago • 105
Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World Paper • 2609.23038 • Published 23 days ago • 63
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 27 days ago • 560
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 24 days ago • 153
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training Paper • 2609.26774 • Published 20 days ago • 56
RULER: Instance-aware Rubric Rewards for SVG Generation Paper • 2609.25270 • Published 21 days ago • 104
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 21 days ago • 55
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 25 days ago • 139
Region-Level Policy Optimization for Fine-grained MLLM Perception Paper • 2609.19745 • Published 25 days ago • 41
Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts Paper • 2609.05661 • Published Sep 4 • 41