1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
Paper • 2609.24432 • Published • 16
Natural Language Processing
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
Video-Based Reward Modeling for Computer-Use Agents