Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding Paper • 2610.07342 • Published 6 days ago • 20
When to Switch: Reliable Action-Chunk Extension for Vision-Language-Action Models Paper • 2610.05719 • Published 6 days ago • 32
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper • 2609.37969 • Published 12 days ago • 43
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 12 days ago • 115
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures Paper • 2609.29429 • Published 17 days ago • 29
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation Paper • 2609.30221 • Published 17 days ago • 47
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 18 days ago • 56
FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation Paper • 2609.27657 • Published 18 days ago • 9
ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning Paper • 2609.22323 • Published 25 days ago • 10
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents Paper • 2609.23986 • Published 20 days ago • 29
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 23 days ago • 138
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 24 days ago • 115