-
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Paper • 2607.07508 • Published • 31 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 11 -
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Paper • 2012.13255 • Published • 6
Collections
Discover the best community collections!
Collections including paper arxiv:2503.14476
-
RL + Transformer = A General-Purpose Problem Solver
Paper • 2501.14176 • Published • 28 -
Towards General-Purpose Model-Free Reinforcement Learning
Paper • 2501.16142 • Published • 31 -
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Paper • 2501.17161 • Published • 127 -
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
Paper • 2412.12098 • Published • 4
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Training language models to follow instructions with human feedback
Paper • 2203.02155 • Published • 26 -
Llama 2: Open Foundation and Fine-Tuned Chat Models
Paper • 2307.09288 • Published • 252 -
The Llama 3 Herd of Models
Paper • 2407.21783 • Published • 119
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 149
-
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
Paper • 2603.12180 • Published • 65 -
Flow-OPD: On-Policy Distillation for Flow Matching Models
Paper • 2605.08063 • Published • 102 -
Normalizing Trajectory Models
Paper • 2605.08078 • Published • 15 -
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
Paper • 2605.08029 • Published • 12
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 222 -
Reinforced Attention Learning
Paper • 2602.04884 • Published • 30 -
Learning to Reason in 13 Parameters
Paper • 2602.04118 • Published • 6 -
LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters
Paper • 2405.17604 • Published • 4
-
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Paper • 2509.07980 • Published • 106 -
Robot Learning from a Physical World Model
Paper • 2511.07416 • Published • 32 -
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
Paper • 2511.06805 • Published • 13 -
GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
Paper • 2511.17592 • Published • 122
-
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520 -
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
Paper • 2510.07499 • Published • 49 -
Improving Context Fidelity via Native Retrieval-Augmented Reasoning
Paper • 2509.13683 • Published • 8 -
Multimodal Iterative RAG for Knowledge-Intensive Visual Question Answering
Paper • 2509.00798 • Published • 1
-
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Paper • 2607.07508 • Published • 31 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 11 -
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning
Paper • 2012.13255 • Published • 6
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Proximal Policy Optimization Algorithms
Paper • 1707.06347 • Published • 12 -
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Paper • 2305.18290 • Published • 71 -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Paper • 2402.03300 • Published • 149
-
RL + Transformer = A General-Purpose Problem Solver
Paper • 2501.14176 • Published • 28 -
Towards General-Purpose Model-Free Reinforcement Learning
Paper • 2501.16142 • Published • 31 -
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Paper • 2501.17161 • Published • 127 -
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
Paper • 2412.12098 • Published • 4
-
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
Paper • 2603.12180 • Published • 65 -
Flow-OPD: On-Policy Distillation for Flow Matching Models
Paper • 2605.08063 • Published • 102 -
Normalizing Trajectory Models
Paper • 2605.08078 • Published • 15 -
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
Paper • 2605.08029 • Published • 12
-
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
Paper • 2602.10693 • Published • 222 -
Reinforced Attention Learning
Paper • 2602.04884 • Published • 30 -
Learning to Reason in 13 Parameters
Paper • 2602.04118 • Published • 6 -
LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters
Paper • 2405.17604 • Published • 4
-
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Paper • 2509.07980 • Published • 106 -
Robot Learning from a Physical World Model
Paper • 2511.07416 • Published • 32 -
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
Paper • 2511.06805 • Published • 13 -
GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
Paper • 2511.17592 • Published • 122
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Paper • 2503.14476 • Published • 148 -
Training language models to follow instructions with human feedback
Paper • 2203.02155 • Published • 26 -
Llama 2: Open Foundation and Fine-Tuned Chat Models
Paper • 2307.09288 • Published • 252 -
The Llama 3 Herd of Models
Paper • 2407.21783 • Published • 119
-
Less is More: Recursive Reasoning with Tiny Networks
Paper • 2510.04871 • Published • 520 -
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
Paper • 2510.07499 • Published • 49 -
Improving Context Fidelity via Native Retrieval-Augmented Reasoning
Paper • 2509.13683 • Published • 8 -
Multimodal Iterative RAG for Knowledge-Intensive Visual Question Answering
Paper • 2509.00798 • Published • 1