Library · 18-reinforcement-learning

RL & robot learning

TitlePeerLink
Asynchronous Methods for Deep Reinforcement Learning (A3C)◦ preprintarxiv.org/abs/1602.01783
Trust Region Policy Optimization (TRPO)✓ peerarxiv.org/abs/1502.05477
Proximal Policy Optimization Algorithms (PPO)— unrefarxiv.org/abs/1707.06347
Soft Actor-Critic (SAC)◦ preprintarxiv.org/abs/1801.01290
Addressing Function Approximation Error in Actor-Critic Methods (TD3)◦ preprintarxiv.org/abs/1802.09477
Continuous control with deep reinforcement learning (DDPG)◦ preprintarxiv.org/abs/1509.02971
Rainbow: Combining Improvements in Deep RL◦ preprintarxiv.org/abs/1710.02298
IMPALA: Scalable Distributed Deep-RL✓ peerarxiv.org/abs/1802.01561
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero)◦ preprintarxiv.org/abs/1911.08265
When to Trust Your Model: Model-Based Policy Optimization (MBPO)✓ peerarxiv.org/abs/1906.08253
Mastering Atari Games with Limited Data (EfficientZero)◦ preprintarxiv.org/abs/2111.00210
Temporal Difference Learning for Model Predictive Control (TD-MPC)◦ preprintarxiv.org/abs/2203.04955
TD-MPC2: Scalable, Robust World Models for Continuous Control◦ preprintarxiv.org/abs/2310.16828
Mastering Diverse Domains through World Models (DreamerV3)◦ preprintarxiv.org/abs/2301.04104
Dream to Control: Learning Behaviors by Latent Imagination (Dreamer)◦ preprintarxiv.org/abs/1912.01603
Deep RL in a Handful of Trials using Probabilistic Dynamics Models (PETS)◦ preprintarxiv.org/abs/1805.12114
Conservative Q-Learning for Offline RL (CQL)◦ preprintarxiv.org/abs/2006.04779
Offline RL with Implicit Q-Learning (IQL)◦ preprintarxiv.org/abs/2110.06169
Off-Policy Deep RL without Exploration (BCQ)◦ preprintarxiv.org/abs/1812.02900
Is Conditional Generative Modeling all you need for Decision-Making? (Decision Diffuser)◦ preprintarxiv.org/abs/2211.15657
Decision Transformer: Reinforcement Learning via Sequence Modeling◦ preprintarxiv.org/abs/2106.01345
Generative Adversarial Imitation Learning (GAIL)◦ preprintarxiv.org/abs/1606.03476
A Reduction of Imitation Learning to No-Regret Online Learning (DAgger)◦ preprintarxiv.org/abs/1011.0686
ALVINN: An Autonomous Land Vehicle in a Neural Networkproceedings.neurips.cc/paper/1988
Model-Agnostic Meta-Learning (MAML)◦ preprintarxiv.org/abs/1703.03400
Efficient Off-Policy Meta-RL via Probabilistic Context Variables (PEARL)✓ peerarxiv.org/abs/1903.08254
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control◦ preprintarxiv.org/abs/2307.15818
RT-1: Robotics Transformer for Real-World Control at Scale◦ preprintarxiv.org/abs/2212.06817
OpenVLA: An Open-Source Vision-Language-Action Model◦ preprintarxiv.org/abs/2406.09246
Octo: An Open-Source Generalist Robot Policy◦ preprintarxiv.org/abs/2405.12213
π₀: A Vision-Language-Action Flow Model for General Robot Control◦ preprintarxiv.org/abs/2410.24164
Open X-Embodiment: Robotic Learning Datasets and RT-X Models◦ preprintarxiv.org/abs/2310.08864
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ALOHA / ACT)◦ preprintarxiv.org/abs/2304.13705
Mobile ALOHA: Bimanual Mobile Manipulation via Whole-Body Teleoperation◦ preprintarxiv.org/abs/2401.02117
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation◦ preprintarxiv.org/abs/2410.07864
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion◦ preprintarxiv.org/abs/2303.04137
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments◦ preprintarxiv.org/abs/2605.30280
RoboBrain 2.0 Technical Report◦ preprintarxiv.org/abs/2507.02029
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots◦ preprintarxiv.org/abs/2503.14734
Gemini Robotics: Bringing AI into the Physical World◦ preprintarxiv.org/abs/2503.20020
π0.5: a Vision-Language-Action Model with Open-World Generalization◦ preprintarxiv.org/abs/2504.16054
CogACT: A Foundational VLA Model for Synergizing Cognition and Action◦ preprintarxiv.org/abs/2411.19650
SpatialVLA: Exploring Spatial Representations for Vision-Language-Action Models◦ preprintarxiv.org/abs/2501.15830
TinyVLA: Towards Fast, Data-Efficient VLA Models for Robotic Manipulation◦ preprintarxiv.org/abs/2409.12514
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics◦ preprintarxiv.org/abs/2506.01844
WorldVLA: Towards Autoregressive Action World Model◦ preprintarxiv.org/abs/2506.21539
Towards Generalist Robot Policies: What Matters in Building VLA Models (RoboVLMs)◦ preprintarxiv.org/abs/2412.14058
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT)◦ preprintarxiv.org/abs/2502.19645
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective◦ preprintarxiv.org/abs/2507.01925
Domain Randomization for Transferring Deep NNs from Simulation to the Real World◦ preprintarxiv.org/abs/1703.06907
Solving Rubik's Cube with a Robot Hand◦ preprintarxiv.org/abs/1910.07113
RMA: Rapid Motor Adaptation for Legged Robots◦ preprintarxiv.org/abs/2107.04034
Curiosity-driven Exploration by Self-supervised Prediction (ICM)— unrefarxiv.org/abs/1705.05363
Exploration by Random Network Distillation (RND)◦ preprintarxiv.org/abs/1810.12894
Mastering Chess and Shogi by Self-Play (AlphaZero)◦ preprintarxiv.org/abs/1712.01815
Hindsight Experience Replay (HER)✓ peerarxiv.org/abs/1707.01495
QT-Opt: Scalable Deep RL for Vision-Based Robotic Manipulation◦ preprintarxiv.org/abs/1806.10293
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan)◦ preprintarxiv.org/abs/2204.01691
PaLM-E: An Embodied Multimodal Language Model◦ preprintarxiv.org/abs/2303.03378
A Generalist Agent (Gato)◦ preprintarxiv.org/abs/2205.06175
RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation◦ preprintarxiv.org/abs/2306.11706
Code as Policies: Language Model Programs for Embodied Control◦ preprintarxiv.org/abs/2209.07753
VoxPoser: Composable 3D Value Maps for Robotic Manipulation◦ preprintarxiv.org/abs/2307.05973
Eureka: Human-Level Reward Design via Coding LLMs◦ preprintarxiv.org/abs/2310.12931
3D Diffusion Policy (DP3)◦ preprintarxiv.org/abs/2403.03954
Genie: Generative Interactive Environments◦ preprintarxiv.org/abs/2402.15391
Learning to predict by the methods of temporal differences✓ peerdoi.org/10.1007/BF00115009
Human-level control through deep reinforcement learning (DQN)✓ peerdoi.org/10.1038/nature14236
Reinforcement learning in the brain✓ peerdoi.org/10.1016/j.jmp.2008.12.005
Reinforcement Learning and Optimal Control
The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL◦ preprintarxiv.org/abs/2607.19749
The Mechanism Matters: When Knowledge Graphs Help Reinforcement Learning◦ preprintarxiv.org/abs/2607.19616
Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning— unrefarxiv.org/abs/2608.05111
Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning◦ preprintarxiv.org/abs/2607.24996
AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning◦ preprintarxiv.org/abs/2608.01980
DARLING: Detection Augmented Reinforcement Learning with Non-Stationary Guaranteesarxiv.org/abs/2604.16684