Library · 18-reinforcement-learning
RL & robot learning
| Title | Peer | Link |
|---|---|---|
| Asynchronous Methods for Deep Reinforcement Learning (A3C) | ◦ preprint | arxiv.org/abs/1602.01783 |
| Trust Region Policy Optimization (TRPO) | ✓ peer | arxiv.org/abs/1502.05477 |
| Proximal Policy Optimization Algorithms (PPO) | — unref | arxiv.org/abs/1707.06347 |
| Soft Actor-Critic (SAC) | ◦ preprint | arxiv.org/abs/1801.01290 |
| Addressing Function Approximation Error in Actor-Critic Methods (TD3) | ◦ preprint | arxiv.org/abs/1802.09477 |
| Continuous control with deep reinforcement learning (DDPG) | ◦ preprint | arxiv.org/abs/1509.02971 |
| Rainbow: Combining Improvements in Deep RL | ◦ preprint | arxiv.org/abs/1710.02298 |
| IMPALA: Scalable Distributed Deep-RL | ✓ peer | arxiv.org/abs/1802.01561 |
| Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (MuZero) | ◦ preprint | arxiv.org/abs/1911.08265 |
| When to Trust Your Model: Model-Based Policy Optimization (MBPO) | ✓ peer | arxiv.org/abs/1906.08253 |
| Mastering Atari Games with Limited Data (EfficientZero) | ◦ preprint | arxiv.org/abs/2111.00210 |
| Temporal Difference Learning for Model Predictive Control (TD-MPC) | ◦ preprint | arxiv.org/abs/2203.04955 |
| TD-MPC2: Scalable, Robust World Models for Continuous Control | ◦ preprint | arxiv.org/abs/2310.16828 |
| Mastering Diverse Domains through World Models (DreamerV3) | ◦ preprint | arxiv.org/abs/2301.04104 |
| Dream to Control: Learning Behaviors by Latent Imagination (Dreamer) | ◦ preprint | arxiv.org/abs/1912.01603 |
| Deep RL in a Handful of Trials using Probabilistic Dynamics Models (PETS) | ◦ preprint | arxiv.org/abs/1805.12114 |
| Conservative Q-Learning for Offline RL (CQL) | ◦ preprint | arxiv.org/abs/2006.04779 |
| Offline RL with Implicit Q-Learning (IQL) | ◦ preprint | arxiv.org/abs/2110.06169 |
| Off-Policy Deep RL without Exploration (BCQ) | ◦ preprint | arxiv.org/abs/1812.02900 |
| Is Conditional Generative Modeling all you need for Decision-Making? (Decision Diffuser) | ◦ preprint | arxiv.org/abs/2211.15657 |
| Decision Transformer: Reinforcement Learning via Sequence Modeling | ◦ preprint | arxiv.org/abs/2106.01345 |
| Generative Adversarial Imitation Learning (GAIL) | ◦ preprint | arxiv.org/abs/1606.03476 |
| A Reduction of Imitation Learning to No-Regret Online Learning (DAgger) | ◦ preprint | arxiv.org/abs/1011.0686 |
| ALVINN: An Autonomous Land Vehicle in a Neural Network | proceedings.neurips.cc/paper/1988 | |
| Model-Agnostic Meta-Learning (MAML) | ◦ preprint | arxiv.org/abs/1703.03400 |
| Efficient Off-Policy Meta-RL via Probabilistic Context Variables (PEARL) | ✓ peer | arxiv.org/abs/1903.08254 |
| RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | ◦ preprint | arxiv.org/abs/2307.15818 |
| RT-1: Robotics Transformer for Real-World Control at Scale | ◦ preprint | arxiv.org/abs/2212.06817 |
| OpenVLA: An Open-Source Vision-Language-Action Model | ◦ preprint | arxiv.org/abs/2406.09246 |
| Octo: An Open-Source Generalist Robot Policy | ◦ preprint | arxiv.org/abs/2405.12213 |
| π₀: A Vision-Language-Action Flow Model for General Robot Control | ◦ preprint | arxiv.org/abs/2410.24164 |
| Open X-Embodiment: Robotic Learning Datasets and RT-X Models | ◦ preprint | arxiv.org/abs/2310.08864 |
| Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ALOHA / ACT) | ◦ preprint | arxiv.org/abs/2304.13705 |
| Mobile ALOHA: Bimanual Mobile Manipulation via Whole-Body Teleoperation | ◦ preprint | arxiv.org/abs/2401.02117 |
| RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation | ◦ preprint | arxiv.org/abs/2410.07864 |
| Diffusion Policy: Visuomotor Policy Learning via Action Diffusion | ◦ preprint | arxiv.org/abs/2303.04137 |
| Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments | ◦ preprint | arxiv.org/abs/2605.30280 |
| RoboBrain 2.0 Technical Report | ◦ preprint | arxiv.org/abs/2507.02029 |
| GR00T N1: An Open Foundation Model for Generalist Humanoid Robots | ◦ preprint | arxiv.org/abs/2503.14734 |
| Gemini Robotics: Bringing AI into the Physical World | ◦ preprint | arxiv.org/abs/2503.20020 |
| π0.5: a Vision-Language-Action Model with Open-World Generalization | ◦ preprint | arxiv.org/abs/2504.16054 |
| CogACT: A Foundational VLA Model for Synergizing Cognition and Action | ◦ preprint | arxiv.org/abs/2411.19650 |
| SpatialVLA: Exploring Spatial Representations for Vision-Language-Action Models | ◦ preprint | arxiv.org/abs/2501.15830 |
| TinyVLA: Towards Fast, Data-Efficient VLA Models for Robotic Manipulation | ◦ preprint | arxiv.org/abs/2409.12514 |
| SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics | ◦ preprint | arxiv.org/abs/2506.01844 |
| WorldVLA: Towards Autoregressive Action World Model | ◦ preprint | arxiv.org/abs/2506.21539 |
| Towards Generalist Robot Policies: What Matters in Building VLA Models (RoboVLMs) | ◦ preprint | arxiv.org/abs/2412.14058 |
| Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT) | ◦ preprint | arxiv.org/abs/2502.19645 |
| A Survey on Vision-Language-Action Models: An Action Tokenization Perspective | ◦ preprint | arxiv.org/abs/2507.01925 |
| Domain Randomization for Transferring Deep NNs from Simulation to the Real World | ◦ preprint | arxiv.org/abs/1703.06907 |
| Solving Rubik's Cube with a Robot Hand | ◦ preprint | arxiv.org/abs/1910.07113 |
| RMA: Rapid Motor Adaptation for Legged Robots | ◦ preprint | arxiv.org/abs/2107.04034 |
| Curiosity-driven Exploration by Self-supervised Prediction (ICM) | — unref | arxiv.org/abs/1705.05363 |
| Exploration by Random Network Distillation (RND) | ◦ preprint | arxiv.org/abs/1810.12894 |
| Mastering Chess and Shogi by Self-Play (AlphaZero) | ◦ preprint | arxiv.org/abs/1712.01815 |
| Hindsight Experience Replay (HER) | ✓ peer | arxiv.org/abs/1707.01495 |
| QT-Opt: Scalable Deep RL for Vision-Based Robotic Manipulation | ◦ preprint | arxiv.org/abs/1806.10293 |
| Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan) | ◦ preprint | arxiv.org/abs/2204.01691 |
| PaLM-E: An Embodied Multimodal Language Model | ◦ preprint | arxiv.org/abs/2303.03378 |
| A Generalist Agent (Gato) | ◦ preprint | arxiv.org/abs/2205.06175 |
| RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation | ◦ preprint | arxiv.org/abs/2306.11706 |
| Code as Policies: Language Model Programs for Embodied Control | ◦ preprint | arxiv.org/abs/2209.07753 |
| VoxPoser: Composable 3D Value Maps for Robotic Manipulation | ◦ preprint | arxiv.org/abs/2307.05973 |
| Eureka: Human-Level Reward Design via Coding LLMs | ◦ preprint | arxiv.org/abs/2310.12931 |
| 3D Diffusion Policy (DP3) | ◦ preprint | arxiv.org/abs/2403.03954 |
| Genie: Generative Interactive Environments | ◦ preprint | arxiv.org/abs/2402.15391 |
| Learning to predict by the methods of temporal differences | ✓ peer | doi.org/10.1007/BF00115009 |
| Human-level control through deep reinforcement learning (DQN) | ✓ peer | doi.org/10.1038/nature14236 |
| Reinforcement learning in the brain | ✓ peer | doi.org/10.1016/j.jmp.2008.12.005 |
| Reinforcement Learning and Optimal Control | — | |
| The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL | ◦ preprint | arxiv.org/abs/2607.19749 |
| The Mechanism Matters: When Knowledge Graphs Help Reinforcement Learning | ◦ preprint | arxiv.org/abs/2607.19616 |
| Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning | — unref | arxiv.org/abs/2608.05111 |
| Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning | ◦ preprint | arxiv.org/abs/2607.24996 |
| AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning | ◦ preprint | arxiv.org/abs/2608.01980 |
| DARLING: Detection Augmented Reinforcement Learning with Non-Stationary Guarantees | arxiv.org/abs/2604.16684 |