An evidence-checked map of how physical-robot reinforcement learning reduces—or relocates—sample, robot-time, wall-clock, human, compute, hardware, and engineering cost.
这是一个经过证据核验的真机强化学习资料库,追踪方法如何减少或转移样本、机器人时间、日历时间、人工、计算、硬件和工程成本。
Literature cutoff / 文献截止:2026-07-31. This snapshot contains 36 included works: 34 physical-robot core studies and 2 measurement/context papers.
Canonical repository: https://github.com/liaohr9/awesome-real-world-rl-efficiency
- 中文详细综述
- English LaTeX survey
- Review methodology
- Evaluation protocol
- Research gaps
- Data schema
- Data: papers · mechanisms · quantitative evidence · claims
- Audits: atomic time ontology · claim-to-evidence map · tier rationales · lifecycle grid · hardware roles · zero-demo basis
- Validation report · Contributing
The bottleneck is not just “number of samples.” It is the resource vector required to reach a declared physical-task target:
C = {
physical steps / episodes,
active robot-hours,
elapsed wall-clock,
reset and operator time,
demonstrations and prior data,
compute,
parallel hardware,
engineering,
safety, wear, and maintenance
}
The coordinates are overlapping diagnostic views, not additive cost terms. For example, active robot-hours can already contain reset/recovery time. They may be summed only after an explicit inclusion matrix removes overlap and a declared price/utility model supplies weights. A fleet may shorten calendar time without reducing aggregate robot-hours; offline RL may reduce target-task exploration while inheriting hundreds of hours of physical data; a learned reset may reduce manual resets without making reset time or failure disappear. For that reason this repository is an evidence atlas, not a leaderboard (CL-EFF-001, CL-EFF-003, CL-EFF-023).
| Ledger | Frozen size | Meaning |
|---|---|---|
| Quantitative evidence | 424 rows | 408 standard rows (34 × 12) plus 16 supplemental quantities; 162 rows are not_reported. |
| Time ontology | 88 rows | Atomic task/phase records. Reported phase duration is available for 27/34 works; strict time_to_declared_target is available for 14/34 and NR for 20/34. |
| Lifecycle cost | 272 rows | Complete 34 × 8 channel grid: 230 NR, 35 source-located quantitative components, and 7 qualitative/incomplete proxies. No work reports a full lifecycle total. |
| Hardware roles | 36 rows | Learner, helper/reset, fleet-collector, and evaluation-only roles; helper hardware is not hidden inside learner count. |
| Tier and zero audits | 36 + 19 rows | One E-tier rationale per included work and one source/scope record per explicit zero-demo claim. |
| Claim evidence | 509 links | Every one of 25 synthesis claims resolves to typed evidence; count claims carry executable filters and full membership. |
reported_phase_duration and time_to_declared_target answer different questions. A run, data-collection program, or model-training job may have a duration without reporting the first time a declared physical target was reached. The strict field is available only when target, evaluator/checkpoint, phase boundary, and task scope can all be identified; “available” does not imply preregistration.
- Original technical research with an accessible full paper.
- At least one physical-robot task experiment for a core record.
- RL, or an RL-enabling mechanism, is central to the reported task-learning system.
- The work directly measures, changes, or targets physical interaction, robot-hours, wall-clock, compute, reset, human input, reward engineering, prior data, safety, or wear.
- Two explicitly labeled context papers may define the measurement problem; they cannot support a physical-effectiveness claim.
- Simulation-only results as evidence that a mechanism works reliably on hardware.
- Perception, planning, or language components without task-level robot-policy learning.
- Search snippets or project-page claims used as final numerical evidence.
- Public code treated as independent physical reproduction.
- Missing quantities silently treated as zero.
R0–R4describe physical-robot evidence/deployment intensity.E0–E4describe sample/time-cost claim strength.RandEanswer different questions: long autonomous operation is not automatically sample-efficient.- No core work in this snapshot reaches
E4independent cross-team cost replication (CL-EFF-011).
Full definitions are in the data schema.
The taxonomy groups systems by the resource transformation they perform, not by algorithm name. A paper can occupy several Level-3 nodes; the complete many-to-many mapping is in mechanism_matrix.csv.
- L1 — Reduce marginal online interaction / 减少目标在线交互
- L2 — Direct online update efficiency
EFF-M01Off-policy replay and actor-critic reuse
- L2 — Model-based learning
EFF-M02Black-box dynamics and policy searchEFF-M03Online latent world-model planning
- L2 — Representation and goal learning
EFF-M04Contrastive representations and imagined goals
- L2 — Constrained policy improvement
EFF-M05Residual policy around an engineered controller
- L2 — Task decomposition
EFF-M06Scheduled auxiliary intentions
- L2 — Trial-limited policy search
EFF-M07One-step or few-episode black-box optimization
- L2 — Direct online update efficiency
- L1 — Compress elapsed time / 压缩日历时间
- L2 — Parallel physical actors
EFF-M08Homogeneous robot-fleet collection
- L2 — External physical automation
EFF-M09Teacher or helper robot
- L2 — Fast online systems
EFF-M10Overlapped collection and learning on local accelerators
- L2 — Parallel physical actors
- L1 — Remove reset and recovery downtime / 减少复位与恢复停顿
- L2 — Task reciprocity
EFF-M11Forward/backward or multi-task self-reset
- L2 — Environment/task cycling
EFF-M12Pseudo-resets and goal cycles
- L2 — Learned recovery
EFF-M13Reset or recovery policy from demonstrations/data
- L2 — Scripted recovery and fixtures
EFF-M14Reorientation, reels, gravity, launchers, or actuated bins
- L2 — Task reciprocity
- L1 — Shift learning out of the target online phase / 将学习移出目标在线阶段
- L2 — Offline real-robot experience
EFF-M15Offline pretraining followed by limited online fine-tuningEFF-M16Large-scale offline policy learning
- L2 — Simulation and prior controllers
EFF-M17Simulation-initialized or repertoire-guided transfer
- L2 — Learned world models
EFF-M18Policy improvement inside a learned video world
- L2 — Offline real-robot experience
- L1 — Improve exploration information / 提高探索信息量
- L2 — Demonstration bootstrap
EFF-M19Initial task demonstrations
- L2 — Live supervision
EFF-M20Human interventions and corrective actions
- L2 — Autonomous data diversification
EFF-M21Play, failure data, and continuous multi-task collection
- L2 — Reward and progress supervision
EFF-M22Learned success/reward classifiers
- L2 — Demonstration bootstrap
- L1 — Adapt instead of relearn / 快速适应而非重新学习
- L2 — Dynamics adaptation
EFF-M23Meta-learned dynamics updated online
- L2 — Damage/behavior adaptation
EFF-M24Repertoire-guided Bayesian adaptation
- L2 — Policy adaptation
EFF-M25Teacher-mediated real-world adaptation
- L2 — Dynamics adaptation
| Claim | Evidence-bounded finding | Required limit |
|---|---|---|
CL-EFF-002 |
Minute- or hour-scale physical learning has been demonstrated on several engineered manipulation, locomotion, and driving tasks. | Each result uses its own task, target, reset, demonstrations, and upstream assets; it is not a minute-scale lifecycle claim. |
CL-EFF-004 |
Robot fleets improve calendar throughput. | They do not eliminate aggregate physical interaction, hardware, maintenance, or staffing cost. |
CL-EFF-005 |
Offline RL can avoid target-deployment exploration. | It inherits collection, demonstration, failure-data, labeling, and curation cost. |
CL-EFF-009 |
Reset-free and continuous systems reduce particular manual reset operations. | They may still depend on reversible tasks, recovery policies, fixtures, batteries, or occasional rescue. |
CL-EFF-010 |
Adaptation can take only minutes after an upstream prior exists. | It cannot be ranked directly against from-scratch learning without C_pretrain and reuse count K. |
CL-EFF-012 |
15/34 core works lack a clear evaluation-trial denominator. | Cross-paper success-rate comparisons are therefore often underdetermined. |
CL-EFF-017 |
In the complete 272-cell lifecycle grid, 230 cells are NR, 35 contain source-located quantitative components, and 7 contain only qualitative/incomplete proxies; no work reports a full lifecycle total. |
This is a reporting-gap finding, not evidence that any named system is unsafe. |
CL-EFF-023 |
The corpus does not support a universal efficiency ranking or pooled effect size. | Local comparison remains possible when task, target, budget, system boundary, assets, and repetitions match. |
| Visible reduction | Typical displaced resource | What must be reported |
|---|---|---|
| Fewer target online transitions | Offline robot data, simulation, pretrained policy/model | Provenance, collection robot-hours, human-hours, compute, reuse count |
| Shorter wall-clock | Fleet size, accelerators, overlapping compute, maintenance | Per-robot hours, actor count over time, duty cycle, learner lag |
| Fewer manual resets | Reverse task, recovery data, fixture/helper robot | Reset attempts, conditional success, duration, rescue fallback |
| Safer/faster exploration | Demonstrations or live intervention | Active and standby operator-minutes, discarded attempts, informational role |
| Few-minute adaptation | Repertoire, meta-training, teacher policy/system | C_pretrain, C_adapt, transfer support, amortization over K deployments |
| Lower-dimensional RL search | Base controller, reward, detector, task fixture | Engineering person-hours, reuse scope, sensitivity to specification errors |
| Less physical policy optimization | World-model data and GPU training | Real-data origin, GPU-hours/energy, model-bias failures, corrective physical trials |
These are cost channels, not accusations: the moved cost may be worthwhile and reusable. The requirement is to expose it (CL-EFF-006, CL-EFF-007, CL-EFF-020, CL-EFF-024).
Each work appears once below under its primary browsing mechanism. The many-to-many taxonomy retains all secondary mechanisms.
- SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning (2024) — replay-based fast online learning; R3/E2 · Project · Code
- Learning to Walk via Deep Reinforcement Learning (2019) — direct model-free locomotion learning; R3/E1 · Project
- Learning Visual Robotic Control Efficiently with Contrastive Pre-training and Data Augmentation (2022) — contrastive representation plus demonstrations; R3/E2 · Project
- Sample-efficient Reinforcement Learning in Robotic Table Tennis (2021) — trial-limited black-box optimization; R3/E1
- DayDreamer: World Models for Physical Robot Learning (2023) — online latent world-model learning; R3/E2 · Project · Code
- Learning by Playing: Solving Sparse Reward Tasks from Scratch (2018) — scheduled auxiliary intentions; R3/E1
- Visual Reinforcement Learning with Imagined Goals (2018) — imagined latent goals; R3/E2 · Project · Code
- Data Efficient Reinforcement Learning for Legged Robots (2020) — black-box dynamics/policy search; R3/E1
- Black-Box Data-efficient Policy Search for Robotics (2017) — few-episode model-based policy search; R3/E2 · Code
- Residual Reinforcement Learning for Robot Control (2019) — residual correction around an engineered controller; R3/E3
- Demonstrating a Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning (2023) — fast replay-based locomotion learning; R3/E3 · Project · Code
- MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation (2024) — online visuo-motor world model; R3/E2 · Project · Code
- Continuously Improving Mobile Manipulation with Autonomous Real-World RL (2025) — autonomous collection with helper/procedural infrastructure; R4/E1 · Project
- FastRLAP: A System for Learning High-Speed Driving via Deep RL and Autonomous Practicing (2023) — fast local online learning and autonomous practice; R3/E3 · Project · Code
- Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation (2018) — fleet-scale physical data collection; R4/E2
- Scaling Up Multi-Task Robotic Reinforcement Learning (2022) — multi-robot, multi-task fleet learning; R4/E2 · Project
- Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids (2025) — teacher-robot-mediated physical adaptation; R3/E3 · Project · Code
- One Demonstration Is Enough for Real-World Robotic Reinforcement Learning (2026) — learned recovery plus single-demonstration bootstrap; R3/E3 · Project · Code
- Fully Autonomous Real-World Reinforcement Learning with Applications to Mobile Manipulation (2022) — pseudo-reset and task cycling; R4/E2 · Project
- REBOOT: Reuse Data for Bootstrapping Efficient Real-World Dexterous Manipulation (2023) — learned reset/recovery from prior data; R4/E3 · Project
- Reset-Free Reinforcement Learning via Multi-Task Learning: Learning Dexterous Manipulation Behaviors without Human Intervention (2021) — multi-task self-reset; R4/E1 · Project
- Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning (2023) — forward/backward learning with recurring reset support; R3/E1 · Project
- Don't Start From Scratch: Leveraging Prior Data to Automate Robotic Reinforcement Learning (2023) — prior-data reuse plus limited online learning; R3/E3 · Project
- World-Gymnast: Training Robots with Reinforcement Learning in a World Model (2026) — policy learning inside a learned world model; R3/E3 · Project · Code
- PlayWorld: Learning Robot World Models from Autonomous Play (2026) — autonomous play data and learned video world; R3/E3 · Project
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets (2020) — offline initialization followed by online fine-tuning; R3/E2 · Project · Code
- COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning (2021) — reuse of prior physical experience; R2/E2 · Project · Code
- Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials (2023) — offline pretraining plus few-trial adaptation; R3/E3 · Project · Code
- Q-Transformer: Scalable Offline Reinforcement Learning via Autoregressive Q-Functions (2023) — large-scale offline physical-robot policy learning; R2/E2 · Project
- Real World Offline Reinforcement Learning with Realistic Data Source (2023) — physical dataset and human/robot cost accounting; R2/E2 · Project
- Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (2025) — demonstrations plus online intervention; R3/E3 · Project · Code
- Real-world Reinforcement Learning from Suboptimal Interventions (2025) — learning from corrective/suboptimal interventions; R3/E3 · Project · Code
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning (2019) — meta-learned dynamics adaptation; R2/E2 · Project
- Robots that Can Adapt Like Animals (2015) — repertoire-guided damage adaptation; R2/E3
These sources define real-world RL constraints and measurement concerns. Their R0/E0 label means they must not be used as physical-effectiveness evidence.
- Challenges of Real-World Reinforcement Learning (2019) — problem-definition context; R0/E0
- Challenges of Real-World Reinforcement Learning: Definitions, Benchmarks and Analysis (2021) — benchmark/measurement context; R0/E0 · Code
A credible new result should report:
- exact robot, task, initial-state distribution, success event, target threshold, and trial denominator;
- physical transitions, episodes/trials, control frequency, action repeat, and active robot-hours;
- elapsed wall-clock with reset, recovery, update, evaluation, idle, and maintenance boundaries;
- robot count over time, aggregate robot-hours, duty cycle, and learner/actor overlap;
- demonstrations, discarded attempts, labels, interventions, resets, rescue, monitoring, and operator-minutes;
- prior-data provenance, simulation/model/controller assets, compute duration, GPU count, and reuse/amortization rule;
- reset/recovery attempts, coverage, success, latency, fallback, and unrecoverable failures;
- collisions, safety aborts, damage, wear, consumables, downtime, and failed development runs;
- at least three independent seeds where feasible, including failures, with performance versus steps, robot-hours, wall-clock, and operator-hours;
- both a target-task marginal ledger and a full lifecycle ledger.
The operational template and comparison gates are in docs/evaluation_protocol.md.
NRis not zero. It means the value was not located in the verified source. Numeric zero is retained only when explicitly source-reported.
Context papers are not physical-robot reliability evidence. They support definitions and measurement design only.
No universal ranking. Tasks, success thresholds, action rates, resets, parallelism, prior assets, and lifecycle boundaries differ too much for a pooled leaderboard.
Code is not reproduction. The 15/34 official-code count records availability at verification time, not independent physical replication.
make validate
python3 scripts/test_validator.py
make paper
make cleanValidation is offline and uses only the Python standard library. Online link health is deliberately separate so the core check is deterministic.
Please read CONTRIBUTING.md. A paper suggestion needs an original-source locator, physical-evidence classification, resource-boundary extraction, and corresponding updates to every affected ledger. Claims that only improve a headline proxy without exposing displaced cost will not be merged as lifecycle conclusions.
Repository text, tables, and scripts are available under the MIT License. Paper contents and linked external code remain under their original terms. Citation metadata is in CITATION.cff.