Hebero / GPU-parallel robot learning

Forty tasks.
One policy.

A GPU-parallel framework for heterogeneous multi-task reinforcement learning.

/ HETEROGENEOUS TASKS. SHARED LEARNING. Explore the benchmark
Heterogeneous tasks
40
All four LIBERO suites
End-to-end training
78.5k
Steps / second · 8 × L20 GPUs
Visual IW-ABC
93.5%
Mean success · Hebero
Physical Piper policy
82.5%
66 / 80 trials · Four tasks

Forty tasks.
One training
process.

Hebero brings LIBERO’s scenes, assets, and success predicates into one GPU-parallel simulator, renderer, rollout buffer, and learner.

A common interface supports state and visual policies. Randomized initial states test how reliably the policy solves each of the 40 known tasks.

4 suitesState + RGB7D actions
Parallel environments
Training in parallelShared policy · Heterogeneous scenes
Evaluating the policyParallel task execution

PARALLEL ENVIRONMENT SCALING

More parallel
experience.

Increasing parallel replicas improves endpoint success within the shared training-time budget in this experiment.

Evaluation success versus parallel environment count, and training reward versus wall-clock hours. On one L20 GPU, success rises from 5.5% at 40 environments to 76.0% at 4,000.
40–4,000 environments on one 48 GB L20 GPU · Checkpoints near a 16-hour training budget.

Task diversity, in motion.

Selected successful rollouts
Selected successes from state and visual policies. Playback is slowed for viewing.
01

Long

Sequences of interactions.

02

Object

Variation in manipulated objects.

03

Spatial

Different spatial relationships.

04

Goal

Different goals in shared scenes.

A shared foundation.
Adaptive guidance.

Demonstration-Guided Policy Optimization provides a common training stack for controlled comparisons of how demonstrations enter learning.

Hebero defines the benchmark. DGPO supplies the reference training framework. IW-ABC is the selected recipe within it.

DGPO / SHARED TRAINABILITY STACKFixed across reference learners
Demonstration-start resetsTracking rewardPrivileged criticTask-conditioned PPO
Learner-specific interface
Select a learner ↓
Demo entry method

ABC · Live action MSE loss

Adds a mean squared error (MSE) loss between the live policy mean and the time-aligned demonstration action to the PPO loss. Its weight β adapts to task success and training progress, balancing teacher guidance with self-exploration.

  • BC → PPO
    Actor initialization

    Demo → initial policy

  • DAPG
    Likelihood regularization

    PPO loss + λ · NLL

  • RFCL
    Reverse–forward curriculum

    Reverse resets → forward coverage

  • ABCSelected
    Live action targets

    PPO loss + β · MSE

Each entry adds to the same shared DGPO stack.

THE SELECTED RECIPE

IW + ABC

A / TASK IMPORTANCE WEIGHTING

Give lagging tasks
more weight.

IW increases the PPO loss weight for tasks below the current multi-task success mean. Rollout allocation stays fixed; the learning signal changes.

Balances learning across tasks
B / ADAPTIVE BEHAVIOR CLONING

Relax guidance
as learning progresses.

ABC uses each task’s success and elapsed training to regulate teacher guidance, balancing imitation with the policy’s own exploration and reward-driven improvement.

Balances teacher guidance and PPO
IW-ABC in motion. A conceptual illustration of how task weighting and adaptive demonstration guidance shape the policy’s learning trajectory.
FIGURE 4 / A CLOSER LOOKHow task weights evolve

Guidance fades. Task priority stays.

As demonstration guidance relaxes, IW continues giving lagging tasks more weight. These training dynamics show how the two mechanisms work together in visual IW-ABC.

Scroll sideways to compare all three panels →Open original PDF
Three heatmaps across 40 tasks and 0–20,000 PPO iterations: training success, raw importance weights, and adaptive behavior-cloning coefficients. Lagging tasks retain higher importance weights, while all BC coefficients reach 0.1 by the dashed line at 12,500 iterations.
aLearning progress
Each row is one task. Darker blue indicates higher training success.
bContinued task priority
Blue indicates higher raw IW. Lagging tasks retain priority after BC annealing ends.
cRelaxing teacher guidance
Teal fades as BC weakens. Every task reaches the 0.1 floor by 12.5k iterations.
Visual IW-ABC, training seed 2; 100-iteration averages. Success is a training moving average. IW and BC are reconstructed from logged success; IW is shown before minibatch normalization. Dashed lines mark the end of ABC annealing.

Broader capability.
Compact policies.

IW-ABC improves both mean and long-horizon success within the shared DGPO stack. The visual recipe covers 38 of 40 tasks at ≥80% success.

STATE IW-ABC

90.1%

Mean success

0.32M actor parameters
VISUAL IW-ABC

93.5%

Mean success

0.40M trainable actor parameters
+ 5.52M frozen Tiny ViT frontend
PPO-based references under the shared DGPO stack
LearnerIWMean SR ↑Long SR ↑Coverage ↑SR-AUC ↑
PPONo50.8 ± 3.110.0 ± 0.020 / 4045.2
BC → PPONo47.5 ± 2.620.0 ± 0.018 / 4040.8
ABCNo75.5 ± 3.640.0 ± 0.029 / 4071.9
IW-PPOYes44.9 ± 2.520.0 ± 0.018 / 4041.3
IW-DAPGYes68.5 ± 3.030.0 ± 0.027 / 4062.4
IW-RFCLYes66.4 ± 2.830.0 ± 0.026 / 4060.1
IW-ABCYes90.1 ± 3.870.0 ± 10.035 / 4082.1
Vis IW-ABCYes93.5 ± 2.681.9 ± 11.538 / 4083.3

Success rates and SR-AUC are percentages. Final checkpoints after 30,000 PPO iterations; 50 demonstrations per task; three independent training seeds. Coverage counts tasks with ≥80% success. SR-AUC summarizes success over the training horizon. These are known-task evaluations under randomized initial states.

BEYOND THE SIMULATOR

One Piper policy.
Four physical tasks.

82.5%66 successful trials out of 80
01 / Train jointly in parallel simulation

02 / DEPLOY WITH FIXED WEIGHTS

From parallel practice
to a physical robot.

A separate state-input policy is jointly trained on four RoboTwin Piper tasks, using 50 demonstrations per task. It transfers to the physical Piper without weight updates.

The physical interface supplies task identity and robot, object, and target states in the simulation observation format.

Four tasks. One set of weights.

01 / PIPER

Click bell

20 physical trials
20 / 20
02 / PIPER

Shake bottle

20 physical trials
15 / 20
03 / PIPER

Move pill bottle onto pad

20 physical trials
13 / 20
04 / PIPER

Place container on plate

20 physical trials
18 / 20

Representative physical rollouts, shown at the supplied recording speed. Success counts summarize the full 20-trial evaluation per task.

Explore Hebero.

Anonymous repositories for code, simulation assets, and demonstrations.