Real2Gym˙

REAL DATA · INTERACTIVE GYMS · REUSABLE SKILLS

Real2Gym: Building Gyms from Videos,
Bringing Skills to Robots

Kerui Ren1,2*Yingxiang Xu1,3*Kaiwen Song1,4Linning Xu5Bo Dai6Mulin Yu1†Tao Lu1†
1Shanghai Artificial Intelligence Laboratory2Shanghai Jiao Tong University3Zhejiang University4University of Science and Technology of China5The Chinese University of Hong Kong6The University of Hong Kong
* Equal contribution.† Corresponding authors.

TL;DR Turn human and robot videos into visually aligned, physically interactive gyms. Let agents explore, learn from feedback, and bring reusable skills to real robots.

01 / REAL2SIM

From real observations to interactive worlds.

24 tasks from DROID and EgoDex. Inspect the scene. Follow the demonstration.

02 / SCENE AUGMENTATION

One scene. More possibilities.

Four scenes, six independent variations each.

03 / AGENT EXECUTION

Observe. Write code. Act.

Explore 18 tasks, one operation stage at a time.

04 / REAL ROBOT

From simulation to the physical world.

Four tasks. GPT-6 Astra and Real2Gym. Two camera views per method.

05 / WHY REAL2GYM?

What changes beyond direct GPT-6 Astra?

  1. 01 / REAL2SIM

    Geometry first.
    Refine through execution.

    Pi3X + SAM2 initialize object-wise point clouds to constrain scene geometry and layout. An iterative self-questioning and correction process refines local discrepancies, producing high-fidelity scenes aligned with fine-grained manipulation.

    Object-level initialization → local correction → aligned interaction
  2. 02 / AGENT POLICY

    Decide in stages.
    Improve through experience.

    SAM3 and GraspNet improve object localization and grasp generation. Each decision generates code for an entire operation stage, coordinating multiple actions and checks. Successes and failures become reusable skills that guide later execution.

    Perception → stage code → feedback → updated skills
06 / METHOD

Build the gym. Evolve the agent.

Real2Gym reconstruction, execution-based refinement, augmentation and agent skill evolution
Overview of Real2Gym. Reconstruct scenes from real observations, refine alignment through motion and physics, augment the environments, and extract reusable skills from agent execution.
07 / QUALITATIVE RESULTS

Aligned scenes. Effective execution. Reusable experience.

Qualitative comparison of Real2Sim reconstruction
Real2Sim reconstruction across robot and human demonstrations. Each column shows one scene; rows compare GPT-5.6 Sol, GPT-6 Astra, Real2Gym, and the source video. The examples highlight differences in object geometry, spatial layout, camera viewpoint, and task-relevant details, with Real2Gym more closely preserving the demonstrated scene and interaction setup.
Qualitative comparison of zero-shot agent execution
Zero-shot execution on a bowl-stacking task. Frames progress from left to right, with the human demonstration above the three agent rollouts. Real2Gym completes the ordered stack, while GPT-6 Astra and GPT-5.6 Sol encounter placement or grasping failures highlighted in red.
Qualitative examples of skill reuse
Skill reuse in cabinet manipulation and adapter placement. Paired rows compare initial exploration with later execution using extracted skills. The cabinet example contrasts repeated door-opening attempts with a complete storage-and-closing sequence; the adapter example shows fewer repeated alignment attempts. Frames are selected excerpts rather than equally spaced execution timestamps.
CITATION

Build on Real2Gym.

@software{real2gym,
  author = {Ren, Kerui and Xu, Yingxiang and Song, Kaiwen and Xu, Linning and Dai, Bo and Yu, Mulin and Lu, Tao},
  title = {{Real2Gym: Building Gyms from Videos, Bringing Skills to Robots}},
  year = {2026},
  url = {https://github.com/real2gym/Real2Gym}
}