γ-Art

Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining

γ-Art: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining

Jicong Ao1†‡ Shuhan Jiang1,2† Yuling Zhong1,3 Yanwen Liu1,3 Yuhan Gao1,4 Jiangyuan Zhao1,5 Yang Zhang1,6 Shiqiang Zhu2 Chenjia Bai1,7* Xuelong Li1*
1Institute of Artificial Intelligence (TeleAI), China Telecom
2Zhejiang University
3Technical University of Munich
4Harbin Institute of Technology
5Shanghai Jiao Tong University
6Tsinghua University
7Gamma Robotics (γ)
†Equal contributions      ‡Project Lead      *Corresponding authors

01 Overview

Large-Scale Synthetic Pretraining for Articulated Manipulation

Articulated-object manipulation demands precise contact with functional parts while satisfying joint constraints, making high-quality real-world demonstrations difficult to collect at scale. γ-Art addresses this bottleneck with γ-Sim, an articulation-aware IsaacLab-based simulation platform for asset annotation, skill and motion generation, and efficient demonstration collection. On top of it, agentic task generation creates articulated-object tasks with varied complexity and temporal horizons, while a distributed synthesis system—shown in the accompanying videos as γ-Hive—scales trajectory generation and dataset production across parallel workers. Together, these components produce γ-Art-Data, with 1,004,584 demonstrations spanning 44 task types, 5 robot setups, and 2,507 articulated objects. A VLA model pretrained on this synthetic data shows competitive simulation performance and transfers zero-shot to real-world articulated-object manipulation, demonstrating the value of scalable synthetic supervision for contact-rich tasks.

Agentic task generation

Natural-language requests become structured task definitions, simulation assets, skills, and execution settings.

Simulation framework

γ-Sim composes assets, scenes, skills, camera layouts, and randomized physical settings into executable tasks.

γ-Hive

A distributed production layer scales trajectory collection, visual rendering, and dataset conversion across parallel workers.

Sim-to-real transfer

Representative drawer, sliding-door, bucket-lifting, and dispenser-pressing trials on RM75, AC1, and R1Pro.

02 Platform

γ-Sim

γ-Sim platform overview showing its asset library, annotation and motion generation, parallel simulation, platform capabilities, and data production pipeline
γ-Sim platform overview: asset library, annotation and motion generation, parallel simulation, domain randomization, and data production.

γ-Sim is a versatile IsaacLab-based simulation platform designed for articulated-object manipulation. It turns natural-language task descriptions into executable configurations and combines asset management, motion generation, parallel simulation, domain randomization, and data production into one platform.

The platform supports seven embodiments—Franka Emika Panda, RM75, AC1, R1Pro, Flexiv Rizon 4S, Marvin M6s, and Unitree H1-2—and assembles them with configurable camera layouts. Five representative setups are used in γ-Art-Data.

Its asset library contains 9,116 rigid objects across 1,071 categories and 2,507 articulated objects across 88 functional categories. γ-Sim also integrates 1,452 visual scenes and more than 10,000 high-resolution 4K indoor images generated with Qwen-Image to broaden scene and illumination coverage. Hunyuan-World 3DGS scenes and Hunyuan3D assets are additionally used for sim-to-real task synthesis.

Constraint-aware Cartesian trajectories, collision-aware GPU planning, and lower-velocity contact establishment are specialized for sustained articulated interactions such as doors, drawers, knobs, and appliance controls. An asynchronous plan-and-execute design scales collection across parallel environments.

Asset Library

γ-Sim builds executable task instances from a broad library of rigid and articulated objects, robot embodiments, and generated assets for sim-to-real task synthesis.

Embodiment Support

Representative robot platforms and configurable embodiment setups supported by γ-Sim.

AC1 dual-arm robot embodiment
AC1
Franka Research 3 robot embodiment
Franka Research 3
Unitree H1-2 humanoid robot embodiment
Unitree H1-2
AgileX Piper robotic arm embodiment
AgileX Piper
Universal Robots embodiment
Universal Robots
ARX-AC1 dual-arm robot embodiment
ARX-AC1
Dualarm RM75 robot embodiment
Dualarm RM75
Dualarm Franka Panda robot embodiment
Dualarm Franka Panda
Galaxea R1Pro humanoid robot embodiment
Galaxea R1Pro
Marvin M6S dual-arm robot embodiment
Marvin M6S
Flexiv Rizon 4S robot embodiment
Flexiv Rizon 4S

Manipulable Object Support

γ-Sim builds task instances from a broad library of rigid objects, articulated objects, robot embodiments, and generated assets for sim-to-real task synthesis.

9,116Rigid objects
1,071Rigid object categories
2,507Articulated objects
88Functional articulated categories
Oven articulated object asset
Oven
Refrigerator articulated object asset
Refrigerator
Microwave articulated object asset
Microwave
Water dispenser articulated object asset
Water dispenser
Nightstand articulated object asset
Nightstand
Wardrobe articulated object asset
Wardrobe
Toilet articulated object asset
Toilet
Window articulated object asset
Window
Juicer articulated object asset
Juicer
Toolbox articulated object asset
Toolbox
Desk lamp articulated object asset
Desk lamp
Faucet articulated object asset
Faucet

Scene Library

γ-Sim integrates 1,452 visual scenes and more than 10,000 high-resolution 4K indoor images, spanning bathrooms, bedrooms, kitchens, and living rooms.

1,452Visual scenes
10,000+4K indoor images
4KImage resolution
Multi-SourceGN0, SDGScene, InteriorGS, etc.
Bathroom scene asset example
Bathroom scene
Bedroom scene asset example
Bedroom scene
Kitchen scene asset example
Kitchen scene
Alternative kitchen scene asset example
Kitchen variation

Asset Annotation

Rigid and articulated assets are annotated with geometry, category, and interaction metadata so they can be instantiated consistently across simulation tasks.

Rigid object annotation pipeline from RGB and depth rendering through AnyGrasp
Rigid object annotation

Rigid assets are organized by category and visual variants, with stable geometry and appearance metadata for placement, material randomization, and scene composition.

Articulated object annotation with part and joint interaction metadata
Articulated object annotation

Articulated assets expose movable parts and joint semantics, including links, axes, limits, and interaction regions for constraint-aware motion generation.

Domain Randomization

Domain randomization increases demonstration diversity while preserving articulated-part semantics through spatial, physical, and visual variation:

  • Object poses
  • Robot initial configurations
  • Camera extrinsics
  • Part-wise friction and density
  • Joint stiffness and damping
  • Part-wise material and texture
  • Local and global illumination

03 Data Synthesis

Synthetic Data Scaling

γ-Art connects task authoring to large-scale demonstration production. A task-generation agent first grounds natural language in executable simulation configurations; distributed workers then decouple trajectory rollout, multi-view rendering, and dataset conversion so failed trajectories do not consume the rendering budget. At scale, each 8-GPU server produces approximately 144 hours of data per day.

Data synthesis workflow from agentic task generation to distributed demonstration production
The data-synthesis workflow: agentic task generation produces executable configurations, and distributed workers turn them into validated γ-Art-Data demonstrations.

Agentic Task Generation

A task-generation agent converts a natural-language manipulation request into a grounded, executable simulation configuration. The agent reasons over the asset catalog, part-level manipulation annotations, and examples of valid task files rather than creating simulator calls directly from unconstrained text.

  1. Unified task specification One structured task describes the scene, objects, and robot embodiment, the skill sequence, initial state, and success criterion, plus rendering, physics, randomization, and parallelism settings.
  2. Scene grounding An asset-retrieval tool maps the instruction to a robot setup, target object, background scene, and auxiliary assets; it also resolves robot configurations and physically plausible object states and poses.
  3. Skill-program construction Part-level joint and interaction annotations guide composition of the 23 reusable manipulation skills into a parameterized high-level program.
  4. Simulation setup generation The agent selects lighting, rendering quality, and domain randomization, then scales environment parallelism to the expected task horizon; humans can optionally refine the result in the simulator UI.

Distributed Demonstration Synthesis

Large-scale synthesis separates CPU-bound planning from GPU-intensive rendering. A scheduler coordinates three workload types across shared GPU and CPU resources, while quality checks remove unstable or unsuccessful rollouts before they reach the rendering stage.

  1. Trajectory collection Workers execute up to 500 parallel environments with a shared motion planner; jerk and configuration-deviation checks trigger early resets, and only successful trajectories are retained.
  2. Visual rendering Stored seeds enable deterministic replay. Tiled Isaac Lab cameras perform parallel multi-view rendering while observations are incrementally encoded into HDF5.
  3. Dataset conversion CPU workers stream HDF5 results into LeRobot format, map action and observation keys across embodiments, encode multi-camera videos, and remove redundant fields.
  4. Distributed scheduling A shared storage pool connects stages; the scheduler supports FIFO and priority dispatch, splits long jobs into subtasks, monitors status, and retries failed tasks.

04 Dataset

γ-Art-Data

1,004,584
Trajectories
44
Task types
2,507
Articulated Objects
5
Robot setups
23
Manipulation skills
500M+
Frames

γ-Art-Data is a large-scale synthetic dataset produced by γ-Sim for articulated-object manipulation. It contains 1,004,584 trajectories, approximately 500 million frames, and more than 4,600 hours of manipulation data across five robot setups.

It covers 44 task types and 23 reusable skills, with 2,507 articulated objects spanning 23 functional categories in the collected trajectories and 88 categories in the full asset library. Examples include cabinets, drawers, windows, refrigerators, ovens, dishwashers, dispensers, and microwave controls. Each demonstration contains multi-view observations, proprioceptive states, language instructions, and action trajectories in a standardized LeRobot v3.0 format.

The dataset includes atomic, composite, and long-horizon tasks (69.4%, 20.9%, and 9.7%), plus variation in object placement, camera parameters, robot initial configuration, physical properties, materials, and illumination. These controlled variations target the visual, geometric, and contact-related sources of the sim-to-real gap.

Trajectories by Robot Setup

1.0M+trajectories

Task Complexity Mix

44task types
Distribution of trajectories across articulated object categories
Articulated object category coverage
Distribution of manipulation skills
Manipulation skill distribution
Distribution of task durations
Task duration distribution

05 Results

Experimental Results

Simulation Benchmark Evaluation

96.8%LIBERO average success rate
28.8%RoboCasa365 average score
+2.1 ptsAdvantage on LIBERO
+11.9 ptsAdvantage on RoboCasa365

The model borrows the PRTS architecture: a Qwen3-VL-4B backbone with a DiT-based action expert trained with flow matching. For controlled comparisons, the CRL objective is removed. γ-Art-VLA uses γ-Art-Data for first-stage behavior-cloning pretraining, then is post-trained on the downstream benchmark; γ-Art-VLA(w/o S1) skips first-stage pretraining. On standard post-training settings, γ-Art-VLA reaches 96.8% average success on LIBERO and 28.8% on RoboCasa365.

On LIBERO, γ-Art-VLA exceeds γ-Art-VLA(w/o S1) by 2.1 points (96.8% vs. 94.7%), surpasses π0 (94.2%) and Qwen3-VL-π (94.7%), and approaches π0.5 (96.9%) and γ-Art-VLA(Real) (97.8%), which uses real-world pretraining. On RoboCasa365, γ-Art-VLA outperforms the strongest listed baseline, GR00T N1.5 (23.9%), despite using synthetic pretraining only. γ-Art-VLA also exceeds π0.5 by 11.9 points on RoboCasa365 and demonstrates an earlier performance advantage during post-training, while γ-Art-VLA(w/o S1) reaches 89.9% at the same point.

LIBERO Benchmark

Average success rate across the LIBERO benchmark suites (%).

MethodAverage
π094.2
Qwen3-VL-π94.7
π0.596.9
γ-Art-VLA (Real)97.8
γ-Art-VLA (w/o S1)94.7
γ-Art-VLA96.8

RoboCasa365 Benchmark

Success rate by task split (%); bold values indicate the best result in each column.

MethodAtomic SeenComposite SeenComposite UnseenAverage
π034.66.11.114.8
π0.539.67.11.216.9
GR00T N1.651.19.41.721.9
GR00T N1.550.714.82.723.9
γ-Art-VLA (w/o S1)43.39.44.420.0
γ-Art-VLA52.823.17.528.8

Zero-Shot Sim-to-Real Transfer

76.7%overall success across 12 real-world tasks
85.0%Average success on AC-1
84.4%Average success on R1Pro
+26.7 ptsAdvantage on real-world tasks

A key result is zero-shot sim-to-real transfer: models are post-trained with synthetic task demonstrations only, without task-specific real-world demonstrations, and then evaluated on physical robots. The evaluation covers 12 articulated-object tasks across a dual-arm RealMan RM75, R1Pro, and ARX AC-1. Each platform uses three RGB cameras, including one center camera and one wrist camera on each arm. We synthesize 500 demonstrations per task in approximately two hours and measure success over 15 independent trials.

γ-Art-VLA achieves 76.7% overall success, compared with 50.0% for γ-Art-VLA(w/o S1), 59.4% for π0.5, and 69.4% for γ-Art-VLA(Real). It matches or outperforms γ-Art-VLA(w/o S1) on every matched task; the largest gains are 46.7 points on RM75 bucket-handle lifting and AC-1 drawer pulling. Performance also scales positively as synthetic task demonstrations increase from 1 to 500 episodes.

Real-world evaluation across three robot platforms and twelve tasks; success rates are measured over 15 trials.

RM75 · Dispenser Press — Zero-shot Sim-to-Real
AC1 · Drawer Pull — Zero-shot Sim-to-Real
AC1 · Sliding Door — Zero-shot Sim-to-Real
R1Pro · Microwave Door Closing — Zero-shot Sim-to-Real

Performance scaling with synthetic demonstrations

Increasing the task-specific synthetic dataset from 1 to 500 episodes produces a positive success-rate trend. At 500 episodes, γ-Art-VLA exceeds 80% on each of the five scaling tasks shown below; constrained actions such as drawer pulling and knob rotation continue to benefit from additional demonstrations.

R1Pro success rates as synthetic demonstrations increase from 1 to 500 episodes
R1Pro · Bucket lifting, microwave closing, and toaster pressing
AC1 success rates as synthetic demonstrations increase from 1 to 500 episodes
AC1 · Drawer pulling and microwave knob rotation

Citation

@article{smart2026,
  title={SMART: Zero-Shot Sim-to-Real Articulated Object Manipulation via Large-Scale Synthetic Pretraining},
  author={Ao, Jicong and Jiang, Shuhan and Zhong, Yuling and Liu, Yanwen and Gao, Yuhan and Zhao, Jiangyuan and Zhang, Yang and Zhu, Shiqiang and Bai, Chenjia and Li, Xuelong},
  year={2026},
  url={https://arxiv.org/abs/2610.07652}
}