The dataset viewer should be available soon. Please retry later.
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Imagined Data
This repository hosts imagined interaction data generated by world models across different environments, tasks, and data sources. Data are organized into separate subdatasets, with additional types of imagined data to be added over time.
The repository contains the RoboTwin2.0 and RealWorld subdatasets. Storage formats, field definitions, and state and action semantics are documented separately for each subdataset.
Dataset Index
| Subdataset | Directory | Contents |
|---|---|---|
| RoboTwin2.0 | RoboTwin2.0/ |
Imagined interaction segments for 50 RoboTwin 2.0 tasks |
| RealWorld | RealWorld/ |
Imagined interaction segments for nine real-world Aloha tasks, initialized from 18 recorded data sources |
RoboTwin2.0
The directory structure, field definitions, array shapes, and reading examples below apply to the data under RoboTwin2.0/.
The RoboTwin2.0 subdataset contains imagined interaction segments generated by a world model for 50 RoboTwin 2.0 tasks. Data are organized by task. Each HDF5 file stores one fixed-length chunk: 21 observation frames, 20 action steps, and the corresponding rewards and episode-ending flags.
Observations include RGB images from three camera viewsβhead, left wrist, and right wristβalong with joint states, gripper states, and end-effector poses for both arms. Frame 0 is the initial observation of the chunk; frames 1β20 are subsequent observations generated by the world model. Each file represents an interaction segment and should not be treated as a complete episode.
Directory Structure
RoboTwin2.0/
βββ adjust_bottle/
β βββ chunk_<uid>.hdf5
β βββ ...
βββ handover_block/
β βββ chunk_<uid>.hdf5
β βββ ...
βββ place_dual_shoes/
β βββ ...
βββ ... # 50 task directories in total
The task directory name identifies the task, whereas the task field inside each file contains the natural-language instruction for that sample.
File Schema
Each file contains 18 HDF5 datasets and a root attribute named source_identity. All dataset paths in the table below are relative to the file root. The root attribute is documented separately below.
| Field | Shape | Dtype | Description |
|---|---|---|---|
task |
() |
UTF-8 string | Natural-language task instruction for the current chunk |
obs/head_cam/rgb |
(21, 240, 320, 3) |
uint8 |
RGB images from the head camera |
obs/left_wrist_cam/rgb |
(21, 240, 320, 3) |
uint8 |
RGB images from the left wrist camera |
obs/right_wrist_cam/rgb |
(21, 240, 320, 3) |
uint8 |
RGB images from the right wrist camera |
obs/left_joint_pos |
(21, 6) |
float32 |
Commanded positions of the six left-arm joints (rad) |
obs/right_joint_pos |
(21, 6) |
float32 |
Commanded positions of the six right-arm joints (rad) |
obs/left_gripper |
(21,) |
float32 |
Commanded opening fraction of the left gripper |
obs/right_gripper |
(21,) |
float32 |
Commanded opening fraction of the right gripper |
obs/left_ee_pose |
(21, 7) |
float32 |
Left-arm end-effector pose: [x, y, z, qw, qx, qy, qz] |
obs/right_ee_pose |
(21, 7) |
float32 |
Right-arm end-effector pose: [x, y, z, qw, qx, qy, qz] |
action/left_joint_pos |
(20, 6) |
float64 |
Raw target position commands for the six left-arm joints (rad) |
action/right_joint_pos |
(20, 6) |
float64 |
Raw target position commands for the six right-arm joints (rad) |
action/left_gripper |
(20,) |
float64 |
Raw action values for the left gripper |
action/right_gripper |
(20,) |
float64 |
Raw action values for the right gripper |
rewards |
(20,) |
float32 |
Rewards associated with the 20 action steps |
terminations |
(20,) |
bool |
Task or environment termination flags |
truncations |
(20,) |
bool |
Truncation flags, for example due to time limits |
dones |
(20,) |
bool |
terminations OR truncations |
Images
Image arrays use the (T, H, W, C) layout, with RGB channel order and pixel values in 0β255. The three views are aligned at each frame index, with a resolution of 240 pixels in height and 320 pixels in width.
RGB images are stored as pixel arrays using lossless HDF5 gzip compression at level 4, with an HDF5 storage chunk shape of (1, 240, 320, 3). Reading a dataset returns image arrays directly; no additional JPEG or PNG decoding is required.
Joint States, Gripper States, and End-Effector Poses
The joint and gripper states of both arms can be combined into a 14-dimensional vector in the following order:
[left_joint_1, ..., left_joint_6, left_gripper,
right_joint_1, ..., right_joint_6, right_gripper]
obs/*_joint_pos and obs/*_gripper store the control targets used as the policy state. Row 0 is the initial state of the chunk. For the subsequent 20 rows, obs/*_joint_pos[t+1] is taken from the target joint angles in action/*_joint_pos[t], and obs/*_gripper[t+1] is obtained by clipping the corresponding gripper action to [0, 1]. Values are stored in the dtypes listed above and have not been normalized using the policy's normalization statistics. These fields do not represent measured joint or gripper feedback after robot motion.
For RoboTwin, the policy input state should use the joint positions and gripper states described here, not the end-effector poses.
Gripper values follow the convention 0 = closed, 1 = open. action/*_gripper retains the raw policy output, which may be slightly below 0 or above 1. Commands used for imagined execution clip gripper values to [0, 1], so a gripper action may differ from the gripper state in the next frame. Joint actions specify target positions, not joint increments.
The first three components of obs/*_ee_pose are positions in the world coordinate frame, expressed in meters. The remaining four components form a quaternion in wxyz order. The pose reference is the URDF link6 joint frame after RoboTwin's joint-coordinate correction, not the tool center point (TCP).
End-effector poses in generated frames come directly from the world model's predictions. During export, they are not replaced with forward kinematics computed from joint commands. Each end-effector pose field contains only the seven pose components; the separate obs/*_gripper fields come from the control targets described above.
Temporal Alignment
For t = 0, ..., 19:
obs[t] -- action[t] --> obs[t + 1]
β
βββ rewards[t], terminations[t], truncations[t], dones[t]
obs[0]is the last conditioning observation for the current chunk. For the first chunk of a branch, it comes from the offline data; for subsequent chunks, it is the final observation of the preceding chunk.obs[1:21]contains the 20 subsequent observations generated for the current chunk.action[t]corresponds to the transition fromobs[t]toobs[t+1]. The reward array and all three flag arrays use the same action index.
Each file therefore contains one more observation than action. The policy generates an action sequence at each chunk boundary; 21 observation frames do not imply 21 policy calls. Files also do not include the world model's complete conditioning history across multiple frames.
Rewards and Episode-Ending Flags
This subdataset stores the sparse binary rewards used during data generation. The first 19 reward entries in each chunk are zero, and the final entry stores the chunk's 0/1 reward. A reward model assigns this reward based on the generated observations; it is not a human annotation for each frame.
terminations indicates termination, and truncations indicates truncation. At every step:
dones = terminations | truncations
The last step of a chunk does not necessarily have done=True: the same branch may continue with another chunk. Use the stored flags to determine episode boundaries rather than inferring them from file boundaries alone.
Root Attribute
The root attribute source_identity in each HDF5 file stores a JSON string containing three provenance fields:
{
"branch_id": "<unique imagined branch identifier>",
"chunk_index": 0,
"transition_id": "<source record identifier>"
}
| Field | Description |
|---|---|
branch_id |
Identifier of the imagined rollout branch to which this chunk belongs |
chunk_index |
Zero-based chunk index within the branch |
transition_id |
Transition identifier in the original export record |
Read this attribute with json.loads(f.attrs["source_identity"]). It is included in the distributed HDF5 files and is not counted among the 18 datasets listed above.
To link chunks, group them by branch_id and sort by chunk_index. Concatenate chunks only when both adjacent chunks are available and their indices are consecutive. Retain the shared boundary observation between adjacent chunks only once.
RealWorld
The directory structure, field definitions, and semantics below apply to data under RealWorld/. This subdataset contains world-model-generated interaction segments initialized from recorded real-world Aloha observations. The generated frames and states are imagined observations, not measurements collected by executing these actions on a physical robot.
Each HDF5 file stores one fixed-length chunk: 21 observation frames, 20 action steps, and the corresponding rewards and episode-ending flags. Frame 0 is the last conditioning observation; frames 1β20 are generated by the world model. A chunk is an interaction segment and does not necessarily represent a complete episode.
Tasks and Directory Structure
The subdataset covers nine logical tasks:
cover_pencapfill_pen_holderfold_towelhandover_bottlematch_bottlesplugsort_blocksstack_blockstidy_the_table
Each task has two recorded sources: a directory named after the task and a corresponding <task>_rollout directory, for 18 sources in total. The _rollout sources are not separate tasks. Source-qualified episode identities are retained during data generation so that episodes with the same index in different sources remain distinct. The published task field preserves the original natural-language instruction for each sample; it is not rewritten from the task directory name.
RealWorld/
βββ cover_pencap/
β βββ rank_00/
β β βββ chunk_<uid>.hdf5
β β βββ ...
β βββ rank_01/
β β βββ ...
β βββ ... # rank_00 through rank_07
βββ fill_pen_holder/
β βββ ...
βββ ... # Nine logical task directories
The path template is RealWorld/<task>/rank_<2 digits>/chunk_<uid>.hdf5. The rank directories partition the files by their export provenance; they do not introduce additional tasks or define episode boundaries. Chunk continuity is determined from source_identity, not from directory order or filename order.
File Schema
Each file contains 16 HDF5 datasets and a root attribute named source_identity. All dataset paths below are relative to the file root.
| Field | Shape | Dtype | Description |
|---|---|---|---|
task |
() |
UTF-8 string | Original natural-language task instruction for the current chunk |
obs/head_cam/rgb |
(21, 240, 320, 3) |
uint8 |
RGB images from the head camera |
obs/left_wrist_cam/rgb |
(21, 240, 320, 3) |
uint8 |
RGB images from the left wrist camera |
obs/right_wrist_cam/rgb |
(21, 240, 320, 3) |
uint8 |
RGB images from the right wrist camera |
obs/left_joint_pos |
(21, 6) |
float32 |
Left-arm joint positions in radians; generated rows are world-model predictions |
obs/right_joint_pos |
(21, 6) |
float32 |
Right-arm joint positions in radians; generated rows are world-model predictions |
obs/left_gripper |
(21,) |
float32 |
Left gripper position in the original device-native numerical scale |
obs/right_gripper |
(21,) |
float32 |
Right gripper position in the original device-native numerical scale |
action/left_joint_pos |
(20, 6) |
float64 |
Absolute target positions for the six left-arm joints (rad) |
action/right_joint_pos |
(20, 6) |
float64 |
Absolute target positions for the six right-arm joints (rad) |
action/left_gripper |
(20,) |
float64 |
Absolute left gripper commands in the original device-native numerical scale |
action/right_gripper |
(20,) |
float64 |
Absolute right gripper commands in the original device-native numerical scale |
rewards |
(20,) |
float32 |
Sparse binary rewards associated with the 20 action steps |
terminations |
(20,) |
bool |
Task or environment termination flags |
truncations |
(20,) |
bool |
Truncation flags, for example due to time limits |
dones |
(20,) |
bool |
terminations OR truncations |
Images
Image arrays use the (T, H, W, C) layout, with RGB channel order and pixel values in 0β255. All three views are aligned at each frame index, with a resolution of 240 pixels in height and 320 pixels in width.
RGB images are stored as pixel arrays using lossless HDF5 gzip compression at level 4, with an HDF5 storage chunk shape of (1, 240, 320, 3). Reading a dataset returns image arrays directly; no additional JPEG or PNG decoding is required.
Joint States and Gripper States
The joint and gripper states form a 14-dimensional vector in this order:
[left_joint_1, ..., left_joint_6, left_gripper,
right_joint_1, ..., right_joint_6, right_gripper]
For this Aloha subdataset, the policy input state consists of the joint positions and gripper positions. The first observation of a branch comes from the recorded native robot state. The subsequent 20 rows in a chunk are the world model's predictions of the native slave-arm joint and gripper state. At the next chunk boundary, the preceding chunk's final predicted state supplies both the policy input and the world model's conditioning history.
The published observation fields are unnormalized values in native joint space. They represent predicted motion outcomes for generated frames and are not constructed by copying action targets. An observation at t+1 may therefore differ from the target command at t. The policy's normalization statistics have not been applied to the stored observation or action arrays.
Gripper values retain the original numerical scale used by the recorded data and robot interface. They are not converted to a normalized 0β1 opening fraction and are not clipped to [0, 1]. The source contract identifies their units as arx_x5_native_position; no conversion to a physical width or an angle is defined in this release. Preserve these values when reading or processing the data.
This subdataset does not contain obs/left_ee_pose or obs/right_ee_pose. The native world-model state is joint space, and no additional end-effector pose fields are synthesized for publication.
Actions
The action fields contain absolute target commands for the 12 arm joints and two grippers. The policy's internal delta-joint representation has already been transformed back to absolute joint targets before export. Joint actions are not increments, and the gripper commands retain their original native numerical scale.
Temporal Alignment
For t = 0, ..., 19:
obs[t] -- action[t] --> obs[t + 1]
β
βββ rewards[t], terminations[t], truncations[t], dones[t]
obs[0]is the last conditioning observation of the current chunk. For the first chunk of a branch, it comes from the recorded data; for subsequent chunks, it is the final predicted observation of the preceding chunk.obs[1:21]contains the 20 subsequent observations predicted for the current chunk.action[t]is the target command for the transition fromobs[t]toobs[t+1]. Rewards and episode-ending flags use the same action index.
Each chunk therefore contains one more observation than action. The policy generates an action sequence at each chunk boundary; 21 observation frames do not imply 21 policy calls. The file does not contain the world model's complete multi-frame conditioning history.
Rewards and Episode-Ending Flags
Rewards follow the sparse binary convention used during generation. The first 19 entries in each chunk are zero; the final entry stores the chunk's 0/1 reward. A frozen reward model scores all 20 generated frames using the three simultaneous camera views and the sample's original task instruction. If any frame has a predicted success probability of at least 0.5, the final reward entry is 1 and the chunk is marked terminated. These rewards are model-assigned labels, not human annotations for each frame.
At every step:
dones = terminations | truncations
A chunk boundary does not by itself imply done=True. Use the stored flags to determine episode boundaries. Recorded successful source episodes have no imposed episode time limit; recorded failed source episodes use their own trajectory length as the time limit, with elapsed time measured from the original anchor. Time-limit checks occur after complete chunks. This timing policy does not change the fixed 21-observation and 20-action file format.
Root Attribute
Each HDF5 file includes a root attribute named source_identity, containing this JSON structure:
{
"branch_id": "<unique imagined branch identifier>",
"chunk_index": 0,
"transition_id": "<source record identifier>"
}
| Field | Description |
|---|---|
branch_id |
Identifier of the imagined rollout branch to which this chunk belongs |
chunk_index |
Zero-based chunk index within the branch |
transition_id |
Transition identifier in the original export record |
Read the attribute with json.loads(f.attrs["source_identity"]). It is an HDF5 attribute, not one of the 16 datasets.
To link chunks, group by branch_id and sort by chunk_index. Concatenate only when both adjacent chunks are available and their indices are consecutive. Keep the shared boundary observation only once. The metadata used during generation are reduced to this three-field attribute in the published files; per-frame reward probabilities, scoring masks, and internal training metadata are not additional fields in the published schema.
- Downloads last month
- 7,161