Environments
SO-101 tasks, with their IDs, state dimensions, success predicates, and reward shape.
SO101-Nexus ships Gymnasium-compatible manipulation tasks for the SO-101 arm.
Import so101_nexus.mujoco for the single-environment IDs or so101_nexus.warp for the GPU-batched counterparts.
Both backends expose the same task contracts, including versioned pick-and-place tasks.
For a runnable loop, see the quickstart. For the observation components themselves, see Observations; for action spaces, see Control modes.
Tasks
| Task | MuJoCo ID | Warp ID | Config class | Max steps | State dim |
|---|---|---|---|---|---|
| Pick and lift an object | MuJoCoPickLift-v1 | WarpPickLift-v1 | PickConfig | 1024 | 31 |
| Pick and return the arm to rest | MuJoCoPickReturn-v1 | WarpPickReturn-v1 | PickReturnConfig | 1024 | 42 |
| Center an object on a target disc | MuJoCoPickAndPlace-v2 | WarpPickAndPlace-v2 | PickAndPlaceV2Config | 1024 | 43 |
| Legacy origin-distance placement | MuJoCoPickAndPlace-v1 | WarpPickAndPlace-v1 | PickAndPlaceConfig | 1024 | 43 |
| Stack cube A on cube B | MuJoCoStackCube-v1 | WarpStackCube-v1 | StackCubeConfig | 1024 | 43 |
| Touch an object on the table | MuJoCoTouch-v1 | WarpTouch-v1 | TouchConfig | 512 | 31 |
| Gaze at an object | MuJoCoLookAt-v1 | WarpLookAt-v1 | LookAtConfig | 256 | 23 |
| Move the TCP a fixed distance | MuJoCoMove-v1 | WarpMove-v1 | MoveConfig | 256 | 22 |
State dim is the width of the default state observation. It changes if you override
observations on the config. all_registered_env_ids() returns all versions for each imported backend.
Seeded resets
reset(seed=...) restarts the environment's random sequence. Subsequent resets without
a seed continue the sequence. The quickstart also shows
how to seed random action sampling.
Warp accepts a scalar seed or one seed per world. Seeds must be integers from 0 through
2**32 - 1. A scalar derives distinct streams for each world. An explicit list assigns
each seed directly, so a world keeps its samples when moved to another batch position.
Repeated entries produce identical random samples for matching episode counters and configuration.
Each Warp world has separate task, robot, wrist-camera, and visualization-camera streams. Only worlds that reset advance their episode counters. Camera randomization does not change task sampling. GPU physics can still differ despite identical random samples. For training reproducibility and checkpoint limits, see Training.
Success conditions
Every task publishes a binary success flag in the info dict.
PickLift succeeds when the object is grasped and its height above the spawn pose
exceeds lift_threshold (default 0.05 m). Both conditions must hold on the same step.
PickReturn requires a grasp of the selected object and lift above lift_threshold.
Every arm joint must be within return_threshold_deg of its configured resting angle,
and every arm joint speed must be below robot.static_vel_threshold.
All conditions must hold on the same step. The gripper joint is excluded from the
posture and speed checks so its opening can match the held object.
Staticness is checked at the current step, without a dwell interval.
PickAndPlace v2 requires the area centroid of the object's projected visual footprint inside the visible target disc. The footprint is the union of projected triangles, not an asset origin, bounding box, or convex hull. The table must support the settled object without significant robot or other-object contact for a continuous interval. The evaluator samples physical support and object motion at every physics substep. It checks the centroid at the current control step. An open gripper alone does not prove release.
PickAndPlace v1 retains the legacy predicate.
The free-joint origin must lie within goal_thresh of the disc center in XY.
Its height must stay less than 10 mm above its initial height.
The object must pass the speed thresholds and is_grasped < 0.5.
The visible disc remains larger than the default acceptance radius.
This version does not verify physical table support.
StackCube succeeds when cube A's horizontal offset from cube B is within
sqrt(2) * cube_half_size + stack_alignment_margin and its height above cube B is within
stack_alignment_margin of exactly 2 * cube_half_size (resting flush on cube B's top
face), cube A is static, the robot is static, and cube A has been released
(is_grasped < 0.5). Unlike PickAndPlace, the robot-static check is required here.
Touch succeeds when the TCP comes within the target object's bounding radius plus
touch_margin (default 0.03 m) of the object center, so the threshold scales with object
size.
LookAt succeeds when the angle between the wrist camera's optical axis and the
direction to the target is at most half the camera's vertical field of view (fov_deg / 2,
or the live cam_fovy / 2 when fov_deg is None). This is exactly the GazeState
observation.
Move succeeds when the TCP's displacement along the move direction reaches at least
target_distance - success_threshold. Perpendicular drift is ignored.
Rewards
Each step returns a scalar roughly in [-1, 1], where 1.0 means the task just finished.
It is a weighted sum of reaching, grasping, task_objective, and a one-time
completion_bonus, and those four weights must sum to 1.0. The component functions live
in so101_nexus.rewards; the weights are set through RewardConfig.
Simple tasks pay the current level. Touch, LookAt, and Move each have exactly one thing
to get right, so the reward is how close the robot is right now: low when far away, rising
smoothly toward 1.0 as it closes in, plus the bonus on success. Distance terms use tanh
shaping rather than a hard cutoff, so reward rises gradually instead of jumping.
Multi-step tasks pay the change in progress. PickLift, PickReturn, PickAndPlace, and StackCube use potential-based shaping: outside the completion bonus, the reward is how much better this step is than the last, not how good the current state looks. Once the robot stops improving, that term drops to exactly zero. This removes the incentive to freeze in a good-looking pose (holding a lifted object for the rest of the episode, or hovering a cube above the target) instead of finishing, because dwelling pays nothing.
Because those terms measure a change, reward can go slightly negative when progress is
genuinely lost, such as dropping an object that was already grasped. That is intentional
and symmetric with the credit earned for making the same progress in the first place.
Success always pays the full 1.0.
Two optional penalties, action_delta_penalty (jerky motion) and energy_penalty (high
effort), default to 0.0. See Configs for every reward field and its
default.
Task details
Training
PickLift, Touch, LookAt, and Move have solved vanilla-PPO baselines on the Warp backend. PickAndPlace is solved by the demo-seeded BC-then-PPO recipe, not by vanilla PPO. StackCube has no baseline yet. Commands, hyperparameters, and per-seed results are in Training.