SO101-Nexus
Environments

Environments

SO-101 tasks, with their IDs, state dimensions, success predicates, and reward shape.

SO101-Nexus ships Gymnasium-compatible manipulation tasks for the SO-101 arm. Import so101_nexus.mujoco for the single-environment IDs or so101_nexus.warp for the GPU-batched counterparts. Both backends expose the same task contracts, including versioned pick-and-place tasks.

For a runnable loop, see the quickstart. For the observation components themselves, see Observations; for action spaces, see Control modes.

Tasks

TaskMuJoCo IDWarp IDConfig classMax stepsState dim
Pick and lift an objectMuJoCoPickLift-v1WarpPickLift-v1PickConfig102431
Pick and return the arm to restMuJoCoPickReturn-v1WarpPickReturn-v1PickReturnConfig102442
Center an object on a target discMuJoCoPickAndPlace-v2WarpPickAndPlace-v2PickAndPlaceV2Config102443
Legacy origin-distance placementMuJoCoPickAndPlace-v1WarpPickAndPlace-v1PickAndPlaceConfig102443
Stack cube A on cube BMuJoCoStackCube-v1WarpStackCube-v1StackCubeConfig102443
Touch an object on the tableMuJoCoTouch-v1WarpTouch-v1TouchConfig51231
Gaze at an objectMuJoCoLookAt-v1WarpLookAt-v1LookAtConfig25623
Move the TCP a fixed distanceMuJoCoMove-v1WarpMove-v1MoveConfig25622

State dim is the width of the default state observation. It changes if you override observations on the config. all_registered_env_ids() returns all versions for each imported backend.

Seeded resets

reset(seed=...) restarts the environment's random sequence. Subsequent resets without a seed continue the sequence. The quickstart also shows how to seed random action sampling.

Warp accepts a scalar seed or one seed per world. Seeds must be integers from 0 through 2**32 - 1. A scalar derives distinct streams for each world. An explicit list assigns each seed directly, so a world keeps its samples when moved to another batch position. Repeated entries produce identical random samples for matching episode counters and configuration.

Each Warp world has separate task, robot, wrist-camera, and visualization-camera streams. Only worlds that reset advance their episode counters. Camera randomization does not change task sampling. GPU physics can still differ despite identical random samples. For training reproducibility and checkpoint limits, see Training.

Success conditions

Every task publishes a binary success flag in the info dict.

PickLift succeeds when the object is grasped and its height above the spawn pose exceeds lift_threshold (default 0.05 m). Both conditions must hold on the same step.

PickReturn requires a grasp of the selected object and lift above lift_threshold. Every arm joint must be within return_threshold_deg of its configured resting angle, and every arm joint speed must be below robot.static_vel_threshold. All conditions must hold on the same step. The gripper joint is excluded from the posture and speed checks so its opening can match the held object. Staticness is checked at the current step, without a dwell interval.

PickAndPlace v2 requires the area centroid of the object's projected visual footprint inside the visible target disc. The footprint is the union of projected triangles, not an asset origin, bounding box, or convex hull. The table must support the settled object without significant robot or other-object contact for a continuous interval. The evaluator samples physical support and object motion at every physics substep. It checks the centroid at the current control step. An open gripper alone does not prove release.

PickAndPlace v1 retains the legacy predicate. The free-joint origin must lie within goal_thresh of the disc center in XY. Its height must stay less than 10 mm above its initial height. The object must pass the speed thresholds and is_grasped < 0.5. The visible disc remains larger than the default acceptance radius. This version does not verify physical table support.

StackCube succeeds when cube A's horizontal offset from cube B is within sqrt(2) * cube_half_size + stack_alignment_margin and its height above cube B is within stack_alignment_margin of exactly 2 * cube_half_size (resting flush on cube B's top face), cube A is static, the robot is static, and cube A has been released (is_grasped < 0.5). Unlike PickAndPlace, the robot-static check is required here.

Touch succeeds when the TCP comes within the target object's bounding radius plus touch_margin (default 0.03 m) of the object center, so the threshold scales with object size.

LookAt succeeds when the angle between the wrist camera's optical axis and the direction to the target is at most half the camera's vertical field of view (fov_deg / 2, or the live cam_fovy / 2 when fov_deg is None). This is exactly the GazeState observation.

Move succeeds when the TCP's displacement along the move direction reaches at least target_distance - success_threshold. Perpendicular drift is ignored.

Rewards

Each step returns a scalar roughly in [-1, 1], where 1.0 means the task just finished. It is a weighted sum of reaching, grasping, task_objective, and a one-time completion_bonus, and those four weights must sum to 1.0. The component functions live in so101_nexus.rewards; the weights are set through RewardConfig.

Simple tasks pay the current level. Touch, LookAt, and Move each have exactly one thing to get right, so the reward is how close the robot is right now: low when far away, rising smoothly toward 1.0 as it closes in, plus the bonus on success. Distance terms use tanh shaping rather than a hard cutoff, so reward rises gradually instead of jumping.

Multi-step tasks pay the change in progress. PickLift, PickReturn, PickAndPlace, and StackCube use potential-based shaping: outside the completion bonus, the reward is how much better this step is than the last, not how good the current state looks. Once the robot stops improving, that term drops to exactly zero. This removes the incentive to freeze in a good-looking pose (holding a lifted object for the rest of the episode, or hovering a cube above the target) instead of finishing, because dwelling pays nothing.

Because those terms measure a change, reward can go slightly negative when progress is genuinely lost, such as dropping an object that was already grasped. That is intentional and symmetric with the credit earned for making the same progress in the first place. Success always pays the full 1.0.

Two optional penalties, action_delta_penalty (jerky motion) and energy_penalty (high effort), default to 0.0. See Configs for every reward field and its default.

Task details

Training

PickLift, Touch, LookAt, and Move have solved vanilla-PPO baselines on the Warp backend. PickAndPlace is solved by the demo-seeded BC-then-PPO recipe, not by vanilla PPO. StackCube has no baseline yet. Commands, hyperparameters, and per-seed results are in Training.

On this page