Skip to main content
At a glance: 20 min · Beginner · Python 3.10+, a Cyberwave API key, a laptop webcam, a cup and a bottle. No robot, no credits. You will build the smallest version of a real robot cell: vision-guided pick-and-sort. A camera looks at your desk. A perception model finds a cup or a bottle and where it is. The robot aims at it, grasps it, lifts it, and drops it in the bin for its class: cups on the left, bottles on the right. The robot is an SO-101 arm running in the free Playground simulation, so you can build and debug the logic before you touch hardware. When you have a real SO-101, you change one line and the same script drives it. The object’s class decides where it goes and its position in the frame decides where the arm aims. A perception model decides, your script controls the arm: the same structure as a sorting station on a production line. See Connect AI to robots.
In the Playground only the arm is simulated. Your real cup stays on the desk. After each pick, move the object out of view by hand and the loop waits for the next one. With a real arm, the gripper closes on the object for real.

Before you start

  • A Cyberwave account and an API token. Create the token under Profile → Access. See API tokens.
  • If your token covers more than one workspace, also set CYBERWAVE_WORKSPACE_ID. Otherwise the first request fails with 400 workspace_required.
  • A cup and a bottle. Both are in YOLO’s default classes, so no training is needed.

Build it

1

Install the SDK with the ML and camera extras

ml installs Ultralytics, which runs YOLO locally. camera installs OpenCV, which reads the webcam. Keep the quotes: zsh treats square brackets as a pattern.
2

Export your credentials

3

Write the sorting loop

Save this as pick_and_sort.py:
4

Run it and watch the arm

The SDK prints an environment link. Open it in your browser. Put a cup in front of the webcam: the arm aims, reaches, grasps, lifts and drops it on the left. Take the cup away, show a bottle: it goes to the right. Move the object left or right in the frame and the arm aims differently. Press Ctrl+C to stop.

Expected output

UUIDs, URLs, pixel positions and angles will differ:
Each object you show is picked once and dropped in the bin for its class. You have a working camera → model → robot loop with a real decision in it.
The “may have been lost” warning is expected on the first run: ensure_attached() has just attached a controller policy to the twin. On later runs the policy is already attached and the warning goes away.

How it works

model.predict(frame, classes=["cup", "bottle"]) returns only those two classes. Each detection has a label, a confidence and a bbox in pixels. The script takes the most confident one. Its label picks the bin; the center of its box picks the aim angle.
aim = (x_center / frame_width - 0.5) * 2 * AIM_RANGE maps the left edge of the image to −45° and the right edge to +45°. It is a straight-line guess that is good enough to see the idea. A real cell replaces it with hand-eye calibration, which turns a pixel into a position in the robot’s frame.
Aim → reach → close → lift → turn to bin → open → home. Each step is one arm.set_joints(pose, degrees=True) call. Positions are in radians unless you pass degrees=True. The pause after each move lets the arm arrive; on real hardware, tune it to your speed.
"playground" is a free kinematic simulation with no cloud instance. "simulation" starts a MuJoCo instance with physics, which uses credits. "live" sends the same commands to a real arm through a paired edge device.
Exactly one controller drives a twin at a time. This call attaches a controller that accepts SDK joint commands, so your script is the one in control. You can take over from the dashboard at any moment.

Make it real

  1. Pair the arm: install the CLI and pair a device.
  2. Mount the camera above the workspace, looking down, so left and right in the image match left and right for the arm. A laptop webcam facing you is mirrored: use frame = cv2.flip(frame, 1) or flip the sign of aim.
  3. Record HOME, REACH, LIFT, the bin angles and the gripper values in the twin’s Saved poses panel and paste them into the script.
  4. Change one line: cw.affect("live"). Keep a hand on the dashboard’s controller switch the first time.

If something goes wrong

  • ModuleNotFoundError: No module named 'cv2': install the camera extra (step 1).
  • No API key found!, 401 or 400 workspace_required: check the two environment variables in step 2.
  • Nothing is detected: light the scene, move the object closer, or lower confidence to 0.35.
  • The arm aims the wrong way: your camera is mirrored. Flip the frame or the sign of aim.
  • The script runs but the arm doesn’t move: make sure cw.affect("playground") comes before the first command and you opened the link the SDK printed.
Full list: Troubleshooting.

Next steps

Replace the rules with a learned policy

Record demonstrations of the same pick, train a VLA, and let it drive the arm.

Control agent

Say “put the cups on the left” and approve the plan before it runs.

Connect AI to robots

The map: AI agents, AI models, control, and how each run trains the next model.

Hand-eye calibration

Turn pixels into robot coordinates for accurate picks.