Three roles a model can play
Perceive
Sensor data in, facts out. Detection, segmentation, pose, depth, captioning, speech-to-text. Something else (your code, a workflow, an agent) decides what the robot does with the result.
Reason
Instructions in, plans out. Language and vision-language models answer prompts and break a goal into steps. They also power the AI agents.
Act
Observations in, robot actions out. VLA models and RL policies produce joint or velocity commands directly, so they run as the robot’s controller.
Modalities: what goes in, what comes out
Each model declares the inputs it accepts. The catalog filters on them.
Outputs come back as typed results in the SDK: detections (boxes, masks, keypoints, oriented boxes), classifications, depth maps, semantic masks, embeddings, text, JSON and images.
Tasks: the job you ask a model to do
A task fixes the prompt and the shape of the answer, so every model that supports it returns the same format. Models that support several tasks let you pick one in the Playground and in the workflow Call Model node.
Output schemas and parsing: Tasks and outputs.
Where models come from and where they run
- The catalog: public models plus your workspace’s own, filterable by input, task, vendor, size and edge runtime. See the model catalog.
- Providers: local or edge runtimes, cloud APIs such as OpenAI and Anthropic, Hugging Face, or your own inference server.
- Your own model: register one behind your API, or upload code for Cyberwave to host. See Bring your own model.
Use a model
Try it in the Playground
Run any catalog model on an image, a frame or a recording, no code.
Call it from a workflow
The Call Model node, with the same tasks as the Playground.
Load it in Python
cw.models.load("yolo26n.pt").predict(frame), locally or in the cloud.Make it the controller
VLA models and RL policies drive the robot directly.