Skip to main content
An observation is a Python dictionary with one model’s camera frames, the robot’s state and a task prompt. relay sends it as is, and the server does all the preprocessing. The keys depend on the model.
image
Main third-person camera.
image
Left wrist camera.
image
Right wrist camera.
float32 array
required
Joint state, shape (32,) or (1, 32). π0.5 writes the state into its prompt, so it rejects any other length. Pad a shorter state with zeros yourself, in the order the checkpoint was trained on.
str
required
The task in plain language, for example "fold the towel".
Images
  • Send uint8 arrays in height × width × RGB order, floats in [0, 1], or JPEG/PNG bytes.
  • You can send any frame size. The server resizes to 224×224, keeps the aspect ratio and pads the short side.
  • You can leave out a camera your robot doesn’t have, but you must send at least one. A missing camera’s slot is masked out rather than filled with a blank frame.
  • Use the same key for the same physical camera every time. The model learns which view each key is.
Actions A (50, 32) float32 array of joint-position targets: 50 future steps, padded to 32 dimensions. Use the first n dimensions, where n matches your state layout.

Sending JPEG frames

For π0 and π0.5, you can send any camera as JPEG or PNG bytes instead of an array. A JPEG is several times smaller than the raw frame, so this cuts upload time on Wi-Fi or a slow uplink.
JPEG is lossy, so actions differ slightly from those for the raw frame. On a fast wired connection, raw arrays are usually just as fast and match exactly.

Request options

You can add these keys to any observation. They change how this one request runs and are not passed to the model as input.
int
Fixes the sampling noise. With the same seed and observation, you get the same actions, which is useful for tests and comparing runs. Without it, every call samples fresh noise.
str
default:"default"
Names the episode or robot this observation belongs to. The first observation of a new session starts the model fresh. See sessions.
dict
The metadata from get_server_metadata() always describes the base model. An adapter fine-tuned for a different robot can expect other cameras or state, and can return a different action width.

Limits

  • An observation can be up to 64 MiB after encoding.
  • Supported array dtypes are integers, unsigned integers, floats and booleans. Object, structured and complex arrays are rejected.