> ## Documentation Index
> Fetch the complete documentation index at: https://docs.relayintelligence.org/llms.txt
> Use this file to discover all available pages before exploring further.

# Observations

> What to send each step, for each model, and the options you can add to any request.

An observation is a Python dictionary with one model's camera frames, the robot's state and a task prompt. `relay` sends it as is, and the server does all the preprocessing. The keys depend on the model.

<Tabs>
  <Tab title="π0 and π0.5">
    ```python theme={null}
    obs = {
        "observation.images.base_0_rgb": front,        # (H, W, 3) uint8
        "observation.images.left_wrist_0_rgb": left,
        "observation.images.right_wrist_0_rgb": right,
        "state": joints,                               # (32,) float32
        "prompt": "pick up the cup",
    }
    actions = policy.infer(obs)                        # (50, 32) float32
    ```

    <ParamField path="observation.images.base_0_rgb" type="image">
      Main third-person camera.
    </ParamField>

    <ParamField path="observation.images.left_wrist_0_rgb" type="image">
      Left wrist camera.
    </ParamField>

    <ParamField path="observation.images.right_wrist_0_rgb" type="image">
      Right wrist camera.
    </ParamField>

    <ParamField path="state" type="float32 array" required>
      Joint state, shape `(32,)` or `(1, 32)`. π0.5 writes the state into its prompt, so it rejects any other length. Pad a shorter state with zeros yourself, in the order the checkpoint was trained on.
    </ParamField>

    <ParamField path="prompt" type="str" required>
      The task in plain language, for example `"fold the towel"`.
    </ParamField>

    **Images**

    * Send `uint8` arrays in height × width × RGB order, floats in `[0, 1]`, or JPEG/PNG bytes.
    * You can send any frame size. The server resizes to 224×224, keeps the aspect ratio and pads the short side.
    * You can leave out a camera your robot doesn't have, but you must send at least one. A missing camera's slot is masked out rather than filled with a blank frame.
    * Use the same key for the same physical camera every time. The model learns which view each key is.

    **Actions**
    A `(50, 32)` `float32` array of joint-position targets: 50 future steps, padded to 32 dimensions. Use the first *n* dimensions, where *n* matches your state layout.
  </Tab>

  <Tab title="GR00T N1.7">
    GR00T takes nested dictionaries and returns one array per action group. The server runs the DROID embodiment (`OXE_DROID_RELATIVE_EEF_RELATIVE_JOINT`).

    ```python theme={null}
    eef = np.zeros((1, 1, 9), np.float32)
    eef[..., 3:] = [1, 0, 0, 0, 1, 0]          # position + 6D rotation (identity)

    obs = {
        "video": {
            "exterior_image_1_left": exterior,  # (1, T, H, W, 3) uint8
            "wrist_image_left": wrist,          # (1, T, H, W, 3) uint8
        },
        "state": {
            "eef_9d": eef,                                        # (1, 1, 9)
            "gripper_position": np.zeros((1, 1, 1), np.float32),  # (1, 1, 1)
            "joint_position": np.zeros((1, 1, 7), np.float32),    # (1, 1, 7)
        },
        "prompt": "pick up the object",
        "session_id": "robot-1",
    }
    actions = policy.infer(obs)
    actions["joint_position"].shape   # (1, 40, 7)
    ```

    <ParamField path="video" type="dict" required>
      One array per camera with shape `(batch, frames, height, width, 3)` and dtype `uint8`. The reference observation sends 2 frames per camera at 256×256.
    </ParamField>

    <ParamField path="state" type="dict" required>
      `eef_9d` (end-effector position plus 6D rotation), `gripper_position` and `joint_position` (7 joints), each `float32` with shape `(1, 1, k)`.
    </ParamField>

    <ParamField path="prompt" type="str" required>
      The task in plain language.
    </ParamField>

    <ParamField path="session_id" type="str">
      Recommended. The server reports `needs_session_id: true` for GR00T. See [sessions](/control-loop#sessions).
    </ParamField>

    **Actions**
    A dict with `eef_9d` `(1, 40, 9)`, `gripper_position` `(1, 40, 1)` and `joint_position` `(1, 40, 7)`, all `float32`. Execute the action group your controller accepts.

    <Warning>
      GR00T frames must be raw `uint8` arrays. The server only decodes JPEG/PNG bytes at the top level of an observation, not inside `video`.
    </Warning>
  </Tab>
</Tabs>

## Sending JPEG frames

For π0 and π0.5, you can send any camera as JPEG or PNG bytes instead of an array. A JPEG is several times smaller than the raw frame, so this cuts upload time on Wi-Fi or a slow uplink.

```python theme={null}
import cv2

ok, jpeg = cv2.imencode(".jpg", frame_bgr, [cv2.IMWRITE_JPEG_QUALITY, 90])
obs["observation.images.base_0_rgb"] = jpeg.tobytes()
```

JPEG is lossy, so actions differ slightly from those for the raw frame. On a fast wired connection, raw arrays are usually just as fast and match exactly.

## Request options

You can add these keys to any observation. They change how this one request runs and are not passed to the model as input.

<ParamField path="seed" type="int">
  Fixes the sampling noise. With the same seed and observation, you get the same actions, which is useful for tests and comparing runs. Without it, every call samples fresh noise.
</ParamField>

<ParamField path="session_id" type="str" default="default">
  Names the episode or robot this observation belongs to. The first observation of a new session starts the model fresh. See [sessions](/control-loop#sessions).
</ParamField>

<ParamField path="sampling_params" type="dict">
  <Expandable title="properties">
    <ParamField path="num_inference_steps" type="int" default="10">
      Denoising steps for the action head (π0 and π0.5). Fewer steps are faster and slightly less precise.
    </ParamField>

    <ParamField path="lora" type="str | dict | false">
      Selects a fine-tuned adapter on π0 and π0.5 servers. Give an adapter name, `{"name": "so100_pick", "scale": 0.8}`, or `false` for the base model. If you leave it out, the server's default adapter applies, if it has one. Only adapters registered on the server can be selected. Naming any other adapter returns an error that lists the ones available.
    </ParamField>
  </Expandable>
</ParamField>

```python theme={null}
obs |= {
    "seed": 0,
    "session_id": "arm-left",
    "sampling_params": {"num_inference_steps": 5, "lora": "so100_pick"},
}
```

<Note>
  The metadata from `get_server_metadata()` always describes the base model. An adapter fine-tuned for a different robot can expect other cameras or state, and can return a different action width.
</Note>

## Limits

* An observation can be up to 64 MiB after encoding.
* Supported array dtypes are integers, unsigned integers, floats and booleans. Object, structured and complex arrays are rejected.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.