> ## Documentation Index
> Fetch the complete documentation index at: https://docs.relayintelligence.org/llms.txt
> Use this file to discover all available pages before exploring further.

# Control loops

> Run a policy on a real robot, episode after episode.

A policy returns a **chunk** of future actions, not a single one: 50 steps for π0 and π0.5, 40 for GR00T. Your loop decides how many of them to execute before asking for the next chunk.

## A complete loop

```python theme={null}
from relay import PolicyClient

REPLAN_EVERY = 10   # actions to execute from each chunk

policy = PolicyClient("pi05")
try:
    for episode in range(10):
        policy.reset()
        while not episode_done():
            actions = policy.infer(get_observation())
            if isinstance(actions, dict) and actions.get("type") == "error":
                raise RuntimeError(actions["message"])
            for action in actions[:REPLAN_EVERY]:
                send_to_robot(action)
finally:
    policy.close()
```

Keep one `PolicyClient` for the whole run. Connecting costs a handshake and possibly a route switch, while `infer` on an open connection costs one round trip.

## Choosing how much of a chunk to execute

The server returns the whole chunk and keeps no action queue, so the trade-off is yours:

| Execute | Behavior |
| - | - |
| The whole chunk | Fewest calls. The robot runs open-loop for the whole chunk and can't react to anything that changes during it. |
| The first few steps, then replan | Reacts quickly to the scene. You need at least one call per replan, so per-call latency must fit inside the time those steps take. |

A common starting point is to execute about the first fifth of the chunk, then replan. Tune it on your robot: replan more often if the arm overshoots or misses moving objects, less often if motion is jerky at chunk boundaries.

<Tip>
  To hide latency, request the next chunk while the robot is still executing the current one. Send the observation from a background thread a few steps before the current chunk runs out.
</Tip>

## Episodes and reset

Call `reset()` before each episode. It clears the server's state for your connection, so the next observation starts the model fresh rather than continuing the last episode.

## Sessions

Every observation belongs to a session, named by its `session_id` key, or `"default"` if you don't send one. The server tracks sessions per connection:

* The first observation of a session it hasn't seen starts that session fresh.
* Later observations in the same session continue it.
* `reset()` forgets every session on the connection.

This means you can start a new episode by switching to a fresh `session_id` instead of calling `reset()`. It also lets several robots, or several arms, share one connection without mixing state:

```python theme={null}
left  = policy.infer(obs_left  | {"session_id": "arm-left"})
right = policy.infer(obs_right | {"session_id": "arm-right"})
```

GR00T reports `needs_session_id: true` in its metadata. Always send a `session_id` to it.

## Connection lifetime

<Warning>
  The server closes a connection that sends nothing for **30 seconds**. If your loop can pause longer, for example while a person resets the scene, close the client and create a new `PolicyClient` when you resume.
</Warning>

A call on a closed connection raises `websockets.exceptions.ConnectionClosed`. To survive network drops in long runs, catch it, reconnect and retry the observation:

```python theme={null}
from websockets.exceptions import ConnectionClosed

def infer(obs):
    global policy
    try:
        return policy.infer(obs)
    except ConnectionClosed:
        policy = PolicyClient("pi05")
        return policy.infer(obs)
```

A new connection starts with no sessions, so the next observation starts the model fresh.

## Latency

Each `infer` call costs one network round trip, the upload of your observation and the model's inference time. To keep it low:

* **Use the direct route.** The client switches to it automatically when the server offers one. It skips Modal's web proxy and saves about 14 ms per call. `WALLE_DIRECT=0` turns it off. Use that only to rule out the direct route when you're debugging.
* **Run close to the servers.** The servers run in Modal's `us-east` region, so round-trip time grows with your distance from it.
* **Send JPEG frames on slow links.** See [Sending JPEG frames](/observations#sending-jpeg-frames).
* **Lower `num_inference_steps`** on π0 and π0.5 if you need faster inference and can accept slightly less precise actions.

Run `python examples/demo.py` to measure p50, p90 and p99 latency from your own network.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.