Skip to main content
A policy returns a chunk of future actions, not a single one: 50 steps for π0 and π0.5, 40 for GR00T. Your loop decides how many of them to execute before asking for the next chunk.

A complete loop

Keep one PolicyClient for the whole run. Connecting costs a handshake and possibly a route switch, while infer on an open connection costs one round trip.

Choosing how much of a chunk to execute

The server returns the whole chunk and keeps no action queue, so the trade-off is yours: A common starting point is to execute about the first fifth of the chunk, then replan. Tune it on your robot: replan more often if the arm overshoots or misses moving objects, less often if motion is jerky at chunk boundaries.
To hide latency, request the next chunk while the robot is still executing the current one. Send the observation from a background thread a few steps before the current chunk runs out.

Episodes and reset

Call reset() before each episode. It clears the server’s state for your connection, so the next observation starts the model fresh rather than continuing the last episode.

Sessions

Every observation belongs to a session, named by its session_id key, or "default" if you don’t send one. The server tracks sessions per connection:
  • The first observation of a session it hasn’t seen starts that session fresh.
  • Later observations in the same session continue it.
  • reset() forgets every session on the connection.
This means you can start a new episode by switching to a fresh session_id instead of calling reset(). It also lets several robots, or several arms, share one connection without mixing state:
GR00T reports needs_session_id: true in its metadata. Always send a session_id to it.

Connection lifetime

The server closes a connection that sends nothing for 30 seconds. If your loop can pause longer, for example while a person resets the scene, close the client and create a new PolicyClient when you resume.
A call on a closed connection raises websockets.exceptions.ConnectionClosed. To survive network drops in long runs, catch it, reconnect and retry the observation:
A new connection starts with no sessions, so the next observation starts the model fresh.

Latency

Each infer call costs one network round trip, the upload of your observation and the model’s inference time. To keep it low:
  • Use the direct route. The client switches to it automatically when the server offers one. It skips Modal’s web proxy and saves about 14 ms per call. WALLE_DIRECT=0 turns it off. Use that only to rule out the direct route when you’re debugging.
  • Run close to the servers. The servers run in Modal’s us-east region, so round-trip time grows with your distance from it.
  • Send JPEG frames on slow links. See Sending JPEG frames.
  • Lower num_inference_steps on π0 and π0.5 if you need faster inference and can accept slightly less precise actions.
Run python examples/demo.py to measure p50, p90 and p99 latency from your own network.