A complete loop
PolicyClient for the whole run. Connecting costs a handshake and possibly a route switch, while infer on an open connection costs one round trip.
Choosing how much of a chunk to execute
The server returns the whole chunk and keeps no action queue, so the trade-off is yours:
A common starting point is to execute about the first fifth of the chunk, then replan. Tune it on your robot: replan more often if the arm overshoots or misses moving objects, less often if motion is jerky at chunk boundaries.
Episodes and reset
Callreset() before each episode. It clears the server’s state for your connection, so the next observation starts the model fresh rather than continuing the last episode.
Sessions
Every observation belongs to a session, named by itssession_id key, or "default" if you don’t send one. The server tracks sessions per connection:
- The first observation of a session it hasn’t seen starts that session fresh.
- Later observations in the same session continue it.
reset()forgets every session on the connection.
session_id instead of calling reset(). It also lets several robots, or several arms, share one connection without mixing state:
needs_session_id: true in its metadata. Always send a session_id to it.
Connection lifetime
A call on a closed connection raiseswebsockets.exceptions.ConnectionClosed. To survive network drops in long runs, catch it, reconnect and retry the observation:
Latency
Eachinfer call costs one network round trip, the upload of your observation and the model’s inference time. To keep it low:
- Use the direct route. The client switches to it automatically when the server offers one. It skips Modal’s web proxy and saves about 14 ms per call.
WALLE_DIRECT=0turns it off. Use that only to rule out the direct route when you’re debugging. - Run close to the servers. The servers run in Modal’s
us-eastregion, so round-trip time grows with your distance from it. - Send JPEG frames on slow links. See Sending JPEG frames.
- Lower
num_inference_stepson π0 and π0.5 if you need faster inference and can accept slightly less precise actions.
python examples/demo.py to measure p50, p90 and p99 latency from your own network.