Skip to main content
relay connects your robot to vision-language-action policies running on WALLE, our GPU inference servers. Each step, you send what the robot sees and feels plus a plain-language task. You get back a chunk of future actions to execute.
You don’t need a GPU, model weights or preprocessing code on the robot. The server resizes images, normalizes state, tokenizes the prompt and runs the model. The robot sends raw frames and joint readings.

Quickstart

Install, configure and get your first action chunk in five minutes.

Observations

The exact keys, shapes and dtypes each model expects.

Control loops

Executing chunks, episodes, sessions and latency.

API reference

Every PolicyClient method, argument and return value.

Models

Choosing a model compares them in more detail.

How it works

relay keeps one WebSocket open for the whole run and sends observations as binary MessagePack, so frames go as raw bytes with no base64 or JSON overhead. On connect, it moves to a direct route to the GPU container when the server offers one. This skips Modal’s web proxy and saves about 14 ms per call.
PolicyClient has the same methods as OpenPI’s WebsocketClientPolicy: infer, reset and get_server_metadata. Code written for OpenPI usually only needs a different import.