View source on GitHub
GET /server_info reports idle_time, the seconds since the last
activity on the server. This is the exact signal runtime-api polls to
decide when a sandbox is idle enough to pause/reap — this example uses the same
signal for your own “has the workspace gone quiet?” check.
On Cloud (the default here), the same response also carries
runtime_idle_timeout_seconds, the platform’s real reap threshold, so you can
see idle_time climbing toward the very number the platform acts on.
One file:
idle_poll.py— start a Cloud sandbox, attach a conversation (no LLM key needed), then poll/server_infountilidle_timecrosses a threshold and declare the agent idle. Pass--localto run against an agent-server you start in Docker instead.
APIs used
Cloud app server — manages the sandbox lifecycle
- Base URL:
https://app.all-hands.dev, auth headerX-Session-API-Key: <OH_API_KEY>. POST /api/v1/sandboxes— start a sandboxGET /api/v1/sandboxes?id=<id>— poll untilRUNNINGPOST /api/v1/app-conversations— attach a conversation (returns a start task)GET /api/v1/app-conversations/start-tasks?ids=<id>— poll for the idDELETE /api/v1/sandboxes/{id}?sandbox_id=<id>— clean up
Agent server — GET /server_info
Read from the sandbox’s AGENT_SERVER exposed URL with its session_api_key.
Returns a ServerInfo object; the fields this example reads:
On Cloud,
runtime-api reaps a sandbox roughly when
idle_time >= runtime_idle_timeout_seconds. This demo uses a much smaller
threshold (--idle-threshold, default 15s) so you can watch idle detection fire
quickly against the same idle_time signal.
idle_time vs. execution_status
They answer different questions — pick per your need:
Use
idle_time when you just want “nothing is happening anymore” without
subscribing to a conversation; use execution_status when you need an
authoritative terminal signal.
The flow (Cloud)
Run it
Flags
What it prints
Running locally without Cloud
The audience for this example is Cloud. If you have no Cloud account, pass--local to start an agent-server in Docker and poll it directly:
--local mode: the script docker runs the
ghcr.io/openhands/agent-server:latest-python image, creates the conversation
directly on the agent-server (POST /api/conversations, which needs an LLM key),
and reads /server_info at http://localhost:8000. Note that
runtime_idle_timeout_seconds is null locally — there is no platform
reaper — so only the idle_time heartbeat is meaningful. Local-only flags:
--llm-api-key, --llm-model, --llm-base-url, --session-key, --image,
--server-port, --container-name.
Notes
- Coarse by design.
idle_timecannot tell you why things went quiet (finished vs. errored vs. stuck vs. simply waiting). It is a heartbeat, not a state machine. That is exactly why the platform uses it for reaping and not for reporting completion. - Threshold choice. Set
--idle-thresholdwell above your longest expected gap between agent actions (LLM latency, long tool calls), or you will declare “idle” mid-run. The platform’s default (runtime_idle_timeout_seconds, ~1200s on Cloud) is deliberately large for this reason. - Full agent-server schema:
<agent-url>/openapi.json.
Related
watch-terminal-state— authoritative per-conversation terminal state over the WebSocket (push)react-to-state-websocket— react to everyexecution_statustransition over the WebSocketstart-sandbox— the sandbox lifecycle this example builds on

