Health Check¶
A workspace can have a health check — a shell command that Klangk runs inside the container at regular intervals to tell whether the service running there is actually healthy. Without one, a running container only proves the container is alive; it says nothing about the process inside it.
Health checks are most useful for service workspaces that combine two other features:
- an Auto-start workspace whose container boots on server startup (before any user connects), and
- a Service Command that launches a long-running process (a dev server, AI gateway, daemon).
In that combination the container is up and the process is launched unattended — the health check is what turns "the container is running" into "the service inside it is actually responding." For an interactive workspace where you're already in the terminal watching the output, a health check adds little; for an auto-started service no one is watching, it's the difference between available and known-good.
Less essential, but still handy: a health check lets a shared
workspace's other members see at a glance (via the status icon and
GET /api/v1/workspaces/{id}/status) whether the service is up,
without opening a terminal.
How it works¶
Klangk polls container health from outside the container using
podman exec. There is no agent baked into the container image and no
extra connection back to the server — the container stays a clean
sandbox that knows nothing about Klangk internals.
For each workspace with a health check configured:
- Every 30 seconds (configurable), Klangk runs your health check
command inside the container as the creating user, with that user's
HOMEset. - Exit code 0 = healthy. Any non-zero exit code, a timeout, or an error counts as unhealthy.
- When the status changes, every connected client gets a
service_healthevent so the UI can update in real time. Because the stream is deltas-only (it fires on transitions, not every poll), a client that connects to an already-unhealthy workspace also receives a one-time snapshot of every health-checked workspace's current status immediately on connect -- so a pure-WS consumer likeklangkc monitorsees steady-state failures right away instead of being blind until the next transition (#1175). Eachservice_healthframe also carries three additive fields:running(falseon the terminal container-death frame, so a consumer watching only this stream learns the service is down instead of seeing "healthy, then silence"),health_checked_at(when the last poll ran), andseq(a per-workspace counter to detect a missed transition across a reconnect). - The current status, the reason it's unhealthy (a tail of the
check's stderr/stdout), and the time of the last check are exposed
via
GET /api/v1/workspaces/{id}/status.
Startup grace period¶
A service that was just launched needs time to boot. To avoid the
very first poll false-flagging a freshly-started service as unhealthy
(e.g. "Gateway not yet ready to accept connections" while an AI gateway
is still coming up), Klangk applies a startup grace window —
modelled on Docker's HEALTHCHECK --start-period:
- For
KLANGK_HEALTH_CHECK_STARTUP_GRACEseconds (default 30) after the service command fires, a failing check does not transition the workspace to unhealthy, broadcast, or log a failure — the blip is treated as expected boot-up noise. - A passing check is still recorded immediately, so a fast-booting service shows up healthy the moment it actually responds rather than after the whole window.
The window is anchored to when the service command actually launches (so it covers the real boot), and falls back to container start for health-checked workspaces that have no service command. After the window elapses, failures are reported normally.
A failing check is not a black box: the tail of its output (stderr
preferred) is logged when the status flips to unhealthy, retained on
the workspace as health_message, and carried on the service_health
event -- so you can see why it's unhealthy without podman exec'ing
in by hand.
The check runs through bash -c "<your command>" — a bash
non-login shell. It sources no startup file: not ~/.profile,
not ~/.bashrc, not /etc/profile.d/*. This is deliberate — the
health check is an operational probe, not a user session (think
Kubernetes liveness probe), so it must stay deterministic and
decoupled from the owning user's interactive shell setup. A slow
nvm load, a broken ~/.profile edit, or a stray read prompt must
never make an unattended 30-second poll flap "unhealthy".
The trade-off: the check command cannot rely on the user's PATH or
environment. It sees only the container's image PATH (so
/opt/klangk/bin and system tools like grep, pgrep, curl
resolve) plus HOME. So:
- Use absolute paths for the binary you're checking — a tool
installed by
setup.sh(under nvm, a venv, a sandbox mount) is on the user's PATH, not the probe's. - For anything non-trivial, point
health-checkat an executable script with a shebang (e.g./openclaw/bin/healthcheck.sh). The script's shebang dictates its interpreter (a script with#!/usr/bin/env bashruns under bash regardless of the probe's shell), you canexportwhateverPATH/env vars it needs, and you can run it locally by hand to test. Both bundled sandboxes ship their check this way (see below). Inline one-liners (pgrep -x …,curl -sf …,test -S …) are fine for trivial checks and run under the probe'sbash.
See The Shell — Startup files for why login-shell exports don't reach the probe.
When checks are skipped¶
Health checks are deliberately conservative:
- No health check configured — nothing is polled; status stays
null(unknown). - Setup not finished — checks are skipped until the workspace's
setup_stateiscomplete. Polling during setup would report false negatives (the service isn't running yet becausesetup.shhasn't installed it). - Container stopped/removed — its health state goes away with the
container. The server first emits one terminal
service_healthframe withrunning=false(andhealthy=false) so a consumer watching only that stream learns the service is down; after that the workspace no longer appears on the stream until it restarts.
Informational only¶
Health status is informational only. Klangk does not automatically restart an unhealthy container. Auto-restart may build on this later.
Seeing health status¶
Health is surfaced in several places so a failing service is hard to miss:
- Web UI — the workspace list shows a health-colored icon (green = healthy, amber = unhealthy, grey = stopped), updated live as checks run. The Settings tab shows the configured check command.
GET /api/v1/workspaces/{id}/status— returnshealth, the failurehealth_message(a bounded tail of the check's stderr/stdout, ornullwhen healthy), plus the time of the last check (health_checked_at).- Server logs — on each transition to unhealthy, the check's output is
logged at
INFO(steady-state unhealthy polls log it atDEBUG). klangkc monitor— stream events and optionally run a command when something changes. This is the automation hook.
Reacting to failures automatically with klangkc monitor¶
klangkc monitor connects to the server and receives the same events
the web UI does — health transitions, container starts/stops, workspace
changes — and can run a command for each one. The event JSON is piped
to the command's stdin, and details are exposed as environment
variables. For service_health events those are:
KLANGK_HEALTHY—true/false.KLANGK_RUNNING—truefor live-container frames,falseon the terminal container-death frame. This lets a command tell "the health check failed" from "the container stopped" without also subscribing tocontainer_status.KLANGK_HEALTH_MESSAGE— the failure reason, when unhealthy.KLANGK_HEALTH_CHECKED_AT— ISO-8601 timestamp of the last poll.KLANGK_HEALTH_SEQ— per-workspace monotonic counter.
Fire a desktop notification when a service goes unhealthy:
klangkc monitor --type service_health -- \
sh -c '[ "$KLANGK_HEALTHY" = false ] && notify-send "klangk" "$KLANGK_HEALTH_MESSAGE"'
Page yourself (or a Slack webhook) on any health change:
klangkc monitor --type service_health --workspace $WS_ID -- \
sh -c 'curl -s -d "klangk health: $KLANGK_HEALTHY" https://hooks.example.com/alerts'
Just watch the stream (pipe to jq):
Because the stream is deltas-only, silence normally means "nothing
changed" — but it's indistinguishable from "the server's health loop
stalled." For a liveness signal, opt into a periodic heartbeat by
sending {"cmd": "subscribe_health_heartbeat", "enabled": true} on
the socket: the server then emits a service_health_heartbeat frame
at the end of every health-loop tick, so if the loop stalls the
heartbeats stop. It's its own event type, so --type service_health
filters it out — drop the filter to observe it.
monitor reconnects automatically — by default forever, with capped
exponential backoff — and refreshes its login token on auth failures,
so it survives server restarts and token expiry as a long-running
daemon. Bound it with --max-reconnects N, or disable reconnect with
--no-reconnect. See klangkc monitor --help.
Setting the health check¶
Web UI¶
Set the health check when creating a workspace, or change it later in the workspace Settings tab.
CLI¶
# Set during creation
klangkc create my-service --health-check 'curl -sf http://localhost:8080/health'
# Change it later
klangkc edit my-service --health-check 'pgrep -f "openclaw gateway"'
# Clear it
klangkc edit my-service --health-check ''
Sandbox config¶
In .klangk-sandbox.yaml (see Sandbox):
This is exactly what the openclaw sandbox
ships. Rather than a bare openclaw health (which would need the
user's nvm PATH + OPENCLAW_HOME from ~/.profile — invisible to
the non-login probe), setup.sh writes /openclaw/bin/healthcheck.sh:
a tiny script with #!/usr/bin/env bash that exports OPENCLAW_HOME
and execs /openclaw/bin/openclaw health by absolute path. The
config points health-check at that absolute path. The hermes
sandbox
does the same. This is the recommended pattern for any non-trivial
check. See The Shell for why the probe
can't see the user's ~/.profile.
Example commands¶
A health check is any command that exits 0 when things are good.
Because the check runs as a non-login bash -c (it sources
nothing — see above), the reliable patterns are either a trivial
inline one-liner using only system tools, or a wrapper script at
an absolute path for anything that needs a sandbox-installed binary
or custom env.
A trivial one-liner (system tools resolve on the image PATH):
# HTTP health endpoint
health-check: curl -sf http://localhost:8080/health
# Process is running
health-check: pgrep -f 'openclaw gateway'
# A unix socket exists
health-check: test -S /tmp/my.sock
# A port is accepting connections
health-check: nc -z localhost 5432
A non-trivial check — point at a script your setup.sh writes, so the
binary path and any env (OPENCLAW_HOME, HERMES_HOME, …) are baked
in and never depend on the user's ~/.profile:
# /openclaw/bin/healthcheck.sh — what the openclaw sandbox ships
#!/usr/bin/env bash
export OPENCLAW_HOME=/openclaw
exec /openclaw/bin/openclaw health
The script's shebang dictates its interpreter, so you can test it
locally (/openclaw/bin/healthcheck.sh) and it behaves identically
when the probe runs it. Avoid inline commands that depend on the
user's PATH (e.g. a bare openclaw health) — they'll work in a
login shell but report perpetually unhealthy under the non-login
probe.
Server tuning¶
The polling interval, the per-check timeout, and the startup grace period are configurable via environment variables on the server:
| Variable | Default | Description |
|---|---|---|
KLANGK_HEALTH_CHECK_INTERVAL |
30 |
Seconds between polls for each workspace. |
KLANGK_HEALTH_CHECK_TIMEOUT |
10 |
Seconds before a single podman exec check is killed and marked unhealthy. |
KLANGK_HEALTH_CHECK_STARTUP_GRACE |
30 |
Seconds after the service launches during which a failing check is not unhealthy (see below). |
podman exec is a local Unix socket call to the Podman API, not a
network round-trip, so running a check every 30 seconds per container
is negligible overhead.
Related¶
- Service Command — the command that usually runs the service being health-checked.
- Auto-start — start service workspaces on server boot.
- Sandbox — the
workspace.health-checkconfig field.