Recovery and fallback
How tasks survive usage limits, exhausted context, crashes, hangs and lost workers — and how to configure the fallback chain.
Outcomes and responses
| Outcome | Response |
|---|---|
| Rate or capacity limit | Record the limit (reset time only if reported), then follow the fallback chain. A limit never fails a task by itself. |
| Context exhausted | Checkpoint, then a fresh session from the checkpoint (up to maxContextResets, default 10) |
| Crash | Restart with backoff, resuming the session if the agent supports it. Three identical failures in a row, or maxRestarts (default 5), → RECOVERY_REQUIRED |
| Hang | Suspected only after no output, no file changes and near-zero CPU for hangTimeoutMs (default 15 minutes); evidence is recorded before the session is stopped |
| Worker lost | Lease expiry → requeue from the checkpoint, or RECOVERY_REQUIRED (onWorkerLost) |
| Question | WAITING_FOR_INPUT until someone answers |
| Not authenticated / not installed | Try another compatible target, else RECOVERY_REQUIRED with instructions |
The fallback chain
fallback.chain is a list of steps tried in order when a limit is hit:
| Step | Effect |
|---|---|
WAIT | Wait for the reset (maxWaitMs optional). Without a reset time, poll from limitPollBaseMs (1 minute) growing to limitPollMaxMs (30 minutes) |
FALLBACK_AGENT | Switch to another agent (agentId optional) |
FALLBACK_PROVIDER | Switch provider (providerId optional) |
FALLBACK_MODEL | Switch model (modelId optional) |
ASK_USER | Wait for a person to decide |
FAIL | Fail the task |
The default chain is [{ "kind": "WAIT" }].
Policy
{
"fallback": {
"chain": [
{ "kind": "FALLBACK_PROVIDER" },
{ "kind": "WAIT", "maxWaitMs": 3600000 },
{ "kind": "ASK_USER" }
]
}
}Add-on models on the worker have their own switch: Keep tasks running when an agent hits its limit.
Network failures
If the control plane is unreachable, running tasks continue; events are buffered on disk and state changes are retried with the same idempotency ids. All recovery paths are covered by end-to-end and chaos tests.