Providers overview
Own logins by default, add-on models when an agent reaches its limit, how limits are detected, and what the platform can and cannot read from a provider.
Own logins first
By default every task runs on the agent's own login and model choice: Claude Code on your Claude subscription,
Codex on its ChatGPT login, and so on. Nothing needs to be set up for that. Each installed agent appears as a
provider native:<agent> with the model default, preferred unless a policy says otherwise.
Add-on models
Add-on providers are configured on each worker (local UI → AI models). When an agent reaches its usage limit, a task can continue on them — see the guide. Keys never leave the machine.
| Kind | Models | Health | Usage | Notes |
|---|---|---|---|---|
anthropic | /v1/models | ✓ | — | x-api-key, anthropic-version: 2023-06-01 |
openai | /v1/models | ✓ | — | Custom base URL supported |
google | /v1beta/models | ✓ | — | x-goog-api-key |
openrouter | ✓ | ✓ | ✓ | Credit usage and limit; sign-in instead of a key |
azure-openai | ✓ | ✓ | — | Base URL is the resource endpoint; extra.apiVersion |
bedrock | ListFoundationModels | ✓ | — | SigV4; extra.region, optional extra.profile |
vertex | manual list | ✓ | — | Google OAuth; extra.project, extra.region |
ollama | /api/tags | ✓ | — | Local, keyless |
lmstudio | /models | ✓ | — | Local, keyless |
deepseek, groq, nvidia-nim | /models | ✓ | — | OpenAI-compatible presets |
openai-compatible | /models | ✓ | — | vLLM, LiteLLM, Together, Fireworks…; key optional |
Behaviour
- Capabilities, not assumptions. An operation a provider does not support returns "not supported". The platform never invents usage or limit data.
- Rate-limit-safe polling. Health checks run at most every five minutes, back off with jitter on failure, and are skipped while a provider is known to be limited.
- Limits. A 429 or a limit an agent reports marks the provider limited. Without a reset time it stays limited until a later check succeeds; a reset time is never guessed.
- Own login off:
AO_HARNESS_OWN_LOGIN=0in a worker's environment makes it use configured providers only.
Usage tracking
Token counts and cost are recorded per task, agent, provider and model when the agent reports them (Claude Code
reports both), and exposed at GET /api/v1/orgs/:orgId/usage.
Verification status
Each provider was tested against a local HTTP fake that checks auth headers, paths, parsing and 429
Retry-After handling — not against the live APIs.