OPENCLAW PLAYBOOK
CTRL+K
INITIATE_PROTOCOL
← Back to Blog

The OpenClaw Model Provider Outage Playbook

By Mira • October 10, 2026 • 7 min read

Your model provider will go down at some point, and the agent on your Mac mini will not wait politely for it to come back. Scheduled jobs keep firing. Each one gets an error, retries, maybe switches to a backup model, and keeps going. By the time you look at your phone the morning briefing has run on a different model than the one you tested it with, and a follow-up email has gone out that you would not have approved.

This page is about planning for that hour before it happens, on one always-on machine running the OpenClaw gateway. Most of it is cheap. The expensive part is deciding, job by job, what you actually want to happen.

What the Fallback Guides Already Cover

The articles that rank for LLM fallback come from gateway vendors and platform teams, and they run long, around 2,200 to 3,200 words each. They agree on the mechanics. Retry first, with backoff, then walk an ordered chain of fallback models. Trip a circuit breaker so you stop hammering a provider that is clearly down. Prefer a backup from a different provider, because a same-provider fallback often fails with the primary. Watch your fallback rate.

All correct. Every one of them also assumes a person, or a customer, is waiting on the other end of the request, so the goal is always the same: get an answer back, from anyone, fast. An agent host has a different problem. Most of its work runs at 3 a.m. with nobody waiting, and some of that work has the authority to send messages or spend money. For that kind of work, a fast answer from a weaker model can be worse than no answer at all.

Decide Per Job: Fall Back or Wait

I sort every scheduled job into one of two buckets, and I write the bucket into the job definition so nobody has to remember it during an outage.

The first bucket can fall back. These jobs read and summarize. A morning briefing, a log digest, a capacity report, the research note that lands in a folder for you to skim later. If the backup model writes a slightly flatter summary, you lose a little quality and nothing else, and you can rerun it once the primary is healthy. The second bucket waits. Anything that sends text to a human outside the host, anything that edits a file other agents depend on, anything that commits code or touches a payment. When the primary is down, these jobs write one line saying they skipped, and they stop.

I would hold the second bucket harder than anything else on this page. The behavior of an agent comes from the model plus its instructions, and the instructions were tuned against one model. Put a different model behind the same AGENTS.md and you are running an untested agent with production credentials. If a waiting job genuinely cannot wait, route its output into the queue described in the human approval gates playbook and let a person look at it first.

Interactive chat is the exception. When you are on Telegram talking to the agent yourself, falling back is fine, because you are the review step. Just make it say so.

Log the Model That Actually Answered

On my host there was once a runner script with a model name pinned in two places while the status file claimed an upgrade had happened months earlier. The configuration drift playbook tells that story. Fallback makes the same confusion happen on purpose, every time a provider blinks, unless you record what ran.

Every run already deserves one line of JSON, as the observability guide recommends. Add two fields to it: the model you asked for and the model that answered. Read the second one from the response or the wrapper that made the call, never from your config.

{"ts":"2026-10-10T03:00:12Z","job":"morning-briefing",
 "model_requested":"primary","model_used":"backup",
 "fallback_reason":"http_503","attempts":4,"done":true}

Once those fields exist, a weekly count of runs where the two differ tells you how often you are really on the backup. If that count is high in a week with no news about an outage, the problem is probably on your side: a quota you outgrew, or a model name the provider retired while your runner script still asks for it. Provider keys from Anthropic or OpenAI generally do not expire, as the credential rotation playbook notes, so a sudden wall of auth errors usually means something else changed. Check billing first.

Building with OpenClaw?

Get the Starter Kit with annotated config, 5 production skills, and deployment checklist.

Grab the Starter Kit →

Retries Are the Outage You Cause Yourself

A single job retrying with backoff is harmless. Twelve scheduled jobs that all fire on the hour, each retrying five times, each spawning a fresh session per attempt, are a process storm on a small machine, and the host capacity playbook explains how that kind of burst ends in processes getting killed. The provider comes back after twenty minutes and your gateway is the thing that is down.

Cap attempts per run, and cap them per host too. A small shared file works: the first job to see a provider error writes a timestamp, and every other job checks that file before it calls the model. If the timestamp is under ten minutes old, waiting jobs skip immediately and fallback jobs go straight to the backup without burning attempts on the primary. That file is your circuit breaker. It takes a few lines of shell and it lives outside the gateway, which matters, because a breaker the agent controls can be talked out of tripping.

When the Provider Comes Back

Recovery is where the second bucket pays you back or bites you. Every waiting job that skipped left a line in the log, and now each of those jobs wants to run, often all at once and often twice: once because you reran it by hand, once because its schedule came around.

This is exactly the case the idempotent jobs playbook exists for. A job keyed on its evidence, with the done marker written last, can be rerun as many times as you like and only acts once. Without that, an outage turns into a duplicate-email problem a day later, and duplicates are harder to apologize for than silence.

Stagger the catch-up. Clear the breaker file by hand, trigger the most important waiting job with openclaw cron run, read its output, then let the schedule handle the rest. And rerun anything from the fallback bucket that a person will actually read, so the record you keep was written by the model you trust.

Rehearse It on a Quiet Afternoon

Point a test agent at a deliberately wrong API key and watch what it does for an hour.

The Short Version

Write fall back or wait into every job. Log the model that answered, from the response. Keep the circuit breaker in a file the agent cannot edit, and make recovery boring with idempotent jobs. If you only do one thing this week, do the bucket labels, because the rest of this page is plumbing and the labels are the decision.

Related Reading

Frequently Asked Questions

Should my OpenClaw agent fall back to another model when the provider is down?

For jobs that only read and summarize, yes. For jobs that send messages, edit shared files, commit code or spend money, have them skip and log instead, because the agent's instructions were tested against the primary model.

How do I know which model actually ran a job?

Record both the requested model and the model that answered in each run's log line, reading the second from the API response or the calling wrapper rather than from your configuration.

Why did my host slow down during a provider outage?

Usually retries. Many scheduled jobs retrying at once, each with a fresh session, can overload a small machine. A shared circuit-breaker file that jobs check before calling the model stops the pile-up.

What should I do when the provider recovers?

Clear the breaker, run the most important skipped job by hand and check it, then let the schedule catch up. Idempotent jobs keep the catch-up from sending anything twice.

Get the free OpenClaw quickstart checklist

Zero to running agent in under an hour. No fluff.