The OpenClaw Upgrade and Rollback Playbook
Most agent outages I have caused were self-inflicted. A new OpenClaw release lands, the changelog looks harmless, somebody types npm install -g openclaw on a Tuesday evening, and by Wednesday morning the briefing never arrived, two cron jobs are silently skipping, and the Slack channel is full of an agent apologizing in a slightly different tone than before. Upgrades break things. This page is the routine I use so that when one does, the damage is measured in minutes.
Pin the Version Before You Need To
A bare npm install -g openclaw installs whatever is newest at that second. Fine for a laptop. On a host that runs your mornings, it means the version you are running depends on when you last happened to touch the terminal, and you cannot roll back to a number you never wrote down.
Run openclaw --version right now and put the answer in a file next to your config. I keep a one-line VERSION file in the same git repo as my skills. Every install after that names a version explicitly (npm install -g openclaw@<version>), and every successful upgrade ends with a commit that changes that one line. The git log becomes an upgrade history with dates, which is the first thing you want when a behavior change shows up three days late and you are trying to work out whether the release caused it.
Read the Changelog Like a Suspicious Person
Release notes are written by people who are proud of the release. Skim past the features. You are hunting for four kinds of lines, and they are not equally dangerous:
- Config schema changes. Renamed keys, new required fields, changed defaults. These are the big one, because a renamed key often fails quietly: the old value is ignored and the new default applies, so the gateway starts cleanly and simply behaves differently.
- Scheduler or cron changes. Anything touching how jobs are timed, retried, or deduplicated. Read these twice if your agent has overnight work.
- Channel integration updates, like a new Slack or Telegram API version. Usually fine. Occasionally they require re-authorizing an app, which you want to know before midnight.
- Node version requirements, which are rare and loud.
If you skip more than one release, read every changelog in between. Skipping versions is where the renamed-key problem compounds.
The Pre-Flight Snapshot
Before touching the binary, take a fresh state snapshot. Same procedure as the nightly job in the backup and restore playbook: SQLite .backup for every state database, a copy of the config directory, an integrity check on the result. Label it with the version you are leaving, something like pre-upgrade-from-<old-version>, so nobody later mistakes it for a regular nightly.
Why bother when last night's backup exists? Because new releases sometimes migrate the state schema on first start. Once that migration runs, the old binary may refuse to open the database, or open it and misread it. A snapshot taken ten minutes before the upgrade is the only clean copy that matches the version you might need to go back to, and last night's copy is missing everything the agent learned today.
Building with OpenClaw?
Get the Starter Kit with annotated config, 5 production skills, and deployment checklist.
Grab the Starter Kit →Pick the Window Your Agent Is Idle
Look at your cron schedule before choosing a time. The obvious slot, late evening after dinner, is often the worst one, because it sits right before the overnight jobs that you will not be awake to watch fail. I upgrade mid-morning on a weekday, after the briefing has shipped and with at least four hours of me being at a keyboard ahead. Boring timing is a feature.
Then the sequence itself is short. Stop the gateway with openclaw gateway stop. Install the pinned new version. Run openclaw doctor and actually read what it says. Start the gateway again with openclaw gateway start --daemon and confirm with openclaw gateway status.
Canaries, Then Real Work
A running gateway proves almost nothing. The process being up and the agent doing its job are separate facts, and the gap between them is exactly where the observability discipline from detecting silent failures earns its keep. Right after the restart, trigger jobs by hand rather than waiting for the schedule:
- One harmless job that posts to a test channel. This confirms the channel tokens still work and the model routing still resolves.
- One job that reads memory and has to use it, for example asking the agent to summarize what it did yesterday. If the answer is vague or generic, the state migration may have gone sideways, and you want to know that now while the pre-flight snapshot is fresh.
- Your single most important scheduled job, run manually, with its real output checked against a known-good example from last week. For me that is the morning briefing, and I compare section headings line by line because formatting drift is often the first visible symptom of a prompt or tool change inside the release.
Then watch the first full scheduled cycle. Check each cron job actually fired the next morning, and fired once. Duplicate runs after an upgrade are more common than skipped ones and much more annoying for whoever receives the emails.
Rolling Back
Decide the rollback trigger before you start, because at 11am with a broken briefing you will be tempted to debug forward for three hours. My rule: if any canary fails and the cause is not obvious within twenty minutes, roll back and investigate later on a spare machine.
The rollback is the upgrade in reverse. Stop the gateway. Reinstall the old version by its pinned number. Restore the pre-flight snapshot over the state directory, since the new release may already have migrated it. Restore the config copy. Start, check status, rerun the same canaries. Revert the VERSION commit, or better, add a new commit noting why, so that three months from now you remember this release was skipped on purpose.
Rehearse it once.
Anything you learned during the canary window (a message the agent handled, a note it saved) is lost when you restore, which is why the window should be short and the host quiet. That trade is almost always worth it. Losing forty minutes of agent memory costs you an apology in a chat thread, while a half-migrated state database that the old binary misreads can cost you a weekend.
When You Run More Than One Agent
Upgrade one host, or one agent, first and let it run a full day before touching the others. If your agents coordinate (the patterns in multi-agent coordination), check that a new-version agent and an old-version agent can still hand work to each other, because message formats occasionally shift between releases. Mixed versions for a day is normal. Mixed versions for a month means the upgrade stalled and nobody wrote down why.
Related Reading
- OpenClaw Backup and Restore: A Disaster Recovery Playbook
- OpenClaw Agent Observability: Detecting Silent Failures
- OpenClaw Cron Jobs: Automated Workflows Guide
- OpenClaw Security Best Practices: Securing Your AI Agent Deployment
Frequently Asked Questions
How often should I upgrade OpenClaw?
For a production agent host, a monthly cadence works well: frequent enough that you never skip many releases at once, rare enough that each upgrade gets the full pre-flight and canary routine. Security fixes are the exception. When a release patches a vulnerability, upgrade within a day or two, using the same steps on a compressed schedule.
Can I downgrade OpenClaw without restoring state?
Sometimes, but you cannot know in advance. If the newer release migrated the state schema on first start, the older binary may fail to open the database or read it incorrectly. Restoring the pre-flight snapshot along with the old version is the only rollback that is reliably safe, which is the reason to take that snapshot minutes before upgrading.
Should I enable automatic updates on my agent host?
No. Automatic updates remove the two things that make upgrades safe: a snapshot taken right before the change, and a person watching the first canary runs. Keep automatic updates for the operating system's security patches if you like, and upgrade OpenClaw itself deliberately, by pinned version, at a time you picked.
What should I check first if something breaks after an upgrade?
Config keys. Compare your config against the release notes for renamed or newly required fields, since a key the new version no longer recognizes is usually ignored without an error. Then run openclaw doctor, then check that each channel integration can still post. Most post-upgrade breakage sits in one of those three places.
Get the free OpenClaw quickstart checklist
Zero to running agent in under an hour. No fluff.