OPENCLAW PLAYBOOK
CTRL+K
INITIATE_PROTOCOL
← Back to Blog

OpenClaw Backup and Restore: A Disaster Recovery Playbook

By Mira • September 27, 2026 • 8 min read

The Mac mini running your agent will die. Maybe the SSD wears out, maybe a power event corrupts the filesystem, maybe you spill coffee into the ventilation slots at 11pm while reorganizing cables. When it happens, the OpenClaw binary is the least of your problems. You can reinstall the software in ten minutes. What you cannot reinstall is the state: the memory your agent accumulated, the configuration you tuned over months, the session history that explains why it behaves the way it does. That state is the machine. This is the playbook for not losing it.

What Is Actually Worth Backing Up

People back up the wrong things. They image the whole disk, feel responsible, and later discover that restoring a 400GB image to recover 2GB of agent state takes four hours and a spare machine they don't own. Start from the other direction. Enumerate what a rebuild would need, in order of pain:

  • Agent state databases. Memory stores, conversation history, job tracking. Usually SQLite files under your OpenClaw data directory. Lose these and your agent wakes up with amnesia: polite, useless, asking you to reintroduce yourself.
  • Configuration. Gateway config, channel tokens for Slack or Telegram, cron schedules, model routing rules. Small files, enormous rebuild cost, because half of it lives in your head and the other half was set once and never written down.
  • Custom skills and scripts. Anything you authored. This should already be in git, which means the backup for this category is called git push. If your skills live only on the host, that is the first thing to fix, tonight.
  • Credentials and API keys. A special case, handled below, because the obvious approach creates a security problem worse than the data loss you were preventing.

Notice what is missing: the application itself, node_modules, model weights you can re-download, OS packages. Reproducible things. Your backup should be small enough that you stop resenting it.

Copying SQLite Without Corrupting It

Here is where most agent backup scripts are quietly broken. OpenClaw state databases typically run in WAL mode, which means the data you care about is split across the main database file, a -wal file, and a -shm file. If your script does a plain cp or rsync of the directory while the gateway is live, you can capture the main file from one moment and the WAL from another. The copy looks fine. The file sizes look fine. It restores to a database that fails an integrity check, or worse, passes the check and silently contains a torn write. You find out during a real recovery, which is the worst possible time to learn your backups were decorative.

The fix is one command. Use SQLite's own backup API instead of copying files:

sqlite3 /path/to/state.db ".backup '/backups/state-2026-09-27.db'"

The .backup command takes a consistent snapshot through the database's own locking, safe against concurrent writes, and produces a single self-contained file with the WAL folded in. No need to stop the gateway. No coordination with the write-ahead log. Every state database gets this treatment, every night, and then your file-level tool (rsync, restic, Time Machine, whatever ships things offsite) works on those clean snapshots instead of live files. If you want the deeper model of why this state matters so much, how OpenClaw memory works covers it.

One more discipline: run PRAGMA integrity_check; against the snapshot right after taking it. Backups that fail validation should page you the same night, at backup time, while the source is still healthy enough to re-copy.

Building with OpenClaw?

Get the Starter Kit with annotated config, 5 production skills, and deployment checklist.

Grab the Starter Kit →

Secrets Do Not Belong in the Tarball

The tempting move is to archive .env files and channel tokens along with everything else, encrypt the archive, and call it done. Two problems. First, the encryption key now has to live somewhere, and if it lives on the same host, your backup is a single point of failure wearing a disguise. Second, a bundle containing both your agent's memory and all its credentials is the most attractive file you own. Whoever gets it gets everything, with context.

Keep secrets in a real password manager and store only references in the backup: a manifest listing which keys exist, where each one is used, and when it was last rotated. Rebuilding from that manifest takes twenty minutes of copy-pasting from the vault. Rebuilding without it takes days of guessing which of your seven Slack apps was the production one. Rotation deserves a mention too: a disaster is the ideal moment to rotate every token, since you are touching them all anyway and you cannot be certain how the old host died.

The 3-2-1 Rule, Agent Edition

The classic rule says three copies, two media, one offsite. For an agent host it compresses into something concrete. Copy one: the live state on the host. Copy two: nightly snapshots on an external drive attached to the same machine, which handles the common case (accidental deletion, botched upgrade, filesystem corruption) with a five-minute local restore. Copy three: an encrypted sync to object storage or a second machine on a different power circuit, which handles the house-fire case. A nightly cron job drives all of it. If you have not set up scheduled jobs yet, your first OpenClaw cron job is the prerequisite, and the backup job is an excellent candidate for it.

Retention: keep seven daily snapshots, four weekly, twelve monthly. State databases are small. Hoard them.

The Restore Drill Is the Whole Point

An untested backup is a rumor. Once a quarter, restore to something that is not the production host: a spare mini, a VM, an old laptop. The drill has a checklist and a timer.

  • Install Node and OpenClaw from scratch. Time this; it is your floor for recovery time.
  • Restore the latest state snapshot. Run PRAGMA integrity_check; on it before letting anything touch it.
  • Re-enter credentials from the vault using only your manifest. Every key you cannot find is a manifest bug. Fix the manifest, not your memory.
  • Run one canary job end to end, something harmless like a health check that posts to a test channel. If the agent wakes up, remembers what it should, and completes a real task, the backup is real.

Write down the elapsed time. That number is your actual RTO, and it is always longer than the number you carry in your head. Mine was 3 hours 40 minutes the first time, mostly spent hunting a Telegram bot token that existed nowhere except the dead machine's environment. The manifest exists because of that afternoon. Pair the drill with the detection machinery from the observability playbook: the same work-product assertions that catch narrated success will also tell you whether a restored agent is genuinely functional or just confident.

One strange sentence deserves its own paragraph: practice the failure while succeeding is still boring.

How Much Loss Can You Tolerate

Nightly snapshots mean you can lose up to a day of agent memory. For most single-host setups that is the right trade. Conversation context from this morning is recoverable by apologizing to the chat. Six months of accumulated memory is not recoverable at any price. If your agent runs jobs whose intermediate state would be expensive to redo (long research pipelines, multi-day builds), snapshot before those jobs start as well as on the nightly timer. RPO is a per-job decision wearing a global costume, and the security side of this picture, file permissions on snapshots, who can read the offsite bucket, encryption at rest, follows the same rules as everything else in securing your OpenClaw deployment.

Related Reading

Frequently Asked Questions

Can I just use Time Machine for my OpenClaw backups?

Time Machine is a fine offsite layer for files, but it copies live SQLite databases the same unsafe way cp does. Keep Time Machine running, and also write nightly .backup snapshots into a directory Time Machine covers. Then your hourly backups are capturing already-consistent files, which is the combination that actually restores cleanly.

Do I need to stop the OpenClaw gateway to back up state?

No. SQLite's .backup command coordinates with the database's own locking and produces a consistent snapshot while the gateway keeps writing. Stopping the gateway for backups creates a nightly gap in every channel your agent serves, and it trains you to skip backups on busy nights, which are exactly the nights with the most new state to lose.

How long should a full rebuild take?

A practiced rebuild onto warm hardware runs about one to two hours: OS and Node, OpenClaw install, state restore, credential re-entry from your manifest, one canary job. If your first drill takes four hours, that is normal. The gap between four hours and ninety minutes is almost entirely missing documentation, and the drill is how you find it.

What is the single most common backup mistake with agent hosts?

Backing up credentials inside the same encrypted archive as the data, with the decryption key stored on the host being backed up. It feels thorough and protects against almost nothing. Secrets go in a password manager, the backup carries only a manifest of references, and the restore drill proves the two halves reconnect.

Get the free OpenClaw quickstart checklist

Zero to running agent in under an hour. No fluff.