The OpenClaw Configuration Drift Playbook
On my host there used to be a markdown file that said which agents existed and which model each one ran. Every agent read it at the start of a session. A person updated it by hand, when they remembered, and for weeks it listed an agent that had been retired as active and had me on the wrong model. Nothing errored. The agents simply believed it.
That is configuration drift. Agent hosts have a second place for it that most infrastructure guides never mention: the text the agents read about themselves. This page covers both, for a single always-on machine running the OpenClaw gateway, a set of scheduled jobs, and some skills.
Where Drift Lives on an OpenClaw Host
Start with ~/.openclaw/openclaw.json. It holds your channels, your model keys (or the ${VAR} references to them), tool settings and usually your cron section. People edit it live, over SSH, at night, to fix one thing. The edit works, nobody commits it, and the copy in git is now a historical document.
The running gateway is a second copy. It read the config and the environment when it started, so a change on disk does nothing until a restart, and a restart you forgot about means the process is running a config that exists nowhere on disk anymore. The credential rotation playbook leans on this same fact from the other direction.
Then the service manager. On a Mac mini that means launchd plists in ~/Library/LaunchAgents, which point at wrapper scripts, which sometimes hardcode a flag or a model name that contradicts what the config says. I have found a model name pinned in two places inside a runner script while the status file claimed an upgrade had happened months earlier. Both were wrong in different ways.
Skills drift less, because most people keep them in git already. The risk there is a skill enabled on the host with openclaw skills enable that was never added to the list you think is canonical.
The last place is prose, and it gets its own section below.
Pick One Source and Generate the Rest
Every serious guide on drift says the same first thing: decide what the source of truth is. I agree, with one addition. Anything that a human or an agent reads to learn the state of the system should be generated from that source, never written by hand next to it.
My fix for the stale agent list was to stop editing it. The facts now live in a small database where every row has to carry evidence (the config file and key it came from, and the date somebody checked). A script renders the markdown from that, and the first lines of the file say so:
# Fleet State (GENERATED, do not hand-edit)
<!-- Rendered by render_fleet_state.py from verified rows only.
Change the source, then re-render. -->You do not need a database. A YAML file in the same repo as your skills is plenty, as long as the readable version is produced from it by a script and the script runs on a schedule. What matters is that a hand edit to the rendered file gets overwritten within a day, which teaches everyone (agents included) where changes actually go.
For openclaw.json itself, the source is the copy in git. Commit it with secrets replaced by environment references, as the backup playbook describes, and treat any live edit as a draft until it is committed.
The Nightly Diff
Detection is a snapshot and a comparison. Once a night, during a quiet hour, I capture what the host is actually doing and diff it against what the repo says it should be doing:
#!/bin/sh
# drift-snapshot: record live state, compare to the repo
OUT=~/drift/$(date +%F)
mkdir -p "$OUT"
openclaw cron list > "$OUT/cron.txt"
openclaw skills list > "$OUT/skills.txt"
openclaw gateway status > "$OUT/gateway.txt"
launchctl list | grep -v com.apple > "$OUT/launchd.txt"
shasum -a 256 ~/.openclaw/openclaw.json > "$OUT/config.sha"
cd ~/agent-config && git diff --no-index --stat \
expected/ "$OUT/" > "$OUT/DIFF" || trueThe expected/ folder holds the same captures from the last time someone approved the state. When the diff is empty, nothing gets sent. When it is not, the file goes to a channel a person reads, with the changed lines inline and no summary on top. Summaries are where an agent decides a difference is probably fine.
Keep the config hash in there even though it tells you nothing about what changed. A changed hash with an unchanged git log means somebody edited the file on the host, and that is the drift I run into most.
Add openclaw doctor to the same run if you can. It will not catch drift on its own, but after an upgrade it is often the first thing to notice a config key the new version stopped reading, which the upgrade playbook warns is usually ignored without any error at all.
Building with OpenClaw?
Get the Starter Kit with annotated config, 5 production skills, and deployment checklist.
Grab the Starter Kit →Prose Is Config Too
None of the drift tools built for servers look at this, and on an agent host it is where the expensive mistakes come from. An agent's behavior is set by its instruction files and by whatever status documents it reads at session start, and those go stale in exactly the way a config file does.
The worst case I have seen was an omission. A dispatch script existed that sent hard problems to a stronger model, it was tested and it worked, and the agent it was built for never called it, because that agent's instruction file did not mention it once. From the outside the agent just seemed less capable than it should have been. Every config file was correct.
Two checks catch most of this, and both fit in the nightly script. First, list every script or skill the agent is supposed to be able to use, and grep its instruction file for each name; anything missing is a capability the agent does not know it has. Second, for every factual claim in a status document (a model name, a count, a threshold), require a pointer to where it was verified, and flag any claim whose pointer is older than the last change to the file it points at.
The memory guide explains how agents store what they learn. Memory drifts too, mostly toward stale warnings. The credential playbook has the example of an agent reporting a broken service for days after it was fixed, and the rule from there applies here: an agent may only state a fact about the system if a check in the same run confirmed it.
Which Side Wins
Some drift tools revert anything that does not match the baseline automatically. I think that is wrong for a one-person agent host. The live edit at midnight was usually a real fix for a real problem, made by someone who meant to commit it later. Reverting it brings the problem back, quietly, on a schedule.
My rule is that the repo is intent and the host is evidence. When they disagree, a person looks at the diff and picks one. If the live change was right, commit it and copy the new capture into expected/. If it was wrong, restore the file from git, restart the gateway with openclaw gateway stop and openclaw gateway start --daemon, and run the most important scheduled job by hand with openclaw cron run to confirm it behaves. Either way the diff should be empty the next night.
A thresholds file deserves extra care. When my host's spawn gate moved from the total process count to the per-user count (the story is in the host capacity playbook), the change had to land in the gate script, the incident notes and the instructions the agents read. Update two of those and you have created drift on purpose.
The Short Version
If an agent reads it, generate it.
Related Reading
- OpenClaw Agent Observability: Detecting Silent Failures
- The OpenClaw Upgrade and Rollback Playbook
- The OpenClaw Credential Rotation and Expiry Playbook
- OpenClaw Backup and Restore: A Disaster Recovery Playbook
Frequently Asked Questions
What is configuration drift in OpenClaw?
It is any difference between what your host is actually running and what your repo or notes say it runs. On an OpenClaw host that covers openclaw.json, the running gateway, cron jobs, launchd services, enabled skills, and the instruction and status files your agents read.
Does editing openclaw.json take effect immediately?
No. The gateway reads its config and environment when it starts, so a change on disk waits for a restart. Until then the running process uses the old values, which is its own form of drift.
Should drift be reverted automatically?
For a small agent host, I recommend against it. Live edits are often real fixes that were not committed yet. Alert on the diff, have a person choose which side is correct, then commit or restore.
How do I stop agents from reading stale status files?
Generate those files from a single source with evidence attached to each fact, mark them as generated at the top, and re-render them on a schedule so hand edits are overwritten.
Get the free OpenClaw quickstart checklist
Zero to running agent in under an hour. No fluff.