The OpenClaw Idempotent Jobs Playbook
One of my morning jobs reads a JSON file called deficiency-report.json, looks for problem categories that have not been resolved yet, and writes a skill for each one. A different job generates that file. Every time the generator ran, it wrote the report fresh, and the fields saying which categories were already handled (signal_resolved and skill_authored) disappeared. The morning job then saw five old problems as brand new.
By early September that had happened on more than seventeen consecutive runs.
Nothing crashed, and that is why it went on so long. This page is about making scheduled agent jobs safe to run twice, or twenty times, on a single host running the OpenClaw gateway and a handful of cron jobs.
What the Usual Advice Covers
Read the well-known guides on reliable cron jobs and they agree on a short core. Make the job idempotent, meaning a second run produces the same result as the first. Add a lock so two copies cannot overlap. Decide on purpose whether a missed run gets caught up or skipped. Google's SRE book goes further and says that, given the choice, they would rather skip a launch than risk launching twice, because a missed run is easier to recover from than a doubled one.
All of that holds for agent jobs. I follow it. What those guides assume, though, is that the code doing the work is the code you wrote, and on an agent host it often is not. The worker is a model reading instructions and a state file, then deciding what to do. If the state file lies, the model acts on the lie with complete confidence, and it is a far more capable actor than a SQL upsert. A dumb script that re-runs inserts a duplicate row. An agent that re-runs writes a second skill for a problem it already solved, and gives it a slightly different name so your dedupe check misses it.
Merge, Never Overwrite
The report bug had a one-line cause. The generator built a new object from what it measured and wrote it to disk, which threw away every field it did not personally produce. Two writers shared one file and only one of them knew it.
The fix is read-then-merge: load the current file, change only the keys you own, write the result. With jq it looks like this:
#!/bin/sh
# merge-write: set one key, keep everything else
STATE=~/state/report.json
tmp=$(mktemp "$STATE.XXXXXX")
jq --arg k "$1" --argjson v "$2" '.[$k] = $v' "$STATE" > "$tmp" \
&& mv "$tmp" "$STATE"The temp file sits in the same directory on purpose. A mv within one filesystem is atomic, so a reader sees either the old file or the new one, never half of each. Write straight to the real path and a job that dies mid-write leaves truncated JSON behind, which the next agent will either choke on or, worse, quietly treat as empty.
Better still is not sharing the file at all. My own recommendation for the report was to move the resolution state out of it entirely, into a file only the resolving job writes. Measurements in one place, decisions in another. If two jobs need to know about each other, one reads and one writes.
Watch your agents here specifically. Most agent write tools replace a whole file, so an agent asked to “mark category X resolved” will often rewrite the JSON from its own memory of what was in it. Give it a merge script as a skill and tell it, in its instructions, that the script is the only allowed way to touch that file.
Write the Done Marker Last
This site has a small example of the right order in its own repo. The script that pings IndexNow when pages change keeps a state file of page hashes, and it only writes that file after the API accepts the submission. If the ping fails, the state stays old, and the next run tries again. Nothing is recorded as done until it is done.
Agent jobs get this backwards all the time, usually because the instructions say something like “update the tracker, then post the summary.” If the post fails after the tracker update, the job believes it succeeded. Reverse it. Do the side effect, confirm it took, then record it.
The flip side is that a crash between the side effect and the record will cause one repeat. For most jobs that is the cheaper failure. For jobs that message people it is not, and those need the next section.
Building with OpenClaw?
Get the Starter Kit with annotated config, 5 production skills, and deployment checklist.
Grab the Starter Kit →Key the Run on the Evidence
A date-based marker (digest-2026-10-06.done) stops a job from running twice in one day. It does nothing for the case that actually bites agent hosts: the same input arriving twice on different days, or new input arriving on a day the marker already exists.
Hash what the job is acting on instead. The IndexNow script uses a sha256 of each page's source as its unit of change, which means an unchanged page is never resubmitted and a changed one always is, regardless of when the script last ran. The same idea works for agents. Before a job drafts a reply to an insight, it hashes the insight's evidence and checks for a marker with that hash. Same evidence, skip. Different evidence, run.
KEY=$(cat "$INPUT" | shasum -a 256 | cut -c1-16)
MARK=~/state/runs/$JOB-$KEY.done
[ -e "$MARK" ] && { echo "$JOB: already handled $KEY"; exit 0; }
# ... do the work, confirm it ...
touch "$MARK"On October 5 I rebuilt one of our proactive jobs around exactly this, and alongside it made its drafts atomic (written to a temp path, then moved into place). Then I tested it against five different interruptions and failures. It held. The point of the interruptions is that you do not know a job is idempotent until you have killed it at the worst moment and run it again.
That skip line should print something. A job that exits silently because work was already done looks identical, from the outside, to a job that never ran, a problem the silent failures guide spends a whole section on.
Locking on a Mac Mini
Most cron guides tell you to wrap the job in flock. macOS does not ship it. Type which flock on a stock Mac mini and you get nothing. You do get /usr/bin/shlock, or you can use the old trick of mkdir, which either creates the directory or fails, atomically:
LOCK=/tmp/$JOB.lock
if ! mkdir "$LOCK" 2>/dev/null; then
echo "$JOB: previous run still going, skipping" >&2
exit 0
fi
trap 'rmdir "$LOCK"' EXITThe weakness is a stale lock after a hard kill, since the trap never fires on SIGKILL. Write the PID into the lock directory and have the next run check whether that PID is still alive before giving up. If you run many jobs, the lock also keeps a slow one from piling up copies of itself, which matters on a host where process count is already the constraint (the capacity playbook covers why).
When the Script Is Missing, the Agent Improvises
On October 1 that same morning job ran in an environment where python3 was not available. Its instructions said to run a Python script. The agent could not, so it read the script, understood what it did, and reproduced its logic by hand against the JSON file with its own read and write tools. It reported this honestly, and that run happened to be correct.
I find this both impressive and alarming. Every guarantee in this page lives in the scripts. An agent reimplementing them by hand keeps none of those guarantees unless it happens to be careful. My rule now is that a job whose safety depends on a script should stop and report when the script cannot run, and the instructions say so in plain words. Improvising is fine for a summary. It is not fine for anything that writes state.
A Rerun Test You Can Do Today
Pick your most important scheduled job. Run it by hand with openclaw cron run, then run it again immediately. Diff the state files before and after the second run, and check every channel it posts to.
If the second run changed anything, you have found your next fix.
Then kill it halfway through a run and start it again. Do this before a scheduler retry or a restore from backup does it for you, because a restore rewinds your state files to last night and every job on the host will then try to redo whatever it did since.
Related Reading
- The OpenClaw Configuration Drift Playbook
- OpenClaw Agent Observability: Detecting Silent Failures
- OpenClaw Parallel Task Execution
- OpenClaw Cron Jobs and Automated Workflows
Frequently Asked Questions
What does idempotent mean for an OpenClaw cron job?
Running the job a second time with the same input changes nothing, and in particular sends nothing. It matters more for agent jobs than for ordinary scripts, because an agent that thinks work is undone will redo it creatively rather than identically.
Why do agent jobs lose fields in shared JSON state files?
Usually because one writer builds a fresh object and overwrites the file, discarding keys that another job added. Agent write tools make this easy, since they tend to replace whole files. Merge only the keys you own and write through a temp file plus mv, or give each job its own file.
How do I lock a cron job on macOS without flock?
Use mkdir on a lock directory, which succeeds or fails atomically, and remove it in an exit trap. /usr/bin/shlock also ships with macOS. Store the PID in the lock so a later run can clear a stale lock left by a killed process.
Should the done marker be written before or after the work?
After the work is confirmed. Writing it first turns any failure into a silent skip. Writing it last means a crash can cause one repeat, so jobs with outward side effects like messages should also key their marker on a hash of the input.
Get the free OpenClaw quickstart checklist
Zero to running agent in under an hour. No fluff.