The OpenClaw Credential Rotation and Expiry Playbook
For about three weeks, one of my reporting jobs ran every morning and produced nothing. The Search Console credential it used had stopped working. Somebody had cleaned up an old Google Cloud project, the OAuth client inside it went with it, and every call came back with deleted_client. The job caught the error, logged it at debug level, and reported a quiet day. Nobody was looking at debug.
Credentials are the part of an agent host that changes without you. You can pin the OpenClaw version and freeze your skills in git, and a key will still die on a schedule someone else set. This page is how I keep track of them now. Most of it comes down to a short file that lists every credential, plus a morning script that tries each one before any real work depends on it.
Expiry Is a Calendar Someone Else Keeps
Every provider has its own rules, and almost none of them tell your agent. A Telegram bot token never expires, but the moment anyone runs /revoke in BotFather the old one is dead and the new one has to reach the host by hand. Slack bot tokens last forever too, until somebody switches on token rotation for the app, after which access tokens expire every twelve hours and your agent had better know how to use a refresh token.
Google is where most of my pain has come from. If your OAuth app is still in “Testing” status on the consent screen, refresh tokens expire after seven days, which is exactly long enough for a new integration to pass every test and then fail the following week. Changing the Google account password also kills any app passwords you made for outbound mail. Apple does the same thing to app-specific passwords.
GitHub is the polite one. Fine-grained personal access tokens make you choose an expiry date and send an email about a week before it arrives, to an inbox that is usually yours and not the agent's.
Model provider keys from Anthropic or OpenAI generally do not expire at all. They get revoked instead, usually by a billing change or a teammate tidying up the console.
Keep a Ledger, Not a Vault Dump
The backup playbook says secrets belong in a password manager, with only references stored alongside your agent's state. That list of references is the ledger, and it does more work than its size suggests. Mine is a YAML file in the same git repo as my skills. It holds no secret values, only facts about them:
- name: slack-bot-ops
env: SLACK_BOT_TOKEN
used_by: [morning-brief, incident-notify]
owner: jascha
issued: 2026-03-14
expires: never
rotates_on: [person-leaves, token-logged, host-lost]
probe: slack_auth_test
- name: github-pat-skills
env: GH_TOKEN
used_by: [skill-publish]
owner: jascha
issued: 2026-06-02
expires: 2026-12-02
probe: github_userThe field people skip is used_by. When a key dies, the first question is which jobs just broke, and when you rotate, the question is which processes need the new value. Without that list you grep the whole home directory at midnight and still miss the cron job that reads its token from a different .env file. The owner field matters too. It names the human who can log into the provider console, which is not always the person running the host.
Probe Every Key Before the Work Starts
Waiting for a real job to fail is how I lost three weeks. A probe is a cheap, read-only call that proves a credential works, and I run all of them at 5:30, half an hour before the first scheduled job. Most providers have an endpoint built for this:
# Slack: returns "ok": true for a live token
curl -s -H "Authorization: Bearer $SLACK_BOT_TOKEN" https://slack.com/api/auth.test
# Telegram: returns the bot's own profile
curl -s "https://api.telegram.org/bot$TELEGRAM_BOT_TOKEN/getMe"
# GitHub: the response header carries the expiry date
curl -sI -H "Authorization: Bearer $GH_TOKEN" https://api.github.com/user \
| grep -i github-authentication-token-expiration
# Anthropic: listing models costs nothing
curl -s https://api.anthropic.com/v1/models \
-H "x-api-key: $ANTHROPIC_API_KEY" -H "anthropic-version: 2023-06-01"Wrap those in a script that reads the ledger, runs each probe, and writes one line per credential with a pass or fail and the HTTP status. Any failure goes to a channel a person reads, through a token that is not the one that just failed. That last part sounds obvious and I got it wrong anyway: my first version sent alerts over the same Slack bot it was testing, so a dead Slack token produced a perfectly silent morning. I send credential alerts over Telegram now, and Telegram's own failures go to email. Schedule the script like any other cron job.
For anything with a known expiry date, have the probe warn at thirty days and again at seven. Thirty days is enough time to rotate on a weekday instead of a Sunday.
Building with OpenClaw?
Get the Starter Kit with annotated config, 5 production skills, and deployment checklist.
Grab the Starter Kit →The Opposite Failure
An agent that remembers a credential problem can keep reporting it long after it is fixed. I had one append “analytics auth is broken” to its daily summary for days after the service account had been repaired and verified, because the warning lived in its context and no check ever contradicted it. The operator corrected it three times. A false alarm repeated daily is worse than no alarm, since it teaches the reader to skim past the very line that will someday be true.
My rule: an agent may only say a credential is broken if a probe in the same run said so, and it has to quote the status code. No probe, no warning.
Rotating Without a Gap
The order matters more than the speed. Issue the new credential first, while the old one still works. Put the new value in the password manager, then into whatever .env or secret store the agent reads. Restart the gateway with openclaw gateway stop and openclaw gateway start --daemon, because a running process keeps the old value in memory and will happily go on using it until the old key is revoked out from under it. Run the probe. Run one real job by hand with openclaw cron run. Only after both pass do you revoke the old credential, and then you update issued and expires in the ledger and commit.
Telegram does not allow overlap, so for bot tokens you just do it fast.
The same sequence covers Slack when you change scopes, since adding a scope means reinstalling the app and the reinstall issues a fresh token. A token with the wrong scopes passes auth.test and then fails on the first chat.postMessage, which is why the hand-run job is part of the sequence and not optional.
When to Rotate
The standard advice is every ninety days. For a single-operator agent host I think a fixed calendar on keys that never expire mostly produces outages you scheduled for yourself, and I rotate on events instead: a person with access leaves, a token shows up in a log or a session transcript, a host is lost or rebuilt, or the agent did something during an incident that you cannot fully account for. Agents paste things into logs far more often than people do, so the “token appeared somewhere it should not” trigger fires more than you would expect. Search your session transcripts for the first few characters of each key (xoxb-, sk-ant-, and the ghp and github_pat prefixes for GitHub) once a month.
There is one exception I keep on a calendar. Any credential that can move money or send messages to people outside your organization gets rotated quarterly regardless, because the cost of a leak there is measured in apologies and the cost of a rotation is fifteen minutes.
The security best practices guide covers how to scope those keys narrowly in the first place, which makes every rotation smaller.
Related Reading
- OpenClaw Agent Observability: Detecting Silent Failures
- OpenClaw Backup and Restore: A Disaster Recovery Playbook
- The OpenClaw Agent Incident Response Playbook
- OpenClaw Security Best Practices
Frequently Asked Questions
Why did my OpenClaw Google integration stop working after a week?
If the OAuth consent screen for your Google Cloud project is still set to “Testing,” Google expires refresh tokens after seven days. Publish the app (internal apps in a Workspace organization are the easy path), then reauthorize once to get a long-lived refresh token.
Do I need to restart the gateway after changing an API key?
Yes. A running gateway reads environment values when it starts, so it keeps using the old key until it restarts. Restart, probe the new key, and only then revoke the old one.
Where should OpenClaw credential alerts go?
Through a different channel than the one being tested. If your probe script alerts over the Slack bot whose token just died, the alert dies with it. Pair channels so each one can report the other's failure.
How often should I rotate OpenClaw API keys?
Rotate on events: someone with access leaves, a key appears in a log or transcript, or a host is lost. Keep a quarterly calendar only for credentials that can spend money or message people outside your organization, and rotate anything with a provider expiry date before it arrives.
Get the free OpenClaw quickstart checklist
Zero to running agent in under an hour. No fluff.