API Key Rotation Best Practices for LLMs
Why API key rotation matters for LLM integrations
If you're building on top of an LLM provider, your API key is the single credential that controls spend, data access, and uptime for your entire integration. A leaked key can run up thousands of euros in usage before anyone notices, and a poorly planned rotation can take down production. The question isn't whether to rotate keys — it's how to do it on a schedule, without downtime, and with a process that survives a real incident.
This guide covers the operational best practices: how often to rotate, how to avoid downtime during rotation, where to store keys, how to scope them, and what to do when a key actually leaks. It applies whether you're calling a model API directly or through a gateway layer.
Rotate on a schedule, not just after an incident
Most teams only think about key rotation after something goes wrong. That's backwards. Treat rotation like TLS certificate renewal: scheduled, automated, and boring.
- Set a fixed interval. 90 days is a reasonable default for production keys; 30 days for anything with broad scope or high spend limits.
- Rotate immediately on role changes. Someone leaves the team, changes projects, or loses device access — rotate any key they had visibility into, not just "their" key.
- Rotate after any exposure, even suspected. A key pasted into a public GitHub issue, a Slack channel with external guests, or a CI log that got shared — rotate first, investigate after.
- Track key age. If your dashboard or secrets manager doesn't show creation date, log it yourself. A key with unknown age is a liability.
Scheduled rotation also forces you to build the automation you'll need in an emergency. If rotating a key is a manual, two-hour process, you won't do it until you're forced to — and during an actual leak, two hours is too slow.
Design for zero-downtime rotation
The biggest reason teams delay rotation is fear of breaking production. Avoid that by designing for overlap from day one.
Use two active keys per environment where the provider allows it. Issue a new key, deploy it alongside the old one, confirm traffic is flowing on the new key, then revoke the old one. Never revoke-then-create — that's a guaranteed outage window.
Decouple the key from the deploy. If rotating a key requires a full redeploy of your application, you've coupled a secret to your release pipeline unnecessarily. Store keys in environment variables or a secrets manager that your app reads at startup or runtime, not hardcoded in build artifacts.
Health-check before cutover. Before revoking the old key, run a synthetic request against the new one and confirm a 200 response and expected payload shape. A simple script works:
curl -s -o /dev/null -w "%{http_code}" https://api.yourprovider.com/v1/messages \
-H "Authorization: Bearer $NEW_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"your-model","messages":[{"role":"user","content":"ping"}]}'
If that returns anything other than 200, don't revoke the old key.
Stagger multi-service rotation. If five services share one key, rotate one service at a time and verify each before moving to the next. Rotating all five simultaneously multiplies your blast radius if something is wrong.
Scope keys tightly, not broadly
Rotation frequency matters less if a single leaked key has unlimited access. Reduce blast radius upfront:
- One key per environment. Dev, staging, and production should never share a key. A leaked dev key shouldn't touch production spend.
- One key per application, where possible. If you run multiple products against the same LLM provider, separate keys mean one compromised app doesn't expose the others.
- Set spend limits per key. If your provider or gateway supports per-key spend caps, use them. A capped key limits damage even before you notice the leak.
- Use descriptive key names/labels. "prod-checkout-service" is far more useful during an incident than "key-4."
This is where a gateway layer helps. SubToAPI issues separate application keys (sub_live_...) per app from a single underlying Claude subscription, so you can scope, label, and revoke individual application keys without touching the others — useful when you want environment- or service-level isolation without managing multiple provider accounts.
Store keys properly — never in code or config files
- No keys in version control. Ever. Use
.gitignorefor.envfiles and add a pre-commit hook or secret scanner (likegitleaksortruffleHog) to catch accidental commits. - Use a secrets manager (AWS Secrets Manager, HashiCorp Vault, Doppler, or even a well-permissioned 1Password vault) rather than plaintext
.envfiles on disk, especially for production. - Restrict read access. Not every engineer needs production key access. CI/CD pipelines should pull secrets at build/deploy time from a vault, not have them checked into pipeline config.
- Log key usage, not key values. Your logs should show which key ID made a request, never the key itself. Redact aggressively — a key visible in a log aggregator is functionally a leaked key.
Automate what you can
Manual rotation is error-prone and gets skipped under deadline pressure. At minimum, automate:
- Expiry alerts — a scheduled job that flags keys approaching your rotation interval.
- Rotation scripts — a single command that generates a new key, updates your secrets manager, and marks the old key for revocation after a grace period.
- Revocation on offboarding — tie key revocation into your team's offboarding checklist, not a separate manual step someone has to remember.
If you're centralizing LLM access for a team, look for a dashboard that exposes key creation and revocation as first-class actions rather than a support ticket. Check /docs for details on how key management works if you're evaluating a gateway approach.
Incident response: what to do when a key leaks
- Revoke immediately. Don't wait to investigate first — a leaked key actively in use is losing money or data every second it's valid.
- Issue a replacement and redeploy using the zero-downtime process above.
- Audit usage logs for the exposure window to check for abnormal request volume or unfamiliar origins.
- Identify the exposure source — commit history, CI logs, a shared document — and fix the root cause, not just the symptom.
- Document it. A short postmortem (what leaked, how, how fast it was caught and fixed) makes the next incident faster to resolve.
Questions
How often should I rotate LLM API keys? 90 days is a reasonable default for production keys, shorter (30 days) for high-spend or broadly-scoped keys. Rotate immediately on any suspected exposure or team member offboarding, regardless of schedule.
Can I rotate an API key without downtime? Yes — issue the new key alongside the old one, verify it works with a health-check request, update your app or secrets manager to use it, confirm traffic has shifted, then revoke the old key. Never revoke before confirming the replacement works.
Where should I store LLM API keys? In a dedicated secrets manager or environment variables injected at deploy time — never in source control or plaintext config files. Restrict read access and scan repositories for accidental commits with a secret-scanning tool.