OAuth Refresh Token Rotation in Worker Architectures
Refresh token rotation improves replay detection, but it changes a refresh token from a reusable credential into shared, replaceable state. In a serverless Worker architecture, concurrent requests may run in different isolates with no shared memory. If two jobs refresh the same grant, one can invalidate the token the other still expects. Correctness depends on coordination and durable storage, not on adding a retry around the token endpoint.
Published September 28, 202614 min readOAuth lifecycle and concurrency
Serialize refreshes per authorization grant
One refresh owner persists the replacement before releasing waiting callers
3 / RefreshSend current token once to authorization server.
4 / PersistAtomically save access and replacement refresh token.
5 / ReleaseWaiting callers read the new credential version.
6 / RecoverClassify timeout, reuse and revoked-grant outcomes.
Keep the token family as durable state
Persist one record per user grant or connected account: encrypted refresh token, access token expiry, granted scopes, provider, token generation/version, refresh status and last successful refresh. Do not key state only by a transient Worker instance. Protect encryption keys separately, restrict reads to the refresh path and never log token values.
Use a lock, lease or single-flight coordinator keyed by the grant. A second request should wait briefly, then reread the record and use the newer access token rather than submitting the old refresh token again. Include a version check when writing so a slow response cannot overwrite a newer token generation.
Make the exchange and persistence boundary explicit
The authorization server’s token endpoint and your database cannot generally participate in one transaction. A failure can occur after the provider invalidates the old token but before your store saves the replacement. If the HTTP response is lost, blindly retrying the old refresh token may look like token reuse and revoke the active token family under replay detection.
Outcome
Safe response
Avoid
Refresh succeeds and write commits
Publish new token version to waiters
Returning stale cached token
Another worker already refreshed
Reread version and use winner
Refreshing the previous token again
Timeout before response is known
Mark outcome uncertain; follow provider recovery policy
Automatic unbounded retry of old token
Invalid grant or reuse signal
Stop refresh, revoke local session and require reauthorization
Looping retries or silently restoring old token
Some providers document grace periods, idempotency behavior or a recovery endpoint; others do not. Implement the provider’s documented contract, set a bounded refresh timeout and represent “outcome uncertain” explicitly. Do not invent a universal retry strategy for an irreversible credential exchange.
Use Worker coordination primitives deliberately
Stateless Workers are useful request routers but cannot provide cross-isolate mutual exclusion with process memory. Use a durable transactional store or a single coordinator keyed by the grant. Cloudflare Durable Objects provide per-object coordination and strongly consistent storage, but requests can be processed out of arrival order; still use explicit versions and state transitions. A relational database with atomic compare-and-swap or row locking can also work if connection and transaction semantics fit your runtime.
Preserve security boundaries
Encrypt refresh tokens at rest, minimize scopes and bind storage to the correct user/tenant. Restrict the callback and refresh endpoints, validate OAuth state/issuer as appropriate and protect browser sessions independently. RFC 9700 recommends sender-constrained refresh tokens or rotation for public clients; token reuse can be a compromise signal, so handle it as a security event rather than a transient server error.
Test the unhappy paths
Simulate two simultaneous refresh requests, worker restart during persistence, provider timeout after accepting the token, database failure, expired access token and replay detection. Verify only one exchange happens, the newest token survives, logs contain no secrets and reauthorization restores the connection cleanly. Also test logout and provider revocation so local deletion does not falsely imply remote revocation.
In summary
Rotation makes refresh-token storage a concurrency problem. Coordinate one refresh per grant, persist replacements atomically with version checks, distinguish known failure from uncertain outcome and treat reuse signals seriously. Stateless compute still needs stateful coordination for credentials whose previous version becomes invalid on use.