A cron expression looks tiny because it describes only when to wake a process. Production cost lives in everything around that trigger: duplicated delivery, overlapping runs, missed windows, timezone interpretation, retries, partial writes and the question nobody asks until an incident: how do we know the job finished correctly?
Published October 7, 202613 min readScheduling, recovery and ownership
A schedule is a request to attempt work
Reliable scheduled work needs control points beyond the clock
TriggerSchedule and timezone.
ClaimLease or unique run key.
ExecuteBounded, checkpointed work.
RecordOutcome and progress.
RecoverRetry, replay or escalate.
Schedulers differ in delivery semantics, delay and retry policy. Some can invoke twice; some can delay or drop a run under platform conditions; a target can reject calls long after the original scheduled time. Treat the invocation as at-least-once unless the platform contract proves otherwise, and make the work idempotent. A unique key such as job name plus logical period prevents a duplicate tick from creating duplicate business effects.
Failure modes worth designing for
Failure mode
Typical symptom
Design response
Duplicate invocation
Two invoices, exports or notifications for one period
Idempotency key, unique constraint and safe upsert
Overlap
Run N+1 starts before N finishes and races shared state
Lease with expiry, concurrency limit or partitioned work
Missed/delayed run
Stale report or skipped synchronization window
Watermark and catch-up query over unprocessed intervals
Partial completion
Some records changed but run is marked failed
Checkpoint batches and make replay safe
Timezone or DST shift
Local business task runs at an unexpected hour
Store UTC instants and define business timezone explicitly
Silent failure
Scheduler says invoked but output is wrong or empty
Validate postconditions and alert on missing outcome
Separate the clock from business intent
“Every day at 02:00” is incomplete when the business operates in a local timezone or daylight-saving rules change. Decide whether the job follows UTC elapsed time or a local calendar boundary. Persist the intended period, not just the process start timestamp. A daily reconciliation should know which business date it owns, even if execution starts late.
Make execution bounded and observable
Set a maximum runtime and batch size. Record scheduled time, actual start, logical period, attempt number, checkpoint, rows processed, outcome and error category. Emit a heartbeat or durable run record before doing irreversible work. Alert on stale completion, not only exceptions: a job can exit successfully while processing zero records because its query window was wrong.
Retries need limits, exponential backoff with jitter where appropriate and a terminal path such as a dead-letter queue or operator review. Do not retry non-idempotent side effects blindly. Define who can safely replay a period and how replay is audited.
When cron should become a workflow
A single stateless cleanup may need only a scheduler and an idempotent handler. Multi-step work with long waits, human review, compensation or durable progress needs a workflow or queue-backed state machine. Keep the scheduler thin: calculate the period and enqueue work; let a worker claim and process it with bounded concurrency.
The cron syntax is the cheapest part. Reliable jobs need logical run identity, idempotent effects, overlap control, catch-up, bounded retries and proof of completion. Design those guarantees first, then select the scheduler that fits.