Home / Blog / Scheduled work
Backend Reliability

The Hidden Cost of Cron Jobs

A cron expression looks tiny because it describes only when to wake a process. Production cost lives in everything around that trigger: duplicated delivery, overlapping runs, missed windows, timezone interpretation, retries, partial writes and the question nobody asks until an incident: how do we know the job finished correctly?

A schedule is a request to attempt work

Reliable scheduled work needs control points beyond the clock
TriggerSchedule and timezone.
ClaimLease or unique run key.
ExecuteBounded, checkpointed work.
RecordOutcome and progress.
RecoverRetry, replay or escalate.

Schedulers differ in delivery semantics, delay and retry policy. Some can invoke twice; some can delay or drop a run under platform conditions; a target can reject calls long after the original scheduled time. Treat the invocation as at-least-once unless the platform contract proves otherwise, and make the work idempotent. A unique key such as job name plus logical period prevents a duplicate tick from creating duplicate business effects.

Failure modes worth designing for

Failure modeTypical symptomDesign response
Duplicate invocationTwo invoices, exports or notifications for one periodIdempotency key, unique constraint and safe upsert
OverlapRun N+1 starts before N finishes and races shared stateLease with expiry, concurrency limit or partitioned work
Missed/delayed runStale report or skipped synchronization windowWatermark and catch-up query over unprocessed intervals
Partial completionSome records changed but run is marked failedCheckpoint batches and make replay safe
Timezone or DST shiftLocal business task runs at an unexpected hourStore UTC instants and define business timezone explicitly
Silent failureScheduler says invoked but output is wrong or emptyValidate postconditions and alert on missing outcome

Separate the clock from business intent

“Every day at 02:00” is incomplete when the business operates in a local timezone or daylight-saving rules change. Decide whether the job follows UTC elapsed time or a local calendar boundary. Persist the intended period, not just the process start timestamp. A daily reconciliation should know which business date it owns, even if execution starts late.

Make execution bounded and observable

Set a maximum runtime and batch size. Record scheduled time, actual start, logical period, attempt number, checkpoint, rows processed, outcome and error category. Emit a heartbeat or durable run record before doing irreversible work. Alert on stale completion, not only exceptions: a job can exit successfully while processing zero records because its query window was wrong.

Retries need limits, exponential backoff with jitter where appropriate and a terminal path such as a dead-letter queue or operator review. Do not retry non-idempotent side effects blindly. Define who can safely replay a period and how replay is audited.

When cron should become a workflow

A single stateless cleanup may need only a scheduler and an idempotent handler. Multi-step work with long waits, human review, compensation or durable progress needs a workflow or queue-backed state machine. Keep the scheduler thin: calculate the period and enqueue work; let a worker claim and process it with bounded concurrency.

Related patterns: queues for small teams, idempotency keys and observability on a budget.

In summary

The cron syntax is the cheapest part. Reliable jobs need logical run identity, idempotent effects, overlap control, catch-up, bounded retries and proof of completion. Design those guarantees first, then select the scheduler that fits.

References