Observability on a Budget: Logs, Metrics, Traces and Useful Alerts
A small backend does not need an enterprise telemetry platform to answer basic production questions. It does need to know whether users can complete important work, where a request failed and whether a background queue is falling behind. Start with a few high-value signals, connect them with a correlation ID and add detail where incidents show it is missing. The budget includes operator attention, storage and cardinality, not only vendor spend.
Published September 28, 202614 min readSignals and actionable alerting
Think in questions, then choose signals
Each signal answers a different operational question
MetricsHow often, how many, how slow?
TracesWhere did this request spend time?
LogsWhat happened at this specific step?
AlertsDoes a human need to act now?
Workflow
Starter signals
Actionable alert
HTTP API
Request count, error rate, latency percentiles
Sustained user-impacting failures or latency budget burn
Background jobs
Queue age, attempts, completion and failure counts
Oldest work exceeds its business deadline
Integrations
Dependency latency, throttles and sync freshness
Critical data is stale past its agreed window
Database
Connection saturation, query latency, pool wait
Requests cannot obtain connections and user flow degrades
Make telemetry correlate without logging everything
Use structured logs with timestamp, service, environment, severity, operation name and trace/correlation ID. Add entity IDs only when they are safe and useful; never place tokens, passwords, full request bodies or sensitive customer details in routine logs. Correlation lets an operator move from an aggregate metric to a trace and then to a small set of relevant events.
Instrument key boundaries first: incoming request, database call, external API, queue publish and worker processing. Automatic OpenTelemetry instrumentation can cover common libraries; add custom spans for meaningful business operations rather than wrapping every line of code. Sample successful high-volume traces while retaining errors and slow requests according to policy.
Keep metrics cheap and bounded
Metrics are most useful for trends and alert conditions, but every unique label combination consumes state and storage. Avoid user IDs, raw URLs, arbitrary error text and unbounded tenant identifiers as metric labels. Prefer normalized route names, status class, operation and a controlled service/environment set. Put high-cardinality detail in sampled traces or searchable logs.
Alert on symptoms, not every possible cause
A page should mean someone needs to act promptly. Alert on user-visible symptoms such as sustained error rate, latency, missed processing deadline or an SLO budget burn. Send lower-urgency warnings to a ticket or dashboard. Include the affected workflow, measured value, time window, owner and a runbook link. Add a duration threshold to avoid paging on brief spikes, and test the full notification path.
Use retention and sampling deliberately
Keep detailed debug logs briefly, preserve security/audit records under their own policy and sample routine traces. Retention should reflect incident investigation needs and data sensitivity. Review ingestion volume after adding instrumentation; one noisy payload field can dominate cost. Establish a monthly budget for bytes, trace volume and alert count, then remove signals nobody uses.
Start with one service-level dashboard
Show request health, background work age, dependency health and the last successful critical operation. Add deployment annotations and links to logs/traces. A useful dashboard answers βis the service healthy for users?β and points toward a diagnosis; it should not be a wall of every metric the code can emit.
In summary
Cost-conscious observability is selective, correlated and tied to decisions. Measure user outcomes, use traces to locate latency, logs for specific events and alerts only when action is needed. Bound metric labels, protect sensitive data, set retention and sampling, and review whether telemetry earns its cost.