“Use a greener region” is not an architecture strategy by itself. First reduce the energy required for useful work; then decide which work can wait or move, while keeping user latency, data residency, availability and recovery constraints explicit.
Published September 28, 202614 min readCloud sustainability
Energy discussions around software often jump from a large data-center headline to a claim about one API. The missing link is measurement. A credible engineering comparison defines the system boundary, the work being delivered, the resources provisioned and the assumptions behind energy or emissions estimates. Without that, a lower server count might simply mean work moved to a different service or became slower for the user.
The International Energy Agency reported rapidly rising electricity demand from data centers in its 2026 analysis, with AI-focused facilities growing especially quickly. That makes efficiency and scheduling an increasingly practical systems concern. It does not mean every workload should chase carbon intensity at the expense of reliability or legal data-location requirements.
Measure per unit of useful work
The Green Software Foundation’s Software Carbon Intensity (SCI) specification expresses emissions as a rate per functional unit. In simplified form, it includes operational energy multiplied by location-based electricity carbon intensity, plus an allocated share of embodied hardware emissions, divided by the selected unit of work.
SCI = (E × I + M) / R
Here, E is the energy inside the defined software boundary, I is the relevant grid carbon intensity, M is allocated embodied hardware impact, and R is the functional unit: for example, one successful API request, one transaction, one processed image or one completed batch. The formal specification has detailed measurement and allocation rules; this compact equation is an explanatory shorthand.
Choose a unit that reflects customer value, not merely infrastructure activity. “kWh per million successful requests” is more useful than “CPU utilization fell” if the service also increased retries or latency. For an AI workload, report energy or estimated emissions per completed task alongside quality, token count, cache behavior and failure rate. Disclose whether figures are measured or modeled.
Illustrative comparison only: relative energy per equal unit of successful work
Unbatched baseline
100 index
Reuse/cache hot data
78 index
Batch + right-size workers
61 index
These values are a teaching example, not benchmark results. Your own chart should come from repeatable runs with the same workload, data, success criteria, service-level objective and measurement boundary.
Efficiency before carbon-aware scheduling
Software can consume less energy for the same function by avoiding unnecessary work: cache stable reads, deduplicate events, batch compatible jobs, reduce serialization and data transfer, select fit-for-purpose models, limit polling, and right-size provisioned capacity. Measure the complete path: application compute, database work, network traffic, storage, retries and any managed services included in the boundary.
Engineering sequence: shrink work, then place flexible work
01 / DefineUseful outputSet functional unit, quality bar, latency objective and system boundary.
02 / MeasureEnergy per unitCollect workload, compute, storage and transfer signals; label measured vs modeled values.
03 / ReduceRemove wasteCache, batch, right-size, reduce retries and choose efficient algorithms or model tiers.
04 / ScheduleShift eligible workUse time or region only when freshness, residency, latency and recovery allow it.
Separate flexible work from interactive work
Interactive requests have user-facing latency and availability budgets. Moving them to chase an hourly grid signal can add network distance, fail over to a less reliable region or breach residency constraints. Better candidates for time-shifting are queued, delay-tolerant jobs: non-urgent analytics, model evaluation, report generation, compaction, backups where recovery objectives permit, and some training workloads.
Make flexibility explicit in the job contract: earliest start, deadline, maximum deferral, allowed regions, data class, retry budget, required capacity and whether the result can be recomputed. The scheduler may optimize only within those constraints. If the signal is stale or the preferred location is unhealthy, fall back to the service objective instead of waiting indefinitely.
Workload
Likely strategy
Guardrail
Interactive API
Optimize code, cache, database and model path; choose a region near users and permitted data.
Latency SLO, residency and availability outrank hourly carbon optimization.
Daily reporting batch
Queue within a defined window and schedule when capacity or energy conditions are favorable.
Hard completion deadline, input freshness and idempotent reruns.
Model evaluation
Batch compatible test cases; schedule flexible suites and use smaller models for screening.
Evaluation validity, reproducibility and release gate timing.
Backup / compaction
Throttle, stagger and avoid competing peaks where practical.
Recovery point objective, retention policy and restore throughput.
Region choice is a multi-objective decision
Cloud-provider guidance recommends considering carbon alongside business requirements, performance and cost. A lower-carbon electricity profile does not automatically make a region the right place for a workload. Check data residency and sovereignty, user latency, service availability, disaster recovery topology, egress, capacity and pricing. Compare like with like, and document the reason for choosing a region.
For a batch workload, time-shifting within an approved region may be simpler than moving data across borders. For a latency-sensitive service, proximity and redundancy can dominate. Carbon-intensity data also varies in resolution and method: annual regional averages are not the same as hourly marginal signals. Record the source, timestamp, geographic boundary and whether the factor is average or marginal.
Build a carbon-aware scheduler that fails open to the SLO
02 / ValidatePolicy gateResidency, security, capacity and service dependency checks.
03 / ObserveSignalsQueue age, grid data freshness, region health and cost.
04 / DecideConstrained placementChoose an allowed start time and region inside the job’s bounds.
05 / VerifyOutcomeRecord completion, retries, energy estimate, deadline and fallback reason.
Use bounded deferral, a maximum queue age, an explicit fallback and a circuit breaker for bad signals. Avoid creating a self-inflicted thundering herd by releasing all deferred jobs at the same “green” time; use jitter and capacity-aware rate limits. Keep the optimization objective observable: successful work completed on time, energy per unit, estimated emissions per unit, latency, failure rate and cost.
What I would build
I would start with one non-urgent workload and a repeatable benchmark. Add a job envelope containing deadline, permitted region set, data classification, idempotency key and flexibility window. Instrument queue time and resource use; calculate a baseline energy-per-output estimate using the same method for every comparison. Then add a policy service that consumes a timestamped carbon-intensity feed, but can only select among pre-approved regions and times.
For cloud-native systems, export the measurements as workload-level metrics and annotate deploys with the boundary and estimator version. A dashboard should compare SCI-like rates and the service objective together, not show a single “green” score. A change is useful only if it reduces impact per equivalent work without silently lowering quality or moving the cost elsewhere.
Failure modes and signals
Failure
Signal
Control
Optimization reduces completed work
Energy per hour falls but failures, backlog or missed deadlines rise.
Normalize by successful functional units and keep the quality/SLO gate.
Carbon feed is stale or mismatched
Scheduler uses old data or the wrong grid region.
Validate timestamp, geographic mapping and maximum signal age; record fallback.
Deferred jobs stampede
Queue depth and throttling spike at the same scheduled interval.
Add jitter, concurrency limits and queue-age-aware dispatch.
Region move breaks residency or recovery
Data crosses a policy boundary or recovery coverage degrades.
Apply policy before optimization; validate replication and restore paths.
Estimate is treated as a meter reading
Modeled figures are reported with unjustified precision.
Label method, assumptions, boundary, uncertainty and factor source.
In summary
Energy-aware architecture is a measurement and scheduling discipline, not a region leaderboard. Define useful work, reduce the resources needed to deliver it, and only then shift flexible workloads within explicit deadlines, residency, latency and recovery constraints. Publish the boundary and method, pair impact metrics with service quality, and make the ordinary SLO the fallback whenever sustainability signals are missing or unsafe.
Editorial note: This article describes engineering methods, not a claim that a particular workload or cloud region has a specific footprint. SCI methodology and provider guidance should be consulted for formal calculations and current service-specific data.