An integration can pass unit tests and still fail when a provider changes a field, revokes a credential or behaves differently from its sandbox. A useful test strategy layers fast deterministic checks with contract verification and a small amount of carefully controlled live evidence; no single layer proves the whole relationship.
Published October 1, 202614 min readContracts, sandboxes and production confidence
Spend test effort where uncertainty lives
Start by separating code you own from behavior another organization owns. Unit tests should dominate because parsing, mapping, retries and idempotency are yours to control. Contract tests then make assumptions explicit between consumer and provider. Provider or sandbox checks exercise a broader boundary, while a tiny monitored production probe can reveal credential, routing or configuration problems that test environments do not reproduce.
A layered portfolio: many fast checks, fewer expensive boundary checks
Few / production probesRead-only, synthetic, rate-limited signals for availability and configuration; never use real customer data.
Some / provider and sandbox testsValidate authentication, representative workflows, provider states and documented sandbox limitations.
More / consumer contract testsRecord the fields and interactions the client actually depends on; verify provider compatibility.
Many / unit and component testsExercise mapping, validation, retries, deduplication, pagination, errors and boundary conditions deterministically.
Make the contract express what the client needs
A contract is not a copy of an entire provider schema. Capture the interactions your consumer relies on: method, path, relevant request fields, status and the minimum response shape. Consumer-driven contract tooling can generate these expectations from consumer tests and let the provider verify them against its implementation. Keep matching rules intentional: exact examples are useful for identifiers and enums, while flexible matchers may be appropriate for timestamps or generated IDs.
Contracts do not prove business semantics, provider uptime or every consumer’s needs. They are executable compatibility checks for stated assumptions. Review contract changes as API changes, publish the version and environment, and require a successful provider verification for the version being deployed.
Do not confuse a sandbox with production
Layer
Good evidence for
Does not prove
Unit
Local transformations, retry decisions and edge cases
Wire compatibility or remote behavior
Contract
Consumer assumptions remain compatible with a verified provider
Availability, latency or real account configuration
Sandbox
Credential flow and representative end-to-end requests
Production data, limits, integrations or identical provider behavior
Production probe
Current reachability and a narrow safe transaction
All workflows, correctness at scale or absence of hidden failures
Provider sandboxes may omit features, use synthetic data or behave differently under throttling. Write those gaps down and cover what you can locally. Never let a sandbox pass create an unqualified “integration tested” claim. Maintain a short evidence record with environment, credential scope, scenarios exercised and known exclusions.
Design tests around failure behavior
For external APIs, test timeouts, connection resets, rate limits, expired credentials, malformed payloads, duplicate delivery, pagination loops and partial success. Assert the outcome users or operators can observe: bounded retries, idempotent writes, dead-letter visibility, useful error codes and preserved source data. Inject clocks and transport adapters so these cases are deterministic rather than dependent on provider uptime.
Use realistic fixtures captured or authored without personal data. Redact tokens, identifiers and payload fields before committing examples. Make contract fixtures small enough that reviewers can understand the promise being made.
Gate releases without turning CI into a live dependency
Run unit and contract checks on every change. Run sandbox suites on a schedule or before a release when the environment is stable, not on every pull request if a vendor outage can block unrelated work. A controlled production smoke check should be read-only where possible, use a dedicated low-privilege account and alert on sustained failure rather than one transient timeout.
Separate a failed test from an unavailable test dependency. Report skipped or blocked external checks explicitly; do not silently convert them into green. Keep secrets in the CI secret store, scope them to the smallest environment and prevent logs from printing request headers or bodies.
Build a useful integration test portfolio
A small team can begin with high-value mapping and idempotency tests, then add contracts for critical consumers, a documented sandbox checklist and a synthetic production probe for the most important read path. Track contract verification age, sandbox success rate, live-probe availability and integration incidents caused by schema drift. These measures describe confidence and operating health, not a universal quality score.
The testing pyramid for integrations is really a map of ownership and uncertainty. Keep most tests fast and deterministic, use contracts for explicit compatibility, treat sandboxes as partial evidence and reserve production checks for narrow, safe signals. The goal is not to simulate the internet in CI; it is to know which failure each layer can catch before a customer does.