The short version
Capacity is a customer journey, not a server count. Rehearse completed purchases with the real mix of browsing, inventory, payment and confirmation traffic, then design for the first dependency that fails.
Trace a purchase all the way through
Begin with a single purchase and map every required step: cart validation, inventory reservation, tax, payment, order persistence and confirmation. A fast catalogue cannot compensate for a saturated payment dependency. Record which team owns each step and how its timeout or error appears to the buyer.
Define success in customer terms: confirmed order, honest pending state or clear failure. Measure completion rate and tail latency for each outcome, not just average response time on an isolated endpoint. A buyer who waits and retries may create a different failure from one who receives an immediate error.
Purchase path
A completed order is a chain of dependencies.
Cart
Buyer starts checkout.
Inventory
Stock is reserved.
Payment
Charge is confirmed.
Order
Buyer sees a durable result.
Decide what should give way first
Google's SRE guidance describes degraded responses and load shedding when systems are overloaded. In a store, product recommendations and live personalization are candidates to degrade before purchase confirmation, provided the customer is not misled about price or stock. That priority is a business choice the team should agree on before the event.
Check whether optional and critical work share worker pools, database connections, queues or third-party quotas. Scaling checkout containers separately does not isolate checkout if both paths still wait on the same saturated inventory table. Protect the bottleneck, not just the service name on the architecture diagram.
Evidence: Google SRE: Handling Overload
Load priorities
Protect the critical path under pressure.
Keep moving
Payment and order confirmation
Degrade carefully
Noncritical search enhancements
Pause first
Recommendations and nonessential refreshes
Include slow dependencies in the rehearsal
Use an authorized environment and the payment provider's test facilities. Model warm-up, an arrival spike and recovery, with browsing and purchasing mixed at credible proportions. Add controlled slowness to dependencies you own or can safely simulate. Watch database connections, queue age, provider errors and order-state transitions as well as CPU.
A retry after a payment timeout must not create another charge or order. Stripe documents idempotency keys for safely retrying its API requests; the surrounding order workflow still needs its own stable operation identity and reconciliation. Test the ambiguous case where the payment succeeded but the confirmation was lost.
Record completed, pending and failed purchases separately. A system that stops accepting orders cleanly may be preferable to one that accepts work it cannot reconcile later. Be explicit about that tradeoff in the runbook.
- Measure the full purchase outcome, not only HTTP status.
- Inject safe delays and timeouts in the test path.
- Verify duplicate-submission handling.
- Test recovery after the burst and after rollback.
Evidence: Stripe: Idempotent requests
Turn the results into an operating plan
A useful report names the first limiting dependency, the symptom an on-call engineer will see and the action that changes the outcome. State what was not tested, such as a provider quota or a production-only integration, rather than turning a successful rehearsal into a guarantee.
Agree on pre-event capacity, cost guardrails, alerts and escalation owners. Schedule another rehearsal when the purchase path changes materially. The goal is a busy period the team can operate with evidence, not a one-time graph that looks reassuring.
Sources & further reading
- Google SRE: Handling Overload
Explains overload, degraded responses and load shedding. The commerce prioritization is our recommendation.
- Stripe: Idempotent requests
Documents Stripe's API retry behavior; it does not guarantee idempotency for an entire merchant order workflow.
Written by Dopstack Technologies
We design and build software, cloud infrastructure and AI workflows. These notes explain engineering decisions; illustrative scenarios are not claims of client results.
Meet the team