RB-203: Checkout Service Database Recovery Runbook
Date: 2026-06-07 Service: Checkout Service Affected Services: Authentication Service, Checkout Service, Billing Service, Inventory Service, Notification Service Incident References: INC-2025-203, INC-2025-204 Primary Dashboard: acme/checkout/db-restore
When To Use
Use this runbook when Checkout Service shows symptoms related to database recovery, including SLO burn, elevated p95 latency, queue backlog, dependency saturation, or customer-visible failures. This procedure is informed by incidents INC-2025-203, INC-2025-204.
Preconditions
- Confirm the incident commander and service owner are assigned.
- Confirm recent deployments and feature flags for Checkout Service.
- Review API health for API-CHECKOUT-01, API-CHECKOUT-02, API-CHECKOUT-03, API-CHECKOUT-04.
- Check whether Checkout Service or Billing Service is seeing downstream impact.
Procedure
- Capture current error rate, latency, saturation, and queue depth.
- Compare current behavior with the baseline dashboard for Checkout Service.
- If retries are amplifying traffic, reduce client retry budgets and enable degraded mode.
- If data repair is required, run a dry-run replay before any write operation.
- Communicate status in Slack and link the active Jira ticket.
Validation
Recovery is complete when customer-facing error rate returns to baseline, API API-CHECKOUT-01 is healthy, and dependent services report no continuing impact.