Mimir

RB-204: Checkout Service Latency Triage Runbook

runbook
checkout
818f929a2c55

runbooks/rb-204-checkout-latency-triage.md

RB-204: Checkout Service Latency Triage Runbook

Date: 2025-07-14 Service: Checkout Service Affected Services: Authentication Service, Checkout Service, Billing Service, Inventory Service, Notification Service Incident References: INC-2025-204, INC-2025-205 Primary Dashboard: acme/checkout/latency-triage

When To Use

Use this runbook when Checkout Service shows symptoms related to latency triage, including SLO burn, elevated p95 latency, queue backlog, dependency saturation, or customer-visible failures. This procedure is informed by incidents INC-2025-204, INC-2025-205.

Preconditions

  • Confirm the incident commander and service owner are assigned.
  • Confirm recent deployments and feature flags for Checkout Service.
  • Review API health for API-CHECKOUT-01, API-CHECKOUT-02, API-CHECKOUT-03, API-CHECKOUT-04.
  • Check whether Checkout Service or Billing Service is seeing downstream impact.

Procedure

  1. Capture current error rate, latency, saturation, and queue depth.
  2. Compare current behavior with the baseline dashboard for Checkout Service.
  3. If retries are amplifying traffic, reduce client retry budgets and enable degraded mode.
  4. If data repair is required, run a dry-run replay before any write operation.
  5. Communicate status in Slack and link the active Jira ticket.

Validation

Recovery is complete when customer-facing error rate returns to baseline, API API-CHECKOUT-01 is healthy, and dependent services report no continuing impact.