Mimir

RB-304: Billing Service Latency Triage Runbook

runbook
billing
ca7dcb5cd8c5

runbooks/rb-304-billing-latency-triage.md

RB-304: Billing Service Latency Triage Runbook

Date: 2026-04-15 Service: Billing Service Affected Services: Authentication Service, Checkout Service, Billing Service, Inventory Service, Notification Service Incident References: INC-2025-304, INC-2025-305 Primary Dashboard: acme/billing/latency-triage

When To Use

Use this runbook when Billing Service shows symptoms related to latency triage, including SLO burn, elevated p95 latency, queue backlog, dependency saturation, or customer-visible failures. This procedure is informed by incidents INC-2025-304, INC-2025-305.

Preconditions

  • Confirm the incident commander and service owner are assigned.
  • Confirm recent deployments and feature flags for Billing Service.
  • Review API health for API-BILLING-01, API-BILLING-02, API-BILLING-03, API-BILLING-04.
  • Check whether Checkout Service or Billing Service is seeing downstream impact.

Procedure

  1. Capture current error rate, latency, saturation, and queue depth.
  2. Compare current behavior with the baseline dashboard for Billing Service.
  3. If retries are amplifying traffic, reduce client retry budgets and enable degraded mode.
  4. If data repair is required, run a dry-run replay before any write operation.
  5. Communicate status in Slack and link the active Jira ticket.

Validation

Recovery is complete when customer-facing error rate returns to baseline, API API-BILLING-01 is healthy, and dependent services report no continuing impact.