RB-509: Notification Service Capacity Scaleout Runbook
Date: 2026-03-06 Service: Notification Service Affected Services: Authentication Service, Checkout Service, Billing Service, Inventory Service, Notification Service Incident References: INC-2026-509, INC-2026-510 Primary Dashboard: acme/notification/capacity-scaleout
When To Use
Use this runbook when Notification Service shows symptoms related to capacity scaleout, including SLO burn, elevated p95 latency, queue backlog, dependency saturation, or customer-visible failures. This procedure is informed by incidents INC-2026-509, INC-2026-510.
Preconditions
- Confirm the incident commander and service owner are assigned.
- Confirm recent deployments and feature flags for Notification Service.
- Review API health for API-NOTIFICATION-01, API-NOTIFICATION-02, API-NOTIFICATION-03, API-NOTIFICATION-04.
- Check whether Checkout Service or Billing Service is seeing downstream impact.
Procedure
- Capture current error rate, latency, saturation, and queue depth.
- Compare current behavior with the baseline dashboard for Notification Service.
- If retries are amplifying traffic, reduce client retry budgets and enable degraded mode.
- If data repair is required, run a dry-run replay before any write operation.
- Communicate status in Slack and link the active Jira ticket.
Validation
Recovery is complete when customer-facing error rate returns to baseline, API API-NOTIFICATION-01 is healthy, and dependent services report no continuing impact.