Mimir

RB-508: Notification Service Backfill and Event Replay Runbook

runbook
notification
014a64479930

runbooks/rb-508-notification-backfill-replay.md

RB-508: Notification Service Backfill and Event Replay Runbook

Date: 2026-02-12 Service: Notification Service Affected Services: Authentication Service, Checkout Service, Billing Service, Inventory Service, Notification Service Incident References: INC-2026-508, INC-2026-509 Primary Dashboard: acme/notification/backfill-replay

When To Use

Use this runbook when Notification Service shows symptoms related to backfill and event replay, including SLO burn, elevated p95 latency, queue backlog, dependency saturation, or customer-visible failures. This procedure is informed by incidents INC-2026-508, INC-2026-509.

Preconditions

  • Confirm the incident commander and service owner are assigned.
  • Confirm recent deployments and feature flags for Notification Service.
  • Review API health for API-NOTIFICATION-01, API-NOTIFICATION-02, API-NOTIFICATION-03, API-NOTIFICATION-04.
  • Check whether Checkout Service or Billing Service is seeing downstream impact.

Procedure

  1. Capture current error rate, latency, saturation, and queue depth.
  2. Compare current behavior with the baseline dashboard for Notification Service.
  3. If retries are amplifying traffic, reduce client retry budgets and enable degraded mode.
  4. If data repair is required, run a dry-run replay before any write operation.
  5. Communicate status in Slack and link the active Jira ticket.

Validation

Recovery is complete when customer-facing error rate returns to baseline, API API-NOTIFICATION-01 is healthy, and dependent services report no continuing impact.