Anomaly Response
Use this playbook to respond to cost anomalies detected by Cost Management. The workflow helps teams triage alerts, confirm whether changes are expected, identify owners, coordinate remediation, and improve detection rules over time.
Intake and Triage
- Centralize anomaly alerts into your preferred channels (ticketing, chat, incident system).
- For each alert, confirm:
- Scope (BU, app, account, customer).
- Time window and magnitude of the anomaly.
- Whether related budget alerts or anomalies exist.
If data health is in question, check Data Health & Freshness before deeper investigation.
Check for Expected Changes
- Look for known events: deployments, load tests, migrations, marketing campaigns, new regions, or services.
- Ask the owning team whether they recently changed configuration, scaling, or data processing patterns.
If the anomaly is fully explained by planned work, document it and, if needed, adjust detector thresholds or exclusions.
Analyze Drivers
For anomalies that are not clearly expected:
- Break down cost by provider, service, region, and usage type.
- Use allocation views to see which applications, environments, or teams drove the change.
- Identify specific resources or services that changed most (for example, storage growth, data transfer, compute bursts).
Capture this analysis in a ticket or an internal doc so it can inform future prevention efforts.
Decide on Response
Depending on root cause, respond by:
- Optimization actions.
- Rightsizing over‑provisioned resources.
- Adjusting or adding commitments where sustained growth is expected.
- Scheduling or shutting down unnecessary resources.
- Design and configuration changes.
- Fixing misconfigurations (for example, logging verbosity, data export paths, scaling thresholds).
- Updating deployment patterns to avoid accidental over‑provisioning.
- Process and governance changes.
- Tightening guardrails (for example, quotas, policies, approval steps for high‑cost changes).
Agree on owners and timelines for each action, and track them through completion.
Communicate and Document
- Summarize key details for stakeholders:
- What happened, when, and where.
- Root cause and business impact.
- Actions taken and any remaining risk.
- If the anomaly affects showback & chargeback statements, add explanatory notes in the relevant period’s communications.
Store learnings in a shared location (for example, a runbook, wiki, or ticket system) to inform future incidents.
Improve Detection and Prevention
After resolving significant anomalies:
- Tune detectors (sensitivity, scopes, exclusions) to reduce noise or catch similar issues earlier.
- Consider adding or updating policies and dashboards that:
- Highlight risky patterns before they turn into anomalies.
- Track optimization backlogs and savings related to anomaly‑driven work.
Over time, this loop will make anomaly response faster, more predictable, and more valuable for the business.