AWS RUNAWAY SPEND
AWS Costs Increasing Fast? Investigate Runaway Spend
Fast AWS spend growth can involve several resources or services. The AWS connector is not production-ready yet. This page documents the cost dimensions, failure modes and unit-economics model CostNerve is designed to support, while the product keeps AWS clearly marked as planned.
What problem does it solve?
- account
- region
- service
- usage type
- resource
- tag
What to check first
- Pin down the first minute/hour where spend velocity changed; avoid comparing only monthly totals.
- Start with Cost by service, region, account and usage type and then break the delta down across the AWS dimensions that actually moved.
- Correlate the inflection with deployments, traffic, retries, schedulers, background jobs and abuse/bot events.
- Keep a before/after record, then use the smallest reversible mitigation so you can measure whether it worked.
Metrics and signals that matter
- Cost by service, region, account and usage type
- EC2/Lambda utilization and request/runtime growth
- Data transfer/NAT/egress changes
- New resources, autoscaling events and commitment coverage
Likely causes
Deployment or configuration regression
A release can change request fan-out, runtime, memory, model choice, logging volume or cache behavior without obvious user-facing breakage.
Traffic, retries or loops
Legitimate growth, bots, retry storms and recursive/background loops can all multiply a normally cheap unit of work.
Billing dimension changed
For AWS, investigate Cost by service, region, account and usage type and EC2/Lambda utilization and request/runtime growth before assuming the total moved for a single reason.
How it works
What to check in the provider dashboard
Open the AWS usage and billing dashboard; this guide does not imply live CostNerve ingestion. Compare account, region, service over equal, complete periods. Keep usage, estimates and invoice amounts separate. Record the scope, currency and last update. Check integration coverage before connecting; use available connectors for supported evidence.
Integration status
The AWS connector is not production-ready yet. This page documents the cost dimensions, failure modes and unit-economics model CostNerve is designed to support, while the product keeps AWS clearly marked as planned.
Worked example with explicit assumptions
Illustrative example, not a provider rate: 12 USD/hour versus a 3 USD/hour baseline means 9 USD/hour of excess spend. If that rate persists for six hours, the additional cost is 54 USD. Recalculate after mitigation; do not treat this scenario as an invoice.
Frequently asked questions
Which AWS signals should I inspect first?
Start with Cost by service, region, account and usage type, EC2/Lambda utilization and request/runtime growth, Data transfer/NAT/egress changes. Compare the same time window before and after the change so volume and unit-cost effects do not get mixed.
How do I know the incident is contained?
Check request volume, concurrency or the affected usage metric after the change. Then reconcile delayed billing for the same scope and currency. Record the action, owner and rollback condition; a quiet alert alone does not prove recovery.
Should uncertain cost be forced into a project?
No. Keep it unallocated until tags, project IDs, resource IDs or another reliable signal justify attribution. False precision produces worse decisions than visible uncertainty.
What should I do before an emergency cost control?
Capture the affected provider/project, current spend velocity, suspected cause and deployment/traffic context. Use a read-only investigation first; any write action should be explicit, scoped, reversible and audit logged.