Decision Call · USD 260
Your cloud bill tripled and the cause is 41 lines of code
60 minute call, 18 August. Reviewed and signed off by Omar Haddad on 20 August.
You asked whether to move off the cloud. Your bill went from USD 31,000 to USD 94,000 a month in seven months while traffic grew 40 percent, and that gap is not a cloud pricing problem. Roughly USD 47,000 a month of it comes from one change deployed in February.
A retry loop added in February retries on any non-200, including on the 404s that are a normal part of your cache-miss path. It generates around 71 million extra cross-region requests a month and the data transfer on them is most of your increase. Fixing it is a day of work. Migrating off the cloud to solve it would be nine months and would not address it.
The question
What was asked, and what was reviewed
As you put it: "The board thinks we should go back to colo. I need to know if they are right."
| Reviewed | Source |
|---|---|
| 12 months of billing, by service and region | Supplied |
| Traffic and request volume, same period | Supplied |
| Architecture diagram | Supplied |
| Deployment history, February to date | Supplied |
| The retry implementation | Read on the call, screen shared |
| The colo quote the board is looking at | Supplied |
Where
Where the increase actually is
The cause
The change, and why it looked harmless
The relevant part, as deployed in February
// added to improve resilience during the eu-west incident
async function fetchWithRetry(url, opts, attempts = 5) {
for (let i = 0; i < attempts; i++) {
const res = await fetch(url, opts);
if (res.status === 200) return res;
await sleep(200 * 2 ** i); // exponential backoff
}
throw new Error('exhausted retries');
}
// The cache-miss path returns 404 by design.
// Every miss is now retried five times, cross-region,
// with the full payload on each attempt. | What was intended | What happens |
|---|---|
| Retry on transient failure | Retries on every non-200, including 404 |
| Rare, during incidents | On every cache miss, which is about 14 percent of requests |
| Small overhead | 5x the requests and 5x the transfer on that path |
| Backoff limits the damage | Backoff spaces them out. It does not reduce the count |
| Would show up in error rates | It does not. Every retry eventually succeeds or 404s normally |
This is a good engineer's change, made during an incident, that has behaved exactly as written and not as intended. It is invisible in every dashboard you have because nothing is failing.
Colo
The migration question, answered properly
Since the board asked, it is worth answering rather than deflecting.
| Stay, after the fix | Move to colo | |
|---|---|---|
| Monthly run cost | ~USD 47,000 | ~USD 38,000 plus staff |
| One-off cost | One day | USD 300,000 plus 9 months |
| Headcount to run it | Current team | 2 to 3 more |
| Which erases the saving | n/a | Yes, roughly |
| Recovery from a region failure | Existing | You would build it |
| Answers the actual problem | Yes | No. The loop would come too |
The last row is the argument. A cost problem caused by an application defect migrates with the application.
Next
In this order
- Today: retry only on 5xx and on network errors, not on 404. One line, and cap attempts at three.
- This week: a billing alert on cross-region transfer, at 20 percent above trailing average. You had no alert on the line that tripled.
- This week: show the board the chart in clause 2. It is a better answer than a migration plan.
- This month: cost attribution by service, so the next one of these is visible in a week rather than seven months.
- This quarter: review the other retry implementations. Where there is one, there are usually three.