A client asked us to redirect a retired domain to its replacement. One of my engineers did it in twenty-six minutes, by hand, in the AWS console, and it worked perfectly for sixteen days. Then a deploy about WordPress plugins put the cloud back the way the code said it should be, and the dead site climbed out of its grave for seventeen more days, until the client's CEO found it on a Saturday morning. Nothing malfunctioned. Every system did precisely what we built it to do. That is the joke, and the lesson, aimed at my own team as hard as at anyone reading this. There are also cats.
The manifest always wins
The client is a live events promoter. We run around thirty production sites for them, more than forty CloudFront distributions once you count staging: a site for every festival and every year it has run, the venue sites, the corporate site, and the ticketing pipeline behind all of them. Those sites sell millions of dollars of tickets a year and send well over a million text messages a month, each carrying a link to one of these domains. A domain serving the wrong page during an on-sale is not a cosmetic bug. It is a revenue event with a timestamp.
A quick tour of how a page reaches a phone, because the story turns on it. The site that builds the page runs in one place: the origin. Nobody's phone talks to it. In front of it sits a content delivery network, a few hundred small data centers called edge locations, and your phone talks to the closest one. The first time an edge is asked for a page it fetches a copy from the origin and keeps it, and for as long as that copy is allowed to live, every later request is answered at the edge without the origin hearing about it. That is why one data center feels fast in Houston and in Frankfurt, and why the site stays up while the origin is slow, sick, or being redeployed. The cat version: the origin is the real cat, and every edge keeps its own copy of the cat for a while. Hold that thought.
We built a platform to run all of this and call it Rabbit. Every repository has a directory called .rabbit with one manifest folder per production domain: the distribution, its origin, the cache rules, the routes, the monitors, all YAML next to the application code. Each domain gets its own branch. Push to the branch and only that domain's infrastructure is planned and applied, from a main branch that fans out to all of them like spokes from a hub. Thirty sites, two codebases, one process. It has run long enough that nobody thinks of it as a rule anymore. It is just how the weather works, and weather can surprise you.
A twenty-six minute favor
At the end of August the client rebranded a venue. New name, new domain, new site, launched on a Monday. At 5:26 PM our project lead typed the kind of request that ends a launch day: can someone please redirect the old domain to the new one?
The old domain still served a WordPress site on Rabbit, behind its own distribution, described by its own manifest. On this platform that redirect is one value in one file. Every distribution tells the edge which domain is preferred, and a small function at the edge answers every other hostname with a 301 to it. Change the preferred domain, open a pull request, promote, done.
That is not what happened. The engineer who picked it up went to the console. His first try, at 5:42, was a CloudFront Function bolted onto the distribution that answered every request with a 302. For seven minutes the old domain served a mix of 302s and errors, which is when the project lead reported that the site "doesn't support a secure connection" and someone suggested propagation. At 5:49 he reached for a pattern from an earlier era, the way you might reach for the flip phone still in a kitchen drawer: an S3 bucket that redirects everything to the new domain, added as a second origin, with the default behavior pointed at it instead of WordPress. At 5:50 he blanked the root object and cleared the edge cache, twice. Four console edits in eight minutes, three redirect implementations, and at 5:52 a message that said, in effect, check it now.
It worked. I have the edge logs. The homepage went from 200s to 301s that evening and stayed that way for two weeks. Launch day ended. Everyone went to dinner.
Twenty-six minutes on launch day
| What the edge logs say about the old domain's homepage | |
|---|---|
| Days redirecting to the new domain, after the console edit | 16 |
| Days serving the dead WordPress page again, after the revert | 17 |
| Browser loads of the dead homepage during those 17 days | 7,213 |
| Distinct devices behind those loads | about 4,600 |
| Of which arrived from the venue's Instagram bio link | 705 |
| People who noticed before the client's CEO did | 0 |
| Time to fix it properly once someone looked | under an hour |
The most dangerous fix is the one that works
A fix that fails gets looked at. A fix that works removes every reason anyone will ever look again. The thread got a thumbs up. Everyone moved on. An S3 bucket nobody would think about again sat in the account being quietly, perfectly correct.
Meanwhile the manifest still said what it had said for a year: WordPress origin, WordPress root object, no bucket. The repository and the cloud now disagreed, and nothing in our process was built to notice. Sixteen days is long enough for a workaround to become the truth in everyone's head. It is not long enough for the YAML to forget.
The deploy that broke it was about something else
On a Tuesday morning in mid September, the same engineer merged a routine promotion into that site's branch. WordPress plugin updates and a workflow tweak, the kind of pull request you merge before your coffee cools. Nothing in it mentioned the old domain, the redirect, or the distribution.
The push triggered the infrastructure apply, as every push does. The apply read the manifest, looked at the live distribution, noticed an origin the manifest had never heard of and a default behavior pointing somewhere it did not sanction, and tidied up. One resource changed, zero errors. Good job, apply.
Four minutes after that merge, the old domain was serving the dead WordPress page again, complete with a stale event listing and a ticket giveaway for a show that had already happened.
One resource, three states
I want to be exact here, because the easy version is wrong. The engineer did not break the redirect and the apply did not malfunction. The system restored the declared state, which is its one job. The redirect was never part of that state. It was a fact about the cloud that the code did not know, and the code was always going to win the next time the two met.
Thirty-three days, zero alarms
Seventeen days of the wrong page. A little over seven thousand homepage loads from about forty-six hundred devices, seven hundred of them people who tapped the link in the venue's Instagram bio and landed on a giveaway for a show that had already happened. The first human to notice was the client's CEO, on a Saturday morning, in an email asking whether we should maybe redirect the old domain to the new one. From where he sat, the request had simply never been done.
Homepage responses per day, browser traffic
We have a synthetic monitor on that domain. It loads the homepage from a few cities every few minutes and checks for a 200. It was green before the redirect, green during it, green after the revert. It would be green today if the page were a photo of a cat. The monitor was measuring the thing we had, not the thing we wanted, and it was extremely confident about it.
And it is not one cat. Each of the monitor's checks lands on a different edge, holding its own copy with its own expiry, filed under the monitor's own browser signature. The monitor was not sampling the site. It was sampling a few of the cats and reporting that the cats were fine.
We also had a nightly scheduled run on that repository, and I had assumed it would catch drift. GitHub runs scheduled workflows only on the default branch, and on Rabbit the default branch is one environment among dozens. The nightly job watched the hub and never looked at any spoke, by design, every night.
The edge logs had the whole story the entire time: every request every edge answered, with status, hostname, path, and whether it served its own copy or fetched from the origin. Nobody was reading them, because nobody had a reason to.
The fix was two lines, and the second one was a surprise
The fix was what it should have been on day one: the preferred domain in the manifest, changed to the new domain. A pull request, two automated reviewers, a human merge, a promotion, an apply, one invalidation. The distribution updated itself through the same path that had reverted it.
I checked the result with a cache-busting request and got a redirect to the wrong URL. The distribution still carried a default root object, a leftover from the days when the site needed one, and CloudFront rewrites the bare path to that object before our edge function ever sees the request. So the homepage redirected to the new domain plus a file name that does not exist there, which answered 404. Every deep path was fine. The one path that mattered was broken.
The logs caught it in two minutes: exactly three requests to the bad URL, one my curl, two my browser. No visitor saw it. A second one-line change blanked the root object, through the same path, and the apply for it reported no changes on the distribution, which is the sound of source control and the cloud finally agreeing.
Two lines of YAML, one of them a surprise
The whole repair, including the detour, took under an hour. The outage it repaired took thirty-three days. That ratio is the part worth sitting with, ideally somewhere without a console.
The edge made all of it harder to see
Now the cats, properly. The edge cache shaped every transition in this story, and it is why "it works for me" is the least useful sentence in web operations.
That distribution holds every page at the edge for at least a day, filed under a key that includes the hostname and the browser's Accept header, which differs between Chrome, Safari, Firefox, and a bot. So every browser family, at every one of several hundred edges, keeps its own copy of the homepage for twenty-four hours. There is no single "the site." There are thousands of copies with thousands of expiry times, and what you see depends on your city and your browser. The design that keeps a site fast and alive while its origin struggles is the design that kept a dead page alive for seventeen days. Availability and staleness are the same feature, seen on different days.
Every edge has its own cat
On launch day this was the first thing the engineer fought. His 302 went live at 5:42, but edges holding a cached 200 kept serving it while the others served the 302, so for eight minutes two people looked at one domain and saw different things. The invalidations at 5:50 and 5:51 made the redirect appear everywhere at once. An invalidation is the one instruction that reaches every edge: throw out your copy, the next visitor fetches a fresh one. That part he did right.
On September 16 nobody invalidated anything, and nothing needed to be. The logs show the busiest edge serving a cached redirect at 9:38 that morning and, three minutes after the apply swapped the origin, fetching the old page again. CloudFront does not promise to drop cached copies when an origin changes, and I would not build on it, but that is what happened. The revert arrived everywhere within the hour, silently, which is worse than creeping in over a day, because a slow revert might have been noticed as something changing.
On October 3 the cache cut the other way. The apply that installed the real redirect changed only a header on the origin. It did not touch the cached pages, and the copy at my edge was already seventeen hours old. Without an invalidation the fix would have surfaced city by city over the next day, and anyone checking in the meantime would have reported, correctly, that it did not work. So the invalidation is part of the fix. On Rabbit the application deploys already invalidate on their own; the infrastructure apply does not, and after this it should, any time a distribution's origin or headers change.
Changing the origin does not change the cats
Then there is the layer past the edge. My broken redirect lived for two minutes and reached exactly two browsers, both mine. Browsers cache a 301 indefinitely, so a browser is one more edge, in your pocket, with no invalidation button. All afternoon my Chrome profile replayed the dead redirect without contacting CloudFront, while an incognito window showed the fix. A browser is the worst instrument for verifying a redirect. Verify at the edge with a request that cannot match a cached copy, then in the logs, where every edge reports what it actually served.
What doing it right looks like at this scale
On any given week one of these domains is the link in a text message going to a hundred thousand phones, the link in an Instagram bio, the URL on a wristband. So "put it in the YAML" is not tidiness. It is the only way the change survives the next deploy, and the next deploy is never more than a plugin update away.
The version that takes about as long as the console edit and does not expire: two lines in the manifest with a comment saying why, a pull request, a promotion, and an apply that keeps making the distribution match on every push forever. Invalidate once. Then do not trust it: pull the access logs for the next hour and count status codes per hostname. If the homepage is not a 301, you are not done, whatever your browser says.
To put the two pictures together: every spoke below is a branch, every branch produces one live distribution, and that distribution is the origin cat from the map. The console edit changed one origin cat, the edges copied it for sixteen days, then the branch that owns that cat ran again, put the original back, and the edges copied that instead.
One hub, thirty spokes
Then close the two gaps that let this run for thirty-three days. Give the scheduled plan every environment branch, not just the hub, and route a non-empty plan to a human, because a change nobody pushed is a hand edit. And rewrite the monitor on the retired domain to assert the redirect, not a 200.
None of that is more work than the console edit. It is the same twenty-six minutes, pointed at the repository.
The platform did its job, which is the uncomfortable part
I am hard on the process here, so it is fair to say what held. When I needed to know exactly what happened in August, CloudTrail had all of it: the bucket at 5:34, four distribution updates with their full request bodies, two invalidations, every timestamp matching the chat thread to the minute. When I needed to know whether my own bad redirect had hurt anyone, the edge logs answered in one query. The platform wrote all of it down without being asked. The failure was never in the record. It was in nobody reading it until it mattered.
For the people who already knew this
This is the section I would normally soften, and I am not going to.
The engineer who made the console edit has been on this team for years. He wrote some of the YAML that reverted his own change. There was no gap in knowledge. There was a gap between knowing a rule and treating it as binding on a Monday afternoon when the request felt small and the console was right there.
I have made the same trade. Everyone who has run infrastructure long enough has. The console is faster for the first ten minutes and more expensive for every minute after, and the bill arrives on a day you did not choose, in a form you do not recognize, addressed to someone else.
Ten minutes faster, forever slower
So, the rule, restated for a team that already knows it: if a resource has a manifest, the manifest is the only place you change it. Not because the console is forbidden, but because the apply will run again, and it will not ask you first. A hand edit on a declared resource is not a fix. It is a scheduled outage with an unknown date.
The second half of the lesson is mine. I assumed the nightly run was checking every environment for drift. It was not, and I had never verified it. An assumption about monitoring that has never been tested is not monitoring. It is a hope with a cron schedule.
What to take home
Put the redirect in the manifest. If the manifest cannot express it, extend the platform, do not route around it.
Plan every environment on a schedule and put a human in front of the output. Make the monitor assert the outcome you want, not the status code you have. A 200 on a retired site is a failure wearing a green badge.
Treat the edge as what it is: hundreds of copies of your site, each with its own clock. Every change to what the origin serves needs an invalidation, and every check of the result belongs in the logs, not in a browser.
And when a request feels like a twenty-minute favor, that is precisely the moment to open the repository instead of the console. The favor is not the change. The favor is making the change stick.