It Worked for Sixteen Days

A client asked us to redirect a retired domain to its replacement. One of my engineers did it in twenty-six minutes, by hand, in the AWS console, and it worked perfectly for sixteen days. Then a deploy about WordPress plugins put the cloud back the way the code said it should be, and the dead site climbed out of its grave for seventeen more days, until the client's CEO found it on a Saturday morning. Nothing malfunctioned. Every system did precisely what we built it to do. That is the joke, and the lesson, aimed at my own team as hard as at anyone reading this. There are also cats.

The manifest always wins

It worked for sixteen days A manifest document on the left connects by a solid apply line into a larger live resource square. A small console-edit square above drops a thin dashed line down into the live resource from the side. The apply line continues through the live resource and lands on an unrelated deploy tick sixteen days along a timeline; the first sixteen-day span is accent blue and the following seventeen-day span, after the revert, is muted grey. manifest live resource console edit apply unrelated deploy 16 days 17 days
One declared configuration, one live resource, and a hand edit that lives only until the next time the two are reconciled. Illustration by svg-forge, a Claude worker.

The client is a live events promoter. We run around thirty production sites for them, more than forty CloudFront distributions once you count staging: a site for every festival and every year it has run, the venue sites, the corporate site, and the ticketing pipeline behind all of them. Those sites sell millions of dollars of tickets a year and send well over a million text messages a month, each carrying a link to one of these domains. A domain serving the wrong page during an on-sale is not a cosmetic bug. It is a revenue event with a timestamp.

A quick tour of how a page reaches a phone, because the story turns on it. The site that builds the page runs in one place: the origin. Nobody's phone talks to it. In front of it sits a content delivery network, a few hundred small data centers called edge locations, and your phone talks to the closest one. The first time an edge is asked for a page it fetches a copy from the origin and keeps it, and for as long as that copy is allowed to live, every later request is answered at the edge without the origin hearing about it. That is why one data center feels fast in Houston and in Frankfurt, and why the site stays up while the origin is slow, sick, or being redeployed. The cat version: the origin is the real cat, and every edge keeps its own copy of the cat for a while. Hold that thought.

We built a platform to run all of this and call it Rabbit. Every repository has a directory called .rabbit with one manifest folder per production domain: the distribution, its origin, the cache rules, the routes, the monitors, all YAML next to the application code. Each domain gets its own branch. Push to the branch and only that domain's infrastructure is planned and applied, from a main branch that fans out to all of them like spokes from a hub. Thirty sites, two codebases, one process. It has run long enough that nobody thinks of it as a rule anymore. It is just how the weather works, and weather can surprise you.

A twenty-six minute favor

At the end of August the client rebranded a venue. New name, new domain, new site, launched on a Monday. At 5:26 PM our project lead typed the kind of request that ends a launch day: can someone please redirect the old domain to the new one?

The old domain still served a WordPress site on Rabbit, behind its own distribution, described by its own manifest. On this platform that redirect is one value in one file. Every distribution tells the edge which domain is preferred, and a small function at the edge answers every other hostname with a 301 to it. Change the preferred domain, open a pull request, promote, done.

That is not what happened. The engineer who picked it up went to the console. His first try, at 5:42, was a CloudFront Function bolted onto the distribution that answered every request with a 302. For seven minutes the old domain served a mix of 302s and errors, which is when the project lead reported that the site "doesn't support a secure connection" and someone suggested propagation. At 5:49 he reached for a pattern from an earlier era, the way you might reach for the flip phone still in a kitchen drawer: an S3 bucket that redirects everything to the new domain, added as a second origin, with the default behavior pointed at it instead of WordPress. At 5:50 he blanked the root object and cleared the edge cache, twice. Four console edits in eight minutes, three redirect implementations, and at 5:52 a message that said, in effect, check it now.

It worked. I have the edge logs. The homepage went from 200s to 301s that evening and stayed that way for two weeks. Launch day ended. Everyone went to dinner.

Twenty-six minutes on launch day

Twenty-six minutes on launch day A time ruler from 5:20 to 6:00 PM. A request in chat at 5:26, an S3 redirect bucket created at 5:34, four console edits between 5:42 and 5:50, two cache clears at 5:50 and 5:51, and a check-it-now at 5:52, after which the ruler fades to grey. 5:20 5:30 5:40 5:50 6:00 request in chat S3 redirect bucket created console edits check it now cache cleared
Launch day, minute by minute, from the chat thread and CloudTrail. Fast, effective, and invisible to the repository.
What the edge logs say about the old domain's homepage
Days redirecting to the new domain, after the console edit16
Days serving the dead WordPress page again, after the revert17
Browser loads of the dead homepage during those 17 days7,213
Distinct devices behind those loadsabout 4,600
Of which arrived from the venue's Instagram bio link705
People who noticed before the client's CEO did0
Time to fix it properly once someone lookedunder an hour

The most dangerous fix is the one that works

A fix that fails gets looked at. A fix that works removes every reason anyone will ever look again. The thread got a thumbs up. Everyone moved on. An S3 bucket nobody would think about again sat in the account being quietly, perfectly correct.

Meanwhile the manifest still said what it had said for a year: WordPress origin, WordPress root object, no bucket. The repository and the cloud now disagreed, and nothing in our process was built to notice. Sixteen days is long enough for a workaround to become the truth in everyone's head. It is not long enough for the YAML to forget.

The deploy that broke it was about something else

On a Tuesday morning in mid September, the same engineer merged a routine promotion into that site's branch. WordPress plugin updates and a workflow tweak, the kind of pull request you merge before your coffee cools. Nothing in it mentioned the old domain, the redirect, or the distribution.

The push triggered the infrastructure apply, as every push does. The apply read the manifest, looked at the live distribution, noticed an origin the manifest had never heard of and a default behavior pointing somewhere it did not sanction, and tidied up. One resource changed, zero errors. Good job, apply.

Four minutes after that merge, the old domain was serving the dead WordPress page again, complete with a stale event listing and a ticket giveaway for a show that had already happened.

One resource, three states

One resource, three states The same live resource drawn three times. Declared in the manifest it has three component rows. Live after a console edit it gains a fourth row at the top in accent blue, an origin redirect added by hand. Live after the next apply that fourth row is gone, shown as a faint dashed grey ghost, because the apply from the manifest reconciled it away. origin: WordPress root object: index preferred domain: old origin: S3 redirect bucket origin: WordPress root object: index preferred domain: old origin: S3 redirect bucket origin: WordPress root object: index preferred domain: old apply declared in the manifest live, after the console edit (day 1 to 16) live, after the next apply (day 17) day 0 day 1 day 17
The same distribution three times: as the manifest declares it, as it was after the console edit added a redirect origin, and as it was after the next apply removed the thing the manifest never knew about.

I want to be exact here, because the easy version is wrong. The engineer did not break the redirect and the apply did not malfunction. The system restored the declared state, which is its one job. The redirect was never part of that state. It was a fact about the cloud that the code did not know, and the code was always going to win the next time the two met.

Thirty-three days, zero alarms

Seventeen days of the wrong page. A little over seven thousand homepage loads from about forty-six hundred devices, seven hundred of them people who tapped the link in the venue's Instagram bio and landed on a giveaway for a show that had already happened. The first human to notice was the client's CEO, on a Saturday morning, in an email asking whether we should maybe redirect the old domain to the new one. From where he sat, the request had simply never been done.

Homepage responses per day, browser traffic

Homepage responses per day, browser traffic Stacked daily counts of homepage responses over 41 days. First the old WordPress page is served as 200s (Aug 24 to 30, peak 2826 on Aug 27). After a console edit on Aug 31, the page answers 301 redirects for sixteen days (Sep 1 to 15, peak 1151 on Sep 1). After an unrelated deploy on Sep 16 the redirect is gone and the old page returns as 200s for seventeen days (Sep 16 to Oct 2), until a fix on Oct 3. 1k 2k 3k Aug 24 Aug 31 Sep 16 Oct 3 console edit unrelated deploy CEO email, fix 16 days redirecting 17 days, old page again 200, old page served 301, redirected
Daily homepage responses for the old domain from the CDN access logs, bots excluded. Grey is the old page served, blue is the redirect. Nothing in this chart needed a person to produce it. It needed a person to look.

We have a synthetic monitor on that domain. It loads the homepage from a few cities every few minutes and checks for a 200. It was green before the redirect, green during it, green after the revert. It would be green today if the page were a photo of a cat. The monitor was measuring the thing we had, not the thing we wanted, and it was extremely confident about it.

And it is not one cat. Each of the monitor's checks lands on a different edge, holding its own copy with its own expiry, filed under the monitor's own browser signature. The monitor was not sampling the site. It was sampling a few of the cats and reporting that the cats were fine.

We also had a nightly scheduled run on that repository, and I had assumed it would catch drift. GitHub runs scheduled workflows only on the default branch, and on Rabbit the default branch is one environment among dozens. The nightly job watched the hub and never looked at any spoke, by design, every night.

The edge logs had the whole story the entire time: every request every edge answered, with status, hostname, path, and whether it served its own copy or fetched from the origin. Nobody was reading them, because nobody had a reason to.

The fix was two lines, and the second one was a surprise

The fix was what it should have been on day one: the preferred domain in the manifest, changed to the new domain. A pull request, two automated reviewers, a human merge, a promotion, an apply, one invalidation. The distribution updated itself through the same path that had reverted it.

I checked the result with a cache-busting request and got a redirect to the wrong URL. The distribution still carried a default root object, a leftover from the days when the site needed one, and CloudFront rewrites the bare path to that object before our edge function ever sees the request. So the homepage redirected to the new domain plus a file name that does not exist there, which answered 404. Every deep path was fine. The one path that mattered was broken.

The logs caught it in two minutes: exactly three requests to the bad URL, one my curl, two my browser. No visitor saw it. A second one-line change blanked the root object, through the same path, and the apply for it reported no changes on the distribution, which is the sound of source control and the cloud finally agreeing.

Two lines of YAML, one of them a surprise

Two lines of YAML, one of them a surprise Two request paths compared. Before: GET slash passes through a root-object rewrite to default-root-object.php, an edge redirect appends that path, and the result is a 404. After: the root-object rewrite box is a removed dashed ghost, the redirect is clean, and the result is a 200. before after GET / edge redirect root object rewrite /default-root-object.php 301 to new domain/default-root-object.php 404 GET / edge redirect root object rewrite /default-root-object.php 301 to new domain/ 200
The first fix sent the homepage to a file that does not exist on the new site. The edge rewrites the bare path to the distribution's default root object before the redirect runs, so the second line of YAML blanks it.

The whole repair, including the detour, took under an hour. The outage it repaired took thirty-three days. That ratio is the part worth sitting with, ideally somewhere without a console.

The edge made all of it harder to see

Now the cats, properly. The edge cache shaped every transition in this story, and it is why "it works for me" is the least useful sentence in web operations.

That distribution holds every page at the edge for at least a day, filed under a key that includes the hostname and the browser's Accept header, which differs between Chrome, Safari, Firefox, and a bot. So every browser family, at every one of several hundred edges, keeps its own copy of the homepage for twenty-four hours. There is no single "the site." There are thousands of copies with thousands of expiry times, and what you see depends on your city and your browser. The design that keeps a site fast and alive while its origin struggles is the design that kept a dead page alive for seventeen days. Availability and staleness are the same feature, seen on different days.

Every edge has its own cat

Every edge has its own cat A low-poly map of the contiguous United States. Six edge locations sit on it, each a small square holding a cat: Dallas (grey, old page), Houston (blue, redirect), Virginia (grey), San Francisco (grey), New York (grey) and Chicago (blue). Dallas and Houston are close neighbours yet hold different cats, so Chrome in Dallas sees a grey cat while Safari in Houston sees a blue one. A monitor off the Atlantic coast checks Virginia, New York and Chicago and finds three grey cats, so it passes. The grey old-page cats look confused, with crossed eyes and a lolling tongue, while the blue redirect cats look composed, so wrong versus right reads from the faces alone. San Francisco Dallas Houston Chicago Virginia New York Chrome in Dallas sees a grey cat Safari in Houston sees a blue cat monitor three grey cats: pass redirect cached old page cached nothing cached yet
Six of the several hundred edge locations, each holding its own cat for up to a day. Grey cats hold the stale page, which is why they look a little lost. Dallas and Houston are 240 miles apart and see different cats. The monitor checks three cities, finds three grey cats, and reports that cats are fine.

On launch day this was the first thing the engineer fought. His 302 went live at 5:42, but edges holding a cached 200 kept serving it while the others served the 302, so for eight minutes two people looked at one domain and saw different things. The invalidations at 5:50 and 5:51 made the redirect appear everywhere at once. An invalidation is the one instruction that reaches every edge: throw out your copy, the next visitor fetches a fresh one. That part he did right.

On September 16 nobody invalidated anything, and nothing needed to be. The logs show the busiest edge serving a cached redirect at 9:38 that morning and, three minutes after the apply swapped the origin, fetching the old page again. CloudFront does not promise to drop cached copies when an origin changes, and I would not build on it, but that is what happened. The revert arrived everywhere within the hour, silently, which is worse than creeping in over a day, because a slow revert might have been noticed as something changing.

On October 3 the cache cut the other way. The apply that installed the real redirect changed only a header on the origin. It did not touch the cached pages, and the copy at my edge was already seventeen hours old. Without an invalidation the fix would have surfaced city by city over the next day, and anyone checking in the meantime would have reported, correctly, that it did not work. So the invalidation is part of the fix. On Rabbit the application deploys already invalidate on their own; the infrastructure apply does not, and after this it should, any time a distribution's origin or headers change.

Changing the origin does not change the cats

Changing the origin does not change the cats Three panels. Before: a grey origin cat feeds five edge locations that each hold a grey cat with its own timer. After the apply changes the origin, the origin cat turns blue but all five edge cats stay grey with the same timers. After an invalidation, every edge cat is thrown out, shown as dashed empty outlines, and the next visitor to Houston fetches a fresh blue cat from the origin. The grey old-page cats look confused, with crossed eyes and a lolling tongue, while the blue redirect cats look composed, so wrong versus right reads from the faces alone. Dallas Houston Virginia LA Frankfurt Dallas Houston Virginia LA Frankfurt Dallas Houston Virginia LA Frankfurt 22h left 9h left 17h left 3h left 14h left 22h left 9h left 17h left 3h left 14h left — next visitor — — — changed before after the apply changes the origin after an invalidation every edge keeps its cat until the timer runs out every cat thrown out at once; the next visitor fetches a fresh one no invalidation: the fix arrives city by city as timers expire, up to a day later
An apply changes what the origin serves. The cats already sitting at the edges stay exactly as they were until their timers run out. An invalidation throws every cat out at once, so the next visitor at each edge fetches a fresh one.

Then there is the layer past the edge. My broken redirect lived for two minutes and reached exactly two browsers, both mine. Browsers cache a 301 indefinitely, so a browser is one more edge, in your pocket, with no invalidation button. All afternoon my Chrome profile replayed the dead redirect without contacting CloudFront, while an incognito window showed the fix. A browser is the worst instrument for verifying a redirect. Verify at the edge with a request that cannot match a cached copy, then in the logs, where every edge reports what it actually served.

What doing it right looks like at this scale

On any given week one of these domains is the link in a text message going to a hundred thousand phones, the link in an Instagram bio, the URL on a wristband. So "put it in the YAML" is not tidiness. It is the only way the change survives the next deploy, and the next deploy is never more than a plugin update away.

The version that takes about as long as the console edit and does not expire: two lines in the manifest with a comment saying why, a pull request, a promotion, and an apply that keeps making the distribution match on every push forever. Invalidate once. Then do not trust it: pull the access logs for the next hour and count status codes per hostname. If the homepage is not a 301, you are not done, whatever your browser says.

To put the two pictures together: every spoke below is a branch, every branch produces one live distribution, and that distribution is the origin cat from the map. The console edit changed one origin cat, the edges copied it for sixteen days, then the branch that owns that cat ran again, put the original back, and the edges copied that instead.

One hub, thirty spokes

One hub, two dozen spokes A main-branch hub at the centre, holding a document glyph, fans promotions outward along thin navy lines to twelve generic site squares arranged on an ellipse, each holding three stacked marks. The lower-right spoke, site_l, is the same navy as the rest. Outside the ring, below and right of site_l, a blue "live" square holds the origin cat in navy outline and is produced from the spoke by a navy "apply" arrow; a dashed blue "console edit" line enters only this live square, from the right, never touching the spoke, because a console edit changes the live distribution and the spoke only finds out at its next apply, when it wins. main branch sitea siteb sitec sited sitee sitef siteg siteh sitei sitej sitek sitel apply console edit live distribution (the origin cat) the spoke only finds out at its next apply, and then it wins
One main branch promoted out to a production branch per domain, each with its own manifest directory, each producing one live distribution: the origin cat the edges copy. A console edit touches that one live cat and none of the branches. The branch only finds out at its next apply, and then it wins. A scheduled job on the hub sees none of the spokes.

Then close the two gaps that let this run for thirty-three days. Give the scheduled plan every environment branch, not just the hub, and route a non-empty plan to a human, because a change nobody pushed is a hand edit. And rewrite the monitor on the retired domain to assert the redirect, not a 200.

None of that is more work than the console edit. It is the same twenty-six minutes, pointed at the repository.

The platform did its job, which is the uncomfortable part

I am hard on the process here, so it is fair to say what held. When I needed to know exactly what happened in August, CloudTrail had all of it: the bucket at 5:34, four distribution updates with their full request bodies, two invalidations, every timestamp matching the chat thread to the minute. When I needed to know whether my own bad redirect had hurt anyone, the edge logs answered in one query. The platform wrote all of it down without being asked. The failure was never in the record. It was in nobody reading it until it mattered.

For the people who already knew this

This is the section I would normally soften, and I am not going to.

The engineer who made the console edit has been on this team for years. He wrote some of the YAML that reverted his own change. There was no gap in knowledge. There was a gap between knowing a rule and treating it as binding on a Monday afternoon when the request felt small and the console was right there.

I have made the same trade. Everyone who has run infrastructure long enough has. The console is faster for the first ten minutes and more expensive for every minute after, and the bill arrives on a day you did not choose, in a form you do not recognize, addressed to someone else.

Ten minutes faster, forever slower

Ten minutes faster, forever slower The request forks into two paths. The console path is a short blue segment that reaches a works square quickly, then a thin blue line that ends abruptly at a next-apply bar with nothing after it. The repository path is a longer navy segment through pull request, review and apply before its own works square, then a solid navy line to the edge with ticks for every apply. the request console repository works works next apply pull request review apply every apply minutes months
Both paths reach a working redirect. Only one of them is still working after the next apply.

So, the rule, restated for a team that already knows it: if a resource has a manifest, the manifest is the only place you change it. Not because the console is forbidden, but because the apply will run again, and it will not ask you first. A hand edit on a declared resource is not a fix. It is a scheduled outage with an unknown date.

The second half of the lesson is mine. I assumed the nightly run was checking every environment for drift. It was not, and I had never verified it. An assumption about monitoring that has never been tested is not monitoring. It is a hope with a cron schedule.

What to take home

Put the redirect in the manifest. If the manifest cannot express it, extend the platform, do not route around it.

Plan every environment on a schedule and put a human in front of the output. Make the monitor assert the outcome you want, not the status code you have. A 200 on a retired site is a failure wearing a green badge.

Treat the edge as what it is: hundreds of copies of your site, each with its own clock. Every change to what the origin serves needs an invalidation, and every check of the result belongs in the logs, not in a browser.

And when a request feels like a twenty-minute favor, that is precisely the moment to open the repository instead of the console. The favor is not the change. The favor is making the change stick.