Hands-On Labs · Guided Drills

Drill — Write a Rollback Plan

Everybody agrees a cutover needs a way back. Almost nobody writes one that survives contact with 04:12 in the morning — because the plan gets written for a version of the night that has already passed by the time anybody needs it. This drill hands you three cutovers that have already gone wrong, at three deliberately different moments. In the first, reverting is still clean and cheap, and the only real question is whether you call it fast enough. In the second you are forty minutes past the point of no return, and “roll back” now means deciding what happens to work that exists in exactly one place. In the third, the rollback itself is what breaks: the old system was quietly dismantled three weeks ago and nobody re-read the runbook. Three different fictional companies, three different sectors, no continuity to carry — nothing here comes from the Brambleside capstone, and nothing you write here feeds it. Give yourself 20 minutes per scenario, and write your plan down before you open its model answer. A rollback plan you thought about but never wrote is not a rollback plan, which is the entire thesis of this page.

☺ Explain it like I’m 10

Moving day. At two in the afternoon the van is half-loaded and you discover the new house has no hot water — you can still ring the driver, turn the van round, and sleep in your old bed tonight. Annoying, cheap, done. At six in the evening the beds are built, the baby is asleep in one of them, and three boxes have been unpacked into cupboards — “going back” now doesn’t mean driving a van, it means waking the baby and re-packing things you can no longer find. And on the third day the landlord has re-let the old house to somebody else, so there is no back to go to at all, however much you would like one. Same house, same van, same family: three completely different plans. The thing that decides which plan you are in isn’t how you feel about it. It’s what time it is.

🐢👺Your hosts for this drill: Timmy the Turtle & Gizmo the Gremlin — Timmy tests the way back before he needs it, times it with a stopwatch, and writes an expiry date on it. Gizmo leans over your shoulder at 04:12 whispering “we’re nearly there, just push through, it’d be a shame to waste the whole weekend.” One of these two is on every real cutover call. On the bad ones, only one of them is talking.
⚠ Before you start — what you need, and what you don’t

You need a text editor and a clock. That is the entire toolchain. There is no cloud account to open, no card to enter, and nothing to provision, break or tear down — rollback plans are written in documents, at desks, in daylight, which is exactly why the ones written at 04:12 are so bad. The three companies below are invented, and so is every figure, timing, price and volume attached to them; they exist to make the arithmetic bite, and none of it should ever be quoted at anyone as a real number. Each scenario is fully independent — different company, different sector, different failure — so you can do one on a lunch break and come back for the others.

How this drill works

☺ Like you’re 10: Three nights that have already gone wrong. You’re the person everyone on the call is waiting on. Write what you’d do — properly, in sentences — then check it.

Each scenario hands you four things: the estate (what the systems are, what depends on what, how big, how critical), the runbook exactly as it was written, the timeline of the night up to the moment it goes wrong, and the precise minute you are standing in. You are the incident lead on the bridge call. For each one, produce the same five answers, in this order:

ROLLBACK DECISION — <system> · <date> · written at <time>

1. THE CALL
   Decision owner:      <one named human, one role — not "the team">
   Deputy:              <the second name, for when the first is unreachable>
   Decision deadline:   <wall-clock minute>
   Derived how:         <the subtraction that produced that minute>
   Chosen path:         REVERT | MITIGATE FORWARD | HOLD, RE-DECIDE AT <time>

2. THE DATA, AT THIS MINUTE
   Exists only on the new system:   <what, how much, since when>
   Exists only on the old system:   <what>
   In flight / queued / physical:   <what is half-done in the real world>
   Reverse path exists?             <yes/no — and has it ever actually been run?>
   Rollback RPO / RTO:              <data lost going back / time to get back>

3. THE STEPS
   <numbered, with a wall-clock time against each, first step first>

4. THE COMMS
   Audience | Told what | By whom | On which channel | When
   <one row per audience — including the ones who are not on the call>

5. THE LINE THAT WAS MISSING
   <the single sentence that should have been in the runbook before
    the night started, and would have made all of this straightforward>

Section 5 is the one that teaches. All three nights below are survivable. Two of them are only painful because of a sentence nobody wrote down in daylight, when writing it would have cost four minutes. That is the skill this drill is actually drilling: not heroics at 04:12, but the four minutes at 11:00 on a Tuesday.

The six words this drill uses precisely

☺ Like you’re 10: Six words people use loosely on a good day and disastrously on a bad one. Pin them down before you start.

TermWhat it means, exactlyWhat goes wrong when it stays vague
RevertAbandon the new system; the old one resumes as the system of record. Available only while the old system is still capable of being authoritative and nothing of value exists solely in the new one.People say “roll back” meaning revert long after revert stopped existing, then burn twenty minutes of a four-hour window discovering that on the call.
Mitigate forward (fix forward)Stay on the new system and repair it in place. The only road once unique data lives in the new place, or the old place is gone.Treated as the coward’s option, so teams attempt a revert they cannot finish, and end up mitigating forward anyway — two hours later, with a worse data state and a tireder team.
Point of no return (PONR)The first moment at which reverting would destroy or orphan data that exists only on the new system. It is a moment, and it belongs in the runbook as a labelled line with a time against it.It gets crossed by accident — often minutes after a green go/no-go gate — and nobody notices until somebody tries to go back.
Decision ownerOne named human with standing authority to stop the cutover without asking anyone. Named in the runbook, awake, on the call, with a named deputy.“We’ll decide together.” At 04:12, six tired people with equal authority produce delay — and delay is itself a decision, usually the wrong one.
Decision deadlineThe latest wall-clock minute at which the decision can still be made and the chosen path still fit. Computed backwards from the first hard downstream commitment, using measured durations.Everyone anchors on the end of the maintenance window, which is almost never the real deadline — in Scenario 1 the two differ by two and a half hours.
Rollback RPO / RTOHow much data you accept losing on the way back, and how long the way back takes. Both measured in a rehearsal, not estimated on the night.“About an hour” turns out to be three — discovered at the point where you have already spent the hour you did not have.
◆ Key idea

A rollback plan is not one plan. It is one plan per phase of the night, and the phases are divided by the point of no return. Before it: revert, and the only hard question is timing. After it: mitigate forward, and the hard question is what happens to the work that exists in exactly one place. If your runbook has a single section headed “Rollback”, you have written one plan for a night that has at least two — and you will find out which one you are in at the worst possible moment.

Scenario 1 — Kestrel Freight: the call you can still make

☺ Like you’re 10: The van can still turn round. Everyone knows it can. The only question is whether anybody says so before it stops being true.

Kestrel Freight is a regional parcel carrier: 14 depots, 420 handheld scanners, roughly 180,000 parcel events a day. Tonight it is rehosting TrackPoint, its shipment-tracking system, and replatforming TrackPoint’s PostgreSQL database onto a managed database service. Logical replication was seeded six days ago and has been running with about two seconds of lag ever since.

SystemWhat it doesRuns onDataWho depends on it
TrackPoint (app)Shipment tracking, depot scanning, ETA calculation3 app VMsEvery depot, all day; the scanner fleet writes into it continuously
TrackPoint DBSystem of record for every parcel eventPostgreSQL 13, primary + standby1.4 TBEverything below reads from it
Scanner fleet420 handhelds across 14 depotsStore-and-forward capable — they queue scans locally and flush when they reconnect
RouteCalcOvernight route planning for the morning trunk runs1 VM, reads a TrackPoint replica40 GBPlanners. Runs 02:00–03:00.
CustomerPortalPublic “where is my parcel” page2 VMs, read-only against TrackPointThe public, 24×7. Reads only — it writes nothing.
EDI-OutNightly manifest files to three large retail customers over SFTP1 VMContractual: manifests must be lodged by 03:00 daily. The run takes 45 minutes.
ConstraintThe number
Maintenance windowSaturday 22:00 → 04:00
Next hard operational eventThe Sunday e-commerce sort begins at 05:30 across all 14 depots
EDI-Out contractual lodgementAll three manifests lodged by 03:00; the run itself takes 45 minutes
Measured revert time45 minutes — rehearsed twice, timed both times: re-enable the old app, restart RouteCalc, re-enable EDI-Out, point the scanner fleet back and let it flush
Scan-write latency baselinep95 260 ms on the existing on-prem database

The runbook, exactly as written

KESTREL FREIGHT · TrackPoint cutover · Saturday

T-7d          Logical replication running; lag verified < 5 s daily
22:00         Comms: depot managers notified; status page updated
22:15         Freeze: EDI-Out disabled, RouteCalc stopped
22:30         Scanner fleet placed into store-and-forward mode
23:00         Application freeze on old TrackPoint; drain in-flight writes
23:10         Replication lag to zero; confirm; stop replication
23:20         Promote managed database; repoint app configuration
23:40         Smoke test: 12 scripted scans across 3 depots
00:10         Validation queries (row counts, last-event timestamps)
00:40         GO / NO-GO
01:00         Scanner fleet repointed to new endpoint; store-and-forward flushed
02:00         EDI-Out re-enabled; RouteCalc restarted
04:00         Window ends

ROLLBACK:     If NO-GO, revert to the old environment.

Read that rollback section again. It is one line, it names nobody, it contains no time, and it does not say when it stops being available. It is also, word for word, what most real runbooks say.

The night, up to the minute you are standing in

TimeWhat happened
22:00Comms sent, status page updated. On schedule.
22:15EDI-Out disabled, RouteCalc stopped.
22:30All 420 scanners confirmed in store-and-forward mode.
23:00Application freeze on old TrackPoint; in-flight writes drained cleanly.
23:14Replication lag reaches zero. Replication stopped.
23:22Managed database promoted; app configuration repointed. Textbook.
23:41Smoke test 1: 11 of 12 scripted scans pass. The twelfth times out.
00:05Smoke test 2: 9 of 12 pass. Scan-write latency p95 = 4.2 s, against a 260 ms baseline.
00:20The DBA finds it: the index-rebuild script was generated from a schema dump taken in February. Two indexes added in March — parcel_event(consignment_id, event_ts) and parcel_event(depot_id, event_ts) — do not exist on the new database at all.
00:35Estimate to fix: 35–50 minutes to build both indexes on a 1.4 TB table, and the build cannot be safely interrupted once started. Add ~15 minutes to re-run validation afterwards. Realistically the build cannot start before 00:40, so the earliest honest re-decision is somewhere between 01:30 and 01:45.
00:36Gizmo, on the call: “The window’s open till four. We’ve got hours.”
⚠ It is 00:36. Everyone is looking at you.

Write the five-section decision for Kestrel. Before you do, do two pieces of arithmetic in the open, because the whole scenario turns on them. First: a driver scans about 200 parcels in a shift. At 260 ms that is around 52 seconds of waiting across the whole shift; at 4.2 s it is fourteen minutes per driver, per shift — across 420 scanners, roughly 98 hours of paid standing-still every day, and every depot’s sort running late. Decide whether that is a defect or an inconvenience before you decide whether to push through it. Second: find the real deadline. It is not 04:00.

Done when: your decision names one human and one deputy; your decision deadline is a wall-clock minute you produced by subtracting measured durations from the first hard downstream commitment, and you can show that subtraction; and Section 5 contains the single runbook line that would have made 00:36 a thirty-second conversation instead of a debate.

Model answer — Scenario 1: the decision, the arithmetic, and the missing line
  1. The real deadline is 01:30, not 04:00 — and that is the whole scenario. Work backwards from the first hard commitment, which is EDI-Out’s contractual 03:00 lodgement, not the end of the window. EDI-Out takes 45 minutes, so it must start by 02:15. A revert takes a measured 45 minutes, so a revert must start by 01:30 for EDI-Out to make its slot. The window closing at 04:00 is real but irrelevant: it is the deadline for finishing, and by the time it binds, three retail customers have already missed their manifests. Gizmo is not lying about the clock. He is pointing at the wrong clock.
  2. Now price the push-forward path against that line, and watch what happens. The index build cannot start before 00:40. Take the pessimistic end of the 35–50 minute estimate, because you are estimating at half past midnight on a night that has already surprised you once: the build finishes at 01:30 — which is exactly the last minute a revert could have started. Validation then runs to about 01:45, and if it fails, a revert begun at 01:45 completes at 02:30, EDI-Out starts at 02:30 and finishes at 03:15, and three contracts are breached. So pushing on does not “use up spare time”: it spends the entire rollback option, down to the minute, on a fix nobody has rehearsed. That is the trade, stated plainly — and once it is stated plainly, nobody on the call argues for it.
  3. The call: NO-GO, made at 00:40 — at the gate the runbook already had — not at 01:30. The 01:30 line is the last legal minute, not the target; treating a deadline as a plan is how you arrive at it with nothing in hand. And note how cheap the revert is at 00:40: the old TrackPoint is untouched and still authoritative, the scanners have never been repointed and are still queueing locally, and not one parcel event exists only on the new database. Rollback RPO is zero. Rollback RTO is 45 minutes, measured twice. The entire cost of turning round tonight is one wasted weekend and a slightly deflated team — which is a rounding error next to 420 scanners at 4.2 seconds on a Monday.
  4. Is there a defensible push? Narrowly, yes — and it is worth writing down, because “always revert” is as thoughtless as “always push”. If the DBA can show, from the actual query plan, that the slow path is served by one of the two missing indexes, and that one index builds in twelve minutes rather than fifty, then a bounded push is legitimate: start at 00:40, done by 00:52, revalidate by 01:07, and you still hold the 01:30 revert line with 23 minutes to spare. The difference between that and Gizmo’s version is not courage. It is that the bounded push is measured against the 01:30 line and announced against it out loud: “we are spending 27 of our 54 remaining minutes; if we are not green by 01:07 we revert without further discussion.” Say the abort condition before you start, or you will not say it at all.
  5. The comms. Note who does not need telling — that is the point of reverting early.
    AudienceTold whatBy whomChannelWhen
    14 depot managersCutover reverted; the Sunday sort runs on the existing system exactly as normal; no action for your teams; scanners flush automatically when they reconnectCutover LeadDepot managers’ group + email, both pre-drafted before the night01:05 — hours before the 05:00 sort brief
    Customer service duty leadPublic tracking page unaffected throughout; no customer-facing comms neededOps on-callPhone01:10
    Three retail EDI customersNothing at all — the manifest lodges on time. That silence is the deliverable.
    Programme sponsorNO-GO called at 00:40; the reason in one sentence; the cost; the proposed new dateCutover LeadEmail, waiting when they wakeby 02:00
    The cutover teamNew date, plus the two rehearsal items that come out of thisCutover LeadRetro, MondayMonday
  6. The lines that were missing. Three, and all three are one sentence each:
    01:00  POINT OF NO RETURN — scanner fleet repointed and store-and-forward
           flushed. From this minute parcel events exist ONLY on the new
           database. Revert is unavailable after this line.
    
    01:30  LAST REVERT START. Revert is measured at 45 min; EDI-Out needs
           45 min and must lodge by 03:00, so it must start by 02:15.
           After 01:30 a revert cannot complete in time and the path is
           mitigate-forward only.
           Decision owner: A. Rahim, Cutover Lead.  Deputy: J. Okafor, Ops Mgr.
    
    00:40  GO/NO-GO — GO requires 12/12 smoke scans AND p95 scan-write
           < 600 ms. Anything less is a NO-GO. No discussion at the gate.
    
    That last one matters as much as the other two: the runbook had a go/no-go gate, and it was useless, because a gate with no pass threshold and no owner is not a gate — it is a moment where everyone looks at each other. The threshold has to be written when you are calm, precisely so that it can be applied when you are not.
  7. And the item for the retro: the index-rebuild script was generated from a February schema dump. The fix is not “be more careful”; it is a pre-cutover step that diffs the schema of source and target and fails the rehearsal if they differ. One command, run a week earlier, in daylight. Every scenario in this drill has one of these — a cheap check, available days in advance, that nobody ran because nothing in the process demanded it.

Scenario 2 — Pemberton Tools: forty minutes of writes

☺ Like you’re 10: The baby is asleep in the new bed and three boxes are already unpacked. You can’t just drive back.

Pemberton Tools is a builders’ and trade merchant: nine branches with trade counters, a web store, and a 60-seat call centre. Tonight it is replatforming OrderDesk — order capture, contract pricing, credit limits — from an on-prem SQL Server 2017 onto a managed database, with the app VMs rehosted alongside it.

SystemWhat it doesRuns onDataWho depends on it
OrderDeskOrder capture, contract pricing, credit limits2 app VMs + SQL Server 2017900 GB60 call-centre agents (8 on Sunday-night duty), 9 trade counters, the web store
PriceEngineContract pricing for 2,300 trade accounts. Called synchronously by OrderDesk on every order line.1 VM + its own database12 GBOrderDesk, on every single line item
WMS-LinkPushes confirmed orders to the warehouse management system every 5 minutes1 small VMThe warehouse night shift, who pick from 04:00 for the 06:00 dispatch
WebStorePublic storefront; checks out through OrderDeskManaged hostingThe public. About 30 orders/hour on a Sunday night.
PaymentsRedirect to the processor’s hosted pageNo cardholder data touches Pemberton systems, which is the only genuinely good news in this scenario
◆ One detail that decides everything below

Sunday 04:00–06:00 is the busiest ordering window of Pemberton’s week. Trade customers phone in the stock they need on their vans for Monday morning, and it has to be picked for the 06:00 dispatch. This is not a quiet slot into which the team released a few cautious users. It is the peak, and the cutover was scheduled to finish just before it — a completely normal, completely reasonable-looking plan, right up until something goes wrong at 04:20.

The runbook, exactly as written

PEMBERTON TOOLS · OrderDesk cutover · Sunday 01:00–05:00

01:00   Comms; web store into maintenance mode
01:30   Freeze old OrderDesk; final transaction-log ship; replication stopped
02:10   Restore verified on the managed instance; row counts matched
02:40   App VMs repointed; internal smoke test (5 scripted orders)
03:30   DNS flipped (TTL 300 s)
03:45   GO / NO-GO
04:00   Web store out of maintenance; duty agents released
05:00   Window ends

ROLLBACK: if NO-GO at 03:45, flip DNS back and re-enable old OrderDesk.

The night, up to the minute you are standing in

TimeWhat happened
01:30Old OrderDesk frozen. Final log ship completed, replication stopped. From this minute the old database is a still photograph.
02:10Restore verified, row counts matched exactly.
03:30DNS flipped. New OrderDesk live.
03:45GO. All 5 scripted smoke orders passed.
03:50Web store taken out of maintenance early — “since it all looks fine”. The first real customer order writes to the new database.
04:00The 8 duty agents released. They start taking Monday van-stock orders immediately; it is the busiest hour of their week.
04:05WMS-Link resumes. Confirmed orders begin flowing to the warehouse night shift, who start picking.
04:20A duty agent queries a price: a long-standing customer’s usual contract rate is not applying.
04:32Cause found. The managed instance was created with a case-sensitive collation. PriceEngine’s lookup joins on account_code, and 214 of the 2,300 trade accounts have lower-case codes. Those lookups now match nothing, and OrderDesk silently falls back to list price — no error, no exception, no alert, no log line. Customers are being overcharged.
04:35The count. 147 orders since 03:50: 121 taken by the duty agents, 26 through the web store. 38 are mispriced, overcharged by a combined $9,340 — a far higher share than the 9.3% of accounts affected, because the lower-case codes all came from a 2019 bulk import of a merged competitor’s account book, and those are disproportionately the van-stock customers who order at four on a Sunday morning. 96 have already been pushed to WMS. 31 are already physically picked and sitting in cages for the 06:00 dispatch.
Fact you need, that nobody says out loudThe consequence
No reverse replication path was ever built (managed database → old SQL Server), and none has ever been tested.“Revert” means the 147 orders exist nowhere afterwards. There is no mechanism to carry them back.
The old SQL Server has been frozen since 01:30.It is internally consistent — and two and a half hours out of date.
31 orders are physically picked: cardboard in a cage with a pick note attached.Reverting a database does not un-pick a pallet. This is the only genuinely irreversible thing in the building tonight.
The window ends at 05:00; branches open Monday at 07:00.Neither is the real deadline. The real deadline is the 06:00 dispatch, 85 minutes away.
⚠ It is 04:35.

Write the five-section decision. Two things to get right before you write a single step. First, name the actual point of no return — it is not 03:45, and the five-minute gap between the two is the whole lesson. Second, your first step is not the fix. Almost everyone’s instinct at 04:35 is to correct the collation, because that is the interesting problem. Resist it, and notice what you would be doing to the count of affected orders while you worked.

Done when: your plan states the minute the point of no return was actually crossed and why; your first step stops the situation getting bigger rather than fixing the bug; all 147 orders are accounted for in writing, including the 96 in WMS and the 31 in cages; and you have said what the 8 duty agents, the warehouse night-shift supervisor, and the 38 overcharged customers are each told, by whom, and when.

Model answer — Scenario 2: the point of no return, and the mitigate-forward plan
  1. The point of no return was 03:50, not 03:45. The go/no-go gate and the point of no return were five minutes apart, and only one of them was written in the runbook. At 03:45 you could still have flipped DNS back and lost nothing at all. At 03:50, the first web-store order landed in a database that has no path back to the old one — and from that minute, revert stopped meaning “undo the cutover” and started meaning “undo the cutover and destroy every order taken since”. Nobody made that decision. It made itself, five minutes after a green gate, because somebody released the web store early on the grounds that things looked fine. Things did look fine. Things look fine for exactly as long as it takes a lower-case account code to come round.
  2. The call: MITIGATE FORWARD. It is not a preference, it is arithmetic. 147 orders exist only on the new system; 96 of them are already in the warehouse system; 31 of those are physically picked; and no reverse path was ever built or tested. So “revert” is really “revert, plus an unbuilt, unrehearsed data-reconstruction project, executed at 05:00 on a Sunday by people who have been awake since Saturday morning, against a 06:00 dispatch”. Rollback RPO would be 147 orders; rollback RTO would be twenty minutes of DNS and app work plus an unbounded reconciliation. Meanwhile the actual defect is small, understood, and has a narrow fix. Decision owner: the Operations Director (the person accountable for the dispatch), with the Head of Trade Sales as deputy. Decision deadline: 05:25 — the latest minute at which corrected pick notes can reach the warehouse and still make the 06:00 van. Derived by subtracting 35 minutes of physical re-pick and paperwork from 06:00.
  3. The trap answer, worth naming so you can recognise it in yourself: revert the application and hand-replay the 147 orders into the old system. It feels like the best of both. It is the worst of both — you take on the entire data-reconstruction problem and a second cutover, at the hour of the night when human error rates peak, with a physical dispatch deadline 85 minutes out. Anyone who proposes it is really proposing “I would like the new system to not be live”, which is a reasonable feeling and a terrible plan.
  4. The steps, in order, with times. Note that four separate things happen before anybody touches the bug.
    TimeStepWhy it is in this position
    04:36Stop the bleed. Web store back into maintenance with an honest banner carrying a time. Duty agents stop entering orders in OrderDesk and take them on the paper pad that already exists for network outages. WMS-Link paused so nothing further reaches the warehouse.Every minute you spend fixing is a minute the affected-order count grows — and it grows fastest right now, in the busiest hour of the week. This is not the fix. It is what makes the fix have a fixed target.
    04:40Bound the damage. One query: every order created since 03:50 where the applied unit price ≠ the contract price for that account. Export to a file, timestamp it, drop it in the incident channel. Expect 147 rows, 38 flagged.That file is now the reconciliation ledger, and nothing gets closed until every row on it is closed. Without it, “we think we got them all” is the best sentence anybody will be able to say on Monday.
    04:45Hold the physical. Phone the warehouse night-shift supervisor: the 31 picked orders are held, not dispatched, pending price correction; the other 65 pushed-but-unpicked are paused.Cardboard on a van is the only truly irreversible thing tonight. It gets handled before the technical fix, and by phone, because a message at 04:45 is a message nobody reads.
    04:50Fix the cause, narrowly. Do not rebuild the database collation at 04:50 — multi-hour, huge blast radius, zero rehearsal. Instead make the lookup deterministic: normalise the account code on both sides of PriceEngine’s join (or add an explicit COLLATE clause to it). Small, understood, and itself revertible in a single deployment.At 04:50 you take the smallest change that removes the defect, not the most correct one. The proper collation correction goes into a daylight change with a rehearsal and a test — and into the incident’s follow-up actions before anyone goes to bed, or it will never happen.
    05:10Verify by comparison. Re-run pricing for the 38 flagged orders; expect exactly 38 corrections and zero movement on the other 109. Hand-check five against the contract records. Then place one fresh test order against each of three lower-case accounts.“It doesn’t error any more” is not verification. The count of things that changed, matching the count of things you expected to change, is.
    05:25Correct and release the physical. Re-price the 38, reprint pick notes for the affected orders among the 31, release the dispatch.This is the step with the 06:00 deadline attached — which is precisely why steps 1–5 were timed backwards from it rather than forwards from 04:35.
    05:35Re-open writes, with a named watcher. Web store out of maintenance, agents back into OrderDesk, and one person doing nothing else but re-running the mismatch query every five minutes for an hour.“We’ll keep an eye on it” is six people each assuming one of the others is. A watcher is a name and a query.
    06:00Re-key the paper orders taken during the freeze, ticking them off against the same ledger.The paper pad saved you at 04:36 and will quietly lose you eleven orders if nobody owns bringing it back in.
    08:00 MonCustomer comms and credits. The 38 customers are contacted before they ever see an invoice, by their own account manager, with the correction already applied. Finance lead owns the ledger to zero.The difference between an incident and a reputational event is almost always who spoke first.
  5. The comms. Six audiences, only two of which are on the call:
    AudienceTold whatBy whomChannelWhen
    8 duty agentsStop entering orders; use the paper pad; here is why in one sentence; expect a call at 05:35Incident leadDuty group message and a phone call to each — 04:36 messages get missed04:36
    Warehouse night-shift supervisorHold the 31 picked orders, do not dispatch, expect release by 05:30Incident leadPhone. Not a message.04:45
    Web-store customersAn honest maintenance banner with a time on it, rather than silence or a lieOps on-callSite banner + status page04:38
    Branch managersOne paragraph: what happened, and the exact line to say if a customer asksOps DirectorMonday morning brief06:45 Mon
    The 38 overcharged customersWhat happened, the corrected value, the credit, and that we rang them firstTheir named account managerPhone, then email confirmationfrom 08:00 Mon
    Finance leadThe ledger file; owns it to zero; credit notes raised before invoicing runsIncident leadEmail + the exported file05:00
  6. The lines that were missing — three, and the third is the cheapest artefact in this entire drill.
    03:50  POINT OF NO RETURN — first customer write to the new database.
           A revert after this minute loses every order taken since.
           From here the only path is mitigate-forward.
    
    03:45  GATE 1 (technical). PASS = 5/5 smoke orders INCLUDING one
           lower-case account code and one 20-line order.
           Then SOAK: 45 min of read-only + synthetic traffic. No real users.
    04:30  GATE 2 (business). Users released only after Gate 2 passes.
    
    PRE-WRITTEN: the "what exists only on the new system since T" query,
           tested during the rehearsal, saved in the runbook, ready to paste.
    
    The soak is the structural fix: the runbook released users fifteen minutes after the gate, and someone moved it earlier still because it looked fine. Put a deliberate, boring gap between “it works” and “people are relying on it”, and put a second gate at the end of the gap. And that pre-written extract query is worth more than either: even when revert is off the table, you can always answer “what exists only in the new place?” in sixty seconds instead of twenty minutes. Almost nobody has one. It takes ten minutes to write, in daylight, a week early.
  7. And the cheap check that was available days in advance: the smoke test used five scripted orders, all against upper-case demo accounts created by the same people who wrote the test. 214 of 2,300 accounts — 9.3% — carried the trait that broke the night, and the test had a zero percent chance of hitting it. A smoke test that doesn’t include one row of each awkward shape in your real data isn’t a smoke test; it is a compile check with ceremony. Sample your actual data for its oddities, and put one of each into the test.

Scenario 3 — Corrin & Slade: the way back was already dismantled

☺ Like you’re 10: The landlord re-let the old house three weeks ago. Nobody told the person holding the plan that says “move back into the old house”.

Corrin & Slade is a payroll bureau: it runs payroll for 380 client companies, around 96,000 employees. Its migration is a multi-wave programme, and PayEngine — the calculation engine — was cut over in Wave 3 on the 12th. Two small weekly payrolls (clients 1–40, hourly-paid, about 4,000 employees) have since run on it cleanly, on the 15th and the 22nd. Tonight is the 24th, and tonight is month-end.

SystemWhat it doesRuns onDataWho depends on it
PayEngineThe payroll calculation engine4 app servers + SQL Server1.1 TB380 client companies
ClientPortalPayslip access, and client uploads: timesheets, starters, leavers, salary changes2 VMs96,000 employees and 380 client payroll contacts
BankFileAssembles and lodges the bank payment submission1 VM with a fixed egress IP allow-listed by the sponsoring bankEvery pay run. Nothing gets paid without it.
StatFileStatutory submissions to the tax authority1 VMFiled per pay run
RefDataStatutory reference tables: tax bands, thresholds, and pension salary-sacrifice bandsTables inside PayEngine’s databasesmallEvery calculation. Updated every April.
⚠ The deadline that owns everything below

Month-end pays 71,000 employees on the 28th. The sponsoring bank requires the submission lodged three working days before payday, so the hard deadline is 17:00 on the 25th. It does not move, it cannot be negotiated, and missing it means 71,000 people are not paid on time. Everything in this scenario is measured backwards from that minute.

What Wave 3’s exit criteria said — and what they caused

Wave 3’s exit criteria contained a line that looked entirely sensible in month one: “Decommission the source environment within 10 days of a clean cutover.” It was there for good reasons — the old hardware was on a lease with a return date, and surrendering the SQL Server Enterprise licences was part of how the new platform was funded. Every one of those tasks completed on schedule, exactly as the plan asked. Here is what each one removed:

DateDecommission task completedWhat it quietly removed
15thBank allow-list updated to BankFile’s new egress IPThe bank now requires 5 working days’ notice to change it back. Nothing can be lodged from the old address.
20thOld SQL Server Enterprise licences surrenderedThe old instance may not lawfully be run in production — at a bureau whose entire product is compliance.
21stOld DNS records released; internal names re-pointedClient uploads now reach only the new ClientPortal.
22nd3 of the 4 old app servers wiped and collected by the lessorOnly a cold VM image of the old primary remains — restorable in roughly 6 hours.
12thBackups of the old environment stopped when it stopped being authoritativeThat cold image contains no data after the 12th. Twelve days of client uploads exist only on the new system.

The night of the 24th

TimeWhat happened
18:00Month-end run starts: 71,000 employees, 380 clients. Expected duration 3.5 hours.
20:40The run completes. The validation report shows 1,410 employees across 19 client companies with incorrect net pay — 240 of them netting to zero.
21:15Cause. Those 19 clients operate salary-sacrifice pension schemes. The RefData copied during migration came from a snapshot taken in March. The April statutory uplift was applied to the old system on 6 April and never re-applied to the new one, so the pension bands are a year out of date, contributions calculate too high, and low earners net to zero.
21:20Someone opens the runbook. The rollback section — written in month one, never revisited — reads, in full: “If the month-end run fails validation, revert to the source environment and run there.”
⚠ It is 21:20 on the 24th. You have 19 hours 40 minutes.

Write four things. (a) Establish in writing whether revert exists at all, with dates — do not assert it, enumerate it. (b) Write the mitigate-forward plan against 17:00 on the 25th, with a decision deadline you derived by subtraction. (c) Write the exit-criteria line that should have fired on the 20th and didn’t. (d) Name what replaces rollback once the source is gone — because “nothing” is the answer most programmes accidentally choose. And before you start: look hard at the shape of the failure. 19 clients are wrong. 361 are correct. There is something in that sentence that saves the deadline, and most people don’t see it for ten minutes.

Done when: your plan does not contain the word “revert” anywhere except in the section explaining why it is unavailable; you have named the date rollback actually expired and the task that should have fired at that moment; you have used the divisibility of the submission to protect the 17:00 deadline; and you can name at least three concrete capabilities that replace rollback once the source environment is gone.

Model answer — Scenario 3: why revert died on the 20th, and the plan that works anyway
  1. (a) Revert died on the 20th — and arguably on the 15th. Enumerate it, in writing, in four lines, because on the call somebody will say “can’t we just go back?” and you need to close that in ninety seconds rather than ninety minutes:
    DimensionStatusVerdict
    DataThe cold image stops at the 12th. Twelve days of timesheets, starters, leavers and salary changes for 380 clients exist only on the new system.Restoring it produces a correct payroll for the wrong month.
    LegalEnterprise licences surrendered on the 20th.Running it in production is a licensing breach — at a compliance bureau.
    PlumbingThe bank’s allow-list points at the new egress IP; changing it back takes 5 working days.Even a perfect payroll on the old system could not be lodged.
    Capacity3 of 4 app servers physically gone on the 22nd; the fourth is a 6-hour restore.Six of your twenty hours, to reach a state that fails on all three counts above.
    The runbook’s rollback section has been fiction for four days. Nobody re-read it because nothing in the process ever prompted anyone to — which is the actual failure here, and it is a process failure, not a person one.
  2. (b) The insight that saves the deadline: a payroll submission is divisible. The 17:00 deadline applies to a submission, not to a company. 361 clients are correct right now, tonight, already calculated. So the plan protects the 361 first and treats the 19 as a bounded, separately-scheduled problem — rather than doing what tired people do at 21:20, which is re-run all 380 blind and turn a 19-client problem into a 380-client one. Half the disasters in payroll, billing and settlement are survivable purely because the obligation divides and somebody remembered that in time to use it.
  3. The plan, with times:
    TimeStep
    21:20 (24th)Declare, and name the owner. The decision owner is the Payroll Operations Director — the person whose signature goes on the lodgement — with the Head of Service Delivery as deputy. Not the migration lead. When the incident is about a regulated submission, the decision owner is whoever is accountable for that submission, and that is worth settling in the runbook rather than at 21:20.
    21:25Freeze and preserve. Do not delete the failed run, and do not re-run over it. It is evidence, and it is also 361 clients’ worth of correct work you would otherwise be throwing away.
    21:30Set the decision deadline by subtraction. The work that must happen after the gate is fixed and cannot be compressed: BankFile assembly and checks = 90 min, lodgement plus bank confirmation = 45 min, so 2 h 15 m. That puts the latest possible gate at 14:45. The bank’s confirmation is not in your control and this is payroll, so take a deliberate 2 h 45 m of contingency and set the gate at 12:00 on the 25th — everything together, or 361 now and the rest supplementary. The overnight work (a 70-minute 19-client re-run, then 60 minutes of reconciliation) has to be finished long before that gate, which is exactly why it starts at 23:30 and not after breakfast. Write the 12:00 where everyone can see it.
    21:40Scope precisely. Which 19 clients, which 1,410 employees, which reference table, which rows, which effective dates. A named list, not “the pension ones”.
    22:00Fix the reference data — and diff the whole set. Import the current statutory bands, then diff all of RefData against the April export, which still exists as a file even though the server that produced it does not. If one table was stale, assume the rest are until proven otherwise. Fixing only the rows you already know about is how you meet the same incident again in June.
    23:30Re-run the 19 clients only, into a parallel run, leaving the 361 untouched.
    01:00 (25th)Validate by comparison, not by completion. Per-employee net pay against last month, expecting movement only where hours, salary or scheme membership actually changed. “The run completed” is not validation; it is the absence of a crash.
    02:00Two of the 19 are still wrong, from an unrelated cause — one client uploaded a stale timesheet file to the new portal on the 19th. Handle those two separately. Do not let two clients hold 378.
    09:00Client comms, before any employee hears it from a colleague. The 19 named payroll contacts, by phone, from their own account manager, with the fix already in and a lodgement time attached.
    12:00The gate. The decision owner calls it: 378 in the main submission, 2 in a supplementary lodged by 15:00.
    15:00Lodge. Bank confirmation received by 15:40 — the 17:00 deadline met with 80 minutes of slack, which exists only because the gate was at 12:00 rather than “when it’s ready”.
  4. (c) The exit-criteria line that should have fired on the 20th. Two sentences, and they belong in the wave template, not in this one runbook:
    WAVE EXIT CRITERIA — addition:
    
      Decommissioning the source environment VOIDS the rollback plan.
      On completion of the FIRST decommission task, the cutover runbook's
      Rollback section must be rewritten as a Forward-Recovery section and
      re-approved by the named decision owner. The wave is not closed
      until that rewrite exists and has been read by the on-call rota.
    
    ROLLBACK PLAN — mandatory header on every rollback plan:
    
      ROLLBACK VALID UNTIL: <date>
      VOIDED BY:            <the specific event>
      LAST TESTED:          <date, and by whom>
    
    A rollback plan without an expiry date will be read, believed, and found to be fiction — at 21:20 on the 24th, by someone who has nineteen hours and forty minutes and had been counting on it.
  5. (d) What replaces rollback when the source is gone. Four things, all cheap, none heroic. The programme that skips them hasn’t “accepted the risk”; it has simply stopped thinking about it.
    CapabilityWhat it is, concretely
    Forward recovery, rehearsedPoint-in-time restore of the new system, timed with a stopwatch at least once, against a realistic data volume. Knowing it takes 40 minutes is worth far more than knowing it is “supported”.
    Detection moved to the frontWhen you cannot go back, your investment shifts from the way back to finding out sooner. Here that is a one-query reference-data diff between the last known-good export and the live tables, as a mandatory pre-run gate on every payroll. It would have caught this on the 13th — eleven days early — for the cost of one query.
    Divisibility, written down in advanceKnow before the night which obligations split: per client, per submission, per region, per batch. Tonight that knowledge is worth 71,000 salaries.
    A rehearsed “run it twice”The ability to safely re-run a subset without disturbing the rest. It is the single most valuable operational capability in any batch system — and the one most likely to be discovered as impossible on the night you need it.
  6. The generalisable sentence, worth stealing: a rollback plan is a perishable asset, and decommissioning the source is what spoils it. Most teams treat the rollback plan as a document written once at the start of a programme. It is closer to a fire extinguisher with a service date on the label — and “we decommissioned the old environment last Tuesday” is the moment the label expires, whether or not anybody looks at it.

The closing question — where does your runbook stop being reversible?

☺ Like you’re 10: Go and find the exact minute in your own plan when turning the van round stops being possible. Write it in. In bold.

Take whichever cutover runbook you actually have — Part 4 of the capstone if you have worked it, or a real one from your own team if you have one — and find the single line at which reverting stops being possible. Write it in, in bold, with a time and an owner against it. Then answer honestly: did you know that line before you read this page, or did you find it just now? Because nearly every rollback disaster in the wild is not “we had no plan”. It is “we had a plan, and it was written for a moment that had already passed”.

Three questions decide whether a rollback plan is a plan or a wish. Apply them to yours:

The testWhat a real answer sounds likeWhat a wish sounds like
Has it been run?“Twice, in the rehearsal. 45 minutes both times, and the second time we found the DNS TTL was 3600, not 300.”“It should take about an hour.” An untested revert is an estimate in a confident font.
Does it have an expiry?“Valid until the source is decommissioned. That’s scheduled for the 20th, and there’s a task on the 20th to rewrite this section.”Undated. Which means it is valid forever, which means it is valid never.
Does it name a person and a minute?“A. Rahim calls it. Deputy J. Okafor. Last revert start 01:30, derived from EDI-Out’s 03:00 lodgement.”“The team will assess and decide.” At 04:12 that is not a decision procedure; it is a description of a silence.
🐢 Timmy’s-eye view

“People think I’m slow because I test the way back. I’m not slow — I’m cheap. Rehearsing a revert costs one afternoon, and it is the only afternoon in the whole programme that buys you the right to be calm at four in the morning. And I’ll tell you the part nobody expects: the rehearsal almost never confirms the plan. It finds the DNS record with the hour-long TTL, or the firewall rule somebody already deleted, or the fact that the old application refuses to start because its licence check phones a server that was switched off last month. Every one of those is a five-minute fix in daylight and a two-hour catastrophe at 04:12. I don’t rehearse the rollback because I expect to use it. I rehearse it because it is the only way to find out that I couldn’t have.”

🐢 Timmy’s rehearsal · 15 min

Right now, on paper, take Scenario 1 and make it worse: the revert itself fails halfway. At 00:50 the team begins reverting; at 01:05 the old TrackPoint application refuses writes, because the database was set read-only during the 23:00 freeze and nobody recorded that step, so nobody knows to undo it. It is now 01:05, your last revert start was 01:30, and you are 25 minutes into a 45-minute revert that has stalled. Write the next ten minutes. Then write the runbook change that prevents it — and notice that the change is not a clever one: it is that every step which alters state gets a matching undo step, written on the same line, at the moment the step is written. A runbook with a “Rollback” section at the bottom will always miss something. A runbook where each step carries its own reversal cannot.

Done when — the whole drill

☺ Like you’re 10: Six things you can point at. If you can’t point at them, you thought about it rather than wrote it.

  1. You have three written decisions, one per scenario, each with all five sections filled in — not notes, sentences.
  2. Every one of the three names one decision owner and one deputy, and a wall-clock decision deadline you derived by subtraction, with the subtraction shown.
  3. You can state, for each scenario, the exact minute the point of no return fell, and for two of them it is not the minute the runbook implied.
  4. Scenario 2’s plan stops the bleed before it fixes the bug, and accounts for all 147 orders including the 96 in the warehouse system and the 31 in cages.
  5. Scenario 3’s plan never proposes a revert, names the date rollback expired, and uses the divisibility of the submission to protect the 17:00 deadline.
  6. You have gone back to a runbook you actually own and written the irreversibility line into it, in bold, with a time and a name.

Milestones

☺ Like you’re 10: Ten boxes. Progress saves in this browser, so you can do one scenario tonight and the rest on Sunday.

0 / 10 steps complete

Setting up

1Copy the five-section template into your editor, three times
One blank template per scenario, headed with the company name. Read the six-word vocabulary table first — the model answers use revert, mitigate forward, point of no return, decision owner, decision deadline and rollback RPO/RTO in exactly the senses defined there.
Done when: you have three empty templates open and a clock you can see.
Concept: Best Practices (Always keep a way back: rollback and runbooks) · Glossary

Scenario 1 — Kestrel Freight

2Find the real deadline, and show the subtraction
It is not 04:00. Work backwards from the first hard downstream commitment using the measured durations you were given, and write the arithmetic out — not the answer, the arithmetic.
Done when: you have a wall-clock minute and a one-line derivation beside it that someone else could check.
Concept: Best Practices (RTO/RPO, the cutover runbook) · Architecture Patterns (cutover & deployment patterns)
3Make the call, and price the alternative honestly
Revert or push? Whichever you choose, write down what the other path costs if it goes wrong, and what abort condition you would announce out loud before starting it.
Done when: your answer contains a named owner, a named deputy, and a sentence beginning “if we are not green by …, we revert without further discussion.”
Concept: Anti-Patterns & Pitfalls (no rollback / backout plan) · Wave Planning
4Write the depot comms — and notice who hears nothing
Five audiences, one row each: told what, by whom, on which channel, at what time. Include the audiences who are asleep and the ones you deliberately do not contact.
Done when: the three retail EDI customers appear on your table with “nothing” in the message column, and you can say why that is the success condition.
Concept: Best Practices (change management, don’t forget the people) · Migration Checklist

Scenario 2 — Pemberton Tools

5Mark the true point of no return, and account for all 147 orders
Name the minute, and say what it was that crossed it. Then split the 147: how many exist only on the new system, how many reached the warehouse system, how many are physically picked, and what each of those three states means for a revert.
Done when: your PONR is 03:50 rather than 03:45, and you can explain the five-minute gap in one sentence.
Concept: Architecture Patterns (parallel run, data patterns) · Data Migration & Data Gravity (cutover)
6Write the mitigate-forward plan — first step stops the bleed
Nine or ten numbered steps with a wall-clock time against each. If your step 1 is the technical fix, delete it and start again: work out what the count of affected orders is doing while you fix, and what the physical dispatch is doing.
Done when: at least three steps happen before anyone touches the defect, and the step with the 06:00 deadline attached has the preceding steps timed backwards from it.
Concept: Best Practices (rollback & runbooks) · Anti-Patterns (inadequate testing & validation)
7Write the comms for six audiences, two of which are not technical
The 8 duty agents, the warehouse night-shift supervisor, the web-store customers, the branch managers, the 38 overcharged customers, and finance. Channel matters as much as message at 04:45.
Done when: at least two rows say “phone, not a message”, and the 38 customers are contacted before they could ever see an invoice.
Concept: Best Practices (change management) · Security, Cost & Resilience

Scenario 3 — Corrin & Slade

8Prove revert is gone — four dimensions, with dates
Data, legal, plumbing, capacity. One line each, each with the date it became true. Assertion is not enough; on the call someone will ask, and you need to close it in ninety seconds.
Done when: you have named the date rollback actually expired, and can defend an earlier date as also defensible.
Concept: The Journey (Operate & Optimize) · Anti-Patterns (no rollback / backout plan)
9Write the 20-hour forward plan, and the exit-criteria line that was missing
A timed plan to the 17:00 lodgement with a decision gate you derived by subtraction — plus the two-sentence wave exit-criteria change, and the mandatory header every rollback plan should carry.
Done when: the word “revert” appears nowhere in your plan except where you explain why it is unavailable, and your rollback-plan header has a “VALID UNTIL” and a “VOIDED BY” field.
Concept: Wave Planning (entry & exit criteria) · Security, Cost & Resilience (resilience & DR)

Take it home

10Find the irreversible line in a runbook you actually own, and write it in
Capstone Part 4, or a real one from work. Find the minute reverting stops being possible, write it in bold with a time and an owner, then apply the three tests: has it been run, does it have an expiry, does it name a person and a minute.
Done when: the line exists in the document, and you know which of the three tests your plan currently fails.
Concept: Best Practices (the cutover runbook) · Migration Checklist · Deep version: Capstone Part 4
🎬 At the Migration Academy
🦊

Foxy: Honest question. We tested the cutover for six weeks. If the testing was any good, why do we need a way back at all?

🐢

Timmy the Turtle: Because testing tells you the things you thought of work. The rollback plan is for the thing you didn’t think of. Pemberton tested five orders and every one used an upper-case account code — the test was fine, the imagination wasn’t.

👺

Gizmo the Gremlin: Or! Hear me out. Push through. We’re so close. Forty more minutes and it’s done, and nobody has to explain a wasted weekend to anyone.

🦉

Professor Owl: Gizmo, you’re quoting the wrong clock again. The window closes at four. The manifest closes at three. The gap between those two numbers is where your forty minutes actually comes from, and it isn’t yours to spend.

🐘

Ellie the Elephant: And once a real order lands in the new database, going back isn’t going back any more. It’s going back and deleting somebody’s work. Those are different sentences and people say them as if they’re the same one.

🐿️

Nutty the Squirrel: Thirty-one of those orders were already picked. Cardboard, in a cage, with a note on it. You can restore a database in twenty minutes. You cannot un-pick a pallet from a terminal.

🦥

Sol the Sloth: And the payroll one is my favourite, in a grim way. Nothing failed. Every task completed on schedule, exactly as planned. The plan just quietly deleted its own escape route on the twentieth and nobody had written down that it would.

🐢

Timmy the Turtle: Which is the whole drill. A rollback plan has an expiry date, an owner, and a minute. Miss any one of those three and what you own isn’t a plan, Foxy. It’s a paragraph.

🐢 Timmy’s checkpoint

1. What is the difference between a go/no-go gate and a point of no return — and at Pemberton Tools, how far apart were they and why did that matter? 2. At Kestrel Freight the window ran to 04:00 and Gizmo said there was plenty of time. What was the real deadline, and how did you compute it? 3. Name the event that expired Corrin & Slade’s rollback plan, and the date. What should have fired automatically at that moment? 4. Your runbook says “restore from backup and repoint DNS.” Give the three questions that decide whether that sentence is a plan or a wish. 5. Once the source system is gone, what capability replaces rollback — name two concrete things you would build to have it.

Check your answers
  1. A go/no-go gate is a decision you schedule: a moment where you assess against a written threshold and choose to continue or stop. A point of no return is a state change you cross, whether or not anyone is looking: the first moment data exists only on the new system, after which “stop” no longer means what it meant a minute earlier. At Pemberton the two were five minutes apart — a green gate at 03:45, the first real customer write at 03:50 — and only the gate was written in the runbook. So the most consequential transition of the entire night happened silently, and the team discovered it at 04:35 while trying to work out whether they could go back. The structural fix is to label the PONR in the runbook with a time, and to insert a deliberate soak between the gate and the moment users start writing, closed by a second gate.
  2. 01:30, and it comes from the first hard downstream commitment rather than the window. EDI-Out must lodge three retail manifests by 03:00 and the run takes 45 minutes, so it has to start by 02:15. A revert takes a measured 45 minutes, so it must start by 01:30. The 04:00 window end is the deadline for finishing, and it is two and a half hours later than the deadline that actually binds — which is why Gizmo’s “we’ve got hours” is technically true and completely wrong. General rule: your rollback deadline is almost never the end of the maintenance window. Find the first commitment that cannot move, subtract the measured recovery time, and that minute goes in the runbook in bold.
  3. The rollback plan was voided by decommissioning the source environment — decisively on the 20th, when the SQL Server Enterprise licences were surrendered, and arguably on the 15th, when the bank’s allow-list moved to the new egress IP and a 5-working-day change was needed to move it back. What should have fired at that moment is a mandatory rewrite: the wave’s exit criteria should state that completing the first decommission task voids the rollback plan and requires the Rollback section to be rewritten as a Forward-Recovery section and re-approved by the decision owner, with the wave not closable until that rewrite exists. Every rollback plan should also carry a header with VALID UNTIL, VOIDED BY and LAST TESTED fields — an undated rollback plan is valid forever, which is another way of saying it is never checked.
  4. (a) Has it been run? — an untested revert is an estimate, and rehearsals almost never confirm the plan; they find the hour-long DNS TTL, the deleted firewall rule, the licence check that phones a decommissioned server. (b) Does it have an expiry? — name the event that voids it, because the source environment will not be there forever and nothing will otherwise tell you the day it stopped being true. (c) Does it name a person and a minute? — one decision owner, one deputy, and a wall-clock deadline derived by subtracting measured durations from the first hard downstream commitment. “The team will assess and decide” is not a decision procedure at 04:12; it is a description of six people waiting for each other.
  5. What replaces it is forward-recovery capability plus earlier detection — you stop investing in the way back and start investing in getting out of trouble in place and in finding out sooner. Two concrete things, from any of these four: a rehearsed point-in-time restore of the new system, timed with a stopwatch against realistic data volumes; a pre-run data-integrity gate such as the reference-data diff that would have caught Corrin & Slade’s stale pension bands eleven days early for the cost of one query; documented divisibility of your obligations, so you know before the night which submissions, batches or client sets can be split to protect a deadline; and a rehearsed “run it twice” procedure that can safely re-process a subset without disturbing the rest. All four are cheap, all four are written in daylight, and a programme that has decommissioned its source without any of them has not accepted a risk — it has simply stopped thinking about one.

Three plans written? That’s the drill. For the theory underneath it, Best Practices covers rollback plans and cutover runbooks directly, Architecture Patterns covers the cutover patterns and the parallel run that make a reversal cheap in the first place, Anti-Patterns & Pitfalls has the no-rollback and big-bang traps all three of tonight’s companies walked into, and Security, Cost & Resilience puts RTO and RPO on a proper footing. The Migration Checklist is the operational version of the same discipline. Want a different single skill? Try Drill — Pick the Right R or Drill — Size the Data Move. Or take all of it into one continuous estate with Move Brambleside — Start Here, where Part 4 makes you write the runbook that this drill has spent three scenarios breaking.