Hands-On Labs · The Capstone · Part 5 of 5

Capstone Part 5 — Operate, Optimize & Modernize

Brambleside Veterinary Group is out of the colo. Every wave has landed, PIMS-Core cut over inside its four-hour Sunday window, and the last clinic stopped noticing weeks ago. And the first full month’s cloud bill is $57,400 against an on-prem run-rate of $42,000 — which is the number the board made the whole project conditional on. This part is where a migration stops being a move and becomes a business: you will work that bill down with six levers you compute yourself, decide which two systems earn modernization in the next twelve months and write down why every other one does not, switch Fort Rusty off without losing a seven-year regulatory obligation, and run the retro. Then you will do the thing that separates an engineer from a spreadsheet — name, in writing, the one line you refuse to optimize.

☺ Explain it like I’m 10

You’ve finished moving house. The boxes are in, the beds are made, everyone sleeps there now. Then the first bill arrives and it’s bigger than the old house — because you rented a van-sized room for a backpack, you’re keeping forty years of old photo albums on the kitchen table instead of in the attic, and the lights are on in three rooms nobody has walked into since you arrived. Today you go room by room and fix all of that. Then you pick the two things worth actually rebuilding — not everything, two — and you make a plan to give the keys back to the old house so you stop paying rent on both. And at the end you write down, honestly, what you’d do differently, including the bit where you looked silly.

🦥🦋Your hosts for this part: Sol the Sloth & Mira the Butterfly — Sol does the slow, careful money math and never rounds in his own favour; Mira turns the creaky things into cloud-native ones, and is the first to tell you which creaky things to leave completely alone. 🐢 Timmy guards the retention obligations, and 👺 Gizmo spends the whole page whispering “just cut the standby, nobody will ever notice.”
⚠ Where you’re arriving from, and what you’ll have when you’re done

Arriving: all sixteen systems are dispositioned (Part 1), the landing zone is built and governed (Part 2), four waves plus Wave 0 have run inside the nine-month calendar (Part 3), and the data is across with PIMS-Core cut over and rolled-back-if-needed (Part 4). Fort Rusty is still powered on, still costing money, and still holding the tapes. Leaving this page: a costed run-rate review that lands under the board’s $42,000/month ceiling with a before-column, a lever-column and an after-column that all reconcile; a 12-month modernization roadmap naming exactly two systems and a strangler-fig first slice for one of them; a decommissioning plan for Fort Rusty with a switch-off date, a power-down order, and every retention obligation re-homed and restore-tested; and a one-page retro. There is no Part 6 — this is where the capstone ends.

What this part assumes, and what it produces

☺ Like you’re 10: Everything is already moved. Nothing new gets carried today — today is about the bill, the plan for what to rebuild, and giving back the old keys.

What Parts 1–4 hand you

Part 5 is the only part that assumes all four earlier parts are behind you, because the whole page is arithmetic on their outputs. If you skipped one, you can still work this page from the tables below — every number you need is restated here — but the exercise will feel like bookkeeping rather than consequences. Here is exactly what each part is assumed to have decided:

FromWhat it decidedWhat Part 5 does with it
Part 1 — DispositionAn R for each of the 16 systems; 5 systems never moved (Retire/Retain), leaving 11 in motionLever 1 is literally Part 1’s decisions, priced. Skip Part 1 and Lever 1 is worth $0.
Part 2 — Landing zoneRegion + in-country DR region, non-overlapping CIDRs, tagging standard with a cost-centre key, budget alerts, one logging destinationThe tagging standard is the only reason the bill below can be split by system at all. Untagged spend is an unattributable bill.
Part 3 — WavesWave 0 + four waves inside nine months; ImageVault seeded by appliance in month 2Tells you which workloads have finished moving (and are therefore safe to commit to) and which have not.
Part 4 — CutoverPIMS-Core cut over in its 4-hour Sunday window; parallel run; rollback with a named point of no returnThe parallel-run capacity is still billing. Deciding when to release it is a Part 5 decision, and it is not the first one you make.
◆ One assumption stated out loud

Part 1’s answer key marks PIMS-Core as the row with the most defensible second answer: Repurchase (the vendor’s hosted SaaS edition) or Replatform (the vendor product on your own cloud VMs with a managed SQL service). This page runs the Replatform branch, because that is the branch Part 4 wrote a SQL-cluster cutover for — the vendor would not commit in writing to keeping the DR copy in-country, and the board’s residency policy covers the DR region too. If your matrix said Repurchase, your PIMS-Core line becomes a per-seat subscription instead of a compute line and Levers 2 and 5 stop applying to it. Say so in your own cost review and carry it through. The method is identical; only the row changes.

The four artifacts this page produces

  1. A — The cost review. The bill, six levers with a number each, and a steady-state run-rate under $42,000/month — plus the line you refuse to cut.
  2. B — The 12-month modernization roadmap. Two systems in, everything else explicitly out, and a strangler-fig outline for the one that earns it.
  3. C — The decommissioning plan for Fort Rusty. Power-down order, evidence captured, obligations re-homed, contracts cancelled with their notice periods respected.
  4. D — The one-page retro. What surprised us, what slowed us down, what changes next time.

All four fit on five sides of paper. That is deliberate: this is the pack you walk a steering committee through, and a steering committee does not read appendices.

⚠ Every number on this page is invented

Brambleside Veterinary Group does not exist, and neither do its bills. Real cloud prices change monthly, differ by region and provider, and are wrapped in commitments, tiers and free allowances that no static page can honestly represent. Nothing here quotes a real per-GB or per-hour rate, and you should never carry one of these figures into a real business case. The skill being practised is the method — the shape of the subtraction, the order you apply the levers in, and knowing which line is not for sale — not the number. When you do this for real, pull live figures from the provider’s own calculator and your own tagged bill, and re-run the same arithmetic.

The estate, restated — only the rows this part needs

☺ Like you’re 10: Here are the numbers you’ll be doing sums with, all in one place so you don’t have to go hunting.

The full fact sheet lives on Move Brambleside — Start Here. Part 5 needs four things from it: the baseline, the systems and their post-move shape, the constraints that still bite after the move, and the calendar.

The baseline the board is measuring you against

The condition the board attached to the whole project was simple and unforgiving: the post-migration steady-state run-rate must be no worse than today’s $42,000/month. Not “roughly.” Not “once we’ve optimized.” Steady state, on a normal month, all-in.

On-prem line$ / monthWhat’s in it
Colo — 8 racks, 40 km away13,800Space, power, cross-connects, remote hands. Contract ends 31 March.
Flagship server room2,400Power, cooling, insurance, rates. HVAC service contract $600 of it. HVAC has failed twice this year.
Licensing, amortized16,300VMware $4,200 · SQL Server Enterprise $6,100 · Oracle DB 12c SE $2,600 · backup software + tape maintenance $1,900 · load-balancer support $700 · Windows/AV/monitoring $800
Refresh fund + support9,500Hardware sinking fund, vendor support contracts, the five-person IT team’s share
Baseline total42,000The ceiling. Everything below is measured against this one number.

The sixteen systems and what happened to each

#SystemSize / shapeChosen RPost-move shape
1PIMS-Core — appointments, notes, billing4 VMs + SQL Server Enterprise cluster, 1.9 TBReplatform4 cloud VMs + managed SQL, in-country, warm standby DR
2ImageVault — imaging archive since 200441 TB, +900 GB/month, 78% untouched in 24 monthsRehost → tierManaged SMB file service (PIMS-Core reads it by UNC path)
3LabBridge — polls the reference lab1 small VM, IPsec VPNRehost1 small VM behind a static egress IP the lab allow-lists
4BramblesideOnline — site, booking, repeat scripts2 VMs + MySQL 5.7 (out of support), 90 GBReplatform2 VMs + managed MySQL 8 + cloud load balancer
5RotaMaster — staff rotaClassic ASP on Win 2012 R2, 90 weekly usersRepurchaseSaaS rota product
6ClaimsFeed — nightly insurance claimsOracle DB 12c SE, 280 GB. Renewal due month 7.ReplatformManaged PostgreSQL. Oracle licence dropped.
7StockCtrl — consumables + controlled-drug registerVendor app on SQL Server, 140 GB, 7-year register retentionRehostCloud VM + managed SQL, vendor-supported
8FileShare — staff file server6.2 TB across 1.1 million files; 40% untouched in 5 yearsRehost hot / Retire cold3.7 TB on a managed file service; 2.5 TB archived
9BackupVault — backup server + tape libraryHolds the 7-year finance retentionReplatformCloud backup + an immutable retention vault
10AD-DC01 / DC02 — AD, DNS, DHCP2 VMs; everything authenticates hereRetain → extendCloud domain controllers authoritative; on-prem pair demoted last
11PrintSrv — dispensing-label printers, 22 clinics1 VM, latency-sensitiveRetainStays on-prem, near the printers. Never moved.
12BI-Reports — SSRS, 40 scheduled reports, 400 GB31 of 40 unopened in 18 months (audit log proves it)Retire 31 / Replatform 99 reports on a managed reporting service
13VetLearn — internal CPD portal, 220 GBVendor SaaS costs less than the VMRepurchaseVendor SaaS
14DevTest — dev/test environments11 VMs, 24×7, 4% average CPU, untouched since 2019Retire 4 / Rehost 77 environments, scheduled
A1AshBook — Ashcombe’s online bookingAlready on a public cloudRetainLeft where it is, integrated
A2Ashcombe PIMS — same vendor, separate tenantVendor-hostedConsolidateFolded into the main PIMS tenant

The constraints that still bite after the move

Artifact A — the cost review

☺ Like you’re 10: The first bill is a shock. Then you go through it line by line and find six ways it was silly, and fix each one, and show your working.

The bill you get if Part 1 never happened

This is April’s bill: every one of the sixteen systems moved at its on-prem shape — same CPU counts, same storage tier, same 24×7, same licences — into the landing zone Part 2 designed. Two lines are not naivety and should not be treated as fat: the landing zone and the DR line are Part 2’s deliberate design outputs, specced tier by tier against stated RTO/RPO targets. Everything else is what happens when you carry the data centre’s decisions across the wire along with the data.

#Line$ / monthWhy it’s that big
1PIMS-Core — 4 VMs + managed SQL + 1.9 TB12,400Sized from the 2019 refresh spec sheet: 4 × 16 vCPU, plus SQL licensing
2ImageVault — 41 TB9,200All 41 TB on a premium SMB file service, because PIMS-Core reads it by UNC path
3FileShare — 6.2 TB, 1.1 M files2,600Everything hot, including 2.5 TB nobody has opened in five years
4DevTest — 11 environments4,30024×7 at 4% average CPU. Nobody switched one off since 2019.
5BramblesideOnline — 2 VMs + managed MySQL 8 + LB1,900Web tier sized for a peak that hasn’t happened since 2018
6ClaimsFeed — Oracle DB 12c SE on a VM3,700$1,100 compute + storage, $2,600 Oracle licence carried straight across
7StockCtrl — VM + managed SQL, 140 GB2,1008 vCPU observed at 11%
8RotaMaster — Win 2012 R2 VM900Includes an extended-support surcharge on a dead OS
9BI-Reports — SSRS VM + 400 GB1,150Running all 40 reports on schedule, including the 31 nobody opens
10VetLearn — VM + 220 GB700A VM that costs more than the vendor’s SaaS
11AD-DC01 / DC026002 × 4 vCPU observed at 3%
12LabBridge — small VM + static egress IP + NAT340Small, and correctly small
13PrintSrv310Moved. It should never have been. Watch this one.
14BackupVault + 7-year retention vault2,900Cloud backup plus immutable long-term retention
15Landing zone & shared services3,800Networking, hybrid link, logging destination, DNS, monitoring, secrets, backup control plane
16DR — per tier, from Part 26,900PIMS-Core warm standby 3,200 · BramblesideOnline + StockCtrl pilot light 1,300 · backup-and-restore for the rest 900 · cross-region storage replication 1,500
17Ashcombe — AshBook + a second PIMS tenant1,800Two tenants of the same vendor product, never consolidated since 2023
18Egress + inter-region replication traffic1,800Cross-region DR replication and the tail of the hybrid link
Total57,400$15,400/month over the ceiling. This is bill shock, and it is entirely self-inflicted.

Anti-Patterns & Pitfalls files this under two names at once — lift-and-shift-everything and ignoring cost → bill shock — and they are the same trap seen from two ends. Nothing in that table is a cloud pricing problem. Every single fat line is an on-prem decision that got a free ride across the wire because nobody was asked to re-justify it. The cloud did not make Brambleside more expensive; it made Brambleside’s existing waste visible and monthly.

▶ Try this before you read on

Cover the next section. Look at that table for ninety seconds and write down, in order, the three lines you would attack first and roughly what you think each is worth. Then read the six levers and see how close you were. Most people get the biggest one wrong — they go for the compute line, because compute is what engineers look at, and miss that the single fattest saving is sitting in storage.

Lever 1 — Retire and Repurchase: banking Part 1’s decisions

☺ Like you’re 10: The cheapest thing to run is a thing you don’t run. Part 1 already found five of those — this is just cashing them in.

This lever costs nothing to pull because Part 1 already did the arguing. That is the point of doing Retire and Retain first: by the time you get to a bill, the decisions are made and priced, and this lever is a data-entry exercise rather than a fight.

MoveBeforeAfterSavingThe evidence that made it uncontroversial
BI-Reports: retire 31 of 40 reports, 9 survive on a managed reporting service1,150290860The SSRS audit log. Not an opinion — a query result.
FileShare: archive the 2.5 TB (40%) untouched in 5 years, serve 3.7 TB hot2,6001,610990Last-accessed timestamps across 1.1 M files
DevTest: retire 4 of 11 environments outright4,3002,7401,560Four had no login and no deploy since 2019
RotaMaster → SaaS rota product900260640Classic ASP, no source control, no tests, author left in 2015
VetLearn → vendor SaaS700180520The vendor’s own price list, compared to the VM
PrintSrv: retained on-prem, never moves3100310Label printers in 22 clinics. Latency is the requirement.
Ashcombe: retain AshBook, fold Ashcombe PIMS into the main tenant1,8007001,100Two tenants of one product, three clinics, one contract
Lever 1−5,980

Note the FileShare line carefully, because it hides the trap Size the Data Move drills into you: moving 440,000 cold files to an archive tier costs a one-off transition charge per object — roughly $1,900 here — which pays back in about two months and would have been invisible if you had reasoned only about terabytes. File count is a cost dimension, not just a transfer-time dimension.

Lever 2 — Rightsizing: stop paying for the 2019 spec sheet

☺ Like you’re 10: You bought a van because one day you might move a sofa. You move a backpack. Get a smaller van.

On-prem, you sized for the peak you might hit over a four-year refresh cycle, because you could not resize a bought machine. In the cloud you can resize on a Tuesday, so sizing from a four-year-old spec sheet is not caution — it is a standing order to overpay. Every provider ships a recommender that reads your own utilization and proposes a size (AWS Compute Optimizer, Azure Advisor, Google Cloud’s Recommender). Use it as an input, never as an authority: it cannot see your seasonality, and Brambleside has a very obvious March.

SystemProvisionedObserved peakResized toBeforeAfterSaving
PIMS-Core cluster4 × 16 vCPU38% at Monday 09:104 × 8 vCPU with burst headroom12,4009,6002,800
StockCtrl8 vCPU11%2 vCPU2,1001,340760
BramblesideOnline web tier2 × 8 vCPU9%2 × 2 vCPU behind the LB1,9001,480420
AD-DC01 / DC022 × 4 vCPU3%2 × 2 vCPU600460140
Lever 2−4,120
⚠ Rightsize after the move, never during it

PIMS-Core is measured at 38% of a machine that was sized for a peak nobody has observed since the estate grew to 22 clinics — but the measurement that matters is Monday 09:10, the busiest ten minutes of the busiest day, not the monthly average. A monthly average would have told you 9% and talked you into 2 vCPU, and the first Monday after that change would have been an outage in 22 waiting rooms. Rightsize on peak, keep burst headroom, and never rightsize inside a wave — you would be changing the machine and the location at the same time and lose the ability to say which one broke it. Best Practices makes the same argument about modernizing during a move; it is the same argument.

Lever 3 — Schedule the non-prod estate 12×5 instead of 24×7

☺ Like you’re 10: Turn the lights off in rooms nobody is in at 3 a.m. on a Sunday.

Seven DevTest environments survive Lever 1, at $2,740/month — $2,400 of that compute, $340 storage that keeps billing whether the machine is on or off. A 12-hour weekday window is 60 of the 168 hours in a week, or 35.7%. The compute portion drops to about $860; storage does not move.

DevTest after Lever 1        2,740   ( compute 2,400 + storage 340 )
Scheduled window             12 h x 5 d  =  60 / 168 h  =  35.7%
Compute at 35.7%             2,400 x 0.357  =  857  ->  860
New DevTest total            860 + 340  =  1,200

Lever 3 saving               2,740 - 1,200  =  -1,540

Two things make this lever stick rather than quietly rot. First, the schedule is a property of the environment, expressed as a tag the Part 2 tagging standard already mandates (schedule=weekday-12x5) and enforced by a scheduler, not by a person remembering. Second, an environment that someone needs at 22:00 can opt out — but opting out sets schedule=always-on with an owner and a review date, so the exception is visible in the same report that produced this saving. An exception you cannot see is not an exception; it is a leak.

Lever 4 — Tier ImageVault’s cold 78% to archive

☺ Like you’re 10: Forty years of old X-rays don’t need to sit on the kitchen table. Put the old ones in the attic — but keep a ladder.

This is the biggest lever on the page and the one people skip, because storage is not where engineers look. ImageVault is 41 TB on a premium SMB file service, chosen because PIMS-Core references studies by UNC path and lifting the path was the fastest way to cut over. That decision was right for the cutover and wrong for the steady state. The audit says 78% of studies have not been touched in 24 months.

ImageVault naive        41 TB, all premium SMB           9,200

After tiering
  hot tier              9 TB on the SMB file service     2,020
  archive tier          32 TB on archive object storage     640
  gateway + retrievals  transparent rehydration + budget    380
                                                        -------
                                                          3,040

Lever 4 saving          9,200 - 3,040  =  -6,160

The mechanism is a lifecycle rule — every cloud has one (S3 Lifecycle, Blob lifecycle management, Object Lifecycle Management) — plus a gateway that keeps the UNC path working so PIMS-Core never learns anything changed. That gateway is not optional garnish; it is the reason this lever does not require touching a vendor application you do not own.

⚠ The clinical exception is not negotiable

Archive storage is cheap to keep and expensive and slow to get back — minutes to hours. In a 24-hour emergency hospital, “your ultrasound will be available in four hours” is not a storage tier, it is a clinical incident. So the demotion rule has two clauses, not one: demote a study when it has been untouched for 24 months and the patient has no open case. The $380 line is not padding — it is a standing retrieval budget plus the gateway, so that pulling a 2009 X-ray on a Tuesday afternoon costs a known amount and works. Write the exception into the rule, not into a wiki page nobody reads at 03:00.

One more consequence worth carrying into the roadmap: ImageVault grows 900 GB/month. Without the rule, the hot tier grows about $200/month every year, forever. With the rule, the hot tier stays roughly flat, because studies age into archive at roughly the rate new ones arrive. A lifecycle rule is not a one-off saving; it is a saving that keeps applying itself.

Lever 5 — One-year commitments, on the steady-state base only

☺ Like you’re 10: Promise to rent the same scooter for a whole year and it’s much cheaper. Just don’t promise it for the scooter you’re about to swap for a bike.

Security, Cost & Resilience lays out the three pricing models — on-demand, committed (reserved instances / savings plans / committed-use discounts), and spot. The art is layering them, and the single rule that keeps a commitment from becoming a liability is this: commit only to a shape you are certain will still exist in twelve months. Applied to Brambleside, that draws a hard line through the estate.

Eligible for a 1-year commitmentCompute baseSaving at ~30%
PIMS-Core compute (post-rightsizing)6,4001,920
StockCtrl compute900270
BramblesideOnline web tier1,000300
AD-DC pair400120
Shared services (log collectors, bastion, gateway)500150
Total eligible base 9,200−2,760
Deliberately excludedWhy not
ClaimsFeedIt replatforms off Oracle onto managed PostgreSQL. Committing to today’s VM shape means paying for a shape you are actively deleting.
DevTestScheduled and spiky. A commitment on a machine that is off 64% of the week is a discount on nothing.
ImageVault storageThe tiering just changed its shape entirely. Wait a quarter, watch where it settles, then look at storage commitments as a separate instrument.
Anything still inside hypercareIf a rollback is still theoretically possible, the workload is not steady state. Commitments start when hypercare closes, not when the cutover finishes.
BramblesideOnline’s managed MySQLThe booking path is first in line for modernization (Artifact B). The web tier is committed; the database deliberately is not.

One year, not three. Three-year commitments are cheaper per month and would look better on this page — and they would lock a nine-month-old estate into decisions made before anyone had seen a full seasonal cycle, including a spring vaccination campaign that has never once run on this infrastructure. Take the smaller discount and keep the option. Revisit at the twelve-month mark, when you have real seasonality data and Artifact B has told you which shapes are about to change.

Lever 6 — Drop the Oracle licence when ClaimsFeed replatforms

☺ Like you’re 10: You were renting an expensive engine to do a small job. Swap it for a normal one and stop renting.

ClaimsFeed is a nightly SFTP job against a 280 GB Oracle Database 12c SE instance, and its licence renewal falls in month 7 — January, two months before the colo dies. That date is not a coincidence in this exercise; it is the whole lever. Renewing means paying $2,600/month for another year on a database you are about to replace. Not renewing means the replatform has a hard deadline with a real number attached, which is the most useful kind of deadline there is.

ClaimsFeed on Oracle 12c SE     compute + storage  1,100
                                Oracle licence     2,600
                                                   -----
                                                   3,700

ClaimsFeed on managed PostgreSQL                     800

Lever 6 saving                  3,700 - 800  =    -2,900

Do not let the number make this look easy. It is a schema migration, an application change, and a re-test of a nightly feed the insurer depends on — Data Migration & Data Gravity is the page to re-read before you scope it. What makes it tractable here is that ClaimsFeed may miss one night: it is the one system in the estate whose downtime tolerance is measured in a whole night rather than in hours, which is why Part 3 could put it where it did. Money and risk lined up on the same system. Take that when it happens; it is rare.

The subtraction, in one place

☺ Like you’re 10: Here’s all six sums stacked up so you can see the big number turn into a small one.

From bill shock to steady state — $/month $42,000 board ceiling 57,400 Naive bill −5,980 L1 retire / repurchase −4,120 L2 rightsize −1,540 L3 schedule −6,160 L4 tier the archive −2,760 L5 commit −2,900 L6 Oracle licence +1,150 on-prem PrintSrv + clinic kit 35,090 Steady state Six levers, −23,460/month. The ceiling is crossed at L4 — the lever nobody looks at.
Naive cloud bill                                     57,400

  L1  Retire & Repurchase (Part 1, banked)           -5,980
  L2  Rightsizing to observed peak utilization       -4,120
  L3  Non-prod scheduled 12x5 instead of 24x7        -1,540
  L4  ImageVault cold 78% tiered to archive          -6,160
  L5  1-year commitments, steady-state base only     -2,760
  L6  Oracle licence dropped at ClaimsFeed replatform-2,900
                                                    -------
  Total levers                                      -23,460

Cloud subtotal                57,400 - 23,460  =     33,940
Residual on-prem (PrintSrv host, 22 clinic print
  appliances, flagship comms rack)                  + 1,150
                                                    -------
STEADY-STATE RUN-RATE                                35,090

Board ceiling                                        42,000
Headroom                                              6,910   (16.5% under)

Three things about that block are worth more than the total. First, the ceiling is crossed at L4 — storage tiering, the lever that lives furthest from where engineers instinctively look. Second, L1 is free: it is Part 1’s homework, cashed. A team that skipped the Retire/Retain pass arrives here needing to find $5,980 from somewhere else, under pressure, with the board already unhappy. Third, the residual on-prem line goes up, and that is correct — PrintSrv was always going to stay, and a cost review that quietly forgets the kit still humming in 22 clinics is not a cost review, it is a sales pitch.

Model answer — the line-by-line reconciliation

Every line of the naive bill, with the levers applied. The Before column sums to 57,400, the Saving column to 23,460, and the After column to 33,940. If your own table does not reconcile on all three, the error is almost always a lever double-counted on a line that two levers touch — PIMS-Core (rightsize then commit) and DevTest (retire then schedule) are where it happens.

LineBeforeLevers appliedSavingAfter
PIMS-Core12,400L2 −2,800 · L5 −1,9204,7207,680
ImageVault9,200L4 −6,1606,1603,040
FileShare2,600L1 −9909901,610
DevTest4,300L1 −1,560 · L3 −1,5403,1001,200
BramblesideOnline1,900L2 −420 · L5 −3007201,180
ClaimsFeed3,700L6 −2,9002,900800
StockCtrl2,100L2 −760 · L5 −2701,0301,070
RotaMaster900L1 −640640260
BI-Reports1,150L1 −860860290
VetLearn700L1 −520520180
AD-DC01 / DC02600L2 −140 · L5 −120260340
LabBridge3400340
PrintSrv310L1 −310 (retained on-prem)3100
BackupVault + retention vault2,900— (see the refusal)02,900
Landing zone & shared services3,800L5 −1501503,650
DR — per tier6,900— refused06,900
Ashcombe1,800L1 −1,1001,100700
Egress + inter-region replication1,80001,800
Cloud subtotal57,40023,46033,940
Residual on-premPrintSrv host + 22 clinic print appliances + flagship comms rack1,150
Steady-state run-rate35,090

The one thing you would not optimize

☺ Like you’re 10: You saved lots of money. Now someone asks for a bit more — and the only bits left are the seatbelts.

You are $6,910/month under the ceiling. This is the moment the CFO says “excellent, find me one more” — and it is the moment the exercise is actually testing. Two seventh levers are sitting right there, both real, both easy, both wrong:

The tempting seventh leverWhat it savesWhat it actually costsVerdict
Drop PIMS-Core’s DR from warm standby to pilot light2,100 / monthMoves PIMS-Core’s recovery time from roughly 20 minutes to roughly 4 hours. The board’s stated tolerance is 4 hours — for a planned, scheduled, 02:00 Sunday outage. An unplanned regional failure at 16:00 on a Monday would consume the entire tolerance with zero slack, in a group that runs a 24-hour emergency hospital.Refuse.
Release the Wave 4 parallel-run capacity two weeks early~5,800 one-off, in one monthThe parallel run is the only mechanism for comparing old and new PIMS-Core output before Part 4’s point of no return. Cutting it does not save money; it removes the evidence that the cutover worked, in exchange for a number on one month’s bill.Refuse.

Write both refusals down, in the pack, with the numbers attached — because an undocumented refusal gets re-litigated every quarter by someone who has only seen the saving. The sentence that works in the room is short: “Both of those are available. Both convert a cost win into an outage risk, and we are already 16% under the number the board set. I will bring you a seventh lever from Artifact B’s roadmap instead.” That is the whole skill. Anyone can find savings; the job is knowing which savings are actually deferred incidents with a discount attached.

◆ Key idea

A cost review with nothing refused in it is not finished. Optimization has a floor, and the floor is made of the things that only look expensive on days when nothing goes wrong. Security, Cost & Resilience puts it as: pick your RTO and RPO first, then buy exactly the resilience they demand. Part 5’s corollary is that once you have bought it, you defend it by name.

Artifact B — the 12-month modernization roadmap

☺ Like you’re 10: Now pick just two old things worth rebuilding properly — and write down why you’re leaving all the others exactly as they are.

Modernization makes the distinction this artifact rests on: migration changes where an app runs; modernization changes how it is built. Nine months of Brambleside’s project was entirely the former. This roadmap is the first time the latter is on the table — and the discipline is not choosing what to modernize, it is choosing what to leave alone and being able to say why in one sentence each.

The four tests a system has to pass to earn a slot

  1. Does the business want something it currently cannot have? Not “would it be nicer” — is there a request, from a named person, that the current design blocks?
  2. Do you own the code? If it is a vendor product, modernization means changing vendors, which is a procurement decision wearing an engineering costume.
  3. Is there a seam? Can you carve off one slice, run it alongside the old path, and switch back in an afternoon? If not, the first piece of work is creating a seam, and that is a separate project.
  4. Would a regulator have to be consulted? If yes, the timeline is theirs, not yours, and it is not a twelve-month item.
▶ Your turn first

Before opening the model answer: go down the sixteen-row table above and mark each system IN or OUT against those four tests, with one sentence each. Give yourself fifteen minutes and allow yourself exactly two INs. The constraint is the exercise — a roadmap with six systems on it is a wish list, and the sixth one never happens.

Model answer — which two, and why every other one is left alone
SystemIn / outThe one sentence
BramblesideOnlineIN — firstThe practice managers have been asking for out-of-hours self-service since 2022, Brambleside owns the code, the load balancer already gives you a routing seam, and no clinical record is written directly by the booking path.
ClaimsFeedIN — secondThe Oracle exit already forced the team into the code and the schema; replacing the nightly SFTP batch with the insurer’s API turns a 3-day settlement into a same-day one, which is working capital, which is a number the board understands.
PIMS-CoreOUTVendor product — “modernizing” it means changing vendors, and that is a contract-renewal conversation, not a sprint.
ImageVaultOUTThe lifecycle tiering was the modernization; a real PACS/DICOM replatform is a two-year clinical programme with regulatory implications.
StockCtrlOUTVendor-supported and holds a regulated controlled-drug register — touching it needs the regulator in the room.
FileShareOUTThe right answer is to keep shrinking it, not to rebuild it. Retire is not a modernization strategy, and it is better than one here.
RotaMaster / VetLearnOUTAlready repurchased — there is nothing left to modernize.
BI-ReportsOUTNine surviving reports on a managed service. Finished. Leave it.
AD / LabBridge / PrintSrv / BackupVaultOUTPlumbing. Modernizing plumbing that works buys you risk and a smaller electricity bill.
AshBook / Ashcombe PIMSOUTConsolidated in Lever 1. Revisit only if the vendor’s roadmap forces it.
DevTestOUTSeven scheduled environments. The next improvement is Infrastructure as Code so they can be destroyed and recreated — cheap, useful, and not a modernization slot.

Two IN, fourteen OUT. If your list has three or more INs, ask which one you would drop if the person who left in month 3 is still not replaced — because they are not, and the roadmap has to survive that.

The roadmap on a page

QuarterWhat happensWhy thenHow you know it worked
Q1 · Apr–JunNothing modernizes. Hypercare closes on the last wave, the rightsizing and lifecycle rules from Artifact A land, commitments are taken, Fort Rusty is decommissioned (Artifact C).An estate that is still settling cannot tell you which change caused which effect. Also, the team is tired.Run-rate under $42,000 for three consecutive months, and the colo invoice reaches zero.
Q2 · Jul–SepBramblesideOnline: build the routing facade, then ship slice 1 (repeat prescription requests) behind it.Smallest slice, no payment path, no clinical write, and a clear success signal.Facade routes 100% of existing URLs with error rate unchanged for 7 days; slice 1 at 100% traffic with completed-request rate at or above the old path’s 30-day baseline.
Q3 · Oct–DecBramblesideOnline slices 2 and 3 (new-client registration, then appointment booking). ClaimsFeed: API pilot against the insurer’s sandbox, running alongside the nightly SFTP batch.Booking is the prize and needs the longest ramp. The claims pilot is parallel-run only — no cutover before the December blackout.Booking completion rate and abandonment at or better than baseline; nightly parity check between API-submitted and SFTP-submitted claims shows zero discrepancies for 30 days.
Q4 · Jan–MarSlice 4 (account / pet profile), old BramblesideOnline app and its VM retired. ClaimsFeed event-driven submission goes live; SFTP batch kept warm as a fallback.The old app is only switched off after a full spring campaign has run through the new one.Old app’s VM deleted from the IaC repo, not just powered off. Claims settling same-day. Nothing routed to the old path for 60 days.

The strangler fig, slice by slice

Architecture Patterns describes the Strangler Fig as replacing one route at a time behind a facade, with the old system live the whole way. Here is what that means concretely for BramblesideOnline. Note the seam is a routing layer keyed on URL path, sitting in front of both the old app and the new service, with an anti-corruption layer translating between the new service’s model and PIMS-Core’s booking API — so the new code never inherits the old app’s data shapes.

SliceWhat moves% of trafficTraffic shiftSuccess measureOld path off
0The facade itself — routing layer in front of the existing app, no behaviour change0%n/aEvery existing URL resolves; error rate and p95 unchanged for 7 daysn/a
1Repeat prescription request~11%5% → 25% → 100% over 6 weeksCompleted-request rate ≥ old path’s 30-day baseline; zero clinician-reported mismatches30 days after 100%
2New-client registration~8%10% → 50% → 100% over 4 weeksRegistration completion rate ≥ baseline; PIMS-Core client records created correctly, checked nightly30 days after 100%
3Appointment booking — the prize~64%5% → 20% → 50% → 100% over 10 weeks, never during the spring campaignBooking completion, abandonment, and nightly write-parity against PIMS-Core60 days after 100%, and only after one full spring campaign
4Account / pet profile~17%10% → 50% → 100% over 4 weeksProfile edits reflected in PIMS-Core within the agreed sync window30 days after 100%
Retire the old app and its VM; remove from the IaC repoNo request has reached the old path in 60 days, proven from the facade’s own logsMonth 12
# Strangler fig — one-page slice template

Slice name:
Seam:                 (where the routing decision is made, and on what key)
Anti-corruption layer:(what it translates, in which direction)
Traffic share today:  __%
Shift schedule:       __% -> __% -> 100%, over __ weeks
Success measure:      (a number, with the baseline it is compared to)
Rollback:             (how traffic goes back, and in how many minutes)
Blackout conflicts:   (which windows this schedule must not cross)
Old path off:         (date, and the evidence required before that date)
Owner:                (named person, not a team)
⚠ “Old path off” is a date and an evidence gate

The single most common way a strangler fig fails is that the fig never strangles: the new path reaches 100% traffic, everyone celebrates, and the old app runs for three more years because nobody owned switching it off. Two rules prevent it. First, every slice carries a switch-off date at the moment it is planned, not afterwards. Second, the old app is deleted from the Infrastructure-as-Code repository, not just powered off — a resource that still exists in code will be recreated by the next apply, and a resource that is merely stopped still bills for storage. This is exactly the discipline The Journey calls “retire the source” and makes the official last step of every wave.

Why nothing was refactored before 31 March — and why that was right

Not one system on Brambleside’s estate was refactored during the nine-month migration. Say that out loud in the steering committee, because someone will otherwise present it as a failure. Three reasons, in order of weight:

  1. The binding constraint was a date, not an architecture. The colo contract ended 31 March with no renewal beyond a punitive month-to-month. A refactor that slipped would not have produced a worse app; it would have produced a $33,120 month-to-month bill and an estate half in and half out.
  2. Changing two things at once destroys your diagnosis. Modernization states the rule plainly: never modernize and migrate the crown jewels at the same time. If PIMS-Core had been re-architected and relocated in the same Sunday window, the 04:12 question in Part 4 would have had no answerable form.
  3. The team lost a person in month 3 and never replaced them. Six became five, two of whom were also on clinic support rota. Capacity was the second binding constraint, and a roadmap that ignores it is fiction.

What makes the deferral honest rather than an excuse is this artifact existing. A refactor postponed with a dated roadmap and named owners is a plan; a refactor postponed with “we’ll look at it later” is how an organization ends up running a lifted-and-shifted 2016 estate in the cloud in 2031, paying cloud prices for data-centre architecture. Anti-Patterns calls that one lift-and-shift and forget, and Artifact B is the specific document that prevents it.

Artifact C — decommissioning Fort Rusty

☺ Like you’re 10: A move isn’t finished until you hand back the old keys and stop paying rent on two houses. And you check the attic first.

Until Fort Rusty is off, Brambleside is paying for both estates and the $42,000 baseline has not gone anywhere — it has been added to. Decommissioning is not the epilogue of a migration; it is the step that realizes every saving the previous four parts earned. It is also the step that is skipped most often, because by this point everyone is tired and nothing about switching off an old server is exciting.

The power-down order

Reverse dependency order, with a gate before each block. Nothing on this list happens until the last wave’s hypercare window has closed — the old estate is your rollback right up to that moment, and The Journey is blunt about it: turning off Fort Rusty too early is a favourite gremlin trick.

StepWhatGate before it
D0Freeze: every Fort Rusty system goes read-only. No changes, no patches, no new data.Last wave’s hypercare closed. Freeze announced to all 22 clinics + the emergency hospital.
D1Capture the evidence (full list below). Nothing is wiped before this is signed off.
D2Re-home the retention obligations and restore-test them.A random 2019 tape restores successfully, witnessed and documented.
D3Application VMs: BI-Reports (SSRS), VetLearn, RotaMaster30 days with zero user sessions, proven from logs
D4BramblesideOnline web tier + the on-prem load balancerDNS TTLs expired; no traffic at the old VIP for 14 days
D5Database VMs: StockCtrl, ClaimsFeed, then PIMS-Core’s SQL clusterFinal backup taken and restored into a scratch instance — a backup you have never restored is a wish
D6FileShareCloud copy has served 100% of reads for 30 days
D7ImageVault NAS — last of the data systems41 TB reconciled by checksum against the cloud copy; 30 days of all reads served from cloud; a random archived study retrieved successfully
D8Demote AD-DC02, then AD-DC01 — do not power them offFSMO roles transferred to a cloud DC; DHCP scopes repointed to cloud DNS; 14-day soak with cloud DCs serving all authentication
D9BackupVault server and the tape libraryOnly after D2’s restore test passed. This is the last contract cancelled.
D10vSphere hosts (8), then the SAN / storagevSphere inventory exported; no VM in a powered-on state
D11Secure erase or certified media destructionCertificate of destruction per drive serial, matched line by line against the asset register
D12Network kit: switches, firewall, VPN concentrator — genuinely lastThe site-to-site tunnel runs through here. Nothing on-prem needs it any more, and the clinic links have been re-terminated.
D13Colo rack hand-back inspection; keys returnedInspection booked 14 days ahead. Racks empty, photographed, cabling removed.
⚠ You do not power off a domain controller

D8 says demote, and the word is doing real work. Powering off a domain controller leaves its object in the directory, its FSMO roles unheld, and its IP still handed out as a DNS server by DHCP — and the failure mode arrives days later as intermittent, unexplainable authentication failures in clinics, which is the worst possible shape for a fault. The sequence is: transfer FSMO roles to a cloud DC, repoint every DHCP scope and static client to cloud DNS, wait fourteen days with cloud DCs serving all authentication, then demote cleanly, then power off. PrintSrv is the one exception on this whole page — it is Retain, it never moves, and it stays humming at the flagship next to the label printers.

The evidence you capture before anything is wiped

You will be asked for something on this list in two years, by an auditor, a regulator, or an insurer, and Fort Rusty will not exist. Capture it while the systems are still running:

The obligations that outlive the data centre

☺ Like you’re 10: Some papers you have to keep for seven years even after you knock the old house down. Find them a new attic, then check you can still read them.

These are Brambleside’s obligations, not a vendor’s. Handing the tapes to someone does not transfer the duty; it only transfers the box.

ObligationSourceDurationNew home after 31 MarchThe test that proves it works
Finance backupsStatutory + company policy7 yearsImmutable archive vault with object lock, in-countryQuarterly restore of a randomly chosen year, documented
Controlled-drug registerRegulator7 yearsSigned immutable export plus live StockCtrl in the cloudAnnual readability test in the regulator’s accepted format — not “the database still opens”
Clinical recordsProfessional body + client dutyPer practice policyPIMS-Core (cloud) + ImageVault archive tierMonthly retrieval of a randomly chosen archived study, timed
Imaging studiesClinicalAs aboveImageVault archive tier, with the clinical exception rule from Lever 4Same monthly retrieval test — and it must beat the clinical response requirement, not just succeed
Employment recordsStatutory6 yearsHR SaaS + the hot FileShare tierSpot check at each annual review
◆ Key idea

A retention obligation is only satisfied when the data is readable, in an accepted format, by someone who is not you. Cheap immutable storage full of a 2019 backup format that needs a piece of software Brambleside cancelled the licence for in month 8 satisfies nothing. That is why D9 — the tape library — is the last contract cancelled and only after a restore test passes, and why the controlled-drug register gets exported to a signed immutable format rather than just being left inside a database.

The contracts, their notice periods, and the deadline behind the deadline

Every one of these has a notice period, and a notice period is a date you have to hit months before the date you were watching. Miss one and the saving simply does not arrive — you keep paying for a thing you already switched off, which is the most annoying possible way to blow a cost review.

Contract$ / monthNoticeLatest notice dateThe catch
Colo — 8 racks13,80090 days, written31 DecemberFalls inside the 20 Dec – 3 Jan blackout. See below — this is the one that bites.
VMware vSphere 7 — 8 hosts4,200Anniversary only, 60 days30 JanuaryAnniversary-only means missing it costs a full further year, not a month
SQL Server Enterprise (cluster)6,100Annual, 30 days28 FebruaryCheck whether any licence is portable to the cloud before cancelling outright
Oracle DB 12c SE2,600Annual renewal month 7, 45 daysMonth 6, week 2 (mid-December)This is Lever 6. Miss the notice and Lever 6 is worth $0 this year.
Backup software + tape library maintenance1,90090 daysExtended by one month, deliberatelyMust outlive the D2 restore test. Cancel this on time and a failed restore leaves you with no reader.
On-prem load balancer support70030 days28 February
Flagship server room HVAC service60030 days28 FebruaryPrintSrv still lives at the flagship — keep some cooling, downgrade the contract rather than cancelling it
Windows / AV / monitoring80030 days28 FebruaryCloud equivalents already billing from month 1 — this is double-paid until cancelled
▶ Work this out before you open it

The colo contract ends 31 March. Given the hand-back inspection needs 14 days’ notice, that the month-to-month extension is 2.4× with a one-month minimum, that the spring vaccination blackout runs two weeks in mid-March, and that the colo requires 90 days’ written notice — what is the real last working day for physical decommission, and what is the real last date to serve notice? Write both down before you look.

Model answer — the deadline, and the three deadlines hiding behind it
  1. 31 March — the contract end date. This is the date everyone quotes and the only one that is not actually a deadline for anyone doing work.
  2. 15 March — racks empty, inspection booked. The hand-back inspection needs 14 days’ notice, and anything still racked on 1 April rolls onto the month-to-month rate at 2.4× ($33,120) with a one-month minimum. So the physical work must be finished, and the inspection booked, by mid-March.
  3. 28 February — the real engineering deadline. 15 March sits inside the two-week spring vaccination campaign blackout. Nothing is touched, moved, powered off, or driven to a landfill during that window. The last working day for physical decommission is therefore the end of February — which also happens to be the latest notice date for four of the eight contracts above, so the last week of February is genuinely busy and should be planned as such in month 5, not discovered in month 8.
  4. 18 December — the date that actually decides whether any of this works. The colo’s 90 days’ written notice must reach them by 31 December, which is inside the 20 Dec – 3 Jan emergency-cover blackout. The blackout is a cutover freeze, not an administrative one — but the person authorised to sign is on leave, and “we’ll email it on the 29th” is how organizations accidentally buy a fourth quarter. So: board resolution in November, notice served and acknowledged in writing by 18 December, before anyone goes anywhere. The Oracle notice (mid-December) lands in the same week, which makes that week a single named milestone in the wave plan rather than two things somebody remembers.

The general lesson, which transfers to every migration you will ever do: the contract end date is never the deadline. Work backwards through hand-back logistics, then blackout windows, then notice periods, and put the earliest resulting date on the plan as the real one. On this estate that walks 31 March back to 18 December — three and a half months earlier than the number on the contract.

Artifact D — the one-page retro

☺ Like you’re 10: Write down what surprised you, what slowed you down, and what you’d change — honestly, including the bit where you looked silly.

Wave Planning closes every wave with a short retrospective and folds the answers straight back into the tooling and the runbook — that is what makes each wave cheaper than the last. This is the same document at project scale: one page, three columns, no blame, and at least one entry that makes you look bad. A retro with no uncomfortable line in it was written for an audience rather than for the next migration.

# Brambleside migration — retro
Date:            Attendees:            Facilitator:

## What surprised us
- (the thing nobody had on the risk register)

## What slowed us down
- (with the number: how many days, and what it cost)

## What we would change
- (a specific change to a document, a template, or an order of operations
   — not "communicate better")

## What we would keep
- (the thing that worked, so it survives the next reorganisation)

## Actions
| # | Action | Owner | By when | Where it lands |
|---|--------|-------|---------|----------------|
| 1 |        |       |         | (runbook / template / wave plan / checklist) |
Model answer — Brambleside’s retro, as it would actually read

What surprised us

  • The colo’s 90-day notice landed inside the Christmas blackout. We served it on 18 December with four days to spare, and only because someone re-read the contract in November. Nobody had put a contractual date on the wave calendar — only technical ones.
  • The ImageVault appliance seed was the easy part. The delta catch-up over the shared 150 Mbps link was not, because the archive kept growing at 900 GB/month while the appliance was in transit. We had budgeted for the seed and not for the tail.
  • PIMS-Core measured at 38% of its provisioned CPU at Monday 09:10 and 9% on average. We nearly rightsized off the average.

What slowed us down

  • LabBridge’s single allow-listed IP. The reference lab’s change process took six weeks, not the two we assumed. It was a two-line change on our side and a six-week queue on theirs. Cost: it very nearly pushed Wave 2.
  • FileShare’s 1.1 million files. Two of our three copy attempts died on per-file metadata, not bandwidth. We had sized the job in terabytes, which told us nothing useful.
  • Three weeks arguing about PIMS-Core’s R in month 1. Doing Retire and Retain first would have shrunk the argument surface from 16 systems to 11 before anyone opened the vendor’s hosted-edition price list.

What we would change

  • Every external dependency with a third-party change process goes into Wave 0, day 1, with a named contact and a chase date — not Wave 0, week 4. LabBridge is now the worked example in the template.
  • Size instances from observed peak, not the spec sheet, from the first wave — we sized from the spec sheet in Waves 1 and 2 and paid for a whole extra rightsizing pass afterwards.
  • Every transfer estimate carries both a byte figure and an object-count figure. Bandwidth arithmetic alone is misleading above roughly a million files.
  • Every environment gets a retire date and an owner at the moment it is created. One “temporary” DevTest environment from Wave 1 is still there. It is on the Q1 list.
  • Contract notice periods go on the same calendar as the cutovers, in month 1.

What we would keep

  • The Retire/Retain-first discipline. Five systems never moved. Lever 1 was worth $5,980/month and cost nothing at bill time because the arguing was done in month 1.
  • The named go/no-go owner with a wall-clock deadline in the cutover runbook. It was used once, at 03:40, and the decision took ninety seconds.
  • Refusing to refactor anything before 31 March. We would make the same call again.

Actions

#ActionOwnerBy whenWhere it lands
1Add a “third-party change lead time” row to the Wave 0 templateMigration leadQ1 week 2Wave plan template
2Add object-count alongside byte-size to every transfer estimateData leadQ1 week 2Migration checklist
3Retire the surviving “temporary” DevTest environmentPlatform engineerQ1 week 6IaC repo (delete, don’t stop)
4Put all contract notice dates on the programme calendar in month 1Finance partnerQ1 week 4Programme calendar
5Document the two refused optimizations with their numbersMigration leadQ1 week 1Cost review, standing section
🎬 At the Migration Academy
🦊

Foxy: Hang on. We moved everything successfully and the bill went up? Fifty-seven thousand against forty-two? What did we do wrong?

🦥

Sol the Sloth: Nothing, yet. We carried Fort Rusty’s decisions across along with its data. Forty-one terabytes on the fastest storage there is, because that was the fastest way to cut over. Eleven test machines running at four percent, all night, every night, since 2019. The cloud didn’t make that expensive. It made it visible.

👺

Gizmo the Gremlin: Easy! Cut the standby copy of PIMS-Core. Two thousand one hundred a month, gone, and nothing will ever happen to it. Probably.

🐢

Timmy the Turtle: The emergency hospital is open at four in the morning, Gizmo. Four hours to recover is what the board agreed for a planned Sunday outage — you want to spend the whole tolerance on an unplanned Monday. No. That line is not for sale, and I’m writing down that you asked.

🦥

Sol: And we don’t need it. The big money was never in the standby. Seventy-eight percent of those X-rays haven’t been opened in two years. Move them to the attic tier, keep a ladder for the ones a vet actually asks for, and that one change is worth three of Gizmo’s.

🦋

Mira the Butterfly: Then — and only then — we rebuild. Two things. The booking path, one slice at a time behind a facade, starting with repeat prescriptions because nothing clinical is written there. And the claims feed, because we’re already inside that code for the Oracle exit.

🦊

Foxy: Only two? We’ve got sixteen systems.

🦋

Mira: We’ve got five people, and one of them left in September. A roadmap with six things on it is a roadmap where the sixth never happens — and usually the fourth and fifth don’t either.

🐿️

Nutty the Squirrel: And nobody’s finished until Fort Rusty is off. I counted eight contracts still billing. The colo needs ninety days’ written notice — which means the letter goes out before Christmas, or we buy another quarter at two-point-four times the rate.

👺

Gizmo: …so the deadline isn’t the deadline.

🦉

Professor Owl: The deadline is never the deadline, Gizmo. It’s the last date on a chain of earlier ones. Work backwards from it and put the earliest date on the plan. That’s the whole trick, and it’s the last thing this capstone has to teach you.

Milestones

☺ Like you’re 10: Tick each box only when the thing is actually written down and the sums actually add up — not because it “sounds right.”

Work these in order; each one uses the output of the one before. Progress saves in this browser.

0 / 12 milestones complete
1Write out the naive bill for your own estate and total it
One row per system plus landing zone, DR, backup and egress. Use the table above as your starting point, or re-derive it if your Part 1 dispositions differed.
Done when: every one of the 16 systems appears exactly once and the column totals.
2Bank Lever 1 — price Part 1’s Retire, Retain and Repurchase decisions
For each system that never moved or was replaced by SaaS, write the before, the after, and the evidence that made it uncontroversial.
Done when: every row cites evidence (an audit log, a timestamp, a price list) rather than an opinion.
3Rightsize from observed peak, and say what your peak moment is
For each surviving compute line, write provisioned size, observed peak, the moment that peak happens, and the new size with headroom.
Done when: no row is sized from a monthly average, and PIMS-Core’s peak moment is named as Monday 09:10.
4Schedule the non-prod estate and make the schedule a tag
Compute the hours ratio, split compute from storage, and write the tag key that carries the schedule plus the opt-out rule.
Done when: your saving only touches the compute portion, and an always-on exception requires a named owner and a review date.
5Tier the imaging archive — and write the clinical exception into the rule
Split hot from archive, budget the gateway and retrievals, then write the demotion rule as a sentence with two clauses.
Done when: the rule says “untouched 24 months and no open case,” and there is a line of budget for getting a study back.
6Draw the commitment line — and list what you deliberately excluded
Name the steady-state base eligible for a 1-year commitment, then write the exclusion list with one reason each.
Done when: nothing still moving, still in hypercare, or on Artifact B’s roadmap appears in the eligible column.
7Total the six levers and reconcile all three columns
Before − levers = after, on every row and on the total. Then add residual on-prem and compare to $42,000.
Done when: your three column totals reconcile and your steady-state run-rate is below the ceiling.
Deep version: the line-by-line answer key above
8Name the thing you refuse to optimize, and price the refusal
Write the tempting seventh lever, what it saves, what it actually costs, and the sentence you say in the room.
Done when: the refusal is in the pack with a number attached, so it doesn’t get re-litigated next quarter.
9Pick exactly two systems for the 12-month roadmap — and rule out the rest in writing
Run all 16 through the four tests. Two IN, fourteen OUT, one sentence each.
Done when: you have exactly two INs and can say which one you’d drop if the vacancy from month 3 stays open.
10Sketch the strangler fig: seam, first slice, traffic shift, switch-off date
Fill the one-page slice template for slice 1, then list slices 2–4 with their traffic shares and blackout conflicts.
Done when: every slice has a switch-off date and an evidence gate, and the old app is deleted from the IaC repo, not just stopped.
Concept: Strangler Fig
11Write the decommissioning plan — order, evidence, obligations, notice dates
Power-down order in reverse dependency order with a gate per step; the evidence list; the retention table with a test per row; the contract table with latest-notice dates.
Done when: no domain controller is “powered off” rather than demoted, the tape contract is the last one cancelled, and the earliest notice date is on the calendar before the nearest blackout.
12Run the retro and walk the whole five-document pack end to end
Fill the retro template, then read all five documents aloud in order as if the steering committee were in the room.
Done when: the retro has at least one entry that makes you look bad, and you can present the pack without looking anything up.

What “done” looks like for Part 5

☺ Like you’re 10: Here’s the checklist you can mark yourself against without asking anyone.

  1. Your steady-state run-rate is below $42,000/month, and your before-column, saving-column and after-column each reconcile to the same three totals.
  2. All six levers have a number you computed from the estate data, and no line has two levers double-counted on it.
  3. You can name the thing you refuse to optimize, the exact monthly number that refusing it costs, and the one sentence you say when the CFO asks for a seventh lever.
  4. Exactly two systems are named for modernization in the next twelve months, each with a first slice and a numeric success measure; every other system has a written one-sentence reason it is being left alone.
  5. The strangler fig’s first slice has a seam, a traffic-shift schedule, a rollback measured in minutes, and a switch-off date that was written when the slice was planned — not afterwards.
  6. Fort Rusty has a switch-off date, a power-down order in reverse dependency order with a gate before each step, and no domain controller on that list is “powered off” rather than demoted.
  7. Every retention obligation has a new home, a duration, and a test that proves the data is still readable in an accepted format by someone who is not you.
  8. Every contract has a notice period and a latest-notice date on a real calendar, and the earliest of those dates falls before the nearest blackout window.
  9. Your retro contains at least one entry that makes you look bad, and at least one action that lands in a template or a runbook rather than in someone’s memory.
  10. You can walk all five documents of the migration pack — disposition matrix, landing-zone one-pager, wave plan, cutover runbook with rollback, and this optimization-and-decommission plan — end to end, without looking anything up.
🐢 Timmy’s checkpoint

1. The naive bill is $57,400 against a $42,000 baseline. Which single lever moves the most money, and why is it the one people skip? 2. Which workloads are eligible for a one-year commitment on 1 April, and name one that is deliberately excluded and the reason. 3. You are $6,910/month under the ceiling and the CFO asks for one more cut. What do you refuse, what does refusing it cost, and what do you say? 4. The colo contract ends 31 March. What is the real last working day for physical decommission, and what are the constraints that pull it earlier?

Check your answers
  1. Lever 4 — tiering ImageVault’s cold 78% to archive, worth $6,160/month, more than rightsizing every server in the estate combined. It is skipped because storage is not where engineers instinctively look: the compute lines have familiar names and obvious knobs, while 41 TB sitting on a premium file service looks like a fact rather than a decision. It was a decision — the file service was chosen so PIMS-Core’s UNC path kept working through the cutover, which was right for the cutover and wrong for the steady state. Note also that it is the lever that carries the run-rate across the $42,000 line.
  2. Eligible: the steady-state compute base only — PIMS-Core compute ($6,400), StockCtrl ($900), BramblesideOnline’s web tier ($1,000), the AD pair ($400) and shared services ($500), $9,200 in total, at about 30% off for one year. Excluded, with reasons: ClaimsFeed, because it is actively replatforming off Oracle and committing to today’s shape means paying for a shape you are deleting; DevTest, because it is now off 64% of the week; ImageVault storage, because Lever 4 just changed its shape and you should watch a quarter first; anything still in hypercare, because a workload you might still roll back is not steady state; and BramblesideOnline’s managed MySQL, because the booking path is first on Artifact B’s roadmap. One year rather than three, because the estate has not yet seen a full seasonal cycle including the spring campaign.
  3. You refuse two things: dropping PIMS-Core’s DR from warm standby to pilot light ($2,100/month) and releasing the Wave 4 parallel-run capacity early (~$5,800 one-off). The first moves recovery time from about 20 minutes to about 4 hours, and the board’s 4-hour tolerance was written for a planned Sunday outage — spending all of it on an unplanned Monday failure, in a group running a 24-hour emergency hospital, converts a cost win into an outage. The second removes the only mechanism for comparing old and new PIMS-Core output before the point of no return. The sentence: “Both are available. Both convert a cost win into an outage risk, and we are already 16% under the number the board set. I’ll bring you a seventh lever from the modernization roadmap instead.” Then write both refusals into the pack with their numbers, so they are not re-litigated next quarter by someone who has only seen the saving.
  4. The real last working day is 28 February. Working backwards: 31 March is the contract end; racks must be empty and the hand-back inspection booked by 15 March, because the inspection needs 14 days’ notice and anything still racked on 1 April rolls onto the 2.4× month-to-month rate ($33,120) with a one-month minimum; and 15 March sits inside the two-week spring vaccination blackout, so nothing may be touched then — which walks it back to the end of February. Behind all three sits the date that actually decides it: the colo requires 90 days’ written notice, due by 31 December, which falls inside the 20 Dec – 3 Jan blackout while the authorised signatory is on leave. So the board resolution is taken in November and the notice is served and acknowledged by 18 December. The transferable rule: the contract end date is never the deadline — work backwards through hand-back logistics, blackout windows and notice periods, and put the earliest resulting date on the plan.

That is the capstone. Over five parts you took sixteen systems belonging to an organisation that had never done this before, decided what each one deserved, built somewhere for them to land, sequenced them across nine months with two blackouts and a hard stop, moved fifty terabytes and cut over the system every clinic touches every minute — and then, today, made it pay, chose what to rebuild, and gave back the keys. The five documents in your pack are the ones a real steering committee actually reads. Step back to Move Brambleside — Start Here to see how the parts fit together, or sharpen the individual muscles in the drills: Pick the Right R, Size the Data Move, and Write a Rollback Plan. For the theory behind this page, re-read Modernization, the cost, FinOps and DR sections of Security, Cost & Resilience, Stop 4 of The Journey, and Best Practices on watching the money — and when you run this for real, take the printable Migration Checklist with you.