Hands-On Labs · Guided Drills

Drill — Size the Data Move

Almost every decision in this course is a judgement call you defend with a sentence. This one isn’t. Whether the data arrives before the deadline is arithmetic — it has a right answer, the right answer does not care how confident anyone in the room sounds, and you can get to it on the back of an envelope in about four minutes. That is exactly why it is worth drilling until it’s automatic: you will be asked this question live, standing at a whiteboard, by someone who has already promised a date. Four scenarios below, each from a different invented organization, each rigged around one specific way the naive calculation lies to you. Nothing here connects to Brambleside or to any other drill — no continuity, no carried-over state, the way a real sizing question arrives with no context either. Budget about eighty minutes for all four, and do them with a calculator and a sheet of paper rather than a spreadsheet, because a whiteboard is where this actually gets asked.

☺ Explain it like I’m 10

Imagine you have to empty a paddling pool with a bucket, but there’s a tap running into it the whole time. If you bail faster than the tap fills, you’ll finish — you can even work out roughly when. If the tap is faster than your bucket, you will never finish, not in a day, not in a year, and no amount of trying harder changes that; you have to turn the tap down or get a bigger bucket. That is the single most important idea on this page. Two smaller ones come with it: sometimes it’s quicker to load the water into a van and drive it across town than to carry it bucket by bucket — and a thousand thimbles of water take much longer to carry than one bucket holding the same amount, because of all the picking up and putting down.

🐘🐦Your hosts for this drill: Ellie the Elephant & Pip the Hummingbird — Ellie counts the boxes and refuses to let one go missing; Pip knows exactly how fast the road is and does the sums out loud before anybody promises a date. 🦥 Sol the Sloth turns up in Scenario 3 for the egress bill, and 🐢 Timmy appears at the end to remind everyone that the copy is not the last step.
⚠ Before you start — what you need, and what you don’t

A calculator and a sheet of paper. That is the whole toolchain. There is no cloud account to open, no card to enter, and nothing to provision anywhere in this drill — every question here is answered with numbers you were given. If you want to put a real tool beside it afterwards, the transfer calculators built into AWS DataSync, Azure Data Box and Google Transfer Appliance (see The Tools Landscape) will show you the same arithmetic wearing a nicer coat — but no step below depends on one, and nothing here should ever cause a bill.

⚠ Every number on this page is invented

Harrowgate, Northmoor, Tessellate and Ardley & Vance do not exist. Every link speed, dollar figure, per-gigabyte rate, change rate and efficiency percentage below was made up to make the exercise bite, and none of it should ever be quoted at anyone as a real price. Real cloud pricing changes monthly, is tiered, varies by region and commitment, and is the one thing you must check against the provider’s own calculator on the day you need it. The skill being drilled here is the method — how you convert the units, find the binding constraint, and show your working — not the number that falls out.

How this drill works

☺ Like you’re 10: Four sums, each with a trick hidden in it, and the answers are at the bottom of each one so you can mark your own homework.

Each scenario gives you a short fact sheet — dataset size, link speed, the window you’re actually allowed to use, the rate at which the data is still growing, and where relevant a per-gigabyte exit charge — and then asks four or five questions. Work them in order, on paper, and write down the number you get before you open the key, even when you’re not confident. The keys are folded away under each scenario and contain every step of the working, with the units converted explicitly, so nobody is ever left wondering where a factor of eight came from.

The four scenarios rise in nastiness and each one is built around a specific failure of intuition:

#ScenarioThe thing it teachesTime box
1Harrowgate Housing Trust — 18 TB document archiveThe window, not the link speed, is your real capacity. And the crossover between wire and appliance is a decision, not just a number.15 min
2Northmoor Genomics — 310 TB and growingThe one that never converges. When the change rate beats the copy rate, there is no completion date at any bandwidth, and no appliance fixes it.25 min
3Tessellate Analytics — 100 TB, cloud to cloudThe exit toll. You pay per byte that leaves, every time it leaves, rehearsals included — and that lands in the business case, not the runbook.25 min
4Ardley & Vance — 900 GB in 4.1 million filesThe file-count trap. Bandwidth arithmetic gives a comfortable answer that is completely wrong, and the delta sync is worse than the first copy.15 min

The arithmetic, once — the method you’ll use four times

☺ Like you’re 10: Learn five small facts now and you’ll never have to look any of this up again.

Everything on this page is built from a handful of conversions and one test. Read this section properly; the four scenarios are just this method applied to progressively meaner inputs.

Four conversion facts

Storage is sold in bytes; links are sold in bits. Almost every wrong transfer estimate in the world is somebody who forgot that, so write these out until they’re automatic:

FactWhy it’s that way
1 TB = 1,000 GB = 1,000,000 MBStorage vendors and cloud bills use decimal terabytes, not the binary 1,024 kind. Use decimal here — and know that the difference is about 10%, which is real but is not what sinks a plan.
1 byte = 8 bits, so 1 MB = 8 MbThis is the ×8 that people drop. A 100 Mbps link does not move 100 MB per second — it moves 12.5 MB per second.
Time (s) = size in megabits ÷ link speed in MbpsThe whole calculation, in one line. Convert the size to megabits first and the rest is division.
1 hour = 3,600 s · 1 day = 86,400 s · 1 week = 604,800 sBecause the answer you actually need is in nights, days or weeks — never in seconds.

The two anchors worth memorizing

🐦 Pip’s two numbers

You only ever need to remember one line, because everything scales from it: 100 Mbps sustained = 45 GB per hour = 1.08 TB per day. Everything else is multiplication. A gigabit is ten times that — 450 GB/hour, 10.8 TB/day. Two useful shorthands fall straight out: GB per hour ≈ Mbps × 0.45, and TB per day ≈ Mbps × 0.0108. Say those in a meeting and you will be the only person in the room who knows whether the date is real.

Check the anchor once so you trust it, then never derive it again. 100 Mbps for one hour is 100 × 3,600 = 360,000 megabits; divide by 8 for megabytes and you get 45,000 MB, which is 45 GB. Over a full day, 45 × 24 = 1,080 GB = 1.08 TB. That’s it.

Sustained rateGB / hourTB / day (24×7)TB / week (24×7)
100 Mbps451.087.56
250 Mbps112.52.718.9
400 Mbps1804.3230.24
500 Mbps2255.437.8
1 Gbps45010.875.6
2 Gbps90021.6151.2
10 Gbps4,500108756

Every row above is a ceiling, computed at 100% efficiency, 24 hours a day, with nothing else on the wire. You will never see one of those numbers in real life. They are the starting point you then take two deductions from — efficiency, and the window.

Deduction one — the efficiency factor

A link rated at 400 Mbps does not deliver 400 Mbps of your files. Protocol overhead, encryption, acknowledgement round trips, checksumming, filesystem behaviour at the source and destination, and the ordinary background traffic of a business all take a cut. A defensible working assumption is 70–80% of the allocated rate for a well-tuned bulk transfer over a decent path, less over a long or lossy one. The important thing is not the exact figure — it’s that you write down which figure you assumed, so that when the real throughput comes in at 62% somebody can point at the line and adjust it, rather than re-litigating the whole plan.

Deduction two — the window is your real capacity

The number that matters is not the speed of the link. It is how much data actually crosses it in a week, which is the sustained rate multiplied by the hours you are permitted to use. A team allowed 400 Mbps for nine hours a night is not running a 400 Mbps migration; averaged across the whole week they are running a 120 Mbps one, and every date they quote should be computed from that. Convert your window to a 24×7-equivalent rate at least once, out loud, in front of the people arguing about the link speed. It ends the argument.

The test that decides everything — does it converge?

◆ The convergence test

Net progress per week = window capacity per week − change per week. If that number is zero or negative, stop calculating a completion date, because there isn’t one — not a long one, not a bad one, none. And the elapsed time to finish is remaining ÷ net progress, never remaining ÷ capacity. Dividing by capacity is the single most common mistake in this whole subject: it always produces a plausible-looking number of weeks, even when the true answer is “never”.

This is the step that gets skipped, and it costs more than any other omission on the page, because the skipped version still returns an answer. Nothing errors. Nobody notices. A date gets written on a slide, four months of effort go into it, and the pile is bigger at the end than it was at the beginning. Do the convergence test first, before you work out anything else, on every single dataset — it takes fifteen seconds and it is the only check here that can tell you the plan is impossible rather than merely slow.

The appliance round trip has five parts, not one

“Ship a box” sounds like a single event and is priced by beginners as one. It is five, and only the middle one is affected by how much data you have:

StageWhat sets its durationTypical shape
1 · Order and deliveryThe provider’s lead time and your address. Nothing you do speeds it up.Days — and it starts on the day someone actually clicks the button, which is usually later than the plan says.
2 · On-site fillYour local disk and LAN speed, not your internet link. The only stage that scales with dataset size.Hours to days. Parallel appliances shorten it proportionally.
3 · Return shippingThe courier. Fixed.Days.
4 · Provider ingestThe provider’s queue. Fixed, and not something you can escalate.Days.
5 · Verify, then catch up the deltaReading everything back to check it, then copying whatever changed while the box was in the post — over the wire.Hours to days, and this is the stage everyone forgets.

Because stages 1, 3 and 4 are constant, the appliance takes roughly the same elapsed time whether you send it 20 TB or 200 TB — which is exactly why it wins on big datasets and loses on small ones. The size at which the two approaches tie is the crossover: appliance round-trip days × net TB per day. Below it, the wire is faster. Above it, the box is. And the crossover is where the arithmetic stops and the judgement starts, because the box also brings chain-of-custody paperwork, physical handling, and a period during which your data exists on a device in a van.

🐢 Timmy’s addition — the copy is not the last step

Every plan on this page needs a line for validation, and validation is another full read of everything you just moved. Checksumming 250 TB off local disk at 500 MB/s takes 500,000 seconds — close to six days — and it is a day-one budget item, not a rounding error you discover on cutover weekend. Row counts, checksums and a sampled byte-for-byte compare are the three that catch different faults (see Data Migration & Data Gravity). A copy nobody verified is a rumour.

The artifact — one sheet, filled in four times

This drill produces a transfer sizing sheet: fifteen lines that answer any version of this question. Copy it onto paper now and fill one in per scenario. If you can produce this sheet from a fact list in under five minutes, you have the skill this drill exists to build.

TRANSFER SIZING SHEET — <dataset name>

 1  Size to move, after Retire        ______ TB  = ______ GB  = ______ Mb   (x1000, x1000, x8)
 2  Rated link                        ______ Mbps
 3  Allocated to this transfer        ______ Mbps
 4  Efficiency assumed                ______ %      -> sustained  ______ Mbps
 5  Throughput                        ______ GB/hour            (= sustained x 0.45)
 6  Usable window                     ______ hours/night  x  ______ nights/week
 7  Window capacity                   ______ GB/night = ______ TB/week
 8  Change rate                       ______ GB/day   = ______ TB/week
 9  NET PROGRESS  (7 - 8)             ______ TB/week      <-- if zero or less, STOP HERE
10  Elapsed to complete  (1 / 9)      ______ weeks
11  Appliance round trip              ______ days   (order + fill + ship + ingest + verify)
12  Crossover size  (11 x line 9/day) ______ TB
13  Egress leaving the source         ______ GB x $______/GB x ______ copies = $______
14  Validation read pass              ______ hours
15  DECISION: __________________________________________________________
    The one number that would change it: ______________________________

Scenario 1 — Harrowgate Housing Trust

☺ Like you’re 10: A straightforward one, to check your sums work — with one thing in it that looks like the answer and isn’t.

Harrowgate Housing Trust is a social landlord with 11,000 homes. It is closing its on-premises file infrastructure and moving a document archive — scanned tenancy agreements, repair photographs, gas-safety certificates — into cloud object storage. The archive is one folder tree of large-ish files; nothing exotic. The network team has been helpful and specific, which is rarer than it sounds and is most of what makes this scenario answerable.

FactValue
Dataset18 TB document archive, average file size around 4 MB
Site link1 Gbps business internet circuit, shared with everything else the Trust does
Allocated to the transfer400 Mbps, and zero during the working day — this is a hard cap set by the network team, not a target
Window21:00–06:00, nine hours, seven nights a week
Efficiency to assume80% of the allocated rate
Change rate12 GB/day of newly scanned documents, added continuously
Appliance optionRound trip of about 13 days end to end, including on-site fill and provider ingest
EgressNone — the source is Harrowgate’s own data centre, so no per-GB exit charge applies

Answer these, on paper, before you open the key:

  1. What is the sustained throughput in GB/hour, and what does one nine-hour night actually move?
  2. Express the window as a 24×7-equivalent rate in Mbps. (This is the number to quote when somebody says “but we have a gigabit”.)
  3. Run the convergence test, then work out how many nights the whole archive takes.
  4. Where is the crossover against the 13-day appliance, and which do you choose? Say why in one sentence.
  5. The network team comes back and says the window is being cut to four hours a night. Redo the answer. Does the decision change?
Check your working — Scenario 1
  1. Sustained throughput. 400 Mbps allocated × 80% efficiency = 320 Mbps sustained. Convert: 320 × 3,600 s = 1,152,000 megabits per hour; ÷ 8 = 144,000 MB = 144 GB/hour. (Or straight from the anchor: 320 × 0.45 = 144.) One nine-hour night therefore moves 144 × 9 = 1,296 GB, which is 1.296 TB per night.
  2. The 24×7-equivalent rate. 1,296 GB per day = 1,296,000 MB × 8 = 10,368,000 megabits, ÷ 86,400 seconds = 120 Mbps. That is the honest description of this transfer. The circuit is a gigabit; the migration is a 120 Mbps migration, and every date must be computed from 120, not from 1,000. Notice how large the gap is — a factor of more than eight — and that none of it is anybody’s fault or a tuning problem. It is simply what a nine-hour window at a 400 Mbps cap means.
  3. Convergence, then duration. Capacity is 1,296 GB per night; change is 12 GB per day. Net progress = 1,296 − 12 = 1,284 GB per night, comfortably positive, so this converges — the delta is consuming 0.93% of a night’s capacity and is genuinely negligible here. Duration = 18,000 GB ÷ 1,284 = 14.02 → 15 nights once you round up and allow one night for the accumulated tail. Just over two weeks. Note the discipline even though it didn’t matter: you divided by net progress, not capacity. Dividing by capacity would have given 13.9 nights. The gap is trivial in this scenario and catastrophic in the next one, and the only way to be reliably right in the second case is to have made it a habit in the first.
  4. Crossover and the decision. The appliance is a fixed ~13 days. The wire moves 1.284 TB of net progress per day, so the two approaches tie at 13 × 1.284 = 16.7 TB. At 18 TB you are just past the crossover: the appliance finishes in about 13 days, the wire in about 15. Choose the wire. The appliance is nominally two days faster, and two days does not move any milestone anyone cares about — while the wire has no chain-of-custody paperwork, no physical handling of tenant records, nothing to lose in a van, no ingest queue you cannot escalate, and it can be paused, resumed and re-run at will. The right sentence is: “the appliance is marginally faster and materially more complicated, and at 18 TB the margin doesn’t buy a date.” The point of computing the crossover is not to obey it — it is to know how much you are giving up when you don’t.
  5. Four-hour window. A night now moves 144 × 4 = 576 GB, net 564 GB. 18,000 ÷ 564 = 31.9 → 32 nights, so nearly five weeks. Against 13 days for the appliance, the decision flips decisively: ship the box. And this is the real lesson of Scenario 1 — the dataset didn’t change, the link didn’t change, the appliance didn’t change. The window changed, and the window was the capacity all along. Which is why line 6 of the sizing sheet is the line to nail down first, and why “how many hours a night do I actually get, and on which nights” is the first question you ask the network team, before you ask them anything about speed.

Scenario 2 — Northmoor Genomics

☺ Like you’re 10: The bucket-and-tap one. Bail carefully, then notice the tap.

Northmoor Genomics is a sequencing bureau: instruments run around the clock producing raw reads and aligned outputs, which land in an on-premises object store. The lease on the building housing that store ends, and the whole archive has to be in the cloud. The migration lead has produced a plan that says “approximately twelve months” and would like you to sanity-check it before it goes to the board on Thursday.

FactValue
Dataset310 TB total, of which 250 TB has not been read in 18 months and 60 TB is the live working set
Site link1 Gbps
Allocated to the transfer600 Mbps
Window00:00–06:00, six hours, five nights a week. The other 138 hours belong to the collaborator sync, which cannot be interrupted without a contract conversation
Efficiency to assume75%
Change rate1.4 TB/day, seven days a week — the instruments do not stop because you are migrating
Appliance optionAvailable; appliances hold roughly 80 TB usable each, and you can run several in parallel on-site
The offer on the tableThe network team will give you the entire 1 Gbps circuit, 24×7, for four weeks — if you can justify it in numbers

Answer these, on paper, before you open the key:

  1. What is the window capacity per week, in TB?
  2. What is the change per week, in TB? Run the convergence test and state the result in one plain sentence.
  3. The plan says twelve months. Where did that number come from, and what is wrong with it?
  4. What sustained in-window rate would you need merely to stand still? What allocated rate is that, at 75% efficiency — and is it available?
  5. Does putting the 250 TB cold pile on appliances fix the problem? Answer carefully.
  6. Write the plan you would actually take to Thursday’s board, with the numbers that justify the four-week circuit request.
Check your working — Scenario 2
  1. Window capacity. 600 Mbps × 75% = 450 Mbps sustained → 450 × 0.45 = 202.5 GB/hour. A six-hour night moves 202.5 × 6 = 1,215 GB = 1.215 TB per night. Five nights a week gives 6.075 TB per week.
  2. Change per week, and the test. 1.4 TB/day × 7 days = 9.8 TB per week. Net progress = 6.075 − 9.8 = −3.725 TB per week. In one plain sentence: “The archive is growing 3.7 TB a week faster than we can copy it, so the online transfer never completes — not late, not slowly, never — and the pile will be about 100 TB bigger after six months of effort than it is today.” Say it exactly like that, with the number, because “it won’t work” gets argued with and “it diverges by 3.7 TB a week” does not.
  3. Where twelve months came from. Somebody divided 310 TB by the weekly capacity of 6.075 TB and got 51 weeks, which rounds to a year, and it looks like a proper answer because it is a proper-looking number with a decimal point in it. It ignores the change rate entirely — it silently assumes the dataset holds still while you copy it. That assumption is invisible, nobody states it out loud, and it is wrong on almost every dataset that is still in use. This is the mistake the whole drill exists to inoculate you against. The tell is structural rather than numerical: any estimate produced by dividing size by capacity, with no line for change rate anywhere on the page, is unsafe regardless of what it says.
  4. The break-even rate. To stand still you must move 9.8 TB inside the 30 hours a week you are allowed (5 nights × 6 h). 9.8 TB = 9,800 GB ÷ 30 h = 326.7 GB/hour, and dividing by the 0.45 anchor gives 725.9 Mbps sustained. At 75% efficiency the allocated rate needed is 725.9 ÷ 0.75 = 967.9 Mbps — which is 97% of the entire 1 Gbps circuit, dedicated to the migration during the exact hours the collaborator sync also needs it, purely to break even and make zero net progress. It is not available, and even if it were it would buy you nothing. That is the moment you stop optimising the copy and go and have a different conversation, with different people, about the window and the data.
  5. Do appliances fix it? No — and understanding why is the point. Appliances change the starting size. They do not change the net progress rate, which is a property of the window and the change rate alone. Put all 250 TB of cold data on boxes and you are left with the 60 TB live set against a net progress of −3.725 TB/week: still negative, still no completion date, just a smaller number that never reaches zero. Subtracting from a quantity that is growing faster than you subtract is not progress. Generalize it: no arrangement of appliances, and no amount of patience, rescues a delta that exceeds the pipe. Fix the sign of net progress first; only then does anything else you do to the plan have meaning.
  6. The plan for Thursday. Two levers, and they only work together — that is the finding.
    • Lever one — the cold pile goes by road. 250 TB of data untouched in 18 months has no business consuming a single second of a contended link. At roughly 80 TB usable per appliance that is 4 units, ordered in week 1 because the lead time is the schedule. Filling them in parallel on-site at around 300 MB/s each takes 250,000,000 MB ÷ 1,200 MB/s ≈ 208,000 s ≈ 58 hours, so under three days of local copying that never touches the internet circuit at all.
    • Lever two — the four-week full circuit, justified in numbers. The whole 1 Gbps at 75% is 750 Mbps sustained = 337.5 GB/hour; run 24×7 that is 337.5 × 168 = 56,700 GB = 56.7 TB per week. Net of the 9.8 TB/week of new data, net progress is +46.9 TB per week — positive, decisively, for the first time. Now the sizing: the full 310 TB would need 310 ÷ 46.9 = 6.6 weeks, which exceeds the four weeks on offer, so asking for the circuit on its own fails. The 60 TB live set needs 60 ÷ 46.9 = 1.28 weeks, about nine days — which fits inside four weeks with nineteen days of margin for re-runs, validation and the inevitable bad night.

    So the ask you write down is: four appliances ordered in week 1 for the 250 TB cold archive, and the full circuit for four weeks to carry the 60 TB live set plus its ongoing delta in about nine days, with the remaining margin spent on validation. Neither lever works alone — the appliances alone leave you diverging, and the circuit alone can’t finish in the window offered. That is worth noticing as a general shape: the answer to a transfer problem is very often two levers that are each individually insufficient, which is exactly why people who evaluate them one at a time conclude the problem is unsolvable.

    And one line for Timmy: verifying 250 TB by reading it back at 500 MB/s is close to six days. Put it on the plan as its own bar, not as a footnote.

Scenario 3 — Tessellate Analytics

☺ Like you’re 10: Getting your stuff into the storage unit is free. Getting it out costs money per box — and you pay again every time you take it out to practise.

Tessellate Analytics runs a customer-facing reporting product and is moving from one public cloud to another to consolidate onto the platform its parent company already uses. The transfer arithmetic here is the easy part; the invoice is not. This is the cloud-to-cloud case, where the constraint changes from time to money.

FactValue
Dataset96 TB in the source cloud’s object storage + a 4 TB managed Postgres = 100 TB
Of which derived22 TB is materialized aggregate tables, recomputable in about 9 hours from 8 TB of raw inputs that are themselves already in the 96 TB
Achievable throughput2 Gbps sustained aggregate, using parallel transfer workers, 24×7 — cloud to cloud, so no nightly window
Change rate500 GB/day
Egress rate$0.062 per GB leaving the source cloud. Invented for this exercise — real providers tier this, and the tiering is itself a lever
Ingress at the destinationFree, as it almost always is. The asymmetry is the whole point
Planned rehearsalsOne full dry run, then the real run, then roughly 5 TB of delta re-syncs
Parallel run6 weeks after cutover, during which the new system reads about 900 GB/day back out of the old cloud
Business caseThe move is projected to save $4,100/month once complete

Answer these, on paper, before you open the key:

  1. How long does the transfer take? Run the convergence test first, out of habit.
  2. What does one full copy cost in egress?
  3. What does the whole programme cost in egress — dry run, real run, deltas, and the parallel run? What number do you actually put in the budget?
  4. Express that as months of payback against the projected saving. What kind of cost is it, and why does that matter more than its size?
  5. Would an appliance help here? Answer in one sentence and say why.
  6. Name the one lever that meaningfully reduces this bill, and compute what it saves.
Check your working — Scenario 3
  1. Duration. 2,000 Mbps × 0.45 = 900 GB/hour = 21.6 TB/day. Change is 0.5 TB/day, so net progress = 21.1 TB/day — positive by a factor of forty-two, and the delta is 2.3% of capacity. Converges trivially. 100 ÷ 21.1 = 4.74 → about 5 days. Run the test anyway, every time: it costs fifteen seconds and it is the only thing standing between you and Scenario 2.
  2. One copy. 100 TB = 100,000 GB × $0.062 = $6,200. This is the number most people quote, and it is the number that is wrong.
  3. The whole programme. Egress is charged per byte that leaves, every time it leaves. The dry run is 100 TB out. The real run is another 100 TB out. The deltas are 5 TB. That is 205 TB = 205,000 GB × $0.062 = $12,710. Then the parallel run: 900 GB/day × 42 days = 37,800 GB × $0.062 = $2,343.60 of egress happening after the migration is nominally finished, which is the line item nobody budgets because it doesn’t feel like migration any more. Total: $15,053.60. Put $20,000 in the budget — a third of contingency, because backups, cross-region replication, snapshot exports and log shipping inside the source cloud may also be billed as egress, and because you will do one more rehearsal than you currently think you will. The general rule: estimate the copies, not the copy. Two-and-a-bit copies is a realistic planning assumption for a migration that is being done carefully, and being done carefully is not optional.
  4. Payback and character. $15,053.60 ÷ $4,100/month = 3.67 months of the entire projected saving, consumed by the exit toll alone, before a single hour of engineering time is counted. That does not kill the move — three-and-a-half months against a permanent saving is a perfectly reasonable trade. What matters is its character: it is one-way, up front, and non-refundable. If the migration is abandoned at 60% for reasons that have nothing to do with data, that money is spent and nothing was bought with it. That asymmetry is why the egress line belongs in the business case on day one rather than in the runbook on cutover week — a cost that cannot be unwound changes how you sequence the decision to commit, not just how you budget for it.
  5. Appliances? No — a physical export from a cloud provider is a niche service, and where it exists you are generally still billed for the data leaving, so the box changes how the bytes travel without changing the toll that made this expensive in the first place. Appliances solve a time problem; this is a money problem, and here time is already fine at five days.
  6. The lever that works: don’t move what you can rebuild. Of the 96 TB, 22 TB is derived — materialized aggregates recomputable in about nine hours from raw inputs you are transferring anyway. Recompute them at the destination instead of shipping them. That removes 22 TB from every copy: 22,000 GB × $0.062 = $1,364 saved per copy, and across the dry run and the real run, $2,728. The new totals are 161 TB of migration egress ($9,982) plus the unchanged $2,343.60 of parallel-run reads = $12,325.60, and the transfer shortens to 78 ÷ 21.1 = 3.7 days. Payback falls from 3.67 months to 3.0 months. Sol’s framing is the one to remember: compute at the destination is almost always cheaper than egress from the source — so before you price a cloud-to-cloud move, go through the dataset and separate what is original from what is merely derived, because you only have to pay to move the originals. And note where that idea came from: it is the Retire pass from the disposition matrix applied to bytes instead of servers. Every terabyte you talk yourself out of moving is discounted twice — once on the transfer, once on the bill.

Scenario 4 — Ardley & Vance

☺ Like you’re 10: Carrying a thousand thimbles takes far longer than carrying one bucket with the same water in it.

Ardley & Vance is a structural engineering practice. Its project share holds drawings, calculation sheets, site photographs and revision history going back to 2009. Compared with everything above it is a rounding error in size — which is precisely why it is the scenario that catches experienced people out.

FactValue
Dataset900 GB — in 4.1 million files
Sustained throughput available300 Mbps, no window restriction
Per-file overhead8 ms per file for the metadata operations — open, stat, create, set attributes and ACLs, set timestamps, close. Invented; the real figure depends on protocol, round-trip latency and how the tool batches, and it is the number worth measuring on a sample before you plan anything
Concurrency availableThe transfer tool supports up to 16 parallel workers
Destination request charge$0.0045 per 1,000 write requests. Invented
The cutover plan says“Final delta sync: ten minutes, it’s only a few hundred megabytes of changes”

Answer these, on paper, before you open the key:

  1. What is the average file size, and what does the bandwidth-only calculation say the copy takes?
  2. What does the metadata cost, single-threaded? Which of the two is the binding constraint?
  3. What happens with 16 workers, and what does that tell you about what to buy — bandwidth or concurrency?
  4. What is the diagnostic that tells you which of the two you are short of, while a transfer is running?
  5. What does the destination’s request charge come to, and what is the honest conclusion about it?
  6. Is the cutover plan’s “ten minutes” right? Show the number, and say what you would change.
Check your working — Scenario 4
  1. Average file size and the naive answer. 900 GB = 900,000 MB ÷ 4,100,000 files = 0.22 MB, so about 220 KB average — small, and small in a way that matters. Bandwidth-only: 300 Mbps × 0.45 = 135 GB/hour, so 900 ÷ 135 = 6.67 hours. “It’s a Saturday morning.” That answer is arithmetically correct and operationally worthless.
  2. Metadata, and the binding constraint. 4,100,000 files × 8 ms = 32,800 seconds = 9.11 hours of pure per-file overhead, single-threaded — and that is time during which the link is largely idle, because the tool is waiting on round trips, not on bandwidth. 9.11 > 6.67, so the transfer is metadata-bound, not bandwidth-bound. Every instinct you built in Scenarios 1 through 3 points at the wrong resource here. The tell is on the fact sheet before you compute anything: whenever average file size is well under about a megabyte, do the file-count calculation first, because at that size the per-file cost dominates the per-byte cost and the answer will not come from the link speed.
  3. Sixteen workers. Per-file overhead parallelizes almost perfectly, because it is latency you are waiting on rather than a resource you are saturating: 32,800 s ÷ 16 = 2,050 s = 0.57 hours, about 34 minutes. Metadata drops below the 6.67-hour bandwidth time and bandwidth becomes the binding constraint again. So what you needed was not a bigger link — it was concurrency, which is a flag on a command line and costs nothing. Diagnose before you buy. A team that responds to “the copy is too slow” by upgrading the circuit here would spend real money and change the completion time by zero.
  4. The diagnostic. Watch files per second alongside bytes per second, and compare each against what you need. Here, finishing in 6.67 hours means sustaining 4,100,000 ÷ 24,000 s ≈ 171 files/second and 37.5 MB/s (= 300 Mbps) simultaneously. If bytes/second is far below the link’s capability while files/second is pinned flat, you are metadata-bound — add workers. If files/second is comfortable and bytes/second is at the ceiling, you are bandwidth-bound — and only then is a bigger pipe the answer. Most transfer tools report both; almost nobody looks at the first one.
  5. Request charges. 4.1 million writes ÷ 1,000 = 4,100 units × $0.0045 = $18.45. Trivial. And that is the honest, slightly deflating conclusion worth internalising: with millions of small files, people worry loudly about the per-request money and quietly ignore the per-request time, when it is the time that costs them the cutover window. The fear is pointed at the wrong one. (Keep an eye on it at a hundred times this file count, and keep an eye on per-request charges for reads in a hot workload — but for a one-off migration of 4.1 million files, $18 is not a line item.)
  6. The cutover plan is wrong, and it is the most dangerous line in the scenario. A delta sync must first work out what changed, which means re-stat-ing all 4.1 million files — the same metadata pass as the first copy, whether or not a single byte has changed. At 16 workers that is 34 minutes of scanning before one changed byte moves; single-threaded, or with concurrency not carried over into the cutover script, it is 9.1 hours and the cutover fails outright. The principle, and it generalizes to every file-based migration you will ever plan: scan time is a function of file count, not of change size. A “tiny” delta sync over a huge tree is not tiny. What to change, in order of preference: (a) use a tool that reads a change journal — an NTFS USN journal, an inotify-style feed, a snapshot-diff — so the scan is proportional to the number of changes rather than the number of files, which turns 34 minutes into seconds; (b) pre-seed and freeze, moving the bulk days early and making the share read-only well before the window, so the final delta is genuinely small and genuinely quick to find; (c) bundle the cold subtrees into archives before transfer, accepting the loss of per-file granularity and per-file restore in exchange for line-rate streaming. And whichever you pick — time a rehearsal of the delta sync on the real tree, because this is a number you can measure for free a week early, and it is the number that decides whether cutover night ends at 02:40 or at 11:00 the next morning.

What the four scenarios share

☺ Like you’re 10: Four different sums, one habit — find the thing that’s actually slowing you down before you try to fix anything.

None of the four was hard arithmetic. Every one of them was a case where the obvious calculation returned a confident, plausible, well-formed number that was wrong, and where the correct answer came from asking which resource was actually binding before dividing anything by anything. Harrowgate’s binding constraint was the window, not the gigabit on the invoice. Northmoor’s was the change rate, which made the completion date not merely wrong but nonexistent. Tessellate’s was money leaving per byte, several times over, long after the transfer was declared finished. Ardley & Vance’s was the number of files, in a dataset small enough that nobody thought to check.

The habit that catches all four is a single reordering: before you divide size by speed, write down what you are short of. Hours in the window, headroom against the change rate, dollars per gigabyte, or operations per second. The division is the last step, not the first — and the sizing sheet above exists to make you fill in the constraints before you are allowed to reach line 10.

◆ The four questions, in order

Ask them in this sequence on every dataset and you cannot produce a Scenario 2 answer by accident: 1. How much is there, after deciding what you refuse to move? 2. How many hours a week am I actually allowed on the wire, and at what sustained rate? 3. Does that beat the change rate — and by how much? 4. What does it cost to leave, times the number of copies I will really make? Only then divide. Anything that appears to be a transfer plan and does not answer all four is an estimate wearing a plan’s clothes.

🐦 Pip’s challenge · going further

Re-run each scenario with one variable moved and watch how far the answer travels. Give Harrowgate a 10 Gbps circuit but keep the nine-hour window — does the appliance decision change, and does the completion date change as much as the link speed did? Halve Northmoor’s change rate to 0.7 TB/day and find the exact window, in hours per week, at which the original plan starts converging. Make Tessellate’s parallel run twelve weeks instead of six and see what it does to payback — then decide whether you would shorten the parallel run to save money, and what you would be trading away if you did. Give Ardley & Vance 40 million files instead of 4.1 million and work out at what file count the archive-bundling option stops being an optimisation and becomes the only viable approach. Then do the one that is genuinely uncomfortable: take a real dataset you are responsible for, and fill in the sizing sheet for it honestly, including line 8.

Milestones

☺ Like you’re 10: Ten ticks. Progress saves in this browser, so you can stop after two scenarios and come back.

0 / 10 milestones complete
1Write the conversion facts and the anchor from memory
TB → GB → MB → Mb, the ×8, seconds in an hour/day/week, and the 100 Mbps anchor. Close this page while you do it.
Done when: you can state 45 GB/hour and 1.08 TB/day at 100 Mbps without looking, and derive 2 Gbps from it in your head.
2Scenario 1 — sustained throughput and the 24×7-equivalent rate
Apply the efficiency factor, get GB/hour and TB/night, then express the nine-hour window as a single Mbps figure.
Done when: you have 144 GB/hour, 1.296 TB/night, and the sentence “this is a 120 Mbps migration on a gigabit circuit”.
3Scenario 1 — nights, crossover, and a defended decision
Divide by net progress, not capacity. Compute the crossover against the 13-day appliance and choose, in one sentence.
Done when: you have 15 nights, a crossover near 16.7 TB, and a decision whose reason is not the number of days.
4Scenario 2 — capacity per week against change per week
Both in TB/week, side by side, before you calculate anything else about this dataset.
Done when: you have 6.075 TB/week of capacity against 9.8 TB/week of change written on the same line.
5Scenario 2 — state the divergence in one plain sentence
Not “it’s too slow”. A number, a unit, and the word “never”. Then say where the plan’s twelve-month figure came from.
Done when: your sentence contains −3.725 TB/week and you can name the exact division that produced the false 51-week answer.
6Scenario 2 — the two-lever plan, justified in numbers
Appliances for the cold pile, full circuit for the live set. Show why each lever fails on its own.
Done when: you can show +46.9 TB/week net on the full circuit, 6.6 weeks for all of it, and ~9 days for the 60 TB live set.
7Scenario 3 — duration, then the egress bill for every copy
Dry run, real run, deltas, and the six weeks of parallel-run reads that happen after everyone thinks it’s over.
Done when: you have ~5 days, $12,710 of migration egress, $2,343.60 of parallel-run egress, and a budget ask above the total.
8Scenario 3 — payback, and the lever that reduces it
Months of saving consumed by the toll; then separate original data from derived data and recompute the bill.
Done when: you have 3.67 months falling to 3.0 months, and can say why a one-way cost changes sequencing and not just budgeting.
9Scenario 4 — bandwidth time against metadata time
Compute both, name the binding constraint, then redo it with 16 workers and name it again.
Done when: you have 6.67 h of bandwidth against 9.11 h of metadata, and 34 minutes at 16 workers — and can say what you would have wasted money on.
10Scenario 4 — kill the “ten-minute delta sync”
Show the scan cost of the final sync and write the fix you would put in the cutover runbook instead.
Done when: you can state scan time is a function of file count, not change size, and name a change-journal, pre-seed-and-freeze or bundling fix.
🎬 At the Migration Academy
🦊

Foxy: Northmoor’s plan says twelve months. That’s a long time, but it’s a number, and the board can plan around a number. What exactly is your objection?

🐦

Pip: My objection is that it isn’t a number, Foxy, it’s a division. Somebody took three hundred and ten terabytes and divided it by what the link moves in a week. That sum assumes the sequencers stop. They don’t stop. They add one-point-four terabytes a day while you copy.

🐘

Ellie the Elephant: Six point zero seven five terabytes a week in, nine point eight terabytes a week out. The pile grows by three point seven every week you work on it. In twelve months you would not be finished — you would be a hundred and ninety terabytes further behind than the day you started.

🦊

Foxy: Then we run it longer.

🐦

Pip: Longer is the one thing that cannot possibly help. This isn’t slow. Slow finishes eventually. This never finishes, at any date, because you are moving away from the finish line the entire time you are running towards it.

👺

Gizmo the Gremlin: Then start it tonight anyway! Something is better than nothing. You can always fix the sums later, when there’s more information.

🦥

Sol the Sloth: Something is worse than nothing here, Gizmo. Six months of a transfer that cannot converge costs six months of contended circuit, six months of engineer attention, and — if this were Tessellate’s move rather than Northmoor’s — six months of egress charges, per byte, for bytes that will all have to leave again anyway. You would be paying, repeatedly, to not arrive.

🐘

Ellie the Elephant: Four appliances for the two hundred and fifty terabytes nobody has read since the year before last. Four weeks of the whole circuit for the sixty that people actually use. Nine days, and then it is done. Neither half works on its own, which is why everyone who checked them one at a time said it was impossible.

🐢

Timmy the Turtle: And then six days of reading it all back to prove it arrived intact, which goes on the plan as its own bar with its own dates. A copy nobody verified is a rumour, and I have never yet met a rumour that held up at four in the morning.

🦊

Foxy: So the skill isn’t doing the division. It’s knowing which four things to write down before you’re allowed to divide anything.

🐦

Pip: That’s the whole drill in one sentence. Now go and do Ardley & Vance, and this time don’t look at the size.

🐢 Timmy’s checkpoint

1. Convert 18 TB into megabits, showing each step, and say how long it takes at 320 Mbps sustained. 2. Write the convergence test as a formula, and say what you do when it comes out negative. 3. Why does the elapsed time for an offline appliance barely change between a 20 TB and a 200 TB dataset — and what is the one stage that does scale? 4. A colleague budgets $6,200 of egress for a 100 TB cloud-to-cloud move. Name three separate reasons the real bill will be higher. 5. You have 900 GB in 4.1 million files and 300 Mbps. Which resource are you short of, and how would you know while the transfer was running?

Check your answers
  1. 18 TB → megabits, and the time. 18 TB × 1,000 = 18,000 GB; × 1,000 = 18,000,000 MB; × 8 = 144,000,000 megabits. At 320 Mbps sustained: 144,000,000 ÷ 320 = 450,000 seconds; ÷ 3,600 = 125 hours of actual transfer time. Note that 125 hours is not the answer to “when will it be done” — spread across a nine-hour nightly window it is 13.9 nights, and against a 12 GB/day change rate it is 15 nights. Transfer hours and elapsed days are different quantities, and confusing them is how a fortnight gets promised as five days.
  2. The convergence test. Net progress = window capacity per period − change per period, both in the same unit over the same period. If it is zero or negative, stop — there is no completion date to calculate, and you must change the window, the bandwidth, or the amount of data before any transfer plan means anything. Completion time is remaining ÷ net progress, never remaining ÷ capacity; the second version always returns a plausible number, including when the true answer is “never”, which is exactly what makes it dangerous rather than merely inaccurate.
  3. Why the appliance is size-insensitive. Four of its five stages — order and delivery, return shipping, provider ingest, and the delta catch-up afterwards — are set by lead times and courier schedules, not by how much data you put on the device. Only stage 2, the on-site fill, scales with dataset size, and that runs at local disk and LAN speed rather than over your internet link, so it is fast and can be parallelized across multiple appliances. That fixed-cost shape is precisely why the appliance loses on small datasets (you wait days to move something the wire would have carried overnight) and wins on large ones (the same days move ten times as much). The tie point is the crossover: round-trip days × net TB per day.
  4. Three reasons the $6,200 is low. (a) Rehearsals. Egress is charged every time bytes leave, so a dry run plus the real run plus deltas is two-and-a-bit copies, not one — around $12,710 on these figures. (b) The parallel run. After cutover the new system keeps reading from the old cloud for weeks, and every one of those reads is billed egress; at 900 GB/day for six weeks that is a further $2,343.60, arriving on an invoice after the project is declared finished. (c) Everything else that leaves. Cross-region replication, backup copies, snapshot and image exports, and log shipping out of the source account can all be billed as egress, and none of them appear in a dataset-size estimate. Any two of those three earns the mark; a fourth worth knowing is that real egress pricing is tiered, so a single blended per-GB rate applied to a large volume can be wrong in either direction — which is the reason to price it against the provider’s own calculator on the day rather than from a number in a training page.
  5. Which resource, and how you would know. You are short of operations, not bandwidth: 4.1 million files at 8 ms of per-file overhead is 9.11 hours single-threaded against 6.67 hours of bandwidth time, so the copy is metadata-bound and a faster link would change the completion time by nothing. You would know while it was running by watching files per second next to bytes per second: bytes/second well under what the link can do while files/second sits pinned flat means add concurrency; files/second comfortable while bytes/second is at the ceiling means you genuinely need more bandwidth. With 16 workers the metadata cost falls to about 34 minutes and bandwidth becomes binding again — the fix was a flag, not a circuit upgrade. And the same file count silently governs the delta sync, because finding what changed means re-stat-ing all 4.1 million files whether or not anything changed at all.

Four sheets filled in? That is the drill. For the theory underneath every calculation on this page — online versus offline, bulk load versus continuous replication, validation, and data gravity — go to Data Migration & Data Gravity; for the exit toll and why it behaves differently from every other migration cost, Cloud-to-Cloud, Hybrid & Repatriation; and for the storage and networking choices that set your window in the first place, Security, Cost & Resilience. Ready for a different single skill? Try Drill — Pick the Right R or Drill — Write a Rollback Plan. To put all of this against one estate with a deadline attached, start at Move Brambleside — Start Here, where Part 4 makes you do exactly this arithmetic for 41 TB of imaging archive on a 150 Mbps link — and then live with the answer.