Interview Prep

Interview Prep

Interviews for a Cloud Migration Engineer or Migration Consultant role probe a narrow, deep set of things: can you count an estate before you move it, can you pick the right R and defend it, can you move a database without losing a row, and can you flip a switch you know how to flip back. This page is a working question bank for exactly those things — 55 technical questions with model answers across all thirteen domains of this course, plus the open-ended scenario rounds, behavioural answers you can slot your own experience into, questions worth asking them, and a plan for the last week.

⚠ What this page is — and what it is not

This is a curated question bank and technique guide, assembled from this course's own curriculum. It is not a transcript of any real interview, and it is not built from a specific employer's job description, hiring rubric, or shared syllabus. The questions are phrased the way an interviewer would phrase them, and the emphases below are this course's judgement about what interviewers commonly probe for this role — not data from a published mark scheme. The model answers are written to teach the idea, not to be recited. Wherever an answer touches on your own work, treat it as a skeleton and swap in your own projects, your own numbers, and your own scars. A shorter honest answer beats a borrowed fluent one, every time.

☺ Explain it like I'm 10

Before a big test, you don't just read the book again — you get someone to ask you questions out loud and see which ones make you go "um." This page is a big pile of those questions, with the answers written underneath so you can check yourself. Cover the answer, say yours out loud, then look.

🦊🦉Your hosts for this topic: Foxy & Professor Owl — Foxy asks the awkward version of every question and won't accept "I basically knew that"; Owl gives the structured answer underneath and shows you the shape to copy. 🐢 Timmy checks that nothing here is asserted more confidently than it deserves.
◆ How to read this page

The technical sections run in course order, domain by domain, so this page lines up with the lessons you already read — and each section links back to the lesson that covers it. Difficulty markers: 🟢 warm-up — you should be fluent and fast; 🟡 medium — a structured answer plus a concrete example; 🔴 hard — a trade-off or a scenario where the interviewer wants to hear you think, not recite. Read a question, answer it out loud before you look, then compare. The gap between what you said and what is written is your actual study list.

What the role actually gets asked

☺ Like you're 10: A migration interview isn't one big test — it's usually three or four smaller ones, and each is checking a different thing about you.

Migration hiring is unusual in one specific way: the job is mostly judgement under constraint. Very little of it is inventing new technology; almost all of it is choosing well among known options, with incomplete information and a deadline. That shapes how the interviews are built. Most processes for this role run through some version of four stages, and each one is genuinely testing something different — the mistake candidates make is answering all four the same way.

Stage 1 — The screen

Usually around half an hour, often with a recruiter or a hiring manager rather than a deep specialist. It is checking vocabulary and shape: can you say what a migration is, name the 7 R's, describe a project you were on without rambling, and give a coherent answer to "what did you personally do." Nobody is trying to catch you out here. The failure mode is the opposite of what candidates expect — not being caught short on depth, but talking for six minutes when ninety seconds was the right length. Have a two-minute version of your own migration story ready, in this order: what the estate was, what the driver was, what your part was, and how it ended.

Stage 2 — The technical deep-dive

The core round, usually with a senior engineer or an architect. It works down through the domains: discovery, the R's, data movement, cutover, landing zone, cost, security. Questions start definitional and get harder until you stop being fluent — that is the point, and reaching your limit is normal and expected. What separates a good answer from a weak one here is almost never the fact itself; it is whether you attach a trade-off and a condition. "Use CDC" is a weak answer. "Use CDC when the source can expose a change log and you need the target within seconds of the source — and be aware it needs more plumbing than a plain bulk copy, so for a small archive I'd just do a one-time load" is a strong one.

Stage 3 — The scenario / design round

An open-ended prompt: "We have a data centre lease ending in nine months and 180 applications. Go." There is no single right answer and the interviewer knows it — they are watching your process. Do you ask what the driver is? Do you ask what the estate looks like before you propose a tool? Do you sequence the work? Do you say how you would verify it, and how you would back it out? The scenario section below gives a reusable structure for these; that structure is worth more than any individual answer in it.

Stage 4 — The behavioural round

Migration work is unusually collaborative and unusually political: you are asking application owners to let you touch systems they are accountable for, on a schedule they did not choose. So behavioural questions here skew towards stakeholder friction, decisions under uncertainty, and what you did when a cutover went wrong. The behavioural section gives STAR skeletons you fill with your own material.

🐢 What "senior" sounds like

The single clearest signal of seniority in a migration interview is the sentence "it depends — and here is what it depends on." A junior answer names one option. A mid-level answer names two and picks one. A senior answer names two, picks one, states the condition that would flip the choice, and then says how they would find out which condition they are actually in. Practise that fourth clause; it is the one most candidates skip.

The topic map — where the questions land

☺ Like you're 10: Not every topic gets asked about equally. Here's a rough map of which ones interviewers dig into hardest — and which lesson to reread for each.

Read this as a study-priority guide. The emphasis column is this course's judgement about how hard each area tends to be probed for a migration-focused role, drawn from the shape of the work itself — not a published weighting from any employer. Treat it as "where to spend Tuesday evening," not as a mark scheme.

DomainEmphasisTypical question shapeReread
FoundationsSteadyDefinitional, plus "why are they moving at all?"What & Why We Move
The 7 R'sHeavyName them, then defend a choice per appThe 7 R's
DiscoveryHeavyWhat it produces, and how you find what's hiddenThe Journey
Wave PlanningHeavyGroup, sequence, size — then justifyWave Planning
Landing ZoneSteadyWhat's in it, and why it comes firstArchitecture Patterns
Data MigrationHeavyCDC, schema conversion, validation, sizingData Migration
CutoverHeavyThe low-downtime recipe, and the rollbackArchitecture Patterns
ModernizationSteadyStrangler fig — and "should we, though?"Modernization
Cost & FinOpsSteadyPricing models, and why the bill went upSecurity, Cost & Resilience
SecuritySteadyShared responsibility, least privilege, complianceSecurity, Cost & Resilience
Resilience & DRSteadyRTO/RPO first, then the four strategiesSecurity, Cost & Resilience
Anti-PatternsSituationalSpot the trap in a plan you're handedAnti-Patterns & Pitfalls
Cloud-to-Cloud & HybridSituationalWhy it's harder than it sounds; lock-inCloud-to-Cloud & Hybrid
◆ Key idea

Four domains carry most of the weight in a migration interview: the 7 R's, discovery, data movement, and cutover. That is not an accident — they are the four places where a wrong decision costs real money or real data. If you only have one evening, spend it on those four, and on being able to say why rather than what.

🦫 Benny's workshop · 15 min

Before you read a single answer below, go down the topic-map table and rate yourself out of three on each domain: 3 = "I could teach it," 2 = "I could answer but I'd wobble," 1 = "I'd be guessing." Write the list down. Then read only the sections you scored 1 or 2, and come back to the 3s the night before. Most people discover their honest weak spot is validation or sizing the data move — the two areas that are pure arithmetic, and therefore the two that are easiest to skip while reading.

Domain 1 · Foundations

☺ Like you're 10: The warm-up questions. They sound easy, which is exactly why a sloppy answer here is expensive — it sets the interviewer's expectation for everything after.

Four questions that almost always open a migration interview. The lesson behind them is What & Why We Move, with the lifecycle in The Journey.

🟢 Q1 · What is a cloud migration, in your own words?

They ask: "Before we get into the detail — how would you define a migration for someone who's never done one?"

Model answer. A migration is moving digital workloads — applications, data and databases, virtual machines, or a whole bundle of all three — from one computing environment to another. The classic case is out of an on-premises data centre and into a public cloud, but that is only one of the directions. You also see on-prem to on-prem hardware refreshes, cloud to cloud between providers, deliberate hybrid setups where some things stay on the ground permanently, multi-cloud, and repatriation where a workload comes back down from the cloud on purpose. I'd say the four things that actually get moved are applications, data and databases, virtual machines, and workloads — where a "workload" is the useful catch-all for an app plus the data and machines it needs to run. When someone says "we migrated forty workloads," that is what they mean.

◈ What they're really checking

Two things. First, precision of vocabulary — do you use "workload," "estate," and "environment" the way practitioners use them, or do you say "stuff." Second, whether you have a one-directional mental model. Candidates who define migration as "moving to AWS" often struggle later with the hybrid and cloud-to-cloud questions, because they never had a frame for them.

⚠ The common wrong answer

"Migration means moving your servers to the cloud." It is too narrow in two ways: it drops the data and application layers, which are the hard parts, and it assumes one direction. Say the general definition first, then name the classic case as an example of it.

🟢 Q2 · Why do organisations migrate? Give me the drivers.

They ask: "What actually pushes a company to do this? It's expensive and risky — why bother?"

Model answer. The common drivers are cost reduction, scalability and elasticity, agility and speed to launch, access to managed and innovative services, security and compliance posture, ageing or end-of-life hardware, a data-centre lease expiring, mergers and acquisitions forcing consolidation, disaster recovery, sustainability, and difficulty hiring people to maintain old infrastructure. Real migrations almost never have one driver — they usually have three, and one of them is the deadline that actually forces the project to happen. The reason this matters practically is that the dominant driver quietly decides your strategy. If the driver is a lease expiring in nine months, you are going to rehost hard and schedule modernization for later. If the driver is agility, rehosting everything gets you almost nothing and you should be replatforming and refactoring where the value is. So the first question I ask on any engagement is "what happens if we don't do this?" — the answer names the real driver.

◈ What they're really checking

Whether you connect driver to strategy, or just recite a list. Anyone can list drivers. The follow-up they want to hear unprompted is: the driver determines which of the 7 R's dominates the portfolio.

⚠ The common wrong answer

"Because the cloud is cheaper." It is not automatically cheaper, and saying so flatly signals you have not seen a real bill. Cost is a genuine driver — but for large, steady, predictable workloads, owning hardware can beat renting it, which is exactly why repatriation exists as a mature strategy. Say "cost can be a driver, and here is when it holds and when it does not."

🟡 Q3 · Explain the shared responsibility model — and how the line moves.

They ask: "Who is responsible for security once we're in the cloud — us or the provider?"

Model answer. Both, along a line that moves depending on the service model. The shorthand is that the provider is responsible for security of the cloud — the physical data centres, the hardware, the hypervisor, the core network and services — and the customer is responsible for security in the cloud: their data, their identities and access controls, their configuration, and their encryption keys. Where the line sits depends on what you consume. On raw infrastructure you patch the guest operating system yourself. On a managed platform service the provider patches far more of the stack. On finished SaaS you are mostly managing your data and who can reach it. But two things never move to the provider's side no matter what you buy: your data, and who has access to it. That is why the majority of real cloud incidents are not a provider's data centre failing — they are a customer leaving a door open on their own side of the line.

◈ What they're really checking

Whether you understand this is a sliding line, not a fixed split. The strong version of this answer names the service model explicitly and gives one concrete example of something that moves (OS patching) and one that never moves (access control).

⚠ The common wrong answer

"The provider handles security, that's what you're paying for." This is the single most dangerous misconception in cloud work, and interviewers listen for it specifically. A close cousin is claiming a service is "compliant" — a provider service can be eligible for use under a framework, but compliance is something you configure and evidence, and for some frameworks it also requires the right contractual agreement in place.

🟡 Q4 · How do you build the business case? Walk me through TCO and ROI.

They ask: "The CFO wants a number. How do you produce one you'd stand behind?"

Model answer. Total cost of ownership is the all-in cost of running something, not the sticker price — for the on-prem side that means hardware and its refresh cycle, power, cooling, floor space or lease, software licences, network, backup, and the staff time spent maintaining it. Return on investment is what you get back against what you spend. The business case compares the TCO of staying put against the expected cloud run-rate plus the one-off cost of migrating, and shows the payback over time. The part that makes a business case honest rather than optimistic is the un-glamorous line items: egress charges to move the data out, training, migration tooling and any professional services, and above all the "double bubble" period where you are paying for both estates at once because the source has not been decommissioned yet. I'd also give a range rather than a single number, and name the two or three assumptions that move the range most — usually right-sizing ratio, commitment coverage, and how long the double-bubble lasts.

◈ What they're really checking

Can you talk to a finance stakeholder without either hand-waving or drowning them. The tell of a good answer is mentioning the double-bubble period unprompted — it is the cost that surprises first-time programmes and the one that proves you have actually run one.

⚠ The common wrong answer

Quoting a percentage saving with no baseline — "migrations typically save thirty percent." Against what? Against a data centre already fully depreciated, the honest answer is often that year one costs more. Tie every number to a stated baseline and a stated assumption, or don't say it.

Domain 2 · The 7 R's

☺ Like you're 10: The seven ways to deal with each app. Interviewers ask about these more than anything else, because picking the right one is most of the job.

This is the highest-frequency domain in a migration interview. You must be able to name all seven cold, and then defend a choice under pushback. The lesson is The 7 R's, and the hands-on drill is Drill — Pick the Right R.

🟢 Q5 · Name the 7 R's and say what each one means.

They ask: "Take me through the 7 R's."

Model answer. Two of them mean you don't move the app at all. Retire — switch it off permanently, because it is unused, redundant, or duplicated by something else. Retain — deliberately leave it where it is for now, because a compliance rule, a vendor contract, or a planned decommission makes moving it wasted work today. The other five are the actual moves. Rehost, nicknamed lift-and-shift — move it unchanged onto cloud infrastructure. Relocate — move an entire virtualised platform at once to a matching cloud service, with no conversion and no app changes; the classic case is a VMware estate to a cloud-hosted VMware offering. Repurchase, nicknamed drop-and-shop — retire your app and buy a ready-made SaaS replacement. Replatform, nicknamed lift-tinker-shift — move it while making a few targeted cloud-friendly changes, most commonly swapping a self-managed database for a managed one. Refactor or re-architect — significantly rewrite it to be cloud-native. If they want the history: the list began as Gartner's five R's in 2011 from Richard Watson, was reshaped into the well-known six R's by Stephen Orban at AWS in 2016 — which is where Retire and Retain were added — and Relocate came later to make seven.

◈ What they're really checking

Fluency and completeness. This is a recall question and they are timing you implicitly. The detail that separates candidates is remembering that two of the seven are "don't move it" — that is the conceptual leap the 2016 revision made, and mentioning it shows you understand the list rather than having memorised it.

⚠ The common wrong answer

Listing five or six and trailing off, or conflating Relocate with Rehost. Also, describing the list as a ranking from worst to best. It is a menu, not a ladder — the "best" R is whichever gives the outcome you need for the least effort and risk, and that is a different answer per application.

🟡 Q6 · How do you decide which R an application gets?

They ask: "You've got 200 apps in a spreadsheet. How do you actually assign strategies?"

Model answer. Per application, from discovery evidence — never one R for the whole estate. I weigh five criteria together: business value (is this growing and important, or coasting?), technical fit (will it run well in the cloud as-is, or does its design fight the platform?), effort and time (how many weeks and people, and is there a hard date?), risk (how bad is it if this breaks?), and licensing and cost (do licences transfer, or is a SaaS replacement cheaper over the life of it?). Then, for speed across a long inventory, I walk each app down a decision path and stop at the first yes: do we still need it at all — no, Retire. Is there a hard blocker — yes, Retain. Does an off-the-shelf product already do this better — yes, Repurchase. Is it part of a block of VMs we want to move together untouched — yes, Relocate. Does it work fine and we mainly need speed out of the data centre — yes, Rehost. Would one or two contained tweaks unlock a big win — yes, Replatform. Is it business-critical and actively held back by its design — yes, Refactor. The path gives you a fast first pass; the five criteria are how you sanity-check the ones that matter.

◈ What they're really checking

Whether you have a repeatable method or you improvise per app. They also want to hear that the input is evidence from discovery, not opinion from a workshop. If you can say "the disposition decision is a data product, and its quality is capped by the quality of the inventory," you are speaking their language.

⚠ The common wrong answer

"We standardise on rehost" or "we refactor everything." Both are portfolio-level answers to an application-level question. The first quietly carries every existing inefficiency into a metered environment; the second is the slowest, priciest and riskiest path applied to apps that do not deserve it.

🟡 Q7 · What's the difference between rehost and relocate? Replatform and refactor?

They ask: "These two pairs get muddled a lot. Draw the lines for me."

Model answer. Rehost and Relocate both leave the application untouched; they differ in the unit of movement. Rehost works machine by machine — you replicate a server and stand it up on cloud infrastructure, and you can do that for one app or a hundred. Relocate moves an entire virtualised platform as a block onto a matching cloud service, with no conversion of the VMs at all — that is why it is fast for very large estates and why it typically has minimal downtime. The trade-off is identical for both: you have changed the address, not the design, so the cloud-native payoff is small.

Replatform and Refactor both change the application; they differ in how deep. Replatform makes contained, targeted changes on the way in — the canonical example is moving the app as-is but swapping its self-managed database for a managed one so nobody patches or backs it up by hand any more. Refactor, or re-architect, significantly rewrites the application around cloud-native patterns: breaking a monolith into services, containerising, going event-driven or serverless. The practical warning I'd add on Replatform is that "tinkering" has a habit of quietly growing into a rewrite — you have to guard the scope boundary deliberately, or a medium-effort R turns into a high-effort one halfway through the wave.

◈ What they're really checking

That you think in terms of unit of movement and depth of change rather than memorised labels. The scope-creep warning on Replatform is the detail that reads as field experience.

🔴 Q8 · When would you argue against refactoring a business-critical application?

They ask: "Our core order system is a monolith and leadership wants it rebuilt cloud-native as part of the migration. Thoughts?"

Model answer. I'd want to separate two questions that are being asked as one: should this application eventually be rebuilt, and should the rebuild happen during the migration. The second one I would push back on. Migrating changes where an app runs; refactoring changes how it is built. Doing both at once means that when something breaks — and on a crown-jewel system something always does — you cannot tell which change caused it, and your rollback has to undo two things instead of one. The rule I'd state plainly is: don't modernize and migrate the crown jewels at the same time.

I'd also test whether refactoring is warranted at all. It earns its cost when the application is genuinely business-critical and its current design is actively blocking scale, speed, or new features, and the team has the skills, and there is a clear payoff someone can name. Miss any one of those and you are buying the slowest, most expensive, highest-risk R for a benefit nobody can articulate. If the driver is a hard date like a lease expiry, the honest recommendation is rehost or replatform now, decommission the source, then schedule the refactor with a real date on it — because the failure mode of "we'll modernize later" is that later never comes.

◈ What they're really checking

Seniority — specifically, whether you will push back on a plan handed to you by someone more senior, and whether you can do it constructively rather than just saying no. Notice the shape of the model answer: it does not refuse, it re-sequences, and it names the condition under which the original plan would be right.

⚠ The common wrong answer

Enthusiastic agreement. "Yes, definitely, we should make it cloud-native" is the answer of someone who has not watched a combined migrate-and-rewrite slip by a year. The mirror-image wrong answer is a flat "never refactor during a migration" — sometimes the whole point of the project is the rebuild, and a rigid rule is not judgement either.

🟡 Q9 · Repurchase gives a high benefit for medium effort. So why isn't everything a repurchase?

They ask: "If buying SaaS is such a good deal, why do we still migrate anything?"

Model answer. Because the benefit belongs to the vendor's engineering, not to yours, and that comes with strings. Repurchase genuinely is the outlier on the effort-versus-benefit curve — you get a well-run, professionally maintained, continuously updated product for medium effort, which is why it sits above the trend line rather than on it. But you pay for it in four ways. You still have to migrate your data into the product, which is its own project. You retrain users and rebuild any integrations. You accept the vendor's model of how the work should be done, which means the process bends to the tool rather than the reverse. And you trade infrastructure lock-in for a different lock-in with a per-seat cost curve that grows with your headcount rather than your usage.

So Repurchase is right when a commercial product already does the job better than your custom-built one — self-hosted email, a homegrown CRM, an internal ticketing tool. It is wrong when the application is the differentiator, because then you would be buying the same capability your competitors can buy.

◈ What they're really checking

Whether you can reason about a strategy that isn't purely technical. The line worth landing is: never repurchase the thing that makes you different from your competitors.

Domain 3 · Discovery

☺ Like you're 10: Counting everything before you pack. Boring, unglamorous, and the single most common thing real migrations get wrong.

Discovery questions are how an interviewer separates people who have planned a migration from people who have read about one. The lesson is the Assess stop in The Journey, the tooling is in The Tools Landscape, and the lab is Part 1 — Discovery & the Disposition Matrix.

🟢 Q10 · What does discovery actually produce, and why can't you skip it?

They ask: "Talk me through the assess phase. What comes out the other end?"

Model answer. Two artefacts, plus one dataset that people forget. The first artefact is the inventory — every server, application and database, with what it is and where it lives. The second is the dependency map — which app talks to which, which apps share a database, what calls that one server everybody forgot about. The dataset people forget is utilisation: how much CPU, memory and storage each thing actually uses over time, which is what right-sizing in the cloud depends on. Without utilisation you end up recreating on-prem machine sizes in a metered environment, and the bill tells you about it a month later.

You cannot skip it because every downstream decision leans on it. The 7 R's disposition is assigned from the inventory. Wave grouping is cut from the dependency map. The business case is priced from utilisation. And the specific failure mode of skipping it is not abstract: you cut over an application and discover in production that it was quietly reading from a database you left behind.

◈ What they're really checking

Whether you name utilisation alongside inventory and dependencies. Most candidates give the first two. The third is what connects discovery to cost, and mentioning it links two domains in one answer.

🟡 Q11 · Agent-based versus agentless discovery — what are the trade-offs?

They ask: "How would you actually collect that inventory?"

Model answer. Agent-based discovery installs a lightweight collector on each machine. It gives you the deepest data — per-process detail, real utilisation, what is actually listening and connecting — but it costs you deployment effort, change approval, and coverage: you can only instrument machines you control and are allowed to touch, which in a large estate is never all of them. Agentless discovery watches from outside: network flow data, hypervisor and configuration-management sources, existing monitoring. It stands up fast and creates almost no friction, but it is shallower and it can miss anything that was quiet during the observation window.

In practice most estates use both — agentless for breadth and fast coverage, agents on the machines that matter or that the flow data leaves ambiguous. The point I'd make either way is that discovery has a time dimension, not just a coverage dimension. A two-week observation window will not see a month-end batch job or a quarterly reporting run. If you plan waves from a fortnight of data, you will find your first quarterly dependency in production. Run the collection long enough to span the estate's own business cycle, and explicitly ask owners what runs on a schedule you might not have observed.

◈ What they're really checking

Whether you understand discovery as a sampling exercise with error bars, rather than a button you press. The time-window point is the one that lands hardest, and very few candidates raise it.

🟡 Q12 · How do you find dependencies the tooling can't see?

They ask: "Your dependency map says this app is standalone. How much do you trust it?"

Model answer. Not fully, and I'd say so. Automated flow analysis finds live network conversations during the observation window, which is most but not all of the truth. The gaps I'd go looking for deliberately: hard-coded IP addresses and connection strings in configuration files and, worse, inside application code; scheduled jobs and batch windows that only run monthly or at quarter end; shared file mounts; DNS names resolving to things nobody has an owner for; outbound calls to third-party or partner endpoints, which matter because your source IP is about to change and someone else's allow-list may not know that; and licence servers, which are a classic silent dependency.

Beyond tooling, two human techniques. First, interview the application owner and ask specifically "what breaks if this box disappears?" rather than "what does it depend on" — people answer the first question much better. Second, where there is genuine doubt and a lower environment exists, block the suspected path there and see what complains. Every one of these is cheaper than finding out at cutover, which is the whole argument: the dependency you don't know about is the one that breaks a wave.

◈ What they're really checking

Whether you treat the tool's output as a hypothesis rather than a fact, and whether you have a concrete list of blind spots rather than a general sense of caution. Mentioning third-party allow-lists and licence servers is a strong signal — both are real, both are common, and neither appears in flow data as anything alarming.

🟡 Q13 · What is portfolio rationalization, and what does it save you?

They ask: "We've got the inventory. What's the next thing you do with it?"

Model answer. Rationalization is going through the inventory application by application and deciding what deserves to move at all, before deciding how anything moves. In practice that means finding the apps nobody uses any more, the ten tools that all do the same job, and the systems that are already scheduled for decommission — and marking them Retire or Retain rather than migrating them. The saving is direct and compounding: every application you don't move is one you never pay to migrate, never pay to run, never have to secure and patch, and never have to fit into a wave. The cheapest migration is the one you skip.

The discipline that makes it safe is verification. "Nobody uses this" is a claim, and it should be backed by real usage data — access logs, connection counts, authentication events over a long enough window — before anything is switched off. And the safe way to retire is to stop the service and keep it recoverable for an agreed period rather than deleting it outright, so that a quarterly user who materialises in week six is an inconvenience rather than an incident.

◈ What they're really checking

That you do rationalization before planning waves, not after moving. The verification point is what distinguishes a confident answer from a reckless one.

⚠ The common wrong answer

"Move everything, then clean up once we're in the cloud." This is the migrating-the-junk trap. Complexity and cost travel with you, cleanup after a move is never prioritised because the pain is gone, and you have spent migration effort — the scarcest resource on the programme — on applications with no future.

Domain 4 · Wave Planning

☺ Like you're 10: Splitting the big move into small batches. Interviewers love this because your answer shows whether you think in blast radius.

Expect at least two of these, and expect one of them to be a curveball where the clean answer isn't available. The lesson is Wave Planning; the lab is Part 3 — The Wave Plan.

🟢 Q14 · What is a migration wave, and why not just move everything at once?

They ask: "Why do people bother with waves? Wouldn't one big weekend be faster?"

Model answer. A wave is a group of related applications and their data, migrated together and validated together before the next group starts. Each wave is a self-contained mini-project with its own plan, its own cutover window, and its own rollback. A migration is really just an ordered sequence of these. The alternative — a big-bang cutover of everything in one event — sounds faster because it has one date on the calendar, but it stacks every risk into a single night with no way to undo one part independently.

Waves buy you three things. They shrink the blast radius, so a failure in wave three doesn't touch the apps already running from waves one and two. They turn the programme into a learning machine: the first wave is where you discover the forgotten firewall rule and the slow database copy, and you fold those lessons into your tooling and runbook so the next wave is cheaper — teams often call this the migration factory, and by wave four the same class of app moves in a fraction of the wave-one time. And they spread the load: nobody can take a hundred applications offline on one night, and no team can execute that safely. The trade is a little more elapsed time for a lot less risk, and that trade is almost always worth making.

◈ What they're really checking

Whether "wave" means "batch" to you, or means "unit of safe progress." The strong framing is: anything you put in a wave must be able to move together, be tested together, and be rolled back together. If it can't do all three, your wave boundary is wrong.

🟡 Q15 · What criteria do you use to group applications into a wave?

They ask: "You've got the dependency map. How do you decide what travels together?"

Model answer. Several criteria weighed together, with one that dominates. The dominant one is dependencies: applications that talk to each other constantly, or share a database, must move in the same wave — otherwise every one of those calls has to cross the link between on-prem and the cloud for the whole gap, which adds latency and can add data-transfer cost. Beyond that: business domain or owner, because apps owned by one team share people and context and are easier to plan and support as a batch; environment, because you move dev and test copies first and rehearse the move on the safe copy; risk and complexity, so the easy ones can go early and the crown jewels late; a shared R strategy, because a wave of pure rehosts can reuse one runbook and one toolchain; compliance and data-residency constraints, so regulated workloads are handled together and their controls are designed once; and shared downtime windows, since apps that can only go offline at the same hour naturally batch.

The technique underneath all of it is to cut where the seams are thin: look for clusters in the dependency map with lots of connections inside and only a few thin links to the outside, and put your wave boundaries there. That minimises the number of cross-wave links you have to bridge temporarily.

◈ What they're really checking

Whether dependencies dominate your grouping or are just one item in a list. Also whether you mention environment — moving dev and test first is the cheapest rehearsal available and candidates routinely leave it out.

🟡 Q16 · How do you sequence and size the waves?

They ask: "Fine — you've got your groups. What order, and how big?"

Model answer. Sequencing first. Before wave one there is really a wave zero: the landing zone and the network connectivity have to exist, because there is nowhere for anything to land otherwise. Wave one is the pilot — the smallest, least-connected, lowest-risk, genuinely non-critical application you can find. Its job is not to move something important; it is to prove the landing zone works, the tooling works, the team knows the runbook, and both cutover and rollback behave. After that, four ordering principles: easy before hard, dependencies before dependents (nothing arrives before the thing it relies on), independent before connected, and crown jewels last — the revenue-critical systems move when the team is most practised and the platform most proven.

Sizing is a balance with a hard floor and a soft ceiling. The floor: a wave must be small enough that you can migrate it, test it, and if necessary roll all of it back inside your available cutover window, with the team you actually have. The ceiling: large enough that you are making visible progress and not dragging a two-hundred-app estate out over three years of one-app waves. Early waves are deliberately small because you are still learning; later waves grow as the factory speeds up. A rule of thumb I'd offer: if a wave needs a war room and an all-nighter, it is too big — split it.

◈ What they're really checking

Whether "wave zero" exists in your model. Candidates who go straight to "wave one is a pilot" without the foundation are describing a pilot with nowhere to land. Also whether your sizing rule is anchored to the rollback window rather than the migration window — that is the constraint that actually binds.

⚠ The common wrong answer

Putting the most important system first "to prove the value fast." This is the skip-the-pilot trap: you learn that your process is broken on your most important system, in front of your most important users. If leadership is pushing for it, the constructive counter is to pilot on a low-risk app in the same architectural class as the crown jewel, so the learning still transfers.

🔴 Q17 · Two chatty applications must move together, but one owner won't sign off in time. What do you do?

They ask: "Your dependency map says these two can't be split. The finance team says their app isn't moving this quarter. Now what?"

Model answer. First I'd measure rather than assume. "Chatty" is a judgement; I want the actual call volume and the actual latency sensitivity, because the cost of splitting them is a function of both. If the traffic is a few hundred calls an hour, splitting is annoying but survivable. If it is hundreds of calls a second on a synchronous path, splitting will show up as user-visible latency within an hour of cutover.

Given real numbers, the options in the order I'd consider them: re-cut the wave boundary so the pair stays together and something else moves instead — the plan should bend before the architecture does. If the date genuinely can't move, bridge the gap deliberately: hybrid connectivity sized for the measured traffic, and be explicit about whether that is a VPN over the public internet or a dedicated private link, because the choice depends on how much data and how predictable the performance needs to be. Model the data-transfer cost of the split period, because cross-boundary chatter is billable in a way it wasn't on-prem. Where the coupling is read-only, a read replica on the cloud side can remove most of the round trips.

The condition I'd insist on regardless of which option wins: the hybrid period is time-boxed with a date and an owner, and it is monitored. The real failure here isn't the split — it's the temporary state that quietly becomes permanent because nobody scheduled its end.

◈ What they're really checking

Whether you can operate when the textbook answer is unavailable. They are also testing whether you reach for measurement before you reach for an opinion — "let me get the call volume first" is a strong opening because it is what you would actually do.

Domain 5 · Landing Zone

☺ Like you're 10: Building the new house — floors, locks, wiring, house rules — before a single box arrives.

Expect these to be framed as "why does this come first?" The lessons are the landing-zone section of Architecture Patterns and the governance half of Security, Cost & Resilience. The lab is Part 2 — The Landing Zone.

🟢 Q18 · What is a landing zone, and what's in one?

They ask: "Define a landing zone for me."

Model answer. A landing zone is a pre-built, secure, well-organised cloud environment that exists before any workload moves, so that everything lands somewhere consistent and governed rather than into a hand-made space. Concretely it defines: the account, subscription or project structure and how it maps to teams and environments; the network layout — address ranges, segmentation, connectivity back to on-prem; identity and access, including roles and how permissions are granted; centralised logging and audit; and guardrails — the automatic policy rules that make the safe choice the default. All of it expressed as infrastructure as code, so it is repeatable, reviewable and rebuildable rather than clicked together once and never reproducible.

The reason it is described as "wave zero" is that it genuinely is a prerequisite: your pilot has nowhere to land without it, and each of the major providers ships blueprints and accelerators precisely because nobody should be inventing this from first principles.

◈ What they're really checking

That your list includes both the technical scaffolding and the governance — a landing zone that is only networking and accounts is missing the point of it. Saying "as code" unprompted matters too, because a hand-built landing zone is just a nicer-looking snowflake.

🟡 Q19 · Preventive versus detective guardrails — what's the difference, and which do you use?

They ask: "How do you stop teams doing dangerous things in the new environment?"

Model answer. Both, for different classes of risk. A preventive guardrail refuses the action outright — a policy that will not let anyone create storage open to the public internet, or launch resources in a region you have not approved. A detective guardrail watches for a bad state and raises an alert or auto-remediates — a scan that notices something became public and locks it back down. The mental model is a fence at the top of the cliff versus an ambulance at the bottom, and mature setups run both.

How I'd choose: preventive for the small set of things that must genuinely never happen, where the cost of a false block is lower than the cost of the event. Detective for everything where a hard block would break legitimate work, where the rule is contextual rather than absolute, or where you are still learning what "normal" looks like. The practical trap is going preventive-heavy on day one, discovering it blocks real delivery, and then having someone grant a broad exception that quietly disables the control for everyone. Start detective, learn what the estate actually does, then promote the stable rules to preventive.

◈ What they're really checking

Whether you understand guardrails as a product with users, not just a policy document. The "start detective, promote to preventive" sequencing is a strong, practical answer that also shows you have watched a control get bypassed.

🟡 Q20 · Why build the landing zone first instead of retrofitting governance later?

They ask: "We're under time pressure. Can't we spin up accounts as we go and tidy up in phase two?"

Model answer. You can, and it is one of the most reliably expensive decisions available. Three reasons. First, arithmetic: fixing identity, network layout and naming on one landing zone is one piece of work; retrofitting it across a hundred applications that have already migrated is a hundred pieces of work, each with an owner who now has other priorities and a change window you have to negotiate. Second, auditability: hand-built accounts diverge — different naming, different permissions, different network shapes — and a snowflake estate cannot be audited or automated, which means every subsequent control costs more to apply. Third, and least obvious: every other migration pattern assumes the landing zone exists. A strangler-fig facade needs a governed network to sit in. A blue-green cutover needs two genuinely consistent environments. Change-data-capture replication needs secure connectivity and logging. Skip the foundation and each of those becomes a bespoke negotiation.

The honest concession is that a landing zone can absolutely be over-engineered into a six-month project that blocks all delivery. The answer to time pressure is a minimum viable landing zone — the account structure, the network, identity, logging and the handful of guardrails you would not ship without — delivered as code so you can extend it wave by wave. "Build it first" does not mean "build all of it first."

◈ What they're really checking

Whether you can defend the principle and acknowledge its failure mode. Candidates who only argue the principle sound doctrinaire; candidates who only concede the time pressure sound like they will build the snowflake. Say both.

🔴 Q21 · How would you structure accounts, subscriptions or projects — and why?

They ask: "Sketch the account structure you'd propose for a mid-sized enterprise migration."

Model answer. I'd start from the principle rather than the diagram: the account — or subscription, or project, depending on the provider — is the strongest isolation and billing boundary the platform gives you, so you spend it where a failure boundary or an accountability boundary genuinely matters. In practice that means separating along three axes. Environment: production isolated from non-production, always, because that separation is what makes a mistake in test survivable. Blast radius and ownership: workloads that fail independently and are owned by different teams belong in different accounts, so a runaway process or a bad permission cannot cross. Platform versus application: shared services — connectivity, centralised logging, shared identity, shared tooling — sit in their own platform accounts, distinct from the accounts where application teams work, because their change cadence and their audience are different.

Above all of it sits an organisation or management tier where org-wide policy and consolidated billing live. Each cloud names the tree differently — organisations and organisational units with accounts underneath, management groups over subscriptions with resource groups inside, or an organisation with folders and projects — but the reasoning transfers cleanly, which is the useful thing to say in an interview where you may not know which provider they use.

Where practice genuinely varies is granularity. Some organisations run an account per application per environment, which gives excellent isolation and a lot of accounts to operate. Others run an account per domain per environment and separate applications inside it, which is cheaper to run and gives coarser blast radius. I would not claim one is universally right; I would ask how many application teams there are, how strong the platform team is, and whether anything in the estate has a regulatory reason to be hard-isolated — and let those answers pick the granularity.

◈ What they're really checking

Provider-neutral reasoning, and whether you know this is a genuinely contested design space. The strongest thing you can do here is refuse to assert a universal answer and instead name the three inputs that decide it.

Domain 6 · Data Migration

☺ Like you're 10: The heavy boxes. Apps you can rebuild from code; data you cannot. This is where a migration actually loses something irreplaceable.

The deepest technical domain in the interview, and the one where hand-waving is most obvious. The lesson is Data Migration & Data Gravity; the drill is Drill — Size the Data Move.

🟢 Q22 · Why is the data the hard part of a migration?

They ask: "People say the data is the difficult bit. Why?"

Model answer. Because of asymmetry of recovery. An application is mostly instructions — lose an app server and you rebuild it from code and redeploy; it is annoying and recoverable. Data is stateful and frequently irreplaceable: lose ten years of a customer's order history and there is nothing to recompile it from. That asymmetry is why data gets its own careful process with its own validation step, rather than being treated as one more thing that gets copied.

The other half is that data can fail quietly. While you copy, it can be lost, duplicated, half-copied, silently corrupted, or changed underneath you — and any one of those can go unnoticed until long after the source is gone. Every data move really answers three questions: is the source still in use while we copy (online or offline), do we copy once or keep copying the changes (bulk load or continuous replication), and how do we prove it arrived intact (validation and reconciliation). Answer those three well and the rest is tooling.

◈ What they're really checking

The words "prove" and "irreplaceable." The three-questions framing is a strong structure to open with because everything else in this domain hangs off it.

🟡 Q23 · Explain change data capture, and why you'd choose it over the alternatives.

They ask: "How do you keep the target database in sync with a live source?"

Model answer. Change data capture reads the database's own transaction log — the redo log, write-ahead log or binary log the engine already writes for its own durability — and streams each insert, update and delete to the target as it happens. That has two properties that matter. It is light on the source, because you are tailing a log the engine maintains anyway rather than adding query load. And it cannot miss a change that occurred between two polls, because there is no polling.

Against the alternatives: repeatedly re-querying tables to find what changed is heavy, needs a reliable modified-timestamp you often don't have, and misses rows that were changed and changed back between polls, plus it typically cannot see deletes at all. Dual-write — having the application write to both stores — looks simpler on a whiteboard and is genuinely risky in practice: if one write succeeds and the other fails, the copies drift apart silently and you may not discover it until reconciliation, after which you have no clean way to decide which side is right. CDC streams from a single source of truth, which is why it is the workhorse behind almost every low-downtime cutover.

The honest caveats: it needs the source to expose a change log and it needs the right privileges to read it; long-running transactions and schema changes on the source during replication need handling; and replication lag is a real number you have to monitor, because your achievable RPO at cutover is bounded by it.

◈ What they're really checking

That you can explain why log-based capture is different in kind from polling, not just that it is "better." Naming replication lag as the thing that bounds your RPO connects this domain to the resilience domain and is a strong senior signal.

⚠ The common wrong answer

Proposing dual-write as the default because "the app already knows about both databases." Interviewers who have been burned by this will push hard on it. If you like dual-write for a specific case, you must be able to say how you detect and repair drift — and once you have built that, you have built most of a reconciliation system anyway.

🟡 Q24 · Homogeneous versus heterogeneous database migration — what changes?

They ask: "We're moving from a commercial engine to an open-source one. How much harder is that?"

Model answer. Substantially harder, and the extra work is almost all in translation and re-testing rather than in moving bytes. A homogeneous migration keeps the same engine on both sides, so the internal structure matches and it is mostly a straight copy. A heterogeneous migration changes engine, which means the schema — tables, columns, data types, keys, constraints — has to be translated into the target's dialect, along with everything procedural: stored procedures, functions, triggers, and any database-specific behaviour the application quietly relies on.

Schema-conversion tooling automates most of it and, importantly, flags what it cannot convert so a human can redesign those parts — some proprietary features simply have no equivalent. The typical sequence is convert the schema first, review and hand-fix the flagged items, then let a managed migration service do the bulk load and the ongoing replication.

The two things I'd budget for that people underestimate: application changes, because SQL dialect differences, data-type edge cases, collation and sort-order differences, and transaction-isolation behaviour can all surface as subtle application bugs rather than as errors; and performance re-tuning, because query plans and index strategy do not transfer between engines. The driver for taking this on is usually licence cost, and that saving is real — but the honest estimate includes the re-testing, not just the copy.

◈ What they're really checking

Whether you know the risk lives in the application, not only in the database. Mentioning collation, isolation behaviour or query-plan differences signals you have actually done one.

🟡 Q25 · How do you prove the data arrived intact?

They ask: "The copy finished with no errors. Are you done?"

Model answer. No — a clean exit code says the tool did not fail, not that the data is correct. Validation has two layers. Row counts per table on both sides catch the loud failures: a table that half-copied, or one that copied twice. But identical counts do not prove identical contents — two tables can have the same number of rows and differ inside. So the second layer is checksums or hashes over the actual content, producing a fingerprint you can compare; that is what catches silent corruption like a mangled character or a truncated field. On very large tables you checksum in chunks or on a sampled basis, so validation doesn't take longer than the copy did.

Beyond the data layer, I'd add application-level validation: run the real queries and the real reports against the new store and compare outputs, because that catches the class of problem where the data is technically identical but behaves differently — sort order, date-time handling, numeric precision, a collation difference that changes which rows a query returns.

And the timing matters as much as the technique. You validate after the initial bulk load, and again after the final sync at cutover, and you keep the source intact and untouched until every check passes. The rule I'd state is: never declare a data migration done on the basis that it looked fine.

◈ What they're really checking

The distinction between counts and content, and the fact that validation happens twice. Adding application-level comparison is the answer that separates a database-focused engineer from a migration engineer.

🟡 Q26 · How do you size a data move? When is shipping a physical appliance faster than the network?

They ask: "We have a few hundred terabytes to move. Over the wire, or on a box?"

Model answer. Do the arithmetic before choosing: transfer time is roughly data size divided by usable bandwidth, and the word usable is doing the work. Link speed is not usable bandwidth — you have to subtract what the business is already using, whatever share you are actually permitted to consume during working hours, protocol and encryption overhead, and the fact that sustained real-world throughput on a long path is meaningfully lower than the headline number. If the honest calculation says weeks or months, and an offline appliance is a fixed few days of shipping regardless of size, the appliance wins — and that crossover arrives much earlier than people expect.

The pattern I'd usually propose for a large estate is both: ship the historical bulk on an encrypted appliance, and use an online transfer service to sync the delta that accumulated while the box was in transit. That way the cutover is a short catch-up rather than a fresh multi-week upload. The appliance is encrypted precisely so that a box lost in transit is not a lost dataset.

Two things I'd add to the estimate that people forget. First, egress: the provider or facility you are leaving typically charges for data going out, and that is a real budget line you should price up front — a cheaper destination does not lower the exit toll on your source. Second, the change rate: if the source is being written to at a high rate while you copy, you are not just moving a static pile, you are chasing a moving target, and your plan has to include how the delta gets reconciled.

◈ What they're really checking

Whether you say "usable bandwidth" rather than "our link is ten gigabit." Interviewers use this question to find out whether you have ever watched a transfer estimate collapse on contact with a production network.

🔴 Q27 · Explain data gravity, and how it changes your plan.

They ask: "You mentioned data gravity earlier — what does it actually mean for the sequencing?"

Model answer. Data gravity is the idea that data behaves like mass: the larger a dataset grows, the more applications and services cluster around it because they need fast local access, and the harder the whole cluster becomes to move. A small database is a shoebox. A hundred-terabyte database with twenty applications reading from it is a planet with an orbit, and you cannot move the planet without moving the orbit.

Three consequences for planning. First, sequencing: you generally plan the data move first and let the applications follow, because the thing you do not want is an application running in the cloud reaching back across a slow link to a database still on-prem — that arrangement is slow, expensive per call, and tends to persist far longer than intended. Second, wave boundaries: gravity is a strong argument for keeping a dataset and its dependent applications in the same wave, which is the same reasoning as the tightly-coupled rule but driven by volume rather than call frequency. Third, strategy: gravity plus egress charges is exactly why cloud-to-cloud moves are sticky — your data has accumulated mass where it lives, and leaving costs money charged by the party you are leaving.

The design lesson I'd draw is that if portability matters to an organisation, the decision that determines it is where the data lives and how portable its format is — far more than which compute service you chose.

◈ What they're really checking

Whether gravity is a metaphor you can repeat or a constraint you can plan around. The strong answer converts it into a sequencing rule and a wave-boundary rule.

Domain 7 · Cutover

☺ Like you're 10: The moment you flip the switch. Everything before it is preparation; everything after it is either relief or a very long night.

If there is one domain where the interviewer is listening for the word "rollback," it is this one. The lessons are the cutover patterns in Architecture Patterns and the per-wave lifecycle in Wave Planning; the lab is Part 4 — The Data Move & the Cutover Runbook.

🟢 Q28 · Walk me through a low-downtime cutover.

They ask: "How do you move a live system with minutes of downtime instead of a weekend?"

Model answer. Three phases, and the trick is that almost all the work happens before the downtime starts. First, bulk load the data to the target while everything keeps running normally — this is the multi-terabyte part and it happens days earlier, quietly, with users unaware. Second, replicate continuously, typically with change data capture, so the target stays within seconds of the source while you test and rehearse. Third, the actual cutover window: stop writes at the source, let the final sync drain, run validation, and only then repoint traffic at the new system.

Because replication already did the heavy lifting, the source and target are nearly identical before you stop writes. The real downtime is just the final sync plus the validation check — minutes rather than hours. And critically, because you stopped writes rather than deleting anything, the old system is sitting there intact and able to take traffic back if validation fails at the last moment. That rollback path is not an optional extra; it is the reason you are allowed to attempt the cutover at all.

◈ What they're really checking

Whether "stop writes, don't delete" is in your description. It is one clause, it takes two seconds to say, and it is the difference between a cutover and a gamble.

🟡 Q29 · Blue-green, canary, rolling — when would you use each?

They ask: "Which cutover pattern would you pick, and why?"

Model answer. They trade different things, so I'd choose by what I am most afraid of.

Blue-green keeps two complete environments and flips all traffic from the old to the new in one router move, with the old one held warm. You pick it when instant rollback is the priority — flipping back is a single operation and takes seconds. The costs are that you pay for two full environments simultaneously, and anything stateful needs real thought, because two environments sharing or diverging on one database is where blue-green gets subtle.

Canary sends a small slice of real traffic — say five percent — to the new version, watches, then widens. You pick it when you want early warning about problems that only appear under real user load and real user data. It is gentler than an all-at-once flip, but the rollout takes longer, and it only works if your monitoring is good enough to detect a problem inside a five-percent sample before you widen.

Rolling replaces instances a few at a time until the whole fleet is on the new version. You pick it for fleets of many identical servers where funding a whole duplicate environment isn't justified. The catch is that old and new run simultaneously for a period, so they must be mutually compatible — including any shared database schema — and rolling back means rolling the change back out node by node rather than one instant flip.

And the fourth option is the honest one: phased — move a few users, features or applications at a time. That is not really a deployment mechanism, it is the wave concept applied at the traffic level, and for a large estate it is usually the outer frame that one of the other three sits inside.

◈ What they're really checking

Whether you attach the pattern to a fear rather than a preference, and whether you raise the stateful-data problem for blue-green and the compatibility problem for rolling. Those two caveats are where the real engineering lives.

🟡 Q30 · What actually flips the traffic?

They ask: "You keep saying 'repoint traffic.' Mechanically, what happens?"

Model answer. Usually one of two things. A DNS change — you repoint the system's name at the new environment's address. The complication is caching: resolvers and clients hold the old answer for the length of the record's time-to-live, so unless you lowered the TTL well before the cutover, some users keep landing on the old system after you have flipped. The standard practice is to drop the TTL to a minute or two a day or more ahead of the window, so caches everywhere have expired the long-lived answer by the time you switch. Even then, some clients cache more aggressively than they should, so a DNS-based cutover means the old system has to keep serving correctly for a period afterwards rather than being switched off at the moment of the flip.

Or a load balancer or traffic manager you control — a router in front that spreads requests across targets. Retarget it and the flip is immediate, with no caches involved anywhere. That is also what makes canary percentages possible: you send five percent of traffic to the new version by weight and then dial it up. And it is exactly what blue-green's "instant rollback" is — the router moving back to blue in one operation.

In practice, on a migration you often need both: DNS to move the public name to the new environment's front door, and a load balancer inside it to do the fine-grained shifting.

◈ What they're really checking

TTL awareness. It is a small, concrete, unglamorous detail, and it reliably separates people who have run a cutover from people who have read about one.

⚠ The common wrong answer

"We change DNS and it's instant." It is not, and the consequence is a split-brain period where some users write to the old system after you thought it was drained — which is precisely the situation that corrupts data during a migration.

🟡 Q31 · What's in a cutover runbook, and what is a go/no-go?

They ask: "Describe the document you'd take into the cutover window."

Model answer. The runbook is the precise, step-by-step script for that specific cutover: every command in order, every check with its expected result, who performs each step and who verifies it, the expected duration of each phase so you can tell mid-window whether you are behind, the communications plan for who gets told what and when, and — in the same document, not a separate one — the exact steps to reverse everything.

The go/no-go is a named decision point, usually just before the irreversible step, where a named person decides whether to proceed on the basis of pre-agreed criteria. The reason it must be pre-agreed is human: at two in the morning, after everyone has been awake for sixteen hours and has sunk-cost feelings about the work, nobody has good judgement about whether "mostly working" is good enough. The criteria have to be written down while everyone is rested and unattached.

Two habits make a runbook real rather than ceremonial. It gets rehearsed, usually against the dev or test copy, which both proves the steps work and — just as valuable — times them. And it gets revised after every wave, so the retrospective's findings become next wave's script. A cutover without a written runbook is a cutover you are improvising at two in the morning.

◈ What they're really checking

That the rollback lives inside the runbook and that go/no-go criteria are set in advance. If you also mention timing each rehearsed step, you are describing something you have actually done.

🔴 Q32 · Cutover completed, validation passed. Six hours later error rates climb. Walk me through it.

They ask: "You're in hypercare. Something's wrong. Talk me through your next thirty minutes."

Model answer. First, scope it before touching anything: what exactly is failing, for whom, and how much — from monitoring, not from the loudest message in the channel. I want to know whether this is functional (something is broken), performance (something is slow under real load), or data (something is wrong in the content), because those three have very different responses.

Second, compare against the rollback trigger we agreed before the cutover. This is the moment that pre-agreed criteria earn their existence: the question is not "does this feel bad enough" but "does this meet the condition we wrote down." If it does, I execute the documented backout, because the source is still intact and still stopped rather than deleted — that is precisely the situation the plan was built for. If it does not meet the trigger, I stabilise: scale out, adjust configuration, rate-limit, roll back a single component if it is isolable — while explicitly keeping the rollback path open rather than taking actions that close it.

Third, communicate on the schedule in the runbook, to the named stakeholders, with what is known and what is not. Silence during hypercare is what turns a technical problem into a trust problem.

The complication I'd raise unprompted, because a good interviewer will ask it anyway: rollback gets harder the longer the new system has been taking writes. Six hours in, reverting means either accepting the loss of those writes, or reverse-syncing them back to the source, or freezing writes while you reconcile. Which of those is acceptable is a business decision, not mine — and it should have been decided before the cutover, as part of defining how long the rollback window stays open. If it wasn't, I'd escalate it as a decision rather than make it quietly.

Afterwards: don't decommission the source. Decommissioning is the last step of a wave, after hypercare ends clean — not the first thing you do to declare victory.

◈ What they're really checking

Composure and sequencing. Notice the model answer never starts with a fix — it starts with scoping and with checking a pre-agreed criterion. Raising the "rollback decays over time" problem yourself is the strongest move available in this question, because most candidates describe rollback as if it were free forever.

Domain 8 · Modernization

☺ Like you're 10: Not just moving the furniture, but rebuilding some rooms. The interviewer usually wants to know whether you know when not to.

The lesson is Modernization; the incremental-replacement patterns are in Architecture Patterns, and the final lab stage is Part 5 — Operate, Optimize & Modernize.

🟢 Q33 · What's the difference between migration and modernization?

They ask: "People use these words interchangeably. Do they mean the same thing?"

Model answer. No. Migration changes where a workload runs; modernization changes how it is built and operated. They are independent choices: you can migrate without modernizing — pick up the app and set it down unchanged — and you can modernize something that is already in the cloud and has been for years. On the 7 R's, two of them are modernization: replatform makes contained cloud-friendly changes, and refactor rebuilds around cloud-native patterns. The other five leave the application's design essentially alone.

Cloud-native is the destination that modernization aims at: an application designed to scale out horizontally, recover from failure automatically, and be updated frequently and safely. The practical value of keeping the two words separate is that it lets you sequence them, and sequencing them is usually the right call — because changing location and design at the same time means you cannot attribute a failure to either one.

◈ What they're really checking

That you treat them as independent axes rather than as two points on one line. The follow-up they are teeing up is almost always Q36 — during or after.

🟡 Q34 · Explain the strangler fig pattern.

They ask: "How do you replace a legacy system that can't be taken down?"

Model answer. You put a facade in front of the old system — a router or gateway that receives every request and decides where to send it. At first it sends everything to the legacy system, so nothing has changed. Then you rebuild one capability at a time in the new environment and reroute just that path to the new implementation. Over time more routes point at the new system until the old one is handling nothing and can be switched off. The name comes from the strangler fig, which grows around a tree until it has replaced it, and Martin Fowler is widely credited with describing and popularising it.

Why you'd choose it: it converts one enormous irreversible step into many small reversible ones, so at every moment you have a working product and a short path back. That is exactly what you want for a large, risky legacy application the business cannot freeze for a two-year rewrite.

The trade-offs are real and worth naming. You run and pay for both systems throughout the transition, which can be a long time. The facade itself becomes critical infrastructure that has to stay fast and available, and it is a new single point of failure you did not have before. And where the new services need to talk to the messy legacy data model, you usually want an anti-corruption layer — a translator that keeps the legacy system's quirks from leaking into and rotting the clean new design. Where a network-level facade doesn't fit, because the thing you are replacing is deep inside one codebase, the in-code equivalent is branch by abstraction: introduce a common interface, build the new implementation behind it, migrate callers gradually, then delete the old one.

◈ What they're really checking

Whether you know the pattern's costs, not just its shape. The anti-corruption layer is the detail that shows you have thought past the diagram.

🟡 Q35 · Should we break the monolith into microservices as part of the migration?

They ask: "While we're moving it anyway, shouldn't we split it up?"

Model answer. Usually not at the same time, and often not to the extent people imagine. The first thing I'd separate is "should this be split eventually" from "should the split happen inside the migration" — and the second is a much higher bar, because combining a location change with an architecture change makes failures ambiguous and rollback compound.

Then I'd test whether splitting is warranted at all. The reasons that justify a service boundary are concrete: two parts genuinely need to be deployed independently, or genuinely need to scale independently, or are owned by teams that are blocking each other. Absent one of those, splitting buys operational overhead — network calls that can fail, versioning across boundaries, distributed tracing, more things to secure and monitor — in exchange for nothing.

Two named failure modes are worth calling out because interviewers listen for them. The distributed monolith: services that were split but still must be deployed together and still change together, which gives you every cost of a distributed system and none of the independence — genuinely worse than the monolith you started with. And premature microservices: carving a small system into many services before anyone understands the domain or has felt real scaling pressure, which optimises a problem you do not have.

The pragmatic recommendation for a migration is usually the cheap modernization first: replatform to managed services. Swapping a self-run database for a managed one removes a large amount of operational toil for a small, contained change — and it is available in the same wave without gambling the cutover on an architectural rewrite.

◈ What they're really checking

Whether you can resist an architecturally fashionable answer. Naming the distributed monolith by name, with the test "if two services must always ship together, they are one service," is the strongest single line in this question.

🟡 Q36 · Modernize during the move, or after it?

They ask: "We want to end up cloud-native. Do we do that on the way in, or once we've landed?"

Model answer. It depends on what is applying pressure, and I'd say what it depends on rather than picking a universal.

Modernizing during the move means the workload arrives already improved and you cut over once. It is the right call when the whole point of the programme is agility rather than exit, when the application is a mess that would be painful to move as-is, and when you have time and skills. The costs: the wave is slower, the risk is higher, and rollback is harder because you changed two things at once.

Modernizing after means lift-and-shift first to get out of the old environment quickly, then reshape once you have caught your breath. It decouples the two risks: you prove the app runs in the cloud, you decommission the source and stop paying for it, and only then do you start changing the design. The costs: you cut over twice, and you run a plainly-moved application for a while.

Match it to the pressure. If a lease ends in a few months, migrate first — the deadline is not negotiable and the architecture is. If there is no hard date and the driver is speed of delivery, doing it in one pass may genuinely be cheaper overall. The rule I would hold to in either case is not to modernize and migrate the crown jewels simultaneously.

And there is one failure mode specific to "modernize later" that has to be managed explicitly: later never comes. Once the pain of the old data centre is gone, the modernization ticket loses its sponsor. So the migrate-first plan is only honest if the modernization work has a date, an owner and a budget attached at the point you defer it — otherwise you have chosen to pay cloud prices for on-prem habits indefinitely.

◈ What they're really checking

Whether you name the "later never comes" failure mode. Recommending migrate-first without it is the answer that produces the estate of idle fixed-size VMs everyone complains about two years on.

Domain 9 · Cost & FinOps

☺ Like you're 10: The meter is always running. Someone has to watch it, and in an interview that someone is being asked whether it could be you.

Cost questions are increasingly common for migration roles, because the post-migration bill is where migrations get judged. The lesson is the FinOps section of Security, Cost & Resilience.

🟢 Q37 · What are the three pricing models, and when do you use each?

They ask: "How do you buy compute in a way that isn't wasteful?"

Model answer. Three shapes, and every cloud has all three under different names. On-demand: pay by the hour or second with no commitment, cancel any time. Highest unit price, and correct for spiky, unpredictable or short-lived workloads. Committed capacity — reserved instances, savings plans, committed-use discounts depending on the provider: you commit to a level of usage for a term, typically one to three years, in exchange for a substantial discount. Correct for the steady, always-on baseline you are confident about. Spot or preemptible: rent spare capacity at the deepest discount, with the condition that the provider can reclaim it at short notice. Correct only for work that can survive being interrupted mid-run — batch processing, CI, rendering, non-urgent analytics.

The art is layering them rather than picking one: commit the baseline you are sure of, cover the unpredictable peaks with on-demand, and run genuinely throwaway interruptible work on spot. The one hard rule is that anything a customer is waiting on does not go on spot.

◈ What they're really checking

That you present them as a portfolio to be layered, not a menu to choose one from. If you also note that the names differ per provider but the three shapes are universal, you have shown provider-neutral thinking without being asked for it.

🟡 Q38 · The bill after our lift-and-shift is higher than the data centre was. Why, and what do you do?

They ask: "This actually happened. Diagnose it."

Model answer. The usual root cause is that the migration rehosted the sizing along with the workload. On-prem, machines were bought for peak load plus headroom for a multi-year refresh cycle, and that headroom was a sunk cost nobody saw monthly. In the cloud that same headroom is metered by the hour, so the over-provisioning that was invisible becomes a line item. On top of that, the common contributors: the source estate never got decommissioned, so you are paying twice; nothing autoscales, so peak-sized capacity runs at three in the morning; non-production environments run twenty-four hours a day; orphaned resources accumulate — unattached disks, forgotten snapshots, idle load balancers, test environments nobody deleted; and data-transfer charges that nobody modelled.

What I'd do, in payback order. Right-size from real utilisation data — this is where the discovery utilisation dataset pays for itself, and if it wasn't collected, gather a few weeks of it now. Kill the orphans and schedule non-production to shut down outside working hours, which is often a surprisingly large win for almost no risk. Finish decommissioning the source. Turn on tagging so spend is attributable, and give the number a named owner. Then, once the shape has stabilised, buy commitments against the steady baseline.

That ordering matters and I'd say so explicitly: committing before right-sizing locks in the waste for one to three years. Commitments are the last step, not the first.

◈ What they're really checking

The sequencing — right-size then commit. It is the single most valuable practical insight in cloud cost management and a large fraction of candidates get it backwards.

🟡 Q39 · What is FinOps, and what actually makes it work?

They ask: "Who owns cloud cost?"

Model answer. FinOps is the practice of treating cloud spend as an ongoing engineering concern shared by the people who create it, rather than as a finance surprise discovered at month end. The shift is that on-prem you bought a server once and it became a fixed cost; in the cloud every running resource is a meter, and the person who can turn it off is an engineer, not an accountant.

What makes it work in practice is unglamorous. Tagging, enforced rather than requested, because untagged spend is an unattributable number and you cannot cut a cost you cannot attribute. Right-sizing as a recurring habit rather than a one-off project. Budgets and alerts so a runaway cost is caught in hours instead of at invoice time. Showback — putting each team's own number in front of them regularly, because visibility changes behaviour more reliably than policy does. And a named owner from day one; "everyone owns cost" reliably means nobody does.

The framing I'd land on: cloud cost is not a price tag, it is a choice you keep making. The same workload can cost triple or a third depending on size, pricing model and region — which is the whole benefit, but only if somebody is watching.

◈ What they're really checking

Whether you understand FinOps as a social practice with technical enablers, rather than as a dashboard. The named-owner point is the one hiring managers care about, because it is the thing they will have to staff.

🔴 Q40 · How do you model TCO honestly, including the parts people leave out?

They ask: "Give me the cost model you'd put in front of a steering committee."

Model answer. I'd build three columns, not two: the true cost of staying, the true cost of running in the cloud, and the one-off cost of getting there — because leaving out the third is how business cases get approved and then blown.

Staying is more expensive than the finance system usually shows, because the hidden components are spread across other budgets: hardware refresh amortisation, power and cooling, floor space or lease, network, backup infrastructure, software licences, and staff time spent on maintenance that would not exist in a managed service. Running in the cloud has to include the things that are easy to omit: data transfer including egress, storage tiers and their retrieval costs, backup and DR capacity, observability tooling, and support plan. Getting there includes migration tooling, any professional services, training, and the double-bubble period where both estates run at once — plus, honestly, the cost of not decommissioning on schedule, which is the most common overrun.

Two things I'd insist on presenting. A range with named assumptions rather than a single number, calling out the two or three assumptions that move it most — usually the right-sizing ratio, commitment coverage, and the length of the double-bubble. And a note on where the answer could legitimately come out the other way: at very large, steady, predictable scale, particularly where infrastructure is close to the core product, owning can beat renting. That is the shape behind the well-known repatriation stories, and a business case that cannot survive that being raised in the room is not a business case.

◈ What they're really checking

Intellectual honesty under commercial pressure. A candidate who presents cloud economics as always favourable is either inexperienced or telling the room what it wants to hear, and interviewers for consulting-flavoured roles are specifically screening for the second.

⚠ On the numbers

If you cite public figures from well-known migration or repatriation stories, cite them as shapes rather than precise facts, and say you would verify against primary sources. Published migration stories get simplified in retelling, and being confidently wrong about someone else's savings number is a worse outcome than not quoting one — see the case-study caveat on Case Studies.

Domain 10 · Security

☺ Like you're 10: Locking the doors while you build, not after the burglars have visited. Security questions in a migration interview are mostly about timing.

The lesson is the security and compliance half of Security, Cost & Resilience; the guardrail mechanics are in Architecture Patterns.

🟢 Q41 · How do you apply least privilege during a migration?

They ask: "Walk me through your access model for the new environment."

Model answer. Least privilege means every person and every program gets exactly the permissions it needs and nothing more. Making that real rather than aspirational comes down to a handful of habits. Prefer roles over long-lived keys: give a workload a role it assumes for short-lived, automatically-rotating credentials instead of a permanent secret, because fewer standing credentials means fewer credentials to steal. Attach permissions to groups and roles rather than hand-crafting them per person, so access is consistent and auditable. Require multi-factor authentication, because a stolen password alone should not be enough. Lock away the root or global-administrator identity, protect it strongly, and do daily work through limited identities. And remember that machines have identities too — pipelines, functions and migration tools all authenticate, and they are routinely the over-privileged ones because nobody thinks of them as users.

There is a migration-specific version of all this worth adding: the tooling itself. Replication agents, discovery collectors and migration services need credentials into both the source and the target, often with broad rights, and those credentials frequently outlive the project. Scope them tightly, time-box them, and make revoking them a step in the wave's own decommissioning checklist.

◈ What they're really checking

Whether you think about the migration tooling's own access, not just the workloads'. That is the answer only someone who has run a migration gives.

🟡 Q42 · Encryption in transit and at rest — what's the interesting decision?

They ask: "How do you handle encryption on a migration?"

Model answer. Encryption at rest protects data sitting on a disk, a backup or in storage, so a stolen drive yields nothing useful. Encryption in transit protects it while it moves, so nobody tapping the link between the source environment and the cloud can read or tamper with it. On a migration both matter in a specific extra place: the replication stream itself, and — if you are using one — the offline transfer appliance, which is encrypted precisely so a box lost in the post is not a lost dataset.

The interesting decision, though, is not whether to encrypt. Most managed storage and database services now encrypt at rest by default, so that question mostly answers itself. The real decision is key management, because whoever holds the key holds the data. Provider-managed keys are the easy default and are appropriate for a great deal of workloads. Customer-managed keys give you control and separation of duties — you can revoke access to the data independently of the storage system — at the cost of owning the key lifecycle: rotation, access policy, backup of key material, and the very real operational risk that losing key access means losing the data. Some regulated contexts effectively require the customer-managed model.

So my answer to "how do you handle encryption" is: encrypt both states as a baseline, then have a deliberate conversation about who holds the keys and what the recovery story is if that access is lost.

◈ What they're really checking

That you move the conversation from encryption to key management. Anyone can say "encrypt everything"; the senior answer is about custody and the failure mode of custody.

🟡 Q43 · How do compliance frameworks land on a migration plan?

They ask: "We're regulated. What does that change about how you'd run this?"

Model answer. The first thing I'd establish is which frameworks apply, because each is tied to a kind of data rather than a kind of company, and one business commonly lives under several at once — personal data of European residents, health records, cardholder data, and a customer-facing assurance report can all apply to the same organisation simultaneously.

What it changes practically: region choice becomes a legal decision rather than a latency one, because data-residency rules can require data to physically stay within a country or region — and that constrains not only where the workload lands but where you replicate backups and where you are permitted to fail over. It changes who may access what, and how you evidence that. It changes retention and deletion behaviour. It changes what you may copy into lower environments, which is where migrations get caught out, because copying production data into a test environment to rehearse a cutover is a completely normal migration activity and a completely abnormal compliance one. And it changes sequencing: regulated workloads are usually grouped into the same wave so their controls get designed once rather than re-argued per application.

The clarification I'd make explicitly is that a provider service being described as eligible or in-scope for a framework does not make your use of it compliant. The tooling can be used compliantly; you still have to configure it correctly, produce the evidence, and — for some frameworks — have the appropriate contractual agreement in place with the provider. Compliance is shared, exactly like security.

◈ What they're really checking

The eligible-versus-compliant distinction, and whether you spot the test-data problem. The second one is migration-specific and is a genuinely common real-world incident.

🔴 Q44 · Where do migrations most often go wrong on security?

They ask: "In your experience, what's the security failure you'd be watching for on this programme?"

Model answer. The structural failure is treating security as a phase rather than a design input — "get it working, we'll lock it down in phase two." Later does not come, because once the workload is live the appetite for a change that only reduces risk evaporates. The fix is that identity, encryption and least privilege are built into the landing zone before wave one, so every wave inherits them rather than needing them retrofitted.

Underneath that, the specific things I'd watch for on a migration in particular, as opposed to cloud work generally. Storage or databases opened up "just to test the replication" and never closed. Broad administrative rights handed out during a cutover window for speed and never revoked. Credentials pasted into runbooks and migration scripts, then committed to a repository. Replication and management endpoints exposed to the public internet because it was faster than arranging private connectivity. Production data copied into a lower environment for cutover rehearsal without masking. The temporary hybrid link that outlives the migration and quietly becomes permanent unmanaged infrastructure. And after cutover, the source systems left running, unpatched, unmonitored and still holding a full copy of the data — which is both a live attack surface and a compliance exposure, and it happens because decommissioning is the step everybody defers.

◈ What they're really checking

Whether you can produce migration-specific security risks rather than generic cloud security advice. The abandoned source system is the strongest item on that list, because it connects security to the decommissioning discipline they will already have asked you about.

Domain 11 · Resilience & DR

☺ Like you're 10: What happens if it all falls over. Two numbers decide everything, and interviewers want to hear you ask for them first.

The lesson is the resilience section of Security, Cost & Resilience; the RTO/RPO framing recurs in Wave Planning and Best Practices.

🟢 Q45 · Define RTO and RPO, and give me an example.

They ask: "What do RTO and RPO mean?"

Model answer. RTO — recovery time objective — is how long a system may be unavailable before the impact becomes unacceptable. RPO — recovery point objective — is how much recent data you can afford to lose, measured in time. Picture the disaster in the middle of a timeline: RPO measures backwards to your last good copy, and RTO measures forwards to being back in service. If you back up every six hours, your RPO is six hours, because that is the most work you could lose.

An example: a payment path might have an RTO of minutes and an RPO of near zero, because both a short outage and any lost transaction are directly and immediately expensive. A marketing site might sit at an RTO of hours and an RPO of a day, and paying for anything tighter would be waste. The important part is that these are business decisions, set per application, before you shop for a solution — they are inputs to the design, not outputs of it. On a migration they also do double duty: each wave's cutover is validated against the app's RTO and RPO targets, so the numbers agreed in the business case become the pass criteria at two in the morning.

◈ What they're really checking

That you say "set by the business, per application, before choosing a solution." Candidates who describe RTO/RPO as properties of a technology have the causality backwards.

🟡 Q46 · Name the DR strategies and what each buys you.

They ask: "What are the options, cheapest to most expensive?"

Model answer. Four classic rungs, each buying a tighter RTO and RPO for more money. Backup and restore: keep backups, rebuild after a disaster. Recovery in hours to days, cheapest. Pilot light: the core pieces — typically the database and the essential configuration — run small and always-on in the recovery site, and you scale everything else up when needed. Recovery in tens of minutes, moderate standing cost. Warm standby: a scaled-down but complete copy runs continuously and can take traffic and then scale. Recovery in minutes, higher standing cost. Multi-site active/active: two full environments serve live traffic simultaneously, so losing one is barely visible. Near-zero recovery, most expensive.

Two things I'd add. First, you do not need the top rung for everything, and a mixed estate is a correct outcome, not an inconsistent one — a hobby-grade internal tool on backup-and-restore beside a payment path on warm standby is exactly right. Second, and more important: test it. A backup that has never been restored is a wish, and a failover that has never been rehearsed is the one that fails at two in the morning. I'd want the recovery drill on a schedule, with the measured time compared against the promised RTO — because that comparison is the only thing that turns the number in the business case into a fact.

◈ What they're really checking

Whether "test the recovery" comes from you unprompted. It is the most-agreed and least-practised idea in the whole discipline.

🟡 Q47 · High availability versus disaster recovery — and how do regions and zones map onto that?

They ask: "Are HA and DR the same thing?"

Model answer. No — they answer different fears. High availability keeps a system running through small, everyday failures: one machine dies, one data centre in the area has a problem, and traffic simply shifts so users barely notice. Disaster recovery is the plan for a large failure that takes out a whole zone or a whole region, where "shift the traffic" is not available because the thing you would shift it to is also gone.

The geography maps onto that directly. A region is a geographic area where a provider runs a cluster of data centres. Inside a region are two or more availability zones — physically separate data centres with independent power, cooling and networking, close enough for fast connections but far enough apart that a single flood or fire should not take out two. Spreading a workload across zones inside one region is high availability. Spreading it across two regions is disaster recovery.

Two consequences for a migration. Regions are also how data residency is honoured, so your DR region is a legal choice as well as an engineering one — you cannot fail over to a region a regulation forbids, and discovering that after you built the DR design is expensive. And the multi-region decision is a cost decision above all: the standing cost of the second site has to be weighed against the cost of the outage it prevents, and for many workloads the honest answer is that multi-zone HA plus tested backups is the right level.

◈ What they're really checking

Clean separation of the two concepts, and the residency constraint on the DR region. The last point — that most workloads do not need multi-region — is a maturity signal, because the naive answer is always "go multi-region."

🔴 Q48 · How would you pick a DR strategy for an application you've just migrated?

They ask: "The app is live in the cloud. Design its DR."

Model answer. I'd work in this order. Start with the business: what are this application's RTO and RPO, agreed by whoever is accountable for it, not inferred by me. Then check what the data layer can actually deliver, because that sets the floor — your achievable RPO cannot be tighter than your replication lag, and if the database replicates asynchronously across regions with a measurable delay, promising a near-zero RPO is a promise the architecture cannot keep. Then check the constraints: residency rules on which regions are permitted, and any dependency that exists in only one place — an on-prem system, a third-party endpoint with an allow-list, a licence server — because a failover that moves the app but not its dependencies is theatre.

Then price it: the standing cost of the chosen rung against the cost of the outage it prevents, expressed in the business's terms — lost revenue per hour, regulatory exposure, contractual penalty. That comparison is what makes the conversation a decision rather than a preference.

Finally, and non-negotiably, write the runbook and schedule the drill, with the measured recovery time compared to the target. I'd also say plainly that if nobody will fund the drill, then the organisation has chosen backup-and-restore in practice regardless of what the design document says — and it is better to say that out loud early than to discover it during an incident.

◈ What they're really checking

Whether replication lag appears as the floor on RPO, and whether you check dependencies before declaring a failover viable. Both are the kind of detail that only shows up after someone has watched a failover test fail.

Domain 12 · Anti-Patterns

☺ Like you're 10: The shortcuts that feel clever today and bill you next quarter. Interviewers often hand you a plan with one buried in it and see if you notice.

The lesson is Anti-Patterns & Pitfalls, with the mirror-image habits in Best Practices.

🟡 Q49 · What anti-patterns do you watch for, and what's the tell for each?

They ask: "What are the ways these programmes usually go wrong?"

Model answer. I'd group them by where they bite. Planning traps, sprung before anyone touches a server: skipping discovery and dependency mapping; migrating the junk because nobody rationalised the portfolio; big-bang cutovers with no phasing; boiling the ocean by scoping the entire estate into one quarter; and skipping the pilot to move the crown jewel first. Foundation traps, baked into what you build: no landing zone, so accounts are hand-made snowflakes; security deferred to phase two; the distributed monolith; premature microservices. Execution traps, which strike at cutover: no rollback plan; validation that is a quick smoke test rather than a data comparison; and underestimating data gravity and egress. Cost and lock-in traps, whose invoice arrives long after the decision: no FinOps and the resulting bill shock; lift-and-shift-and-forget, where the temporary rehosted state becomes permanent; and lock-in that was never a conscious choice.

The useful part is the tells, because in a real meeting you hear the phrase before you see the trap. "We don't have time for discovery" — skipped inventory; ask what this app depends on that we have not mapped. "We'll modernize later" — lift-and-forget; ask what date, owned by whom. "Let's just move it all at once" — big bang; ask what the rollback is if the first batch fails. "Cost is finance's problem" — no FinOps; ask who owns the monthly number. "Security review can wait" — afterthought; ask who can read this data on day one. Each of those questions costs a minute and can save a weekend.

And the pattern behind all of them: the reward is immediate — less work today — and the punishment is delayed. That is why experienced teams still fall for them.

◈ What they're really checking

Whether you have a structure rather than a grab-bag, and whether you can convert each trap into a question you would actually ask in a meeting. The tells are what make this answer sound like experience rather than reading.

🟡 Q50 · Why is a distributed monolith worse than the monolith you started with?

They ask: "We split our monolith into services. Why isn't it better?"

Model answer. Because you have taken on all of the costs of a distributed system and none of its benefits. The costs arrive immediately and unavoidably: calls that used to be in-process are now network calls that can fail, time out, or arrive twice; you now have versioning and compatibility to manage across boundaries; debugging needs distributed tracing because a single user action spans several services; and you have multiplied the things to deploy, secure, monitor and pay for.

The benefit of microservices — the thing you paid all of that for — is independence: the ability to deploy, scale, and change one piece without touching the others. If your services must all be deployed together, and a change in one forces changes in the others, you never received that benefit. So you are strictly worse off than the monolith, which at least gave you one deployment, in-process calls, and a single stack trace.

The test I'd apply is blunt: if two services must always ship together, they are one service. And the way to avoid it is to split along real business boundaries — where the data and the decisions genuinely separate — rather than along technical layers or along the shape of the existing code.

◈ What they're really checking

Whether you can articulate what microservices actually buy, precisely enough to notice when you have not bought it.

🔴 Q51 · Leadership wants the whole estate migrated this quarter and says there's no time for discovery. How do you respond?

They ask: "You've said discovery is essential. The sponsor says it's a luxury. Go."

Model answer. I'd start by finding out what is actually driving the date, because "this quarter" is usually standing in for something concrete — a lease, a contract renewal, a hardware support expiry, a commitment made to a board. The response is completely different depending on which. If it is a real external deadline, the plan should be sequenced around that constraint. If it is an internal ambition, it is negotiable and should be negotiated with evidence rather than opinion.

Then I would not refuse discovery; I would time-box it. Discovery is not binary — a two-week automated collection over the estate gives you an inventory and the bulk of the dependency picture, and that is enough to start. Refusing to move without a perfect map is its own anti-pattern, and it hands the sponsor a reason to work around me.

In parallel I would start the pilot, so there is visible movement while the collection runs. That does two things: it gives the sponsor progress to point at, and it produces the only credible input to a re-forecast — a real measurement of how long one wave actually takes with this team, this tooling and this estate. Then I re-plan from that measurement rather than from either of our opinions.

What I would hold firm on is small and specific, because a short list is defensible where a long one is not: the pilot happens before the crown jewels, every cutover has a written rollback, and data validation is not optional. Those three cost days, not months, and they are the difference between a delayed programme and a recoverable disaster. And I'd say the underlying thing out loud, politely: pressure is exactly when these steps get skipped, and exactly when their absence costs the most.

◈ What they're really checking

Whether you can disagree upward without becoming an obstacle. The structure to notice: understand the real constraint, offer a reduced version rather than a refusal, create evidence, and hold a short non-negotiable list. Candidates who simply restate the best practice fail this question even though their content is correct.

Domain 13 · Cloud-to-Cloud & Hybrid

☺ Like you're 10: Moving between clouds, living in two at once, or moving back down. The trap is assuming that because both ends say "cloud," it must be easy.

The lesson is Cloud-to-Cloud, Hybrid & Repatriation; the connectivity patterns are in Architecture Patterns.

🟢 Q52 · Why is cloud-to-cloud often harder than on-prem to cloud?

They ask: "Both ends are already cloud. Surely that's the easy case?"

Model answer. It is frequently the opposite, and the reason is that the shared word "cloud" hides how different the two environments are internally. The compute layer usually moves cleanly — a virtual machine is a virtual machine, and a container is more portable still. What has to be rebuilt is everything around it. Managed services rarely have a one-to-one twin on another provider; the nearest equivalent behaves differently, which means re-testing and sometimes redesign. Identity and access models are architecturally different per provider, so roles and permissions get rebuilt rather than copied. Networking uses different names, different defaults and different constructs. Infrastructure-as-code written for one provider needs rewriting or a substantial portable rework. Your engineers' hard-won knowledge of one provider's quirks only partly transfers, so there is a retraining cost. And you pay egress to the provider you are leaving — the toll is charged by the source, so a cheaper destination does not reduce it.

The general principle: the more deeply you used one provider's proprietary managed services, the harder the move. A workload built on portable technology travels far more easily than one wired into a specific vendor's distinctive offerings.

◈ What they're really checking

Whether you separate "the compute moves" from "the surroundings get rebuilt." That framing is the whole answer, and it is what stops teams from budgeting a cloud-to-cloud move as if it were a data copy.

🟡 Q53 · Hybrid versus multi-cloud — and what does each cost?

They ask: "Define both, and tell me when you'd recommend them."

Model answer. Hybrid cloud is on-premises infrastructure and public cloud connected so they operate as one system. Multi-cloud is two or more public clouds used deliberately at the same time. Hybrid mixes ground and sky; multi-cloud mixes sky and sky.

Hybrid is often the intended permanent design rather than a half-finished migration. The legitimate reasons are concrete: data that legally must stay on-premises or in a specific place; latency-sensitive workloads that need compute physically close, like a factory floor or a hospital; and gradual migration, where the two halves have to keep working together for months or years. That last one is worth stating explicitly in an interview, because nearly every large migration passes through a hybrid phase even when the end state is all-cloud — while waves are still landing, some apps are up and some are not, and they still have to talk over a private link.

Multi-cloud is chosen for resilience against a single provider's outage, for best-of-breed access to a particular provider's strongest service, and for negotiating leverage and reduced lock-in. The cost is the multi-cloud tax, and it is paid mostly in skills and tooling: separate identity systems to keep aligned, separate monitoring and billing to reconcile, cross-cloud networking and its egress, and engineers who must stay fluent in more than one platform. So the recommendation is: go multi-cloud for a specific named reason where the benefit clearly exceeds that standing cost — and not because teams drifted onto different platforms and someone wrote a strategy afterwards to explain it.

◈ What they're really checking

Precision on the definitions, and whether you name the tax. Also whether you point out that hybrid is usually a transitional state as well as a target state — that observation ties this domain back to wave planning.

🟡 Q54 · How do you think about vendor lock-in?

They ask: "Should we avoid provider-specific services to stay portable?"

Model answer. Lock-in is a trade, not a sin — the mistake is making it unconsciously. Using a provider's proprietary managed services is often exactly the right call: they remove real operational burden, they are usually better engineered than what you would build, and the productivity gain is immediate. What makes it dangerous is when nobody ever asked "how would we leave?" and the proprietary pieces are woven through the code with no boundaries.

So my approach is three things. Make it deliberate: decide, per workload, whether portability has value here, because for most workloads it genuinely does not. Where it does, isolate the proprietary pieces behind clear interfaces, so that the blast radius of a future move is known and bounded rather than "the whole system." And where portability is a first-order requirement, prefer portable technology — containers and open standards travel between environments with far less friction than a deeply platform-native design.

The tension worth naming honestly is that these two directions pull against each other. Re-architecting onto a provider's native services gives you the best fit and usually the best price on that provider, and simultaneously re-anchors you to it. Leaning on portable technology keeps your options open at the cost of not using either provider's strongest managed offerings. There is no free lunch — only an honest per-workload choice, made with the exit cost written down.

◈ What they're really checking

Whether you can hold a genuine trade-off without collapsing it into a slogan. "Avoid lock-in" and "embrace the platform" are both slogans; naming the tension and giving a per-workload decision rule is the answer.

🟡 Q55 · When is repatriation the right call, and how do you argue it without sounding like the cloud failed?

They ask: "Would you ever recommend moving something back on-premises?"

Model answer. Yes, for specific workloads, and I'd frame it as a per-workload decision rather than a verdict on the cloud. The drivers are well understood: cost at large, steady, predictable scale — the cloud's pay-as-you-go model shines for variable demand and is less compelling for a big constant load, especially where infrastructure is close to the core product; performance for heavy or specialised workloads that run more consistently on dedicated hardware; data control where the organisation wants sensitive data fully in-house; and compliance, where a new regulation requires something the available regions cannot satisfy.

How I'd argue it: with a TCO comparison for that workload specifically, including the parts people leave out on the on-prem side — hardware refresh, power, floor space, staff time, and the resilience you would have to rebuild yourself. And I'd point out that repatriation is almost always selective: a few workloads come home and the rest stay in the cloud, which is really hybrid by another name. That framing defuses the political temperature, because the conversation stops being "was the cloud a mistake" and becomes "what is the right home for this particular workload."

If I referenced the well-known public examples, I would use them as shapes rather than facts and say so — the widely-discussed cases involved organisations operating at very large scale where infrastructure was core to the product, and at least one of them deliberately stayed hybrid rather than leaving entirely. The specifics get simplified in retelling, so I'd verify before quoting numbers at anyone.

◈ What they're really checking

Whether you are dogmatic. A migration specialist who cannot imagine a case for moving something back is a specialist whose recommendations are predictable, and predictable recommendations are worth less than judgement. Framing it per workload is what keeps it from sounding like heresy.

Scenario & design questions

☺ Like you're 10: "Here's a messy situation — what would you do?" There's no single right answer, so what they're really watching is how you think, not what you conclude.

Scenario rounds are where migration interviews are won and lost, because the prompt is deliberately underspecified. The failure mode is jumping straight to a solution — naming a tool, or announcing "I'd lift and shift it" — within twenty seconds of hearing the prompt. What the interviewer wants is to watch you convert an ambiguous situation into a structured plan, out loud, while asking for the information you are missing.

◆ The reusable structure — use this on every scenario

Say these steps out loud as you go; the visible structure is half the marks.

  1. Clarify the driver and the constraint. "What happens if we don't do this, and by when?" The driver decides the strategy; the deadline decides the sequencing.
  2. Establish what we know and what we'd have to find out. Name the discovery inputs you need — inventory, dependencies, utilisation, data volumes, change rates — and say explicitly which of them you are going to assume for the sake of the discussion.
  3. State your assumptions out loud. "I'm going to assume a few hundred VMs and a single data centre; tell me if that's wrong." This is not a stalling tactic — it is what lets the interviewer correct you early, which is exactly what they want to do.
  4. Propose an approach, in phases. Foundation first, pilot, then waves. Say what goes in wave one and why.
  5. Name the trade-offs of your choice. What you gave up by choosing this. There is always something.
  6. Say how you'd verify. What "done" means for each phase, and what evidence proves it.
  7. Say how you'd back it out. Every irreversible step gets a reverse step and a decision point.

Steps 6 and 7 are the two most frequently skipped and the two the interviewer is most reliably listening for.

Scenario 1 · The data centre lease ends in nine months

The prompt: "A mid-sized retailer has 180 applications in one data centre. The lease expires in nine months and won't be renewed. Design the programme."

Open with the clarifying questions: Is the date genuinely immovable, or is a short holdover possible at a price? What is the appetite for downtime per application class? How many people do we have, and are the application owners committed or merely aware? Is there any existing cloud footprint or landing zone at all, or are we starting from nothing?

A reasonable shape. The driver here is a hard exit date, not agility — so this is a rehost-dominant programme and I would say that early and defend it. Months one and two: build the landing zone and connectivity, run automated discovery across the estate, and rationalise the portfolio aggressively, because the fastest way to make nine months feasible is to reduce 180 applications to the number that genuinely has to move. Month two also runs the pilot: one small, non-critical, low-dependency application, all the way through cutover and a rehearsed rollback. Months three to eight: waves, sized so each fits a single window with rollback room, starting with the low-risk and independent, ending with the revenue-critical. Month nine is deliberately empty — it is contingency, and if you have not left it you have not planned, you have hoped.

Trade-offs to say out loud. Rehost-dominant means carrying existing inefficiencies into a metered environment, so the bill in month ten will be uncomfortable unless right-sizing is scheduled as an explicit follow-on with a date and an owner. Anything that genuinely cannot be moved in time needs an early, honest answer — a colocation holdover for a handful of systems is a legitimate outcome and is much cheaper than a failed cutover on a crown jewel in month nine.

⚠ The trap in this one

Proposing modernization because it sounds ambitious. With a hard external date, "let's containerise while we're in there" is the answer that misses the deadline. Say explicitly that you are separating the migration from the modernization and putting a date on the second.

Scenario 2 · Move a 40 TB database with under 15 minutes of downtime — and change engine

The prompt: "There's a 40 terabyte commercial database behind the order system. Finance wants off the licence, so the target is an open-source engine. The business will give you fifteen minutes."

Open with the clarifying questions: What is the write rate, and what is the size of the daily change volume — because that, not the 40 TB, determines the cutover. What is the RPO, in the business's words? What does the application do with the database beyond plain queries — stored procedures, triggers, engine-specific SQL? Can the source expose a transaction log for change capture, and do we have the privileges? Is there a lower environment with production-shaped data to rehearse against?

A reasonable shape. This is two projects wearing one hat, and I'd name that: an engine conversion, and a low-downtime move. Take them in that order. First, schema conversion — run the conversion tooling, then hand-fix everything it flags, because the flagged items are exactly the ones with no equivalent on the target. Then the application work, which is the part people underestimate: dialect differences, data-type and precision edge cases, collation and sort order, and transaction-isolation behaviour, each of which shows up as a subtle behavioural difference rather than an error. Then test functionally and under representative load against a converted copy, and only then start the data move: bulk load, then continuous change capture, monitoring replication lag as the number that bounds the achievable RPO.

The fifteen-minute window is then: stop writes, drain the final sync, validate — row counts and chunked checksums plus a small set of application-level comparisons that were pre-written and timed during rehearsal — then repoint and open up. The window is achievable precisely because the 40 TB moved days earlier.

Trade-offs and the honest caveat. Fifteen minutes is a validation budget, not a copy budget, and the validation set has to be designed to fit inside it — you cannot checksum 40 TB in the window, so you agree in advance which checks are gating and which run afterwards. And I would want to know whether fifteen minutes is a hard business limit or an opening position, because a parallel-run period, where the new system answers the same queries in shadow and its answers are compared without being served, is a much stronger way to build confidence in a money-moving system than any amount of pre-cutover testing.

⚠ The trap in this one

Treating it as a data-movement problem. The data move is the well-understood part; the risk is in the engine conversion and the application's dependence on engine-specific behaviour. A candidate who spends the whole answer on replication tooling has misread the question.

Scenario 3 · 600 TB of media archive, and the network you actually have

The prompt: "There's a 600 terabyte media archive to move. The site has a 1 Gbps link that the business also uses. How do you get it there?"

Open with the clarifying questions: How much of that link can we actually have, and during which hours? Is the archive static, or still growing — and at what rate? Is it needed for retrieval during the transfer? What is the retention and access profile, because that decides which storage tier it should land in? And is anything in it subject to residency or retention rules?

A reasonable shape. Do the arithmetic first, out loud, and be explicit that link speed is not usable bandwidth — after the business's own use, protocol overhead and real-world sustained throughput on a long path, the usable share of a contended 1 Gbps link is a fraction of the headline number, and the honest transfer estimate lands in months rather than weeks. That is the case for shipping the bulk on an encrypted offline appliance, whose duration is essentially fixed by logistics rather than by size.

The pattern I'd propose is the combined one: appliance for the historical bulk, then an online transfer service to sync the delta that accumulated while the box was in transit and being ingested, then a final reconciliation before the source is retired. For a dataset this size you would likely need several appliance rotations, which is a scheduling exercise in itself — and each rotation needs its own manifest and validation, not one big check at the end.

Trade-offs and the numbers to raise. Egress is charged by the source you are leaving, so it belongs in the budget before anyone commits. Storage tiering matters at this volume: an archive that is genuinely rarely read belongs in a cold or archive tier, but retrieval from those tiers is slow and charged, so you need the access profile before you choose — putting cold data in a hot tier wastes money every month, and putting warm data in an archive tier produces an unpleasant surprise the first time someone needs it back quickly.

⚠ The trap in this one

Answering "use an appliance" without doing the arithmetic. The appliance is probably right, but the interviewer wants to see the calculation and the words "usable bandwidth" — the reasoning is the answer, and it is the reasoning that transfers to the next dataset.

Scenario 4 · A landing zone for a regulated organisation with residency constraints

The prompt: "A healthcare company is migrating. Some data legally has to stay within a specific region. Design the foundation."

Open with the clarifying questions: Which frameworks apply, and to which categories of data specifically — because the constraint is on data types, not on the company. Which regions are permitted, and does that constraint extend to backups and to the DR region? Who inside the organisation owns compliance evidence, and what do their auditors ask for? Is there an existing on-prem identity system that must remain authoritative? How many application teams will be operating in this environment?

A reasonable shape. Start with data classification, because everything else follows from it — you cannot design residency controls before you know which data is constrained. Then the account structure: production separated from non-production always; regulated workloads separated so their controls are designed once and evidenced once; shared platform services — connectivity, centralised logging, identity — in their own accounts distinct from where application teams work; and an organisation tier above it all for org-wide policy and consolidated billing.

Then the controls. Preventive guardrails on the things that must never happen — resource creation outside the permitted regions, storage exposed publicly — because these are absolute rules and a false block is cheaper than the event. Detective controls with alerting for everything contextual. Centralised, tamper-resistant logging with a retention period that matches the audit requirement, because "we have logs" and "we have logs going back as far as the auditor asks" are different claims. Encryption everywhere, with an explicit decision on key custody, since some regulated contexts effectively require customer-managed keys.

Then the migration-specific controls that are easy to miss: masking or synthetic data for lower environments, because rehearsing a cutover with real production data in a test account is a normal migration activity and an abnormal compliance one; and time-boxed, tightly-scoped credentials for the migration tooling, with revocation as an explicit step in each wave's closure.

What I'd say about DR. The DR region is now a legal choice as well as an engineering one — you cannot fail over to a region the rules forbid, so residency has to be checked before the resilience design, not after.

⚠ The trap in this one

Treating "regulated" as a synonym for "more encryption." The distinctive work is evidence, region constraints, data classification, and test-data handling. Also, do not claim a provider service "is compliant" — say it can be used compliantly given the right configuration, evidence and agreements.

Scenario 5 · A core order system that cannot go down and cannot stay as it is

The prompt: "The order system is a fifteen-year-old monolith. It's brittle, expensive to change, and it takes money every second. Modernize it."

Open with the clarifying questions: What specifically hurts — is it change lead time, scaling, reliability, or cost? Because "it's old" is not a driver and the answer differs for each. Who owns it, and does anyone still understand it? Is there test coverage? What is the actual traffic shape, and which parts of it are hot? And is there a deadline attached, or is this a standing programme?

A reasonable shape. Not a rewrite, and I'd say why in one sentence: a two-year secret rebuild followed by a flip is the highest-risk option available for a system that moves money. Instead, incremental replacement. Put a facade in front — a router that receives every request and initially sends all of it to the existing system, so day one changes nothing. Then extract one capability at a time, starting with something valuable but peripheral rather than with checkout: a read-heavy path is usually the right first extraction because it is reversible and low-consequence. Reroute just that path, verify, repeat. Where the new services must speak to the legacy data model, put an anti-corruption layer between them so the legacy quirks do not leak into and rot the new design.

For anything that computes money — pricing, tax, totals — run it in parallel before you trust it: the legacy system's answer is what the customer sees, the new system's answer is computed and recorded silently, and you compare. Weeks of agreement is evidence; a passing test suite is a hope. Then cut that path over with blue-green or a canary, depending on whether you value instant rollback or early warning more.

Trade-offs to say out loud. You run and pay for both systems for a long time, and "a long time" here means years, not months — that has to be in the business case or the programme will be judged as late. The facade becomes critical infrastructure and a new single point of failure. And the strangler needs a defined end: extractions that stall halfway leave you permanently operating two systems, which is the worst of both, so each extraction needs an owner and a completion criterion.

⚠ The trap in this one

Leading with microservices. The question is about safely replacing a critical system; the number of services at the end is a detail, and proposing a specific decomposition before you understand the domain is the premature-microservices trap in interview form.

Scenario 6 · Consolidating an acquired company onto the acquirer's cloud

The prompt: "We've acquired a startup that runs entirely on one public cloud. We standardise on a different one. Consolidate them."

Open with the clarifying questions: What is the actual business driver — cost, operational simplicity, unified security posture, or a contract? Because that determines how much has to move and how fast. Is there a deadline tied to a commitment or contract renewal? How much of their estate is portable — containers and open technology — versus wired into their current provider's proprietary managed services? How large is the data, and what is the egress exposure? And critically: are their engineers staying, because they are the only people who know the estate?

A reasonable shape. Treat it as a fresh migration, not a copy — re-assess every workload against the 7 R's rather than assuming everything moves. Some of it will be Retire, because acquisitions produce duplicates, and consolidating two ticketing systems or two CI platforms is the cheapest win available. Some will be Retain on the original provider for a defined period, which is a legitimate answer and is much better than forcing an unready workload across a boundary. The portable workloads — containerised services, anything on open technology — move first and cheaply, and give the programme early momentum. The workloads wired into proprietary managed services are the real project, and each needs an individual decision: find the nearest equivalent and accept the behaviour difference, re-architect, or replace.

Around all of it, the non-compute rebuild: identity and permissions rebuilt rather than translated, networking redesigned, infrastructure-as-code rewritten for the target, and monitoring and alerting re-established. Budget the retraining explicitly.

Trade-offs and the people point. Egress on the data is a real and immediate cost charged by the provider being left. But the risk I'd raise first is human: acquisitions lose people, and the engineers who understand the acquired estate are the ones with the most options. Getting the knowledge out of their heads — dependency maps, runbooks, the list of things that are held together with tape — is more urgent than any technical step, and it has a short window.

⚠ The trap in this one

Assuming "cloud to cloud, so it's a copy." The compute moves; identity, networking, managed services and infrastructure code get rebuilt. Say that in the first thirty seconds and the rest of the answer lands.

Scenario 7 · The bill is 40% over the business case

The prompt: "Six months post-migration, cloud spend is forty percent above what we told the board. What do you do?"

Open with the clarifying questions: Forty percent over on what basis — against the original model, or against a model that was built before we knew the estate? Is the overage concentrated or spread? Is spend still growing, flat, or already falling? Is the source estate fully decommissioned? Is anything tagged?

A reasonable shape. First, attribute before you act — without tagging you cannot tell an over-provisioned production fleet from a forgotten test environment, and the two have completely different fixes. So if tagging is absent, that is step zero and it is a days-long job, not a quarter-long one.

Then work in payback order. Finish decommissioning the source, because paying for both estates is the largest and most embarrassing line and it is pure waste. Sweep the orphans — unattached disks, forgotten snapshots, idle load balancers, environments nobody deleted. Schedule non-production to shut down outside working hours. Right-size from real utilisation data, which is where the discovery utilisation dataset earns its keep. Only then buy commitments against whatever baseline has emerged, because committing before right-sizing locks the waste in for the length of the term.

Alongside the technical work, fix the mechanism: named owner for the number, budgets and alerts so the next surprise arrives in hours rather than at quarter end, and showback so each team sees its own consumption. And re-forecast honestly with the board rather than promising the original number back — the original model was built on assumptions that discovery has since corrected, and saying so is more credible than a heroic recovery plan.

⚠ The trap in this one

Reaching for commitments first because the discount is the biggest single lever. It is the biggest lever and it is the last step — buy commitments against a right-sized baseline, or you have bought a one-to-three-year contract for capacity you should not be running.

Behavioural questions, STAR-shaped

☺ Like you're 10: "Tell me about a time when…" These aren't small talk. They're checking whether you're someone the team can survive a bad night with.

Migration work is unusually political for an engineering discipline. You are asking application owners to let you touch systems they are accountable for, on a schedule they did not choose, using tools they did not pick — and then you are asking them to be available at two in the morning. So the behavioural round for this role concentrates on stakeholder friction, decisions under uncertainty, and what you did when something went wrong.

◆ How to use these

Each question below gives a skeleton with [slots in square brackets] for your own material. Fill them from your actual experience — a university project, a home lab, a support role, an internal tool migration all count; the scale matters far less than whether the reasoning is real. Then say the answer out loud once. If you cannot fill a slot honestly, pick a different story rather than inventing one; interviewers ask follow-up questions, and invented stories fall apart in the second or third follow-up, which is a much worse outcome than a modest true story.

STAR: Situation (two sentences of context, no more), Task (what specifically was yours), Action (what you did — most of the answer lives here, in first person singular), Result (what happened, with a number if you have one, and what you would do differently).

B1 · "Tell me about a migration or deployment that went wrong."

What they're testing: whether you can discuss failure without either blaming others or performing excessive contrition — and whether you actually extracted a transferable lesson rather than a resolution to "be more careful."

The skeleton. Situation: "We were moving [system] to [target]. The plan was [approach], and the window was [duration]." Task: "I was responsible for [your specific part — the runbook, the data validation, the cutover coordination]." Action: "During the cutover, [what went wrong] — [what the symptom was and how it surfaced]. I [what you did first: scoped it, checked it against the rollback trigger, escalated, communicated]. We decided to [roll back / stabilise] because [the criterion], and I [what you personally did next]." Result: "[Outcome — including the honest one if it was bad.] Afterwards I [the specific change you made]: [added a check to the runbook / lengthened the discovery window / built a pre-cutover verification step]. On the next [wave / release], [the evidence that the change worked]."

◈ What makes this answer strong

A specific, mechanical lesson. "I learned to communicate better" is a non-answer. "I learned that our go/no-go criteria were written vaguely enough that we argued about them at 2 a.m., so I rewrote them as pass/fail checks with numbers" is an answer — it names a defect, a fix, and a mechanism.

⚠ What sinks it

Choosing a story where the failure was entirely someone else's, or one where nothing actually went wrong ("we were slightly behind schedule but recovered"). Both read as an unwillingness to be examined. Pick something that genuinely cost you.

B2 · "Tell me about a time you pushed back on a plan or a deadline."

What they're testing: whether you will say the uncomfortable thing, and whether you can do it in a way that leaves the relationship intact. For a consulting-flavoured role this is often the single most important question in the round.

The skeleton. Situation: "[Who] wanted [what], by [when], and the concern I had was [the specific risk — not a vague unease]." Task: "I needed to either raise it credibly or accept the risk knowingly." Action: "First I [got evidence rather than arguing from opinion — measured something, ran a small test, priced it]. Then I [how you raised it — to whom, in what forum, framed how]. I offered [the alternative or the reduced version] rather than just objecting, and I was explicit about [what I'd hold firm on and what was negotiable]." Result: "[What was decided — including if you lost.] [What happened as a result.] [What you'd do differently.]"

◈ What makes this answer strong

Evidence, an alternative, and a short non-negotiable list. Also: a story where you pushed back and were overruled, then supported the decision professionally, is often a better answer than one where you won — because it demonstrates you can disagree and commit.

B3 · "Tell me about a difficult stakeholder or a resistant application owner."

What they're testing: whether you interpret resistance as an obstacle or as information. In migration work, the owner refusing to sign off is very often right about something you have not understood yet.

The skeleton. Situation: "[Owner / team] was blocking [the move / the window / the change] for [stated reason]." Task: "I needed their sign-off to [do the thing], on [timeline]." Action: "Rather than escalating first, I [went and asked what they were actually worried about]. It turned out the real concern was [the underlying thing — often not the stated one: a past bad experience, an unmanaged risk, a deadline of their own, or a genuine technical dependency we had missed]. I addressed it by [what you did — a rehearsal, a rollback demonstration, a change to the plan, taking on a task they were dreading]." Result: "[Outcome.] What I took from it was [the general lesson about how to open these conversations]."

◈ What makes this answer strong

Discovering that the stated objection was not the real one, and that finding out cost you one conversation. That is exactly the behaviour a migration lead needs, and it is a story you can tell without making the other person a villain.

⚠ What sinks it

Any version where the resolution is "I escalated to their manager and they were told to cooperate." It may even be true — but as an answer it says you reach for authority before you reach for understanding, and it predicts how you will behave with their stakeholders.

B4 · "Tell me about a decision you made without complete information."

What they're testing: tolerance for ambiguity, and whether you know the difference between a reversible and an irreversible decision. Migration work is permanently short of information; waiting for certainty is itself a failure mode.

The skeleton. Situation: "We had to decide [what] but we didn't know [the missing thing], and finding out would have taken [how long] which we didn't have because [why]." Task: "The call was mine / I had to make a recommendation." Action: "I [what you did to reduce the uncertainty cheaply — a sample, a small test, a conversation with the one person who'd know]. I decided [what] on the basis that [reasoning], and I explicitly noted [what would tell us we'd chosen wrong] and [how we'd reverse it if so]." Result: "[What happened.] [Whether the reversal was ever needed.] [What you learned about when to spend time on certainty and when not to.]"

◈ What makes this answer strong

Naming the reversal path as part of the decision. The mature framing is: for a reversible decision, decide fast and watch; for an irreversible one, spend the time. If you can say which kind yours was, you have answered the question underneath the question.

B5 · "Tell me about a time you found or saved significant money."

What they're testing: commercial instinct. For consulting roles this is close to mandatory, because your value is measured in outcomes the client can see on an invoice.

The skeleton. Situation: "[Where the waste was — an over-provisioned environment, a licence nobody used, a duplicated tool, resources left running.] Nobody had noticed because [why it was invisible — untagged, spread across teams, inside a bigger line item]." Task: "[How you came to be looking — deliberately, or you spotted it doing something else.]" Action: "I [how you quantified it, because an unquantified saving is an opinion]. Then I [what you did — and if it needed someone else's agreement, how you got it]." Result: "[The number, or the honest approximation, with the basis.] [Whether it stayed fixed — and if you put a mechanism in place so it would.]"

◈ What makes this answer strong

The mechanism at the end. Finding waste once is luck; adding the tag policy, the budget alert or the scheduled shutdown so it cannot recur is engineering. Also: give the number with its basis, and say if it is an estimate. Inflated savings numbers are the easiest thing in an interview to probe and puncture.

B6 · "Tell me about learning something technical quickly."

What they're testing: whether you can be dropped onto an unfamiliar estate — and migration consultants are dropped onto unfamiliar estates constantly, sometimes onto a cloud provider they have used least.

The skeleton. Situation: "I needed to be useful on [technology / platform / domain] within [timeframe] because [why]." Task: "Specifically I had to be able to [the concrete capability — not 'learn X' but 'do Y with X']." Action: "I [how you approached it — the actual method: read the primary docs rather than tutorials, built a small thing end to end, found the person who knew and asked specific questions, mapped it onto something I already knew]. The mapping that helped most was [the analogy — e.g. one cloud's identity model onto another's, or infrastructure-as-code onto a config-management tool you knew]." Result: "[What you were able to do, and by when.] [Where you were still weak and how you covered that honestly.]"

◈ What makes this answer strong

Naming what you were still weak at and how you handled it — because the honest version of "I learned it fast" always includes a boundary. For migration roles specifically, the mapping-onto-what-I-knew move is exactly the skill that makes someone useful across three providers, so say it explicitly.

B7 · "Tell me about a technical disagreement with a colleague."

What they're testing: whether you can hold a position on evidence and change it on evidence — and whether you can be wrong out loud.

The skeleton. Situation: "[Colleague] and I disagreed about [the technical question]. Their position was [state it fairly and strongly — this is the part being marked]. Mine was [yours]." Task: "We had to converge because [why it mattered and by when]." Action: "We [how you resolved it — found the data, ran the experiment, wrote both options down with their trade-offs, brought in a third view]. [What you personally did to move it forward rather than dig in.]" Result: "[Which way it went — and if it went their way, say so plainly.] [What the outcome was.] [What it changed about how you handle disagreements.]"

◈ What makes this answer strong

Stating the other person's position generously. If you can articulate their argument better than a caricature of it, you demonstrate that you actually engaged with it — and interviewers notice this within about ten seconds.

B8 · "Tell me about explaining a technical risk to a non-technical audience."

What they're testing: whether you can be trusted in front of a client or an executive. In consulting-shaped roles this is often the deciding question, because it is the one that predicts whether you can be put in a room alone.

The skeleton. Situation: "[Who] needed to make a decision about [what], and the risk was [the technical thing] — which mattered to them because [the business consequence, not the technical one]." Task: "I had to make it real enough to decide on without making it frightening or incomprehensible." Action: "I framed it as [the framing you used — a comparison to something in their world, a number they already cared about, a range of outcomes with likelihoods]. I deliberately left out [the detail you cut], and I made sure they had [the one thing they needed to decide]." Result: "[The decision they made.] [Whether the framing held up afterwards.]"

◈ What makes this answer strong

Saying what you deliberately left out. Anyone can simplify; the skill is knowing which detail changes the decision and which is just interesting to you. Migration gives you natural material here — RTO and RPO are technical terms that translate directly into "how long can we be down" and "how much work can we lose," which is a business conversation.

B9 · "Why do you want to do migration work?"

What they're testing: whether you have thought about the actual shape of this job — which involves a great deal of inventory, coordination, careful arithmetic and two-in-the-morning windows, and comparatively little greenfield building.

The skeleton. Answer honestly and specifically, from one of these directions: you like the fact that the work is finite and verifiable — a wave either landed and validated or it did not; you like the breadth, because a migration forces you across networking, identity, data, cost and the application layer in a way that few roles do; you like being the person who reduces risk rather than the person who adds features; or you came to it from [operations / development / support] and found that [the specific thing] was the part you enjoyed most. Then anchor it to something concrete you have actually done, however small, and be honest about what you have not done yet.

◈ What makes this answer strong

Naming a part of the job that most people find unglamorous and saying you like it — the dependency mapping, the reconciliation, the runbook. It is the most credible signal available, because nobody says that to be impressive.

⚠ What sinks it

"Because cloud is the future." It says nothing, and it is the answer of someone who has confused a migration role with a cloud-architecture role. They are related and they are not the same job — this one is defined by moving things safely under constraint.

🦫 Benny's workshop · 30 min

Pick four of the nine above — the failure one, the pushback one, the stakeholder one, and one other — and write the four slots for each in full sentences. Then read them aloud and time them. Anything over two and a half minutes needs cutting, and the cut almost always belongs in the Situation: you are giving three sentences of context where one would do. Most people find that shortening the setup makes room for the Action, which is the part actually being marked.

Smart questions to ask them

☺ Like you're 10: At the end they'll say "any questions for us?" Having good ones is part of the interview — and it's also how you find out whether the job is any good.

Ask questions whose answers would actually change your view of the role. Migration programmes vary enormously in maturity, and a handful of specific questions will tell you within two minutes whether you are joining a disciplined factory or a rescue operation. Both can be good jobs — but you want to know which one you are signing up for.

About the programme

About the estate

About how the team actually works

About the role itself

⚠ Two to avoid

Questions whose answers are on the careers page, and questions that are really statements about how impressive you are. Also, do not ask all sixteen of these — pick three or four that you genuinely want answered, and follow up on what they say rather than moving down your list.

The last week before the interview

☺ Like you're 10: A short, honest plan for the seven days before. Not "read everything again" — that doesn't work and you know it.

This assumes you have already worked through the course. If you have not, the honest advice is different: read The 7 R's, Wave Planning, Data Migration and Anti-Patterns properly, and skim the rest — four domains understood well beats thirteen skimmed. The plan below is deliberately light on volume and heavy on saying things out loud, because the gap between "I know this" and "I can explain this under mild social pressure" is the gap that loses interviews.

WhenDo thisWhy
Day 7
(one week out)
Take the self-assessment in the topic map — rate all thirteen domains 1–3. Then run Self-Check and note every question you got wrong or guessed.You are building an evidence-based study list instead of re-reading what is already comfortable. Guessed-and-correct counts as wrong.
Day 6Read the Q&A sections for your two weakest domains here, and reread the lessons behind them. Answer each question out loud before reading the model answer.Depth where you are weakest has the highest marginal value. Reading silently feels like studying and produces almost no interview-usable fluency.
Day 5Same again for the next two weakest. Then run one pass of Flashcards for pure recall — the 7 R's, RTO/RPO, the DR rungs, the cutover patterns.The recall items are the ones you must produce instantly. Hesitating on "name the 7 R's" costs you credibility that the rest of the interview then has to rebuild.
Day 4Work two of the scenarios from a blank page, out loud, timed at ten minutes each — using the seven-step structure and forcing yourself to state assumptions.Scenario rounds are pattern-matched to structure. Practising the structure twice is worth more than reading all seven scenarios.
Day 3Write your four behavioural stories in full, then read them aloud and time them. Cut anything over two and a half minutes — and cut it from the Situation.Behavioural answers are where prepared candidates ramble most, because the material is emotionally close and hard to edit while speaking.
Day 2Two things. Rehearse your two-minute "tell me about a migration you worked on" story until it is boring to you. Then pick the three or four questions to ask them you actually want answered.The opening story sets the tone of the whole conversation, and it is the one answer you can fully control. Your questions are the last impression.
Day 1
(the day before)
One light pass over the Glossary and the Migration Checklist. Nothing new. Then stop, at a sensible hour.Cramming new material the night before displaces the material you already have and costs you sleep, which costs you more than the material was worth.
🐢 Timmy's rule for the day itself

Three things, and they are all about the first ninety seconds of each answer. Ask before you answer on anything scenario-shaped — one clarifying question is a signal of seniority, not of ignorance. Say what you don't know, then say how you would find out; "I haven't done that with that specific tool, but the shape of the problem is X and here's how I'd approach it" is a strong answer and a bluff is a fatal one, because the follow-up question always comes. And slow down — nearly everyone speaks too fast in interviews, and the difference between a rushed correct answer and a measured correct answer is much larger than it should be.

◆ Key idea

Migration interviews reward the same thing migrations do: calm, ordered thinking under a deadline, with an explicit way back. Almost every strong answer on this page has the same three-part shape — name the trade-off, pick a side, state the condition that would change your mind. If you internalise nothing else, internalise that shape and apply it to a question this page did not anticipate.

🎬 At the Migration Academy
🦊

Foxy: Right — interview tomorrow. I've memorised all seven R's, the four DR strategies, and every cutover pattern. I'm ready.

🦉

Professor Owl: Good. Now: a client has 180 apps and nine months. Which R?

🦊

Foxy: Er… refactor? It has the highest long-term benefit —

👺

Gizmo: YES. Refactor everything! Say "cloud-native" three times and they'll hire you on the spot! 🤑

🦫

Benny: Nine months, Foxy. I couldn't rebuild eight of those apps in nine months, let alone a hundred and eighty. It's a rehost programme with a modernization plan behind it — and a date on that plan.

🐿️

Nutty: And the first thing you'd actually say is "how many of the 180 still need to exist?" Half my job is deleting things nobody has opened since 2019.

🐢

Timmy: Fact-check: knowing the seven R's is the warm-up question. Knowing which one, why, and what would change your mind is the actual interview.

🦉

Professor Owl: Precisely. They are not testing whether you have read the menu, Foxy. They are testing whether you can order dinner for someone else, on a budget, with an allergy nobody mentioned.

A last word on honesty

☺ Like you're 10: The most useful thing on this whole page is knowing what to say when you don't know something.

Two things are worth saying plainly, because they matter more than any individual answer above.

Do not borrow someone else's experience. Every model answer here is written to teach the underlying idea, not to be recited as if it happened to you. Interviewers ask follow-up questions, and follow-ups on a borrowed story fail in the second or third layer — at which point the interview is over in a way that a modest, honest answer would never have been. If your genuine experience is a home lab, a university project, an internal tool you moved between two servers, or reading this course carefully, say that. "I have not run a migration at that scale, but here is how I would approach it and here is why" is a completely respectable answer, and it is one an interviewer can build on.

Prefer durable answers over specifics that rot. Product names change, services get renamed and retired, pricing moves, and certification tracks are restructured. If you are asked about a specific product and you are not certain of its current state, say what job it does and how you would confirm the current details — that is the answer of someone who has been burned by a renamed service, which is to say the answer of someone experienced. The conceptual layer of this course — the seven strategies, the four lifecycle stops, the wave discipline, the data-movement recipe, the two resilience numbers — is what stays true. Lean on it.

🐢 Timmy's checkpoint

Answer these out loud, from memory, before you close this page. 1. Name all seven R's, and say which two mean you don't move the app. 2. What two artefacts does discovery produce — and what third dataset do most people forget? 3. Describe the low-downtime cutover recipe in three phases, and say which clause makes rollback possible. 4. What do RTO and RPO measure, who sets them, and when? 5. What is the one clause that turns a good technical answer into a senior one?

Check your answers
  1. Retire, Retain, Rehost, Relocate, Repurchase, Replatform, Refactor. Retire (switch it off permanently) and Retain (deliberately leave it where it is for now) are the two where you don't move the application at all.
  2. An inventory (everything you have) and a dependency map (how it's all connected). The forgotten third is utilisation data — real CPU, memory and storage usage over time, which is what right-sizing and therefore the cloud cost model depend on.
  3. Bulk load while everything runs → replicate changes continuously (CDC) → at cutover, stop writes, drain the final sync, validate, then repoint traffic. The clause that makes rollback possible is stop writes rather than delete — the source stays intact and able to take traffic back.
  4. RTO is how long a system may be down; RPO is how much recent data may be lost. Both are set by the business, per application, before you choose a solution — and each wave's cutover is validated against them.
  5. The condition. Naming a trade-off and picking a side is a mid-level answer; adding "and here is what would change my mind, and here is how I'd find out which situation we're actually in" is the senior one.

From here: run the Exam Simulator if you want timed pressure, work the Move Brambleside capstone if you want something concrete to talk about in the interview, or go back to the domain lessons for anything above that made you hesitate. And if a question here has no good answer for your situation yet — that is the study list, not a verdict.