Exam Prep · SRE Foundation (SREF) · Mock Exam · Set 3

SREF Mock Exam · Set 3

Most SREF candidates don't fail on material they've never seen — they fail on material they've seen a dozen times and still mix up under a running clock. This is the third of five practice sittings, and unlike a balanced paper drawn proportionally across all eight syllabus modules, Set 3 spends every one of its 40 questions on exactly four confusable term-pairs: SLA vs. SLO vs. SLI, robust vs. antifragile, error budget vs. error-budget policy, and toil vs. ordinary operational work. These four traps produce more wrong single-best-answers than everything else in the syllabus combined, precisely because each pair sounds like a single idea until an exam stem quietly asks you to tell its two halves apart. Sit this paper closed-book, in one 60-minute block, and mark it against the 65% bar before you read a single explanation.

☺ Explain it like I'm 10

Some spelling tests trip you up with hard words. This test trips you up with easy words that sound almost identical — "dessert" and "desert," "affect" and "effect." You already know both words. The test isn't checking whether you've heard them before; it's checking whether you can tell them apart in the half-second before you circle an answer. Set 3 is that kind of test, except the lookalike words are SLA/SLO/SLI, robust/antifragile, error budget/error-budget policy, and toil/ordinary work. Forty questions, four pairs, no new vocabulary at all — just the same four traps, asked ten different ways each, until picking the right half becomes reflex instead of a coin flip.

🦥🐿️Your hosts for this topic: Sol the Sloth & Nutty the Squirrel — Sol won't let a percentage move until the arithmetic comes out exactly right, and Nutty has never once let two similar-sounding terms blur into each other. Between the two of them, that's the entire skill this paper is built to drill.

Why Set 3 is different from a balanced paper

☺ Like you're 10: Instead of one question about every topic, this paper asks the same four "which word is it, exactly" questions over and over, from every angle, until you stop hesitating.

A real SREF sitting draws its 40 questions across all eight modules of the blueprint, and the exam guide plus Set 1 and Set 2 are built to reflect that spread. Set 3 deliberately does not. Every question below sits inside one of four blocks, and every block is built around a single pair (or trio) of terms that look interchangeable in casual conversation and are worth the entire question if you mix them up on the real paper — there's no partial credit on a single-best-answer exam for "close." The four traps aren't arbitrary: they're the ones Module 2, Module 6, and Module 3 each flag as their single most commonly wrong answer, and this paper simply concentrates all forty of its questions on that exact list instead of sampling it once each.

That concentration is the point, not a shortcut. Over-sampling four traps ten times each is how a "yes, but which one exactly" reflex actually gets built — one exposure to a trap teaches you the answer to that one stem; ten exposures, phrased ten different ways, teach you to recognize the shape of the trap regardless of how the stem dresses it up. If you've already sat Set 1 or Set 2 and know your domain scores were fine everywhere except these four ideas, Set 3 is exactly the paper built for you.

Exam conditions — sit it like the real thing

☺ Like you're 10: One sitting, one clock, no notes, no looking anything up — the real test won't give you any of those either.

One 60-minute block, timer started once. The actual SRE Foundation exam is 40 questions in 60 minutes with a 65% pass mark — 26 correct answers — closed-book, single-best-answer, no live terminal and no partial credit for a well-reasoned wrong option. Sit Set 3 under those same conditions: no notes, no glossary tab, no AI assistant, and no going back to re-read the question blocks below once you've started the clock. If a question stumps you, mark your best guess and move on — the real exam typically lets you flag a question and revisit it before final submission, so practice that discipline here too, rather than freezing on any single stem.

Read the question twice before you read the options. Every trap on this paper is built to hide in the last six words of a stem — "unless the scenario explicitly describes a system that measurably improved," "measured from the SLO, never the SLA," "who is negotiating with the customer." Skimming the stem and jumping straight to pattern-matching an option is exactly how a candidate who knows the material perfectly still loses the point.

⚠ Verify current specifics before you register

The 40-question, 60-minute, 65%-pass-mark format above matches DevOps Institute's published SRE Foundation specifics at the time this page was written, and it's the same figure cited consistently across this course's exam-blueprint pages. Exam formats, pricing, and pass marks do change without much notice, though — confirm the current numbers on the DevOps Institute's own certification page before you book a real sitting. Nothing on this page is a substitute for that.

The four traps this paper concentrates on

☺ Like you're 10: Four pairs of words, one quick reminder of what separates each pair, before the ten questions on it start.

A thirty-second primer on each pair, in the order the paper uses them. If any of these four rows reads as unfamiliar rather than "oh right, that one," stop and read the linked lesson before you sit the timed paper — this page drills recognition speed, it doesn't teach the concept from zero.

PairThe one-line distinctionRead this first if it's shaky
SLA vs. SLO vs. SLISLI = the raw measurement. SLO = the internal target for it. SLA = the external, financially-backed promise, set looser than the SLO on purpose.SLIs, SLOs & error budgets · Module 2
Robust vs. antifragileRobust survives a shock and stays flat. Antifragile actually gains capability because of the shock. They look identical after one event and only diverge across many.Chaos engineering · Module 6
Error budget vs. error-budget policyThe budget is a number — (100% − SLO) × window. The policy is the pre-agreed written rule for what happens once that number hits zero.SLIs, SLOs & error budgets · Module 2
Toil vs. ordinary operational workToil clears all six gates — manual, repetitive, automatable, tactical, no enduring value, O(n) with growth. Miss any one gate and it's ordinary work, engineering work, or overhead instead.Toil & automation · Module 3
SLA vs. SLO vs. SLI contract vs. target vs. measurement Block A · Q1–Q10 Robust vs. antifragile stays flat under shock vs. gains from shock (only diverge over time) Block B · Q11–Q20 Error budget vs. error-budget policy a spendable number vs. the rule for zero Block C · Q21–Q30 Toil vs. ordinary work all six gates must pass vs. failing even one reclassifies it Block D · Q31–Q40

The paper — 40 questions, four blocks of ten

☺ Like you're 10: Read every question, pick one answer, write it down — no checking below until you've finished all forty.

Every question is single-best-answer: four options, exactly one correct, no partial credit. Write your letter for each of the 40 questions before you look anywhere near the answer key. Explanations for the whole paper live in one place, after the scoring section, organized by block — opening them mid-sitting defeats the entire point of a timed diagnostic.

Block A (Q1–Q10) — SLA vs. SLO vs. SLI

Q1. A video-streaming platform's engineers define "the percentage of playback-start requests that begin buffering within 300ms" and instrument it continuously against live traffic. What is this quantity itself, by name?

Q2. The same team then declares: "99.5% of playback-start requests must begin within 300ms, measured over a rolling 28-day window." What is this declaration called?

Q3. The company's contract with a premium-tier customer reads: "99% of playback starts within 300ms, or the customer receives a 10% account credit for the billing period." What kind of statement is this?

Q4. Which of the three terms — SLI, SLO, SLA — is the only one that typically carries direct financial consequences for the provider if it's missed?

Q5. A team sets an internal SLO of 99.9% availability and separately signs an SLA promising the identical figure — 99.9% — to its largest customer. What's the problem with this setup?

Q6. Who is generally responsible for negotiating and signing off on an SLA?

Q7. A postmortem notes: "the checkout service's CPU utilization averaged 72% during the incident." Taken alone, is this figure a usable SLI?

Q8. Put these in the order they must logically be established, first to last: (1) an SLA is negotiated with a customer, (2) an SLI is instrumented, (3) an SLO is set as an internal target.

Q9. A stem describes a service's reliability commitment only as: "the target is 99.95%," with no window, no ratio, and no stated consequence for missing it. What's the single biggest problem with treating this sentence as a complete SLO?

Q10. During an outage, an engineer says: "We're still within our SLA, so there's nothing urgent to do here." Why is that reasoning risky?

Block B (Q11–Q20) — Robust vs. Anti-Fragile

Q11. In Nassim Nicholas Taleb's three-way framework, what is the precise opposite of "fragile"?

Q12. A rock sits in a garden through years of hail, rain, and frost, coming out essentially unchanged — neither shattered nor improved. Which category does the rock illustrate?

Q13. Which response-curve shape, plotted against the size of a stressor, characterizes an antifragile system?

Q14. Rocky the Raccoon kills one production instance of the checkout service. Traffic reroutes cleanly, nobody notices, and no code changes afterward. Strictly speaking, what does this single result demonstrate?

Q15. A chaos program runs the same class of experiment repeatedly over a year. Every time a hypothesis fails, the team root-causes the gap, ships a fix, and re-verifies at a wider blast radius. Over time, the fleet's measured tolerance for that entire class of failure trends upward. What does that trend represent?

Q16. A service sits behind a load balancer with N+1 redundant replicas and automatic failover — losing any single replica leaves it exactly as capable as before. By itself, does this redundancy make the service antifragile?

Q17. Biologists have a specific term for the pattern where a mild dose of a stressor triggers an adaptive overcompensation, leaving an organism better equipped for a larger dose later. What's the term?

Q18. In which book did Nassim Nicholas Taleb formally name and popularize "antifragile"?

Q19. Systems A and B are both hit by the same single moderate failure, and both come through it apparently unharmed. Based on this one event alone, can you tell which one is robust and which is antifragile?

Q20. An engineer describes a newly hardened payment service as "really resilient now" after adding retries and a circuit breaker. In the SREF exam's stricter vocabulary, how should that claim most likely be read, unless the scenario says otherwise?

Block C (Q21–Q30) — Error Budget vs. Error-Budget Policy

Q21. A service has a 99.9% SLO measured over a 30-day window (43,200 minutes). How many minutes of downtime does its error budget allow for that window?

Q22. A different service carries a 99.99% SLO over the same 30-day, 43,200-minute window. What's its error budget?

Q23. A checkout API serves 6,000,000 valid requests over a 30-day window and carries a 99.9% success-rate SLO. How many failed requests does its error budget allow across that entire window?

Q24. What, precisely, is an error-budget policy?

Q25. A team's error budget for the current window is fully consumed, with 10 days still remaining before the window rolls forward. According to a properly implemented error-budget policy, what's the standard next step?

Q26. Why is "just lower the SLO" widely treated as an anti-pattern in response to a spent error budget?

Q27. A stem provides both an SLA figure (99.5%) and an SLO figure (99.9%) for the same service and window, then asks you to calculate the error budget. Which figure should feed the calculation?

Q28. What does a well-written error-budget policy's "exception path" typically exist for?

Q29. A service is currently measuring exactly 99.9% success against a 99.9% SLO, with the window not yet closed. Has this service's error budget been breached, and how much margin remains?

Q30. What does "burn rate" measure, and how does it differ from the error budget itself?

Block D (Q31–Q40) — Toil vs. Ordinary Operational Work

Q31. According to the formal definition, how many properties must a task satisfy, all at once, to count as toil?

Q32. A team encounters a data-corruption bug they've never seen before, spends two days investigating it, and ships a schema-validation fix that prevents it from recurring. Why does this fail to qualify as toil, even though it was operational and unplanned?

Q33. A senior engineer manually reads the diff and approves every production deploy — a task that happens on every single release. Why does this usually fail the toil test, despite being manual and highly repetitive?

Q34. An engineer spends 45 minutes a week updating an on-call handoff spreadsheet and attending a reliability status meeting. How should this work be classified?

Q35. An on-call engineer restarts a hung payment-worker pod three or four times a week; the underlying memory leak causing the hangs has never been root-caused. Does this qualify as toil?

Q36. What does the "O(n) with growth" property of toil actually describe?

Q37. What single prior question determines whether a recurring administrative task — a status meeting, a headcount-planning spreadsheet — should even be run through the six-gate toil test at all?

Q38. According to Google's SRE model, what does the 50% toil cap actually govern, and what's the standard organizational response when a team sustains work meaningfully above it?

Q39. A task is manual and repetitive, and an engineer calls it "extremely tedious." Is that combination, by itself, enough to classify the task as toil?

Q40. A platform engineer hand-runs the same seven CLI commands every time a new microservice team onboards, creating its namespace, IAM role, and CI pipeline — a fully scriptable sequence requiring no real judgment call, triggered reactively by each onboarding request, restoring nothing beyond baseline setup, and growing linearly with team count. What's the single most defensible classification?

Score yourself

☺ Like you're 10: Count how many letters you got right out of forty, turn that into a percentage, and check it against 65 — then look at which of the four blocks actually lost you the points.

Mark against the answer key below only after you've committed to all 40 answers — marking as you go turns a diagnostic into an open-book quiz and tells you nothing about your real readiness. Every question is worth exactly one point; there's no partial credit and no penalty for a wrong guess, so an unanswered question and a wrong one cost you identically. The pass mark mirrors the real SRE Foundation exactly: 26 out of 40, or 65%.

BlockPair testedQuestionsYour score (out of 10)
ASLA vs. SLO vs. SLIQ1–Q10
BRobust vs. antifragileQ11–Q20
CError budget vs. error-budget policyQ21–Q30
DToil vs. ordinary operational workQ31–Q40
TotalAll four pairs40

Because this paper is deliberately narrow, the diagnosis it produces is unusually sharp: a low block score can't be explained away as "that domain just wasn't my focus," because every block only tests one already-familiar pair. A block score under 7/10 means that specific distinction is still genuinely shaky, not unlucky — go straight to the linked lesson before re-sitting rather than re-reading this whole paper.

If Block A is weakIf Block B is weakIf Block C is weakIf Block D is weak
SLIs, SLOs & error budgets, then Module 2 and the SLOs & toil practice bank.Chaos engineering, then Module 6 and the resilience & culture practice bank.Module 2, the error-budget-policy section, plus Google & the error-budget policy for a real worked case.Toil & automation, then Module 3 and Drill — Audit the Toil.

Answer key & explanations

☺ Like you're 10: Every answer, block by block, with a short reason for the right one — read these only after you've already picked all forty.

Open one block at a time, and only after you've scored that block above. Each explanation names the deciding word or property, not just the letter — that's the part worth actually reading, since it's the reasoning that transfers to a stem you haven't seen before.

Block A answers (Q1–Q10) — SLA vs. SLO vs. SLI
  1. B. A continuously instrumented ratio of good events to valid events is the definition of an SLI — the raw measurement itself, before any target is set against it.
  2. A. A stated target value ("99.5%") over a defined window ("28 days") is an SLO — the internal goal set for an SLI, not the SLI itself.
  3. C. A contractual promise with a stated financial consequence (a 10% account credit) for missing it is the definitional signature of an SLA.
  4. A. Only the SLA is external and contractual with financial teeth; an SLO miss triggers the internal error-budget policy, not a customer payout.
  5. C. A well-run SLA is set looser than the internal SLO on purpose, preserving an early-warning window. Identical figures eliminate that margin — every SLO miss instantly becomes a contract breach.
  6. D. SLA negotiation is a business/legal function, conducted with the customer — not an engineering or on-call decision.
  7. B. CPU utilization is a resource-based, internal signal. A good SLI is measured at the boundary the user actually crosses (success, latency), not from internal component health.
  8. C. You instrument the SLI first (you need something to measure), then set the SLO as an internal target against it, then separately negotiate the looser SLA with a customer.
  9. A. A percentage with no stated window is incomplete — the same figure produces wildly different absolute allowances depending on whether it's measured daily or over a rolling 28 days.
  10. B. The SLA is deliberately the loosest of the three commitments. A team can be well within its SLA while already missing its (tighter) internal SLO — the SLO miss is the earlier, more useful signal, and "we're within SLA" ignores it entirely.
Block B answers (Q11–Q20) — Robust vs. Anti-Fragile
  1. C. Antifragile is Taleb's true opposite of fragile. Robust is the neutral middle of the spectrum — unaffected by a shock, neither harmed nor improved — not the far end from fragile.
  2. B. Unchanged capability after repeated stress, with no damage and no gain, is the textbook robust pattern.
  3. C. Antifragile responses are convex: benefit accelerates as the size of the stressor grows, up to a point. Fragile is the concave (accelerating-harm) mirror image.
  4. B. A single clean pass, with no code change afterward, only proves the system tolerated that one fault at that one scope on that one day — nothing about its future tolerance for anything else has changed.
  5. C. The closing loop — failed hypotheses get root-caused, fixed, and re-verified at a wider scope, with tolerance for a whole class of failure trending upward over many events — is precisely the operational signature of antifragility.
  6. B. Redundancy and failover make a system robust against a specific, already-known failure mode. They don't, by themselves, generate new capability against a class of failure nobody has tested yet — that requires the fix-and-reverify loop.
  7. D. Hormesis is the biological term for a mild stressor triggering adaptive overcompensation — the same shape as antifragility, named decades before Taleb coined the software/economics term.
  8. C. Antifragile: Things That Gain from Disorder (2012) is the book that names the concept, the third installment of Taleb's Incerto series after Fooled by Randomness and The Black Swan.
  9. D. Robust and antifragile are indistinguishable from the outside after any single event — both a rock and a muscle are still standing the moment after a stressor. The trajectories only split apart once you look across many repeated events.
  10. B. Unless a scenario explicitly describes a system whose future tolerance measurably increased because of the stress, "resilient" used casually should be read as the weaker, more common claim: robust, not antifragile.
Block C answers (Q21–Q30) — Error Budget vs. Error-Budget Policy
  1. C. (100% − 99.9%) × 43,200 min = 0.1% × 43,200 = 43.2 min. A is the 99.99% figure and B is the 99.95% figure — both plausible-looking distractors from adjacent SLO tiers.
  2. C. (100% − 99.99%) × 43,200 min = 0.01% × 43,200 = 4.32 min. Each additional "nine" divides the allowed budget by roughly 10x.
  3. B. 0.1% × 6,000,000 = 6,000 failed requests, for the entire 30-day window — not per day. A stem that quietly asks for a daily figure from a window-scoped budget (or vice versa) is testing whether you noticed the scope, not whether you can multiply.
  4. C. The policy is the written governance mechanism around the number — trigger, enforcement action, and exception path, agreed by both engineering and business stakeholders before anyone is under pressure. The budget itself is just the number the policy reacts to.
  5. D. Freezing risky launches and redirecting engineering capacity to reliability work is exactly what a pre-agreed error-budget policy exists to enforce once its trigger condition — a fully spent budget — is met.
  6. C. Lowering the SLO moves the target to match the failure instead of fixing the reliability problem, and defeats the entire purpose of having a target the team is actually held to.
  7. D. An error budget is always derived from the SLO — the internal target — never from the SLA, which is a separate, looser, externally negotiated commitment with its own breach condition.
  8. B. The exception path exists for genuinely urgent, low-risk changes that need a named approver's sign-off, not a blanket bypass any engineer can invoke on their own judgment.
  9. C. Sitting exactly at the SLO line means 100% of the allowed error rate has already been consumed — zero margin remains, even though the SLO hasn't technically been breached yet. "Haven't breached" and "have budget left" are not the same claim.
  10. D. Burn rate is a speed — how fast the budget is being consumed, as a multiple of the sustainable rate. The error budget itself is a fixed quantity for the window, not a rate at all.
Block D answers (Q31–Q40) — Toil vs. Ordinary Operational Work
  1. D. Six: manual, repetitive, automatable, tactical, devoid of enduring value, and O(n) with growth. All six must hold at once — the exam's favorite trick is a scenario that satisfies five convincingly and quietly fails the sixth.
  2. C. A novel task fails "repetitive" outright, and a fix that leaves the system durably better fails "no enduring value" — this is engineering project work, however stressful or operational it felt in the moment.
  3. D. Reading a diff for real risk requires exactly the kind of judgment a script can't safely replace, so the task fails "automatable" even though it's manual, repetitive, and happens on every release.
  4. C. Administrative work not tied to operating a live production service — meetings, handoff spreadsheets — is overhead. It never even reaches the six-gate toil test, because the prior question ("is this about running the service?") already sorts it into a third bucket.
  5. D. All six gates hold: hands-on, recurs weekly, a process supervisor could do the restart, purely reactive to a page, restores the pod to exactly where it was before, and restart volume scales with worker count.
  6. D. O(n) with growth means the work's volume scales linearly with fleet size, customer count, or traffic — the property that makes unaddressed toil a compounding problem rather than a fixed cost.
  7. D. "Is this about operating the production service, or not?" is the prior gate. Work that would exist regardless of whether the service is even in production — a status meeting, a training course — is overhead, full stop, with no need to run the six-gate test at all.
  8. C. The 50% cap governs the share of an SRE's time spent on toil, tracked per person per quarter. Sustained overage triggers a staffing/automation-investment response — assigning engineering time to fix the highest-volume sources — not simply "work harder."
  9. B. Manual and repetitive are only two of six required properties. The task still has to be automatable, tactical, devoid of enduring value, and O(n) with growth — failing any one of those remaining four disqualifies it regardless of how tedious it feels.
  10. D. Every gate holds: manual execution, recurs per onboarding, fully scriptable with no real judgment call, reactive to each onboarding event, restores nothing beyond baseline setup, and scales linearly with team count — a clean six-for-six toil classification.
🎬 At the Reliability Watch
🐰

Remy the Rabbit: 33 out of 40! I flew through that. Robust, antifragile, whatever — I just picked the word that sounded more impressive each time.

🐿️

Nutty the Squirrel: Let's see which seven you dropped. ...Four of them are Block B. You picked "antifragile" on three separate questions that were actually describing robust.

🐰

Remy the Rabbit: "Gains from disorder" just sounds like the better answer, though!

🦥

Sol the Sloth: ...That's exactly the trap. Robust and antifragile look identical after one event. The word that "sounds more impressive" is the one the exam is baiting you toward.

🦉

Professor Owl: Which is why this whole paper is one pair, asked ten times, instead of one question about everything. Speed isn't the goal on Block B — noticing "one event" versus "many, trending" is.

🐿️

Nutty the Squirrel: Re-sit just Block B tomorrow, cold, no re-reading the lesson first. If the same three questions trip you again, we'll know it's the concept and not the clock.

🐰

Remy the Rabbit: ...Fine. Slow, then fast. I hate that Sol's always right about this.

After the sitting

☺ Like you're 10: Fix only the block that actually lost you points, then come back and sit the whole paper again in a week to prove it stuck.

Resist re-reading all four lessons just because you sat a hard paper. The entire value of a narrow, concentrated diagnostic like this one is that it tells you exactly which single pair still needs work and which three are already solid — spend your next study session on the weak block's linked lesson and drill, not on a general review of everything. If you cleared 34/40 or better with no block under 7/10, the four traps are functionally reflexive for you now; move on to Set 2 or a full domain-balanced paper and stop spending more time here.

If a specific block stayed weak on a second attempt a few days later, treat that as real signal, not bad luck — go past the primer table above into the full lesson and the matching drill: Drill — SLO & Error-Budget Calculation for Blocks A and C, Drill — Design a Chaos Experiment for Block B, and Drill — Audit the Toil for Block D. Answer triage — SREF covers the general skill of spotting a trap distractor even on pairs this paper didn't cover, and Know it cold — SREF is the fast-recall version of the same four distinctions for a final pass the night before your real exam.

✓ Checkpoint

1. Why does Set 3 deliberately concentrate all 40 questions on four pairs instead of spreading them across all eight SREF modules? 2. In one sentence each, what's the fastest way to misdiagnose a robust result as antifragile, and what actually distinguishes the two? 3. A team's error budget hits zero with time left in the window — name the one popular "fix" that's actually an anti-pattern, and say why. 4. A task is manual, repetitive, and mildly unpleasant — why isn't that automatically enough to call it toil?

Check your answers
  1. Because ten exposures to the same trap, phrased ten different ways, build recognition of the trap's shape — not just the answer to one specific stem — which is what actually transfers to an unfamiliar question on the real exam.
  2. The fastest misdiagnosis is judging antifragility from a single clean pass with no follow-up fix. What actually distinguishes the two is a trend across many repeated stress events: robust stays flat, antifragile measurably improves, and the two are indistinguishable after just one event.
  3. Lowering the SLO to match the failure. It's an anti-pattern because it moves the target instead of fixing the reliability gap, and teaches the organization that the target is negotiable under pressure rather than a real commitment.
  4. Because toil requires four more properties beyond manual and repetitive — automatable, tactical, no enduring value, and O(n) with growth. A task that fails even one of those four (most often because it requires real judgment, or because it produces something lasting) isn't toil, however tedious it feels.

That's the paper. Close every tab except this one, set a 60-minute timer, and let the forty questions above tell you — precisely, block by block — which of the four traps still costs you points under pressure. Score it honestly, fix only what's actually weak, and come back to the full practice question bank or the platform-balanced Set 1 once these four pairs stop being a coin flip.

📝 The five practice papers

Set 1 · Set 2 · Set 3 (you are here) · Set 4 · Set 5. Set 3 is a concentrated trap drill, not a domain-balanced sitting — for a paper weighted across all eight blueprint modules the way the real exam is, sit Set 1 or Set 2 instead. See the exam guide for real exam-day logistics and the study plan for how all five papers fit into a full prep sequence.