SREF Mock Exam · Set 3
Most SREF candidates don't fail on material they've never seen — they fail on material they've seen a dozen times and still mix up under a running clock. This is the third of five practice sittings, and unlike a balanced paper drawn proportionally across all eight syllabus modules, Set 3 spends every one of its 40 questions on exactly four confusable term-pairs: SLA vs. SLO vs. SLI, robust vs. antifragile, error budget vs. error-budget policy, and toil vs. ordinary operational work. These four traps produce more wrong single-best-answers than everything else in the syllabus combined, precisely because each pair sounds like a single idea until an exam stem quietly asks you to tell its two halves apart. Sit this paper closed-book, in one 60-minute block, and mark it against the 65% bar before you read a single explanation.
Some spelling tests trip you up with hard words. This test trips you up with easy words that sound almost identical — "dessert" and "desert," "affect" and "effect." You already know both words. The test isn't checking whether you've heard them before; it's checking whether you can tell them apart in the half-second before you circle an answer. Set 3 is that kind of test, except the lookalike words are SLA/SLO/SLI, robust/antifragile, error budget/error-budget policy, and toil/ordinary work. Forty questions, four pairs, no new vocabulary at all — just the same four traps, asked ten different ways each, until picking the right half becomes reflex instead of a coin flip.
Why Set 3 is different from a balanced paper
☺ Like you're 10: Instead of one question about every topic, this paper asks the same four "which word is it, exactly" questions over and over, from every angle, until you stop hesitating.
A real SREF sitting draws its 40 questions across all eight modules of the blueprint, and the exam guide plus Set 1 and Set 2 are built to reflect that spread. Set 3 deliberately does not. Every question below sits inside one of four blocks, and every block is built around a single pair (or trio) of terms that look interchangeable in casual conversation and are worth the entire question if you mix them up on the real paper — there's no partial credit on a single-best-answer exam for "close." The four traps aren't arbitrary: they're the ones Module 2, Module 6, and Module 3 each flag as their single most commonly wrong answer, and this paper simply concentrates all forty of its questions on that exact list instead of sampling it once each.
That concentration is the point, not a shortcut. Over-sampling four traps ten times each is how a "yes, but which one exactly" reflex actually gets built — one exposure to a trap teaches you the answer to that one stem; ten exposures, phrased ten different ways, teach you to recognize the shape of the trap regardless of how the stem dresses it up. If you've already sat Set 1 or Set 2 and know your domain scores were fine everywhere except these four ideas, Set 3 is exactly the paper built for you.
Exam conditions — sit it like the real thing
☺ Like you're 10: One sitting, one clock, no notes, no looking anything up — the real test won't give you any of those either.
One 60-minute block, timer started once. The actual SRE Foundation exam is 40 questions in 60 minutes with a 65% pass mark — 26 correct answers — closed-book, single-best-answer, no live terminal and no partial credit for a well-reasoned wrong option. Sit Set 3 under those same conditions: no notes, no glossary tab, no AI assistant, and no going back to re-read the question blocks below once you've started the clock. If a question stumps you, mark your best guess and move on — the real exam typically lets you flag a question and revisit it before final submission, so practice that discipline here too, rather than freezing on any single stem.
Read the question twice before you read the options. Every trap on this paper is built to hide in the last six words of a stem — "unless the scenario explicitly describes a system that measurably improved," "measured from the SLO, never the SLA," "who is negotiating with the customer." Skimming the stem and jumping straight to pattern-matching an option is exactly how a candidate who knows the material perfectly still loses the point.
The 40-question, 60-minute, 65%-pass-mark format above matches DevOps Institute's published SRE Foundation specifics at the time this page was written, and it's the same figure cited consistently across this course's exam-blueprint pages. Exam formats, pricing, and pass marks do change without much notice, though — confirm the current numbers on the DevOps Institute's own certification page before you book a real sitting. Nothing on this page is a substitute for that.
The four traps this paper concentrates on
☺ Like you're 10: Four pairs of words, one quick reminder of what separates each pair, before the ten questions on it start.
A thirty-second primer on each pair, in the order the paper uses them. If any of these four rows reads as unfamiliar rather than "oh right, that one," stop and read the linked lesson before you sit the timed paper — this page drills recognition speed, it doesn't teach the concept from zero.
| Pair | The one-line distinction | Read this first if it's shaky |
|---|---|---|
| SLA vs. SLO vs. SLI | SLI = the raw measurement. SLO = the internal target for it. SLA = the external, financially-backed promise, set looser than the SLO on purpose. | SLIs, SLOs & error budgets · Module 2 |
| Robust vs. antifragile | Robust survives a shock and stays flat. Antifragile actually gains capability because of the shock. They look identical after one event and only diverge across many. | Chaos engineering · Module 6 |
| Error budget vs. error-budget policy | The budget is a number — (100% − SLO) × window. The policy is the pre-agreed written rule for what happens once that number hits zero. | SLIs, SLOs & error budgets · Module 2 |
| Toil vs. ordinary operational work | Toil clears all six gates — manual, repetitive, automatable, tactical, no enduring value, O(n) with growth. Miss any one gate and it's ordinary work, engineering work, or overhead instead. | Toil & automation · Module 3 |
The paper — 40 questions, four blocks of ten
☺ Like you're 10: Read every question, pick one answer, write it down — no checking below until you've finished all forty.
Every question is single-best-answer: four options, exactly one correct, no partial credit. Write your letter for each of the 40 questions before you look anywhere near the answer key. Explanations for the whole paper live in one place, after the scoring section, organized by block — opening them mid-sitting defeats the entire point of a timed diagnostic.
Block A (Q1–Q10) — SLA vs. SLO vs. SLI
Q1. A video-streaming platform's engineers define "the percentage of playback-start requests that begin buffering within 300ms" and instrument it continuously against live traffic. What is this quantity itself, by name?
- A. An SLA
- B. An SLI
- C. An SLO
- D. An error budget
Q2. The same team then declares: "99.5% of playback-start requests must begin within 300ms, measured over a rolling 28-day window." What is this declaration called?
- A. An SLO
- B. An SLI
- C. An SLA
- D. A burn rate
Q3. The company's contract with a premium-tier customer reads: "99% of playback starts within 300ms, or the customer receives a 10% account credit for the billing period." What kind of statement is this?
- A. An SLO
- B. An SLI
- C. An SLA
- D. An error budget
Q4. Which of the three terms — SLI, SLO, SLA — is the only one that typically carries direct financial consequences for the provider if it's missed?
- A. The SLA
- B. The SLI
- C. The SLO
- D. All three carry equal financial consequences
Q5. A team sets an internal SLO of 99.9% availability and separately signs an SLA promising the identical figure — 99.9% — to its largest customer. What's the problem with this setup?
- A. There's no problem; matching the two figures shows strong confidence in the service
- B. The SLA should be tighter than the SLO to better protect the customer
- C. The SLA should be looser than the SLO, or the team loses its early-warning margin — every SLO miss becomes a contract breach simultaneously
- D. The SLI needs to be renegotiated, not the SLA
Q6. Who is generally responsible for negotiating and signing off on an SLA?
- A. The on-call SRE handling the affected service that week
- B. Whichever engineer happens to be on call when the contract is due
- C. The engineer who built the service's monitoring dashboard
- D. Business or legal stakeholders, negotiating directly with the customer
Q7. A postmortem notes: "the checkout service's CPU utilization averaged 72% during the incident." Taken alone, is this figure a usable SLI?
- A. Yes — any percentage tracked over time automatically qualifies as an SLI
- B. No — CPU utilization is a resource-based signal, not a measurement of user-facing success or latency
- C. Yes, but only once it's measured over a rolling window
- D. No, but only because it was measured during an incident rather than steady state
Q8. Put these in the order they must logically be established, first to last: (1) an SLA is negotiated with a customer, (2) an SLI is instrumented, (3) an SLO is set as an internal target.
- A. 1, 2, 3
- B. 3, 2, 1
- C. 2, 3, 1
- D. 2, 1, 3
Q9. A stem describes a service's reliability commitment only as: "the target is 99.95%," with no window, no ratio, and no stated consequence for missing it. What's the single biggest problem with treating this sentence as a complete SLO?
- A. It's missing the measurement window — the same percentage means something very different measured over one day versus a rolling 28 days
- B. 99.95% is too high a target for any real service to hold
- C. It should have been phrased as an SLA instead of an SLO
- D. There's no problem — the percentage alone fully defines the SLO
Q10. During an outage, an engineer says: "We're still within our SLA, so there's nothing urgent to do here." Why is that reasoning risky?
- A. SLAs are legally unenforceable, so the statement is meaningless either way
- B. The team may already be missing its (tighter) internal SLO — the real early-warning signal that reliability work is overdue — well before any SLA breach occurs
- C. SLAs can't be evaluated while an incident is still active
- D. It isn't risky at all; meeting the SLA is the only threshold that actually matters operationally
Block B (Q11–Q20) — Robust vs. Anti-Fragile
Q11. In Nassim Nicholas Taleb's three-way framework, what is the precise opposite of "fragile"?
- A. Resilient
- B. Redundant
- C. Antifragile
- D. Robust
Q12. A rock sits in a garden through years of hail, rain, and frost, coming out essentially unchanged — neither shattered nor improved. Which category does the rock illustrate?
- A. Antifragile
- B. Robust
- C. Fragile
- D. None of these categories apply to inanimate objects
Q13. Which response-curve shape, plotted against the size of a stressor, characterizes an antifragile system?
- A. A perfectly flat line across the whole range
- B. Concave — damage accelerates as the stressor grows
- C. Convex — benefit accelerates as the stressor grows, up to a point
- D. A straight downward line at constant slope
Q14. Rocky the Raccoon kills one production instance of the checkout service. Traffic reroutes cleanly, nobody notices, and no code changes afterward. Strictly speaking, what does this single result demonstrate?
- A. Nothing meaningful can be concluded from a single chaos experiment
- B. The checkout service is robust to that specific fault, at that specific scope, on that specific day — nothing about its future tolerance has changed
- C. The checkout service is now provably antifragile
- D. The checkout service is fragile, because it required a fallback path at all
Q15. A chaos program runs the same class of experiment repeatedly over a year. Every time a hypothesis fails, the team root-causes the gap, ships a fix, and re-verifies at a wider blast radius. Over time, the fleet's measured tolerance for that entire class of failure trends upward. What does that trend represent?
- A. Toil, since the team keeps repeating the same experiment
- B. Nothing new — a single passing test already proved this months ago
- C. The operational signature of antifragility: capability improving as a direct result of repeated stress
- D. Robustness, since the system keeps recovering each time
Q16. A service sits behind a load balancer with N+1 redundant replicas and automatic failover — losing any single replica leaves it exactly as capable as before. By itself, does this redundancy make the service antifragile?
- A. No — redundancy actually makes systems more fragile over time
- B. No — it makes the service robust against that specific known failure mode; it doesn't, by itself, generate new capability against untested failure classes
- C. Yes, but only once the replica count exceeds N+2
- D. Yes — redundancy is the textbook definition of antifragility
Q17. Biologists have a specific term for the pattern where a mild dose of a stressor triggers an adaptive overcompensation, leaving an organism better equipped for a larger dose later. What's the term?
- A. Homeostasis
- B. Entropy
- C. Attrition
- D. Hormesis
Q18. In which book did Nassim Nicholas Taleb formally name and popularize "antifragile"?
- A. Thinking, Fast and Slow
- B. The Black Swan
- C. Antifragile: Things That Gain from Disorder
- D. Fooled by Randomness
Q19. Systems A and B are both hit by the same single moderate failure, and both come through it apparently unharmed. Based on this one event alone, can you tell which one is robust and which is antifragile?
- A. Yes — antifragile systems always recover measurably faster than robust ones
- B. Yes — only an antifragile system needs a failover mechanism at all
- C. No — the robust/antifragile distinction is purely theoretical and can never be observed
- D. No — robust and antifragile systems look identical after a single event; the difference only shows up across many repeated stress events over time
Q20. An engineer describes a newly hardened payment service as "really resilient now" after adding retries and a circuit breaker. In the SREF exam's stricter vocabulary, how should that claim most likely be read, unless the scenario says otherwise?
- A. As a claim of fragility, since retries and circuit breakers add complexity
- B. As a claim of robustness — the service now survives a known failure mode, which isn't the same as proving it gains from stress
- C. As meaningless, since "resilient" has no defined technical meaning at all
- D. As a claim of antifragility, since "resilient" implies gaining from stress
Block C (Q21–Q30) — Error Budget vs. Error-Budget Policy
Q21. A service has a 99.9% SLO measured over a 30-day window (43,200 minutes). How many minutes of downtime does its error budget allow for that window?
- A. 4.32 minutes
- B. 21.6 minutes
- C. 43.2 minutes
- D. 432 minutes
Q22. A different service carries a 99.99% SLO over the same 30-day, 43,200-minute window. What's its error budget?
- A. 0.432 minutes
- B. 21.6 minutes
- C. 4.32 minutes
- D. 43.2 minutes
Q23. A checkout API serves 6,000,000 valid requests over a 30-day window and carries a 99.9% success-rate SLO. How many failed requests does its error budget allow across that entire window?
- A. 60,000
- B. 6,000
- C. 600
- D. 600,000
Q24. What, precisely, is an error-budget policy?
- A. A dashboard that displays the current burn rate
- B. A tool that automatically recalculates the SLO whenever the budget runs low
- C. A pre-agreed, written governance mechanism, signed off by engineering and business stakeholders in advance, specifying the trigger, the enforcement action, and an exception path once the budget is spent
- D. A synonym for the error budget itself — the two terms describe the same thing
Q25. A team's error budget for the current window is fully consumed, with 10 days still remaining before the window rolls forward. According to a properly implemented error-budget policy, what's the standard next step?
- A. Continue shipping normally, since the window will reset on its own regardless
- B. Immediately reassign the engineers who were on call during the incidents that spent the budget
- C. Lower the SLO so the team is back within its now-smaller allowed error rate
- D. Freeze new feature releases and redirect engineering capacity toward reliability work until the budget recovers or the policy's exception process is invoked
Q26. Why is "just lower the SLO" widely treated as an anti-pattern in response to a spent error budget?
- A. It always requires renegotiating the SLA at the exact same time
- B. It has no real downside at all — it's actually the preferred first response
- C. It moves the target to match the failure instead of fixing the underlying reliability problem, and quietly teaches the organization that targets are negotiable under pressure
- D. Lowering a published SLO is technically impossible once it's set
Q27. A stem provides both an SLA figure (99.5%) and an SLO figure (99.9%) for the same service and window, then asks you to calculate the error budget. Which figure should feed the calculation?
- A. The average of the two figures
- B. The SLA figure, because it's the externally binding one
- C. Whichever figure is smaller
- D. The SLO figure — an error budget is always derived from the internal target, never the external contractual promise
Q28. What does a well-written error-budget policy's "exception path" typically exist for?
- A. Automatically restoring full release velocity the instant the incident that spent the budget is resolved
- B. Genuinely urgent, low-risk changes — such as a security patch — that require sign-off from a named approver rather than an ad hoc argument with whoever's on call
- C. Letting any engineer bypass the freeze whenever they personally judge a change to be safe
- D. Extending the measurement window until the budget refills on its own
Q29. A service is currently measuring exactly 99.9% success against a 99.9% SLO, with the window not yet closed. Has this service's error budget been breached, and how much margin remains?
- A. It hasn't been breached, and roughly half the budget remains
- B. It hasn't been breached, and the full budget remains untouched
- C. It has been breached, and zero margin remains — sitting exactly on the SLO line means the full budget for the window has already been spent
- D. The question can't be answered without also knowing the SLA figure
Q30. What does "burn rate" measure, and how does it differ from the error budget itself?
- A. Burn rate is set by the SLA figure, not the SLO
- B. Burn rate and error budget are the same measurement, just expressed in different units
- C. Burn rate only applies to request-based SLOs, never time-based ones
- D. Burn rate is the speed at which the budget is being consumed, expressed as a multiple of the sustainable rate — the error budget is the total spendable quantity, not the speed of spending it
Block D (Q31–Q40) — Toil vs. Ordinary Operational Work
Q31. According to the formal definition, how many properties must a task satisfy, all at once, to count as toil?
- A. Three
- B. Four
- C. Five
- D. Six
Q32. A team encounters a data-corruption bug they've never seen before, spends two days investigating it, and ships a schema-validation fix that prevents it from recurring. Why does this fail to qualify as toil, even though it was operational and unplanned?
- A. It actually does qualify as toil, since it was operational work performed under pressure
- B. It fails "tactical," because bugs are always planned in advance
- C. It fails "repetitive" (never seen before) and "no enduring value" (the fix leaves the system better than it started) — this is engineering project work
- D. It fails "manual," because a machine performed the investigation
Q33. A senior engineer manually reads the diff and approves every production deploy — a task that happens on every single release. Why does this usually fail the toil test, despite being manual and highly repetitive?
- A. It fails "O(n) with growth," since deploy frequency never actually changes
- B. It fails "repetitive," since every diff is technically different content
- C. It fails "tactical," since deploys are scheduled, planned events
- D. It fails "automatable" — reading a diff for real risk is exactly the kind of judgment call a script can't safely replace
Q34. An engineer spends 45 minutes a week updating an on-call handoff spreadsheet and attending a reliability status meeting. How should this work be classified?
- A. Engineering project work, because it supports the on-call process
- B. Toil, because it is manual and recurring
- C. Overhead — administrative work that isn't operating the production service, so it never even reaches the six-gate toil test
- D. Toil, but only once the meeting runs longer than an hour
Q35. An on-call engineer restarts a hung payment-worker pod three or four times a week; the underlying memory leak causing the hangs has never been root-caused. Does this qualify as toil?
- A. No — because the underlying cause is still unknown, the task counts as ongoing investigation, not toil
- B. Yes, but only once it happens more than five times in a single week
- C. No — restarting a pod is too trivial an action to count as "real" operational work
- D. Yes — it's manual, recurs weekly, is fully automatable by a process supervisor, is purely reactive to a page, restores the system to exactly where it started, and its volume scales with worker count
Q36. What does the "O(n) with growth" property of toil actually describe?
- A. The task must be completed within n hours of being triggered
- B. The task requires exactly one engineer to perform it, never more, regardless of scale
- C. The task takes exactly n minutes to complete, no matter how large the system is
- D. The volume of the work scales linearly with fleet size, customer count, or traffic — twice the fleet means roughly twice the manual work
Q37. What single prior question determines whether a recurring administrative task — a status meeting, a headcount-planning spreadsheet — should even be run through the six-gate toil test at all?
- A. Is the task documented in a runbook anywhere?
- B. Does a senior engineer or a junior engineer usually perform it?
- C. Does the task take longer than 30 minutes to complete?
- D. Is this work about operating the production service, or not? If not, it's overhead, full stop — no gate-checking required
Q38. According to Google's SRE model, what does the 50% toil cap actually govern, and what's the standard organizational response when a team sustains work meaningfully above it?
- A. It caps the total headcount of the SRE team; the response is an immediate hiring freeze
- B. It caps the number of pages one individual can receive per week; the response is to mute non-critical alerts
- C. It caps the share of an SRE's time spent on toil, averaged over a period such as a quarter; the standard response is to treat sustained overage as a staffing/automation-investment problem — assigning engineering time to automate the highest-volume sources, not simply asking the team to work harder
- D. It caps the number of services one team can support; the response is to shut down the lowest-priority one
Q39. A task is manual and repetitive, and an engineer calls it "extremely tedious." Is that combination, by itself, enough to classify the task as toil?
- A. Yes, as long as the engineer performing it genuinely finds it unpleasant
- B. No — the task must also be automatable, tactical, devoid of enduring value, and scale with growth (O(n)); failing even one of those remaining four gates disqualifies it, no matter how tedious it feels
- C. No — tedium is actually a required seventh property, on top of the standard six
- D. Yes — "manual and repetitive" is the complete, sufficient definition of toil on its own
Q40. A platform engineer hand-runs the same seven CLI commands every time a new microservice team onboards, creating its namespace, IAM role, and CI pipeline — a fully scriptable sequence requiring no real judgment call, triggered reactively by each onboarding request, restoring nothing beyond baseline setup, and growing linearly with team count. What's the single most defensible classification?
- A. Engineering project work, because setting up a CI pipeline is inherently valuable
- B. Overhead — onboarding is an administrative process, not operational work on a live service
- C. Not toil — it fails "automatable," since provisioning an IAM role always requires human judgment
- D. Toil — it clears all six gates: manual, repetitive, automatable (no real judgment required), tactical (reactive to each onboarding), no enduring value (baseline setup only), and O(n) with team count
Score yourself
☺ Like you're 10: Count how many letters you got right out of forty, turn that into a percentage, and check it against 65 — then look at which of the four blocks actually lost you the points.
Mark against the answer key below only after you've committed to all 40 answers — marking as you go turns a diagnostic into an open-book quiz and tells you nothing about your real readiness. Every question is worth exactly one point; there's no partial credit and no penalty for a wrong guess, so an unanswered question and a wrong one cost you identically. The pass mark mirrors the real SRE Foundation exactly: 26 out of 40, or 65%.
| Block | Pair tested | Questions | Your score (out of 10) |
|---|---|---|---|
| A | SLA vs. SLO vs. SLI | Q1–Q10 | |
| B | Robust vs. antifragile | Q11–Q20 | |
| C | Error budget vs. error-budget policy | Q21–Q30 | |
| D | Toil vs. ordinary operational work | Q31–Q40 | |
| Total | All four pairs | 40 |
Because this paper is deliberately narrow, the diagnosis it produces is unusually sharp: a low block score can't be explained away as "that domain just wasn't my focus," because every block only tests one already-familiar pair. A block score under 7/10 means that specific distinction is still genuinely shaky, not unlucky — go straight to the linked lesson before re-sitting rather than re-reading this whole paper.
| If Block A is weak | If Block B is weak | If Block C is weak | If Block D is weak |
|---|---|---|---|
| SLIs, SLOs & error budgets, then Module 2 and the SLOs & toil practice bank. | Chaos engineering, then Module 6 and the resilience & culture practice bank. | Module 2, the error-budget-policy section, plus Google & the error-budget policy for a real worked case. | Toil & automation, then Module 3 and Drill — Audit the Toil. |
Answer key & explanations
☺ Like you're 10: Every answer, block by block, with a short reason for the right one — read these only after you've already picked all forty.
Open one block at a time, and only after you've scored that block above. Each explanation names the deciding word or property, not just the letter — that's the part worth actually reading, since it's the reasoning that transfers to a stem you haven't seen before.
Block A answers (Q1–Q10) — SLA vs. SLO vs. SLI
- B. A continuously instrumented ratio of good events to valid events is the definition of an SLI — the raw measurement itself, before any target is set against it.
- A. A stated target value ("99.5%") over a defined window ("28 days") is an SLO — the internal goal set for an SLI, not the SLI itself.
- C. A contractual promise with a stated financial consequence (a 10% account credit) for missing it is the definitional signature of an SLA.
- A. Only the SLA is external and contractual with financial teeth; an SLO miss triggers the internal error-budget policy, not a customer payout.
- C. A well-run SLA is set looser than the internal SLO on purpose, preserving an early-warning window. Identical figures eliminate that margin — every SLO miss instantly becomes a contract breach.
- D. SLA negotiation is a business/legal function, conducted with the customer — not an engineering or on-call decision.
- B. CPU utilization is a resource-based, internal signal. A good SLI is measured at the boundary the user actually crosses (success, latency), not from internal component health.
- C. You instrument the SLI first (you need something to measure), then set the SLO as an internal target against it, then separately negotiate the looser SLA with a customer.
- A. A percentage with no stated window is incomplete — the same figure produces wildly different absolute allowances depending on whether it's measured daily or over a rolling 28 days.
- B. The SLA is deliberately the loosest of the three commitments. A team can be well within its SLA while already missing its (tighter) internal SLO — the SLO miss is the earlier, more useful signal, and "we're within SLA" ignores it entirely.
Block B answers (Q11–Q20) — Robust vs. Anti-Fragile
- C. Antifragile is Taleb's true opposite of fragile. Robust is the neutral middle of the spectrum — unaffected by a shock, neither harmed nor improved — not the far end from fragile.
- B. Unchanged capability after repeated stress, with no damage and no gain, is the textbook robust pattern.
- C. Antifragile responses are convex: benefit accelerates as the size of the stressor grows, up to a point. Fragile is the concave (accelerating-harm) mirror image.
- B. A single clean pass, with no code change afterward, only proves the system tolerated that one fault at that one scope on that one day — nothing about its future tolerance for anything else has changed.
- C. The closing loop — failed hypotheses get root-caused, fixed, and re-verified at a wider scope, with tolerance for a whole class of failure trending upward over many events — is precisely the operational signature of antifragility.
- B. Redundancy and failover make a system robust against a specific, already-known failure mode. They don't, by themselves, generate new capability against a class of failure nobody has tested yet — that requires the fix-and-reverify loop.
- D. Hormesis is the biological term for a mild stressor triggering adaptive overcompensation — the same shape as antifragility, named decades before Taleb coined the software/economics term.
- C. Antifragile: Things That Gain from Disorder (2012) is the book that names the concept, the third installment of Taleb's Incerto series after Fooled by Randomness and The Black Swan.
- D. Robust and antifragile are indistinguishable from the outside after any single event — both a rock and a muscle are still standing the moment after a stressor. The trajectories only split apart once you look across many repeated events.
- B. Unless a scenario explicitly describes a system whose future tolerance measurably increased because of the stress, "resilient" used casually should be read as the weaker, more common claim: robust, not antifragile.
Block C answers (Q21–Q30) — Error Budget vs. Error-Budget Policy
- C.
(100% − 99.9%) × 43,200 min = 0.1% × 43,200 = 43.2 min.A is the 99.99% figure and B is the 99.95% figure — both plausible-looking distractors from adjacent SLO tiers. - C.
(100% − 99.99%) × 43,200 min = 0.01% × 43,200 = 4.32 min.Each additional "nine" divides the allowed budget by roughly 10x. - B.
0.1% × 6,000,000 = 6,000failed requests, for the entire 30-day window — not per day. A stem that quietly asks for a daily figure from a window-scoped budget (or vice versa) is testing whether you noticed the scope, not whether you can multiply. - C. The policy is the written governance mechanism around the number — trigger, enforcement action, and exception path, agreed by both engineering and business stakeholders before anyone is under pressure. The budget itself is just the number the policy reacts to.
- D. Freezing risky launches and redirecting engineering capacity to reliability work is exactly what a pre-agreed error-budget policy exists to enforce once its trigger condition — a fully spent budget — is met.
- C. Lowering the SLO moves the target to match the failure instead of fixing the reliability problem, and defeats the entire purpose of having a target the team is actually held to.
- D. An error budget is always derived from the SLO — the internal target — never from the SLA, which is a separate, looser, externally negotiated commitment with its own breach condition.
- B. The exception path exists for genuinely urgent, low-risk changes that need a named approver's sign-off, not a blanket bypass any engineer can invoke on their own judgment.
- C. Sitting exactly at the SLO line means 100% of the allowed error rate has already been consumed — zero margin remains, even though the SLO hasn't technically been breached yet. "Haven't breached" and "have budget left" are not the same claim.
- D. Burn rate is a speed — how fast the budget is being consumed, as a multiple of the sustainable rate. The error budget itself is a fixed quantity for the window, not a rate at all.
Block D answers (Q31–Q40) — Toil vs. Ordinary Operational Work
- D. Six: manual, repetitive, automatable, tactical, devoid of enduring value, and O(n) with growth. All six must hold at once — the exam's favorite trick is a scenario that satisfies five convincingly and quietly fails the sixth.
- C. A novel task fails "repetitive" outright, and a fix that leaves the system durably better fails "no enduring value" — this is engineering project work, however stressful or operational it felt in the moment.
- D. Reading a diff for real risk requires exactly the kind of judgment a script can't safely replace, so the task fails "automatable" even though it's manual, repetitive, and happens on every release.
- C. Administrative work not tied to operating a live production service — meetings, handoff spreadsheets — is overhead. It never even reaches the six-gate toil test, because the prior question ("is this about running the service?") already sorts it into a third bucket.
- D. All six gates hold: hands-on, recurs weekly, a process supervisor could do the restart, purely reactive to a page, restores the pod to exactly where it was before, and restart volume scales with worker count.
- D. O(n) with growth means the work's volume scales linearly with fleet size, customer count, or traffic — the property that makes unaddressed toil a compounding problem rather than a fixed cost.
- D. "Is this about operating the production service, or not?" is the prior gate. Work that would exist regardless of whether the service is even in production — a status meeting, a training course — is overhead, full stop, with no need to run the six-gate test at all.
- C. The 50% cap governs the share of an SRE's time spent on toil, tracked per person per quarter. Sustained overage triggers a staffing/automation-investment response — assigning engineering time to fix the highest-volume sources — not simply "work harder."
- B. Manual and repetitive are only two of six required properties. The task still has to be automatable, tactical, devoid of enduring value, and O(n) with growth — failing any one of those remaining four disqualifies it regardless of how tedious it feels.
- D. Every gate holds: manual execution, recurs per onboarding, fully scriptable with no real judgment call, reactive to each onboarding event, restores nothing beyond baseline setup, and scales linearly with team count — a clean six-for-six toil classification.
Remy the Rabbit: 33 out of 40! I flew through that. Robust, antifragile, whatever — I just picked the word that sounded more impressive each time.
Nutty the Squirrel: Let's see which seven you dropped. ...Four of them are Block B. You picked "antifragile" on three separate questions that were actually describing robust.
Remy the Rabbit: "Gains from disorder" just sounds like the better answer, though!
Sol the Sloth: ...That's exactly the trap. Robust and antifragile look identical after one event. The word that "sounds more impressive" is the one the exam is baiting you toward.
Professor Owl: Which is why this whole paper is one pair, asked ten times, instead of one question about everything. Speed isn't the goal on Block B — noticing "one event" versus "many, trending" is.
Nutty the Squirrel: Re-sit just Block B tomorrow, cold, no re-reading the lesson first. If the same three questions trip you again, we'll know it's the concept and not the clock.
Remy the Rabbit: ...Fine. Slow, then fast. I hate that Sol's always right about this.
After the sitting
☺ Like you're 10: Fix only the block that actually lost you points, then come back and sit the whole paper again in a week to prove it stuck.
Resist re-reading all four lessons just because you sat a hard paper. The entire value of a narrow, concentrated diagnostic like this one is that it tells you exactly which single pair still needs work and which three are already solid — spend your next study session on the weak block's linked lesson and drill, not on a general review of everything. If you cleared 34/40 or better with no block under 7/10, the four traps are functionally reflexive for you now; move on to Set 2 or a full domain-balanced paper and stop spending more time here.
If a specific block stayed weak on a second attempt a few days later, treat that as real signal, not bad luck — go past the primer table above into the full lesson and the matching drill: Drill — SLO & Error-Budget Calculation for Blocks A and C, Drill — Design a Chaos Experiment for Block B, and Drill — Audit the Toil for Block D. Answer triage — SREF covers the general skill of spotting a trap distractor even on pairs this paper didn't cover, and Know it cold — SREF is the fast-recall version of the same four distinctions for a final pass the night before your real exam.
1. Why does Set 3 deliberately concentrate all 40 questions on four pairs instead of spreading them across all eight SREF modules? 2. In one sentence each, what's the fastest way to misdiagnose a robust result as antifragile, and what actually distinguishes the two? 3. A team's error budget hits zero with time left in the window — name the one popular "fix" that's actually an anti-pattern, and say why. 4. A task is manual, repetitive, and mildly unpleasant — why isn't that automatically enough to call it toil?
Check your answers
- Because ten exposures to the same trap, phrased ten different ways, build recognition of the trap's shape — not just the answer to one specific stem — which is what actually transfers to an unfamiliar question on the real exam.
- The fastest misdiagnosis is judging antifragility from a single clean pass with no follow-up fix. What actually distinguishes the two is a trend across many repeated stress events: robust stays flat, antifragile measurably improves, and the two are indistinguishable after just one event.
- Lowering the SLO to match the failure. It's an anti-pattern because it moves the target instead of fixing the reliability gap, and teaches the organization that the target is negotiable under pressure rather than a real commitment.
- Because toil requires four more properties beyond manual and repetitive — automatable, tactical, no enduring value, and O(n) with growth. A task that fails even one of those four (most often because it requires real judgment, or because it produces something lasting) isn't toil, however tedious it feels.
That's the paper. Close every tab except this one, set a 60-minute timer, and let the forty questions above tell you — precisely, block by block — which of the four traps still costs you points under pressure. Score it honestly, fix only what's actually weak, and come back to the full practice question bank or the platform-balanced Set 1 once these four pairs stop being a coin flip.