Exam Prep · SREF · Practice Bank · Modules 1–3

Practice · SRE Principles, SLOs & Toil

Twenty-four single-best-answer questions covering the three modules the rest of the SREF blueprint quietly leans on: SRE Principles & Practices (Module 1), Service Level Objectives & Error Budgets (Module 2), and Reducing Toil (Module 3). This is a practice bank, not a mock — no 60-minute clock, no interleaving with the other five modules — so it drills these three until the vocabulary and the arithmetic stop needing deliberation: where SRE actually came from and how it differs precisely from DevOps, the SLI/SLO/SLA chain plus the formula that turns an SLO into a spendable number of minutes, and the six-property test that decides whether a recurring task legally counts as toil. Answer cold, in the exam's own four-option, single-best-answer format — every question explains not just which option is correct but which specific trap each wrong option was built from. This is the first of three module-grouped banks; Monitoring & SRE Tools and Anti-Fragility & Organizational Impact cover the remaining five modules.

☺ Explain it like I'm 10

Picture three practice rounds before the real spelling bee. Round one is about where the whole competition came from and exactly what rules it follows — no vocabulary words yet, just the shape of the game and how it's different from a similar-looking game next door. Round two hands you a target number and asks you to do a small bit of subtraction and multiplication that tells you precisely how many mistakes you're allowed before a buzzer sounds. Round three is a checklist game: someone describes a chore, and you run it through six yes-or-no questions to decide whether it counts as the annoying, automatable kind of work — or something else wearing the same disguise. Get good at all three rounds on their own, and the real spelling bee — which mixes all eight rounds together with a sixty-minute clock running — stops feeling like an ambush.

🦉🦥🦫Your hosts for this bank: Professor Owl, Sol the Sloth & Benny the Beaver — Owl frames what SRE actually is and exactly how it differs from DevOps before a single number appears, Sol works the error-budget arithmetic slowly and gets it precisely right instead of fast and wrong, and Benny runs every toil-shaped scenario through his own six-gate checklist before letting it keep the label.

What this bank drills, and how it's built

☺ Like you're 10: Same four-option, one-right-answer format as the real test — just three topics at a time instead of all eight mixed together.

The SRE Foundation (SREF) is a closed-book, 40-question, 60-minute, single-best-answer multiple-choice exam with a 65% pass mark, drawn from eight blueprint modules DevOps Institute doesn't publish a per-module weighting for. Every question below is written in that exact register — four options, exactly one correct, no partial credit for a well-reasoned wrong answer — because the format itself is part of what you're rehearsing, not just the content. Click an option to lock it in: it turns green or red immediately, and the explanation underneath names the specific reasoning, including which trap each wrong option was built from. The chips below the meta bar filter by module, the progress bar and counter track the current attempt, and Shuffle / reset reshuffles both the question order and each question's own option order, so passing a second time by remembering "the third one was green" doesn't work. Your attempt count and best score per module are kept in this browser only.

The three modules, and why they're grouped together

☺ Like you're 10: You can't reason about the budget without the frame from module one, and module three protects the engineering time module two's budget assumes you actually have.

The grouping isn't arbitrary — these are the three modules most people find genuinely sequential, in a way the exam's other five modules mostly aren't. Module 1 establishes what SRE is and how it precisely differs from DevOps and traditional operations, which is the frame every later module quietly assumes. Module 2 turns that frame into an actual number: an SLO target, and the error budget it produces once you subtract that target from 100% and multiply by a window. Module 3 is the operational discipline that number exists to protect — toil is exactly the kind of recurring, budget-draining, engineer-hour-consuming work that a healthy error-budget policy is trying to keep in check, which is why "reduce toil" and "spend the budget deliberately, not by accident" are two ends of the same idea.

One chain, not three unrelated topics 🦉 MODULE 1 · Professor Owl SRE Principles & Practices Sets the frame SRE runs on — origin, and the DevOps relationship 🦥 MODULE 2 · Sol the Sloth SLOs & Error Budgets Turns the frame into a spendable number — the SLI/ SLO/SLA chain 🦫 MODULE 3 · Benny the Beaver Reducing Toil Protects the engineering time that number assumes the team has can't reason about a budget without this toil is what the budget protects against

How to work this bank

☺ Like you're 10: Go in cold, read every explanation even when you got it right, and write down what kind of mistake it was — not just which letter you should've picked.

Reading a practice bank once and nodding along teaches you almost nothing — the format only pays off if you use it the way it's meant to be used.

  1. Go in cold. Close Know It Cold — SREF and the blueprint pages before you start a chip. If you've got a reference tab open while you answer, you're rehearsing search, not recall — a skill the real closed-book sitting gives you zero opportunity to use.
  2. Commit to an answer before reading the explanation, even when you're guessing. A guess you got right for the wrong reason is worth catching now, while it's free.
  3. Read every explanation, including the ones you got right. Half the value here is the sentence naming why the tempting wrong option looked correct — that sentence is the actual skill a single-best-answer exam is testing.
  4. Log the module of every miss, not the question. Three chips, three possible diagnoses. Two misses in "Reducing Toil" and none elsewhere is a specific reading assignment; one miss scattered across each chip is closer to noise.
  5. Re-sit the weak chip a day later, without re-reading the explanations first. Getting it right the second time for the same reason as the first time is genuinely learned; getting it right because you remember which letter was green last time hasn't taught you anything.
⚠ This bank measures recall, not this page

Everything above this heading and below the quiz is teaching material you're allowed to read as many times as you like. The 24 questions themselves are the one part of this page that only works if you treat it like the real thing: no notes, no scrolling back up to the schematic mid-question, no second look at the toil blueprint page while you're mid-way through the toil chip. If a question stumps you, that's the bank doing its job — note it, finish the chip, then go read.

The bank — 24 questions

☺ Like you're 10: Pick an answer. It turns green or red immediately and tells you exactly why.

Eight questions per module, in the exam's own four-option, single-best-answer shape. The chips filter by module; "All 24" runs the full set interleaved, which is closer to what a real sitting actually feels like once you've drilled each module on its own first.

Module
0 / 0 answered · 0 correct

The error-budget arithmetic, worked longhand

☺ Like you're 10: One formula and one table of minutes-per-window — memorize both and almost any error-budget question turns into simple multiplication.

Module 2's arithmetic only requires two things held cold: the formula, and how many minutes live inside whatever window a question hands you. This is the exact same table used on Know It Cold — SREF — reproduced here longhand, Sol-style, because a bank drilling this arithmetic should show its work rather than assume it.

Error budget = (100% − SLO) × window, in the window's own units

Step 1 — convert the window to minutes (30-day month = 43,200 min, unless told otherwise)
Step 2 — subtract the SLO from 100% to get the allowed failure rate
Step 3 — multiply the failure rate by the window
SLOAllowed failure rateBudget / 30-day month (43,200 min)Budget / year
99% (two nines)1%432 min (~7.2 hours)~3.65 days
99.9% (three nines)0.1%43.2 min~8.76 hours
99.95%0.05%21.6 min~4.38 hours
99.99% (four nines)0.01%4.32 min~52.56 min
99.999% (five nines)0.001%~26 sec~5.26 min

Notice the pattern that makes this table reconstructible on exam day even if one cell blanks out: each additional nine divides the allowed downtime by roughly 10x. Hold 99.9% → 43.2 minutes a month cold, and you can derive 99.99% by dividing by ten and 99% by multiplying by ten, without having memorized every row as an isolated fact.

◆ The trap worth over-learning

A service measuring exactly its SLO — not below it, precisely on the line — has already spent its entire error budget for the window, with zero margin left. That reads as "we're fine" and is functionally identical to having already gone slightly under: a properly implemented error-budget policy treats it as spent, full stop, and the release-freeze policy applies exactly as if the team had missed the target. And whichever figure a stem hands you, the formula always runs on the SLO, never the SLA — the SLA is a separate, deliberately looser, external commitment that a question may hand you specifically to see whether you plug the wrong number in.

To drill this arithmetic against a timer rather than reading it, run Drill — SLO & Error-Budget Calculation, and see the policy actually bite in Google & the error-budget policy.

Distinctions this bank keeps returning to

☺ Like you're 10: Here are the lookalike pairs the questions above kept testing, side by side, so the difference stops needing to be re-derived every time.

If you can produce the right-hand column from memory without looking, these three modules are close to done.

These get confusedThe difference in one line
SLI vs SLO vs SLASLI is the raw measurement; SLO is the internal target for it; SLA is the external, contractual promise — deliberately set looser than the SLO, and the only one of the three with financial teeth.
What feeds the error-budget formulaAlways the SLO. The SLA may appear in the same stem, but plugging it into the formula produces a number that matches neither the team's real target nor the contract's own breach condition.
Toil vs overheadAsk one prior question first: is this work about operating the live production service, or not? If not — a status meeting, a headcount spreadsheet — it's overhead, and the six-gate toil test never even applies.
Toil vs engineering project workA novel task, or one that leaves a lasting fix behind, fails "repetitive" or "no enduring value" — it's project work, the deliberate opposite of toil, even when it was unplanned and stressful.
The 50% rule — ceiling, not targetIt caps operational work at roughly half a team's time; it doesn't mandate spending exactly half. Sustained overage is fixed by pushing work back, freezing launches, or adding headcount — not by "working harder."
SRE vs DevOpsDevOps (CALMS) is a cultural interface with no mandated mechanism; SRE is one concrete class that implements it — and SRE's practice at Google predates the word "DevOps" by roughly six years.
The "automatable" gateA task can be manual, recurring, and frequent and still fail toil entirely if it requires a genuine judgment call — a script can't safely replace real risk assessment, whatever the task's other five properties look like.

If you missed these, read this

☺ Like you're 10: Find the chip you dropped marks in, read the page next to it, then go re-sit that chip alone.

Nothing tested above is examined anywhere on this site without also being taught somewhere on this site. Take a miss one module at a time — the middle column is the exam-blueprint version of the idea, the right column is where it's taught at full depth.

🦥 Sol's drill · 10 min

Cover the distinctions table above. For each of the seven rows, say the right-hand column out loud from memory, in one sentence, no hedging. Whatever you can't produce cleanly is your actual reading list for the next hour — and it will be shorter, and far more specific, than "review everything again."

🎬 At the Reliability Watch
🦫

Benny the Beaver: Nineteen out of twenty-four. Not bad for a cold run.

🦉

Professor Owl: Which module took the five you dropped, Benny?

🦫

Benny the Beaver: Toil, mostly. I keep calling something toil the second it's manual and repetitive, and I stop checking there.

🦥

Sol the Sloth: That's not five separate mistakes, then. That's one habit costing you five times — you're skipping "automatable" and "no enduring value" because the first two gates already felt convincing.

🦊

Foxy: So run all six every time, even the ones that feel obvious?

🦫

Benny the Beaver: Especially the ones that feel obvious. That's exactly where I stopped checking.

🐢

Timmy the Turtle: And the SLO arithmetic — any of those five losses arithmetic, or all vocabulary?

🦫

Benny the Beaver: All vocabulary, actually. The one arithmetic question I got right without even rushing it.

🦥

Sol the Sloth: Then that's the honest use of this page — it just told you exactly where to spend the next hour, and it isn't where you assumed.

Where to go next

☺ Like you're 10: Three foundational modules are done — two more grouped banks and a full timed mock stand between here and the real exam.

Modules 1–3 are one third of the blueprint. Work the other two grouped banks the same way — cold, chip by chip, explanations read even when you were right — and only then sit a full mock, since a mock is a measurement and measurements are wasted on material you haven't drilled yet.

✓ Checkpoint

1. Chronologically, which came first — Google's SRE practice or the term "DevOps" — and by roughly how many years? 2. In one sentence each, distinguish an SLI, an SLO, and an SLA, and say which of the three the error-budget formula always runs on. 3. A service's measured attainment sits exactly on its SLO for the window — does the team still have error budget left? 4. Name all six properties a task must clear to count as toil. 5. A task is manual, weekly, and requires a genuine judgment call each time — is it toil, and which gate decides it? 6. What's the standard response when an SRE team sustains operational work meaningfully above the 50% ceiling?

Check your answers
  1. Google's SRE practice came first — Ben Treynor Sloss began building it around 2003, roughly six years before the term "DevOps" was coined at devopsdays Ghent in 2009.
  2. An SLI is the raw measurement (a ratio of good events to valid events); an SLO is the internal target set for that SLI over a defined window; an SLA is the external, contractual promise with stated consequences for a miss. The formula always runs on the SLO, never the SLA.
  3. No — sitting exactly on the SLO line means the entire error budget for that window is already spent, with zero margin remaining. It's functionally identical to having gone slightly under, and a properly implemented error-budget policy treats it the same way.
  4. Manual, repetitive, automatable, tactical, no enduring value, and O(n) — scales linearly — with the service's growth. All six must hold; failing even one disqualifies the task.
  5. Not toil — it fails the automatable gate. A genuine judgment call is exactly the kind of step a script can't safely replace, regardless of how manual and frequent the task otherwise looks.
  6. Push the excess toil back to the owning product team, freeze new toil-generating launches, or add headcount — not simply work harder. The 50% figure is a ceiling that signals a structural problem, not a target to grind through.