Exam Prep · SREF · Mock Exam · Set 2

SREF Mock Exam · Set 2

This is the second of five full-length SREF practice papers, and every structural fact from Set 1 carries over unchanged: the same eight blueprint modules, five questions from each, one unbroken 60-minute sitting, and the same 65% pass mark as the real SRE Foundation exam. What changes is the shape of the question itself. Set 1 tested whether the vocabulary had actually stuck — definitions, formulas, the classic classify-this-scenario calls the blueprint pages already walked you through. Set 2 assumes that vocabulary is already in your head and asks a harder question of it: a team, an engineer, or a VP does something specific, and you have to decide what a correctly-run SRE practice does next — not just name the term for what just happened. Almost every stem below hands you a short situation instead of a slogan, in the same shape the real exam's harder half tends to use. If you haven't sat Set 1 yet, sit it first: Set 2 spends its entire budget on application, on the assumption that recall is no longer the thing being measured.

☺ Explain it like I'm 10

Set 1 was a vocabulary quiz — "what does this word mean?" Set 2 is a fire drill. Nobody asks you to define "smoke detector." Instead, the alarm goes off, there's smoke in one specific room, and you have thirty seconds to say what you actually do next. You already know every term the drill uses — the drill was never testing the words. It's testing whether you can apply them fast enough, to a described situation, without freezing, and without reaching for the first plausible-sounding action instead of the one that's actually correct.

🦊🐢Your hosts for this topic: Foxy & Timmy the Turtle — Foxy digs past the surface symptom in every scenario to what's actually being tested, and Timmy insists the option you pick actually matches the mechanism at work, not just the one that sounds the most confident.

Where Set 2 sits in your SREF prep

☺ Like you're 10: Set 1 taught you the words. Set 2 hands you a situation and asks what you actually do about it — same words, harder use of them.

Because DevOps Institute doesn't publish a confirmed per-module weighting for the real exam's 40 questions, this paper keeps Set 1's working assumption: five questions drawn evenly from each of the eight modules, spread across the sitting rather than grouped by topic, so a domain you drop points in here is telling you precisely where to go back and reread. What's different is the stem itself. Instead of "what is toil" or "what's the difference between an SLI and an SLO," a Set 2 question reads more like: an engineer notices X, a manager proposes Y, a customer complains about Z — and asks what a correctly-run SRE practice does about it. That framing tests something Set 1 structurally cannot: whether you can hold a described situation in your head, spot which mechanism it's actually invoking, and reject the option that merely sounds right in favor of the one that matches the mechanism.

This is the natural next step after Set 1, not a harder version of the same paper. If Set 1 came back weak in a specific module, fix that module's content gap first — a scenario paper won't teach you a term you never learned, it'll just show you a new way of not knowing it. If Set 1 came back strong, Set 2 is where you find out whether "I know the definition" has actually turned into "I'd reach for the right response under pressure," which is a different skill and the one the real exam's harder half is built to test. Set 3 is a different exercise entirely — a concentrated drill on four specific confusable term-pairs rather than a domain-balanced sitting — worth sitting once you know which of those four pairs, if any, are still shaky. The SREF study plan lays out where all five sets fit against your remaining prep time.

Sit it like the real thing

☺ Like you're 10: Real quiz rules: no notes, no browser tab, one hour, go — and read the whole situation before you even look at the four options.

The SRE Foundation is closed-book: no notes, no browser tab, no this course, no colleague, no AI assistant, once the clock starts. Sit Set 2 the same way — close every other tab, put your phone away, and treat the forty questions below exactly as you'd treat them in a locked-down exam environment. The one habit worth adding for this specific paper: read the full scenario before you look at the options. A definitional question can be pattern-matched from its first few words; a scenario question usually hides its deciding detail in the last sentence of the stem — how much budget is left, who's already been notified, what's already been tried — and skimming straight to the answer choices is exactly how a candidate who knows the material still picks the wrong one.

⚠ Format specifics can change — verify before you book

Everything on this page — the 40-question count, the 60-minute duration, the 65% pass mark, and the assumption of no per-module weighting — reflects this course's own certifications page, itself sourced from DevOps Institute's published materials at a point in time. Certification bodies revise format details without much notice. Confirm the current specifics on DevOps Institute's own SRE Foundation page before you register for the real thing — treat this paper as calibrated to that format, not as a substitute for checking it yourself.

The same guessing assumption from Set 1 carries over: most closed-book, single-best-answer foundation-level certifications of this shape don't deduct extra for a wrong answer beyond simply not scoring it. If that holds for the real SREF exam (verify it yourself before you sit the real thing), never leave a question blank — eliminate what you can, commit to your best remaining option, and move on. A considered guess on a scenario question is still a guess grounded in "which of these four actually addresses the mechanism," which is worth far more than a coin flip.

Your 60-minute budget

☺ Like you're 10: Same clock as Set 1 — sixty minutes, forty questions, five blocks of eight — but budget a few extra seconds per question to actually read the situation first.

The block structure is identical to Set 1 on purpose, so your pace on the two papers is directly comparable: five 8-question blocks, one question from every module in each block, with a short settle-in at the start and a flag-review sweep at the end. What's worth adjusting mentally, not on the clock itself, is where the 90-second average actually goes — a one-line definitional stem reads in two seconds, but a three-sentence scenario needs real reading time before you can even start eliminating options. Don't let that reading time come out of nowhere at the end of a block; it's already priced into the same twelve minutes.

60 minutes · 40 questions · 65% to pass settle 2m Block 1 Q1–8 · 12 min Block 2 Q9–16 · 12 min Block 3 Q17–24 · 12 min Block 4 Q25–32 · 12 min Block 5 Q33–40 · 10 min flag 2m 0 12 24 36 48 58 60 Every block still holds one question from every one of the eight modules — but every stem is now a short scenario, worth reading in full before you look at the options.

If a question runs past about 90 seconds without you being confident, don't stall: eliminate the options that clearly don't address the mechanism the scenario is testing, commit to your best remaining pick, and — since there's ordinarily no penalty for a wrong answer — move on. As with Set 1, there's no folded solution to peek at mid-paper; each question's explanation sits directly underneath it, so the only discipline required is not opening it before you've committed.

The paper — 40 questions in exam order

☺ Like you're 10: Forty short situations, five blocks of eight, one from every module in every block — read the whole thing, pick the response that actually fits, then check yourself.

Read each stem in full, pick your answer, and only then open the explanation underneath it. Each question is tagged with its blueprint module in parentheses so you can total your score by module afterward — on the real exam you won't get that label, so once you've sat this paper the first time, consider re-reading the stems with the tags covered to see how many you can still place correctly on content alone.

Block 1 — Q1–8 (minutes 0–12)

Q1 (Module 1 — SRE Principles & Practices). A new VP announces that the ops team will be renamed "Site Reliability Engineering" this quarter — same staffing, same responsibilities, same change process, just a new name on the org chart. What's the correct SRE assessment of this change?

Check the answer

B. The mechanisms — SLOs, an enforced error-budget policy, a toil ceiling — are what make SRE falsifiable and auditable. A title change with none of them attached is the single most common failure mode this discipline warns about, whatever a postmortem template or a reporting line says on paper.

Q2 (Module 2 — SLOs & Error Budgets). A checkout service holds a 99.9% availability SLO over a rolling 30-day window (about 43 minutes of budget). With 12 days left in the window, the team has already used roughly 40 of those minutes. What should happen next, under a properly implemented error-budget policy?

Check the answer

B. Forty of roughly forty-three minutes spent with twelve days still open is exactly the trigger condition a written error-budget policy exists for. Lowering the SLO to match the failure (C) is the classic anti-pattern; there's no reason given to roll back an unrelated year-old release (D).

Q3 (Module 3 — Reducing Toil). An engineer manually resizes a Kubernetes node pool by typing the same three commands, roughly twice a week, whenever traffic climbs. Which property makes this toil rather than legitimate engineering work?

Check the answer

B. Manual, repetitive, automatable, tactical, no enduring value, and scaling with traffic — all six gates hold. Neither access method, time of day, nor task duration is part of the formal definition.

Q4 (Module 4 — Monitoring & SLIs). An on-call engineer is paged 40 times in one week; only 3 of the pages required any real action, and the rest self-resolved within a minute. What's the correct first response?

Check the answer

C. A 37-out-of-40 self-resolve rate is a textbook alert-design problem, not a staffing problem — splitting bad alerts across more people (D) just spreads the fatigue instead of fixing its source.

Q5 (Module 5 — SRE Tools & Automation). A team needs to monitor a legacy application that can't be modified to expose a Prometheus-compatible /metrics endpoint. What's the standard way to bring it into a Prometheus-based monitoring stack?

Check the answer

B. The exporter pattern is exactly what exists for unmodifiable third-party or legacy systems — it translates whatever the app already exposes (a log file, JMX, a proprietary API) into a format Prometheus can scrape, with no change to the application itself.

Q6 (Module 6 — Anti-Fragility & Learning from Failure). Before running a chaos experiment that kills a random instance of a production payment service, what should the team define first?

Check the answer

B. Chaos engineering is a controlled, hypothesis-driven experiment, not unplanned failure injection — a steady-state hypothesis and a bounded blast radius are the two things that make it safe to run against production at all.

Q7 (Module 7 — Organizational Impact of SRE). One central SRE team now supports 40 product teams and has become the bottleneck for every reliability review and every production readiness check. What organizational shift does SRE practice typically recommend at this scale?

Check the answer

B. A purely centralized model doesn't scale linearly with the number of teams it supports; distributing ownership through embedded or platform-style models is the standard structural fix once one central team becomes the ceiling on everyone else's velocity.

Q8 (Module 8 — SRE, Other Frameworks & the Future). An organization running a traditional ITIL-style Change Advisory Board (CAB) — where every change, regardless of size or risk, must be approved at a fixed weekly meeting — wants to start adopting SRE practice. What's the key shift in how change risk gets governed?

Check the answer

B. SRE doesn't remove change governance — it replaces a fixed, calendar-gated committee with a continuous, risk-proportionate one driven by the remaining error budget, which is also the direction ITIL 4 itself moved in.

Block 2 — Q9–16 (minutes 12–24)

Q9 (Module 5 — SRE Tools & Automation). A critical page goes unacknowledged for 15 minutes because the primary on-call's phone was silenced, and there was no fallback configured. Which tool feature would have prevented the silent miss?

Check the answer

A. An escalation policy with a timeout is specifically the guardrail against exactly this failure mode — a single unreachable responder silently swallowing a page. Retrying the same unreachable phone (D) does nothing to route around the actual failure.

Q10 (Module 7 — Organizational Impact of SRE). During a severe, company-wide outage touching five different teams, the response turns chaotic: three engineers independently try to roll back three different services with no shared picture of what's already been tried. What's missing?

Check the answer

B. Uncoordinated parallel action during a multi-team incident is precisely the failure mode incident command exists to prevent — one person directing the response and maintaining a shared picture, rather than everyone independently guessing.

Q11 (Module 2 — SLOs & Error Budgets). A monitoring system pages the on-call engineer because, at the current failure rate, the service is on pace to burn its entire 30-day error budget within the next two hours. What does this fast-burn alert mean, and what's the right response?

Check the answer

B. A fast-burn alert is exactly what a high burn-rate multiplier is designed to catch — severe enough right now that waiting for the slow-burn signal to fire would mean discovering the exhausted budget only after it's already gone.

Q12 (Module 8 — SRE, Other Frameworks & the Future). A company already has a platform engineering team building a self-service internal developer platform. Leadership asks whether the organization still needs SRE as a separate discipline. What's the accurate answer?

Check the answer

C. The two disciplines answer different questions — how easy is it to build and ship, versus how reliable is what's running — and most mature organizations staff both rather than treating them as substitutes.

Q13 (Module 1 — SRE Principles & Practices). An on-call log shows a team spent 65% of the last quarter on paging, tickets, and manual restarts, and only 35% on engineering work. Per Google's toil-ceiling guidance, what's the correct next step?

Check the answer

D. 65% sustained above the 50% ceiling is treated as a staffing and automation-investment signal, not a performance issue and not something to tolerate silently until it reaches 90%.

Q14 (Module 6 — Anti-Fragility & Learning from Failure). A team runs a "game day" simulating a full regional cloud outage. Two hours in, the exercise reveals that the documented failover runbook is badly out of date and would have caused a second outage if followed literally. What's the correct reaction?

Check the answer

C. Finding a broken runbook in a drill, at no real cost, is the exercise working exactly as intended — the alternative is discovering it live, during an actual regional outage.

Q15 (Module 4 — Monitoring & SLIs). A dashboard shows CPU and memory both comfortably under 50%, yet users report the service feels slow. Which category of signal is most likely being missed?

Check the answer

A. Low raw utilization can coexist with a saturated queue or exhausted connection pool — the four golden signals are distinct dimensions precisely because a system can look healthy on three of them and be starving on the fourth.

Q16 (Module 3 — Reducing Toil). A team wants to reduce the recurring toil around certificate rotation. What's the correct order of priority?

Check the answer

C. A runbook (B) is a real rung on the automation ladder, but it's an early one — it makes manual execution consistent, it doesn't remove the human from the loop, which is what full automation actually does.

Block 3 — Q17–24 (minutes 24–36)

Q17 (Module 6 — Anti-Fragility & Learning from Failure). A postmortem's "five whys" chain ends at "...because the on-call engineer typed the wrong command." The facilitator says this isn't a complete root cause yet. What should the team ask next?

Check the answer

B. Stopping at a human action is the single most common way a five-whys chain falls short — the systemic gap that let a routine typo reach production is the fixable cause, and it's what's still unasked.

Q18 (Module 1 — SRE Principles & Practices). A postmortem for a database outage ends with the line: "Priya applied an untested migration under time pressure and should be more careful in the future." What's wrong with this postmortem from a blameless standpoint?

Check the answer

C. A calm tone or a single mention doesn't make a postmortem blameless — what matters is whether it stops at an individual's action or keeps asking what system-level gap let that action reach production.

Q19 (Module 8 — SRE, Other Frameworks & the Future). A team starts using an ML-based anomaly-detection tool that automatically correlates alerts across services and proposes a likely root cause during an incident. What's the SRE-appropriate way to treat its suggestion?

Check the answer

B. AIOps-style correlation tools speed up triage without replacing verification — a correlated suggestion is a hypothesis, not a diagnosis, until it's checked against real evidence.

Q20 (Module 3 — Reducing Toil). A team reports that toil has crept up because every new microservice they onboard requires the same six manual steps to wire up monitoring, logging, and alerting. What's the most durable fix?

Check the answer

B. Rotating who performs a manual task (A) doesn't reduce the toil, it just spreads it — a self-service template removes the six steps from the critical path entirely.

Q21 (Module 7 — Organizational Impact of SRE). A team is asked to push a service from 99.9% to 99.999% availability ("five nines"). What's the correct framing to bring into that conversation?

Check the answer

C. Reliability has diminishing returns and rising marginal cost — the correct conversation weighs that cost against real user and business need, not an assumption that more is unconditionally better.

Q22 (Module 2 — SLOs & Error Budgets). A customer complains the product is "always down," despite the internal dashboard showing the service comfortably inside its 99.9% quarterly SLO. Which distinction most likely explains the mismatch?

Check the answer

B. A comfortable SLO with an unhappy customer is usually an SLI-coverage problem — the measurement isn't capturing the specific journey or failure class the customer is actually hitting.

Q23 (Module 5 — SRE Tools & Automation). A team applies the same Terraform configuration twice in a row with nothing changed in between. What should happen the second time, and why does that property matter operationally?

Check the answer

A. Idempotent applies are what let infrastructure-as-code run unattended and repeatedly — a no-op plan on the second run is the correct, expected behavior, not an edge case.

Q24 (Module 4 — Monitoring & SLIs). A request that touches eight microservices is intermittently slow, but every individual service's own dashboard looks healthy in isolation. What's the correct tool for finding where the time is actually going?

Check the answer

C. When every service looks healthy alone but the end-to-end request is slow, the problem is almost always in the gaps between services — exactly what a per-request trace, not a per-service dashboard, is built to reveal.

Block 4 — Q25–32 (minutes 36–48)

Q25 (Module 2 — SLOs & Error Budgets). A product VP insists error budgets should be set unilaterally by engineering, with no product or business input, "because reliability is purely an engineering concern." What's the correct objection?

Check the answer

B. An SLO target sets the tradeoff between shipping speed and reliability spend — a decision with real business consequences on both sides, which is exactly why it's negotiated jointly rather than set unilaterally by either side.

Q26 (Module 4 — Monitoring & SLIs). An engineer adds a unique user_id label to a latency histogram "for better debugging," and within a day the metrics backend runs out of memory and starts dropping data. What happened?

Check the answer

C. A label with unbounded unique values multiplies the time-series count, not the label count — this is the textbook cardinality-explosion failure, and the fix is moving that detail to logs or traces instead of resizing the metrics store.

Q27 (Module 6 — Anti-Fragility & Learning from Failure). A team has never run a failure-injection experiment before and is worried about causing a real outage on day one. What's the recommended way to begin?

Check the answer

C. Chaos maturity is built progressively — small, bounded experiments first, wider blast radius only once confidence and tooling justify it — not by starting at company-wide scope on day one.

Q28 (Module 1 — SRE Principles & Practices). Leadership wants engineers to "just be more careful" as the primary reliability strategy for the coming year, with no change to review process, testing, or automation. What's the core SRE objection?

Check the answer

B. "Be more careful" with no supporting process or automation change asks humans to substitute for a system fix — SRE's core move is to treat the recurring failure as an engineering problem instead.

Q29 (Module 3 — Reducing Toil). Which of the following is the best example of toil actually worth tolerating in the short term rather than automating right away?

Check the answer

A. The "repetitive" gate is what disqualifies toil-worth-tolerating from toil-worth-automating — a genuinely one-off task never repeats, so the automation investment never pays itself back, unlike B, C, or D, all of which recur indefinitely.

Q30 (Module 8 — SRE, Other Frameworks & the Future). A team's systems have grown more distributed — more microservices, more managed cloud dependencies, more third-party APIs — and their metrics-only monitoring stack increasingly can't answer "why" a specific request failed. Which broader shift does this describe, and what's the recommended direction?

Check the answer

C. This is the monitoring-to-observability shift in a nutshell: a fixed set of pre-chosen metrics can only answer questions you thought to ask in advance, while observability is built to answer the question you didn't anticipate.

Q31 (Module 7 — Organizational Impact of SRE). During an active incident, the on-call SRE realizes the "outage" is actually being caused by an ongoing credential-stuffing attack overwhelming the auth service. What's the correct organizational response?

Check the answer

A. Scaling capacity treats the symptom of an attack, not the attack itself — once the cause is identified as adversarial, the incident belongs jointly to reliability and security response, not to reliability alone.

Q32 (Module 5 — SRE Tools & Automation). A recurring incident type — "disk fills up on node X, evict pods, resize the volume" — has a well-tested, low-risk runbook that's been executed by hand successfully more than 20 times. What's the correct next step for that runbook?

Check the answer

B. Twenty clean executions of a low-risk, well-understood runbook is exactly the evidence base that justifies moving it up the automation ladder to unattended remediation.

Block 5 — Q33–40 (minutes 48–58)

Q33 (Module 8 — SRE, Other Frameworks & the Future). A company's cloud bill has grown faster than its traffic, and finance asks the SRE team why maintaining "five nines everywhere" behaves like a black-box cost center. Which emerging cross-discipline practice is most directly relevant here, and what's its core idea?

Check the answer

A. FinOps is specifically the practice of bringing cost accountability into engineering decision-making in real time — exactly the missing piece when a reliability target is set without anyone pricing what it costs to hold.

Q34 (Module 3 — Reducing Toil). After building three automations this quarter, a team's toil-ticket count drops from 200 to 40. Two engineers report they're now spending that freed-up time entirely on ad hoc, low-priority feature requests instead of reliability project work. What's the correct read of this situation?

Check the answer

B. The entire point of cutting toil is to grow the engineering half of the split — letting the freed capacity drift into unrelated, unmanaged ad hoc work quietly defeats that purpose and can seed a new toil source.

Q35 (Module 5 — SRE Tools & Automation). A platform team's Grafana dashboards are updated by hand-editing them directly in the production UI, with someone "supposed to" export the JSON back into the team's Git repository afterward — a step that routinely gets forgotten. What's the correct fix?

Check the answer

A. Dashboards-as-code removes the manual "remember to export" step by making the repository, not the live UI, the source of truth — the same discipline applied to every other piece of production configuration.

Q36 (Module 2 — SLOs & Error Budgets). A checkout flow depends on three internal services — cart, payments, and inventory. The team defines one end-to-end SLO for "checkout succeeds" instead of three separate SLOs for each service. What's the main advantage of this composite approach?

Check the answer

B. Three individually healthy-looking services can still combine into a broken user journey — a composite SLO measures the thing the user actually experiences instead of three separate numbers that can each look fine in isolation.

Q37 (Module 4 — Monitoring & SLIs). A team standardizes on the RED method (Rate, Errors, Duration) for every service's dashboards and the USE method (Utilization, Saturation, Errors) for the resources underneath. During an incident, RED looks fine on the affected service but a downstream queue is badly saturated. Why does keeping both views matter here?

Check the answer

B. This exact split — clean RED, saturated resource underneath — is the textbook case for why service-level and resource-level views are tracked separately rather than assumed to move together.

Q38 (Module 7 — Organizational Impact of SRE). A manager says, "We already do blameless postmortems — we hold the meeting after every incident." A reviewing SRE points out this alone doesn't prove the culture is actually blameless. What additional evidence would actually demonstrate it?

Check the answer

C. A meeting existing on the calendar proves a process was followed, not that it's safe to be honest inside it — the real test is whether engineers volunteer the uncomfortable details without fear of career consequences.

Q39 (Module 6 — Anti-Fragility & Learning from Failure). Six months after a major outage, the postmortem's tracked action items are still open — they keep losing priority to feature work. What does this reveal, and what's the correct organizational fix?

Check the answer

B. A correct root cause with no completed fix is "postmortem theater" — the paperwork is right, the system is exactly as vulnerable as before, and the durable fix is protected prioritization, not more documentation.

Q40 (Module 1 — SRE Principles & Practices). A candidate paraphrases Ben Treynor Sloss's definition of SRE as "what happens when you ask a software engineer to design an operations function." A team lead asks what that framing actually implies day to day. What's the correct operational implication?

Check the answer

C. The operational payoff of "software engineer designs the operations function" is that recurring manual work gets treated as a bug to fix in code — reviewed, tested, version-controlled — not repeated by hand indefinitely.

Score yourself

☺ Like you're 10: Count how many letters you got right out of forty, turn it into a percentage against 65%, then look at which module — and which kind of mistake — actually cost you the points.

Total your correct answers out of 40 for your raw percentage against the 65% pass mark (26 correct or more). Then total each module separately out of its five questions, exactly as you did for Set 1 — a 30/40 spread evenly across eight modules and a 30/40 built from seven near-perfect modules plus a zero on one are very different results on the real exam, even though they score identically here.

ModuleQuestions in this paperYour scoreIf under 4/5, go here
1 · SRE Principles & PracticesQ1, Q13, Q18, Q28, Q40/5SRE Principles & Practices
2 · SLOs & Error BudgetsQ2, Q11, Q22, Q25, Q36/5Service Level Objectives & Error Budgets
3 · Reducing ToilQ3, Q16, Q20, Q29, Q34/5Reducing Toil
4 · Monitoring & SLIsQ4, Q15, Q24, Q26, Q37/5Monitoring & Service Level Indicators
5 · SRE Tools & AutomationQ5, Q9, Q23, Q32, Q35/5SRE Tools & Automation
6 · Anti-Fragility & Learning from FailureQ6, Q14, Q17, Q27, Q39/5Anti-Fragility & Learning from Failure
7 · Organizational Impact of SREQ7, Q10, Q21, Q31, Q38/5Organizational Impact of SRE
8 · SRE, Other Frameworks & the FutureQ8, Q12, Q19, Q30, Q33/5SRE, Other Frameworks & the Future
Total40 questions/4065% (26/40) to pass

A scenario paper produces a third kind of miss that Set 1 mostly couldn't. Sort what you got wrong into three piles. A genuine content gap means you didn't know the underlying mechanism at all — reread the linked module. The "sounds confident" trap means you knew the material but picked the option that read as decisive, cautious, or authoritative rather than the one that actually matched the mechanism the scenario was testing — that's not a knowledge gap, it's a habit of trusting tone over substance, and the fix is simply sitting more scenario papers and asking "does this option address the actual cause, or does it just sound like the kind of thing a careful person would say?" before you commit. A misread means the scenario contained a specific detail — how much budget was left, who'd already been notified, what role was already assigned — that you skimmed past; the fix there is the "read the whole stem first" habit from earlier on this page, not more content review.

For deeper, module-specific reps beyond a full 40-question sitting, this course's practice banks split the same ground into focused sets: Practice · SRE Principles, SLOs & Toil, Practice · Monitoring & SRE Tools, and Practice · Anti-Fragility & Organizational Impact — plus the full SREF Practice Questions bank and Answer Triage for the general skill of ruling out an option that merely sounds right, which is exactly the trap this paper spends most of its forty questions on.

🎬 At the Reliability Watch
🐢

Timmy the Turtle: Score?

🦊

Foxy: 34. But six of my misses cluster in exactly one place — Organizational Impact.

🐢

Timmy the Turtle: Which one, specifically?

🦊

Foxy: The team-topology bottleneck question. I picked "hire more into the same centralized team," because it sounded like the safe, cautious answer.

🐢

Timmy the Turtle: That's the trap on this whole paper. A scenario question always has one option that sounds careful and conservative and is actually just more of the thing that's already broken.

🦝

Rocky the Raccoon: Meanwhile I got the chaos-engineering one wrong — I picked "no planning needed, that's the whole premise." In my defense, that's basically my entire personality.

🐢

Timmy the Turtle: Which is exactly why blast radius exists, Rocky. Even you plan the boundary before you touch anything.

🦥

Sol the Sloth: The pattern's the same in both cases, though — you both answered fast, on vibes, instead of checking whether the option actually matched the mechanism.

🦉

Professor Owl: Which is the whole difference between Set 1 and Set 2. Set 1 asked what a word means. Set 2 asks whether you'd reach for the right lever under a running clock — and confidence isn't a substitute for checking.

✓ Checkpoint

1. What's the one thing that actually changed between Set 1 and Set 2, given that the module count, questions per module, timing, and pass mark all stayed identical? 2. Name one of the three common ways a candidate gets a scenario question wrong, beyond simply "never learned the term." 3. Why does a scenario stem typically deserve a slower first read than a straight definitional one, even though the clock still averages 90 seconds a question? 4. If your module tally shows one module noticeably weaker than the rest, what should you do before sitting Set 3 or Set 4?

Check your answers
  1. The structure stayed identical to Set 1 — eight modules, five questions each, one 60-minute sitting, 65% pass mark. What changed is the question's shape: almost every stem now describes a short situation and asks what a correctly-run SRE practice does next, instead of asking you to define or classify a term directly.
  2. Any of: a genuine content gap (never learned the underlying mechanism); the "sounds confident" trap (picking the option that reads as decisive or cautious rather than the one that actually matches the mechanism); or a misread (missing a specific stated detail in the scenario, such as time remaining, who's already been notified, or a role that's already assigned).
  3. Because the trap in a scenario question is usually hidden in one specific detail inside the situation, not in the wording of the options — skimming the setup and jumping straight to pattern-matching an answer is exactly how a candidate who knows the material still loses the point.
  4. Re-read that module's linked blueprint page, work the matching practice bank, and re-sit just that module's five questions cold in a day or two before moving on — fixing one soft module now is cheaper than carrying it forward into Set 3 or Set 4.

That's the paper. Set the timer once, keep every scenario's full detail in view before you touch the options, and let the forty questions above show you whether Set 1's vocabulary has actually turned into judgment. Score it by module, fix only what's soft, and when you're ready, Set 3 is a completely different kind of paper — the same four confusable term-pairs, over and over, until they stop being a coin flip.

📝 The five sets

Set 1 · Set 2 (you are here) · Set 3 · Set 4 · Set 5. Sets 1 and 2 are both domain-balanced — five questions per module, 40 total — with Set 1 skewed toward recall and Set 2 skewed toward scenario judgment; Set 3 breaks from that format entirely to drill four confusable term-pairs intensively instead of sampling all eight modules. See the SREF study plan for how to sequence all five against your remaining prep time, and the SREF exam for the real thing's day-of logistics.