SREF Mock Exam · Set 2
This is the second of five full-length SREF practice papers, and every structural fact from Set 1 carries over unchanged: the same eight blueprint modules, five questions from each, one unbroken 60-minute sitting, and the same 65% pass mark as the real SRE Foundation exam. What changes is the shape of the question itself. Set 1 tested whether the vocabulary had actually stuck — definitions, formulas, the classic classify-this-scenario calls the blueprint pages already walked you through. Set 2 assumes that vocabulary is already in your head and asks a harder question of it: a team, an engineer, or a VP does something specific, and you have to decide what a correctly-run SRE practice does next — not just name the term for what just happened. Almost every stem below hands you a short situation instead of a slogan, in the same shape the real exam's harder half tends to use. If you haven't sat Set 1 yet, sit it first: Set 2 spends its entire budget on application, on the assumption that recall is no longer the thing being measured.
Set 1 was a vocabulary quiz — "what does this word mean?" Set 2 is a fire drill. Nobody asks you to define "smoke detector." Instead, the alarm goes off, there's smoke in one specific room, and you have thirty seconds to say what you actually do next. You already know every term the drill uses — the drill was never testing the words. It's testing whether you can apply them fast enough, to a described situation, without freezing, and without reaching for the first plausible-sounding action instead of the one that's actually correct.
Where Set 2 sits in your SREF prep
☺ Like you're 10: Set 1 taught you the words. Set 2 hands you a situation and asks what you actually do about it — same words, harder use of them.
Because DevOps Institute doesn't publish a confirmed per-module weighting for the real exam's 40 questions, this paper keeps Set 1's working assumption: five questions drawn evenly from each of the eight modules, spread across the sitting rather than grouped by topic, so a domain you drop points in here is telling you precisely where to go back and reread. What's different is the stem itself. Instead of "what is toil" or "what's the difference between an SLI and an SLO," a Set 2 question reads more like: an engineer notices X, a manager proposes Y, a customer complains about Z — and asks what a correctly-run SRE practice does about it. That framing tests something Set 1 structurally cannot: whether you can hold a described situation in your head, spot which mechanism it's actually invoking, and reject the option that merely sounds right in favor of the one that matches the mechanism.
This is the natural next step after Set 1, not a harder version of the same paper. If Set 1 came back weak in a specific module, fix that module's content gap first — a scenario paper won't teach you a term you never learned, it'll just show you a new way of not knowing it. If Set 1 came back strong, Set 2 is where you find out whether "I know the definition" has actually turned into "I'd reach for the right response under pressure," which is a different skill and the one the real exam's harder half is built to test. Set 3 is a different exercise entirely — a concentrated drill on four specific confusable term-pairs rather than a domain-balanced sitting — worth sitting once you know which of those four pairs, if any, are still shaky. The SREF study plan lays out where all five sets fit against your remaining prep time.
Sit it like the real thing
☺ Like you're 10: Real quiz rules: no notes, no browser tab, one hour, go — and read the whole situation before you even look at the four options.
The SRE Foundation is closed-book: no notes, no browser tab, no this course, no colleague, no AI assistant, once the clock starts. Sit Set 2 the same way — close every other tab, put your phone away, and treat the forty questions below exactly as you'd treat them in a locked-down exam environment. The one habit worth adding for this specific paper: read the full scenario before you look at the options. A definitional question can be pattern-matched from its first few words; a scenario question usually hides its deciding detail in the last sentence of the stem — how much budget is left, who's already been notified, what's already been tried — and skimming straight to the answer choices is exactly how a candidate who knows the material still picks the wrong one.
Everything on this page — the 40-question count, the 60-minute duration, the 65% pass mark, and the assumption of no per-module weighting — reflects this course's own certifications page, itself sourced from DevOps Institute's published materials at a point in time. Certification bodies revise format details without much notice. Confirm the current specifics on DevOps Institute's own SRE Foundation page before you register for the real thing — treat this paper as calibrated to that format, not as a substitute for checking it yourself.
The same guessing assumption from Set 1 carries over: most closed-book, single-best-answer foundation-level certifications of this shape don't deduct extra for a wrong answer beyond simply not scoring it. If that holds for the real SREF exam (verify it yourself before you sit the real thing), never leave a question blank — eliminate what you can, commit to your best remaining option, and move on. A considered guess on a scenario question is still a guess grounded in "which of these four actually addresses the mechanism," which is worth far more than a coin flip.
Your 60-minute budget
☺ Like you're 10: Same clock as Set 1 — sixty minutes, forty questions, five blocks of eight — but budget a few extra seconds per question to actually read the situation first.
The block structure is identical to Set 1 on purpose, so your pace on the two papers is directly comparable: five 8-question blocks, one question from every module in each block, with a short settle-in at the start and a flag-review sweep at the end. What's worth adjusting mentally, not on the clock itself, is where the 90-second average actually goes — a one-line definitional stem reads in two seconds, but a three-sentence scenario needs real reading time before you can even start eliminating options. Don't let that reading time come out of nowhere at the end of a block; it's already priced into the same twelve minutes.
If a question runs past about 90 seconds without you being confident, don't stall: eliminate the options that clearly don't address the mechanism the scenario is testing, commit to your best remaining pick, and — since there's ordinarily no penalty for a wrong answer — move on. As with Set 1, there's no folded solution to peek at mid-paper; each question's explanation sits directly underneath it, so the only discipline required is not opening it before you've committed.
The paper — 40 questions in exam order
☺ Like you're 10: Forty short situations, five blocks of eight, one from every module in every block — read the whole thing, pick the response that actually fits, then check yourself.
Read each stem in full, pick your answer, and only then open the explanation underneath it. Each question is tagged with its blueprint module in parentheses so you can total your score by module afterward — on the real exam you won't get that label, so once you've sat this paper the first time, consider re-reading the stems with the tags covered to see how many you can still place correctly on content alone.
Block 1 — Q1–8 (minutes 0–12)
Q1 (Module 1 — SRE Principles & Practices). A new VP announces that the ops team will be renamed "Site Reliability Engineering" this quarter — same staffing, same responsibilities, same change process, just a new name on the org chart. What's the correct SRE assessment of this change?
- A. It's a genuine SRE transformation, since the team is already the one carrying the pager
- B. It's traditional operations wearing a new job title — without SLOs, an enforced error-budget policy, and a toil ceiling, a name change alone doesn't create the discipline
- C. It qualifies as SRE as long as the team starts writing a postmortem after every incident, blameless or not
- D. It qualifies as SRE as long as the renamed team now reports into engineering instead of IT
Check the answer
B. The mechanisms — SLOs, an enforced error-budget policy, a toil ceiling — are what make SRE falsifiable and auditable. A title change with none of them attached is the single most common failure mode this discipline warns about, whatever a postmortem template or a reporting line says on paper.
Q2 (Module 2 — SLOs & Error Budgets). A checkout service holds a 99.9% availability SLO over a rolling 30-day window (about 43 minutes of budget). With 12 days left in the window, the team has already used roughly 40 of those minutes. What should happen next, under a properly implemented error-budget policy?
- A. Nothing — only the calendar-year total error budget is ever a real constraint
- B. The near-exhausted budget should trigger the pre-agreed policy: pause risky launches and prioritize reliability work until the window rolls over and the budget resets
- C. The SLO should quietly be lowered to 99% so the dashboard stops showing red
- D. The team should roll the entire service back to the version running a year ago
Check the answer
B. Forty of roughly forty-three minutes spent with twelve days still open is exactly the trigger condition a written error-budget policy exists for. Lowering the SLO to match the failure (C) is the classic anti-pattern; there's no reason given to roll back an unrelated year-old release (D).
Q3 (Module 3 — Reducing Toil). An engineer manually resizes a Kubernetes node pool by typing the same three commands, roughly twice a week, whenever traffic climbs. Which property makes this toil rather than legitimate engineering work?
- A. It requires an SSH key to execute
- B. It's manual, repetitive, fully scriptable with no real judgment call, and produces nothing that outlasts the moment — next week's traffic spike needs the exact same three commands again
- C. It only happens during business hours
- D. It only counts as toil once a single instance takes longer than 30 minutes
Check the answer
B. Manual, repetitive, automatable, tactical, no enduring value, and scaling with traffic — all six gates hold. Neither access method, time of day, nor task duration is part of the formal definition.
Q4 (Module 4 — Monitoring & SLIs). An on-call engineer is paged 40 times in one week; only 3 of the pages required any real action, and the rest self-resolved within a minute. What's the correct first response?
- A. Nothing needs to change about the alerting — the engineer just needs to filter noise faster
- B. Mute all pages for the remainder of the on-call week
- C. Treat this as alert fatigue caused by over-sensitive or poorly designed alerts, and redesign the alerting itself — for example adding a "for" duration or multi-window burn-rate logic — rather than expecting a human to keep absorbing the noise
- D. Add a fourth engineer to the rotation so the same 40 pages get split three ways instead of two
Check the answer
C. A 37-out-of-40 self-resolve rate is a textbook alert-design problem, not a staffing problem — splitting bad alerts across more people (D) just spreads the fatigue instead of fixing its source.
Q5 (Module 5 — SRE Tools & Automation). A team needs to monitor a legacy application that can't be modified to expose a Prometheus-compatible /metrics endpoint. What's the standard way to bring it into a Prometheus-based monitoring stack?
- A. Skip monitoring for this application entirely, since Prometheus can't observe anything it wasn't built to instrument
- B. Deploy or reuse an exporter — a small adapter process that translates the application's native metrics into the Prometheus exposition format so it can be scraped normally
- C. Manually copy the numbers from the app's admin UI into a spreadsheet on a fixed schedule
- D. Migrate the whole organization off Prometheus onto a push-only metrics system
Check the answer
B. The exporter pattern is exactly what exists for unmodifiable third-party or legacy systems — it translates whatever the app already exposes (a log file, JMX, a proprietary API) into a format Prometheus can scrape, with no change to the application itself.
Q6 (Module 6 — Anti-Fragility & Learning from Failure). Before running a chaos experiment that kills a random instance of a production payment service, what should the team define first?
- A. Nothing in advance — deliberately introducing failure with no planning is the entire premise of chaos engineering
- B. A steady-state hypothesis based on real metrics, and a deliberately bounded blast radius, so the experiment can be halted immediately if the impact exceeds that small, pre-agreed scope
- C. A public announcement telling every customer exactly when and how the system will be made to fail
- D. Individual written sign-off from every engineer at the company
Check the answer
B. Chaos engineering is a controlled, hypothesis-driven experiment, not unplanned failure injection — a steady-state hypothesis and a bounded blast radius are the two things that make it safe to run against production at all.
Q7 (Module 7 — Organizational Impact of SRE). One central SRE team now supports 40 product teams and has become the bottleneck for every reliability review and every production readiness check. What organizational shift does SRE practice typically recommend at this scale?
- A. Keep hiring into the same fully centralized team, indefinitely, with no change to the review structure
- B. Shift toward embedding SRE practice into product teams — for example through a platform/enablement model or embedded SREs — so reliability ownership scales with the number of teams instead of funneling every review through one bottleneck
- C. Eliminate the SRE function outright, since the current structure clearly isn't working
- D. Require every one of the 40 product teams to get the CEO's personal sign-off before any deploy
Check the answer
B. A purely centralized model doesn't scale linearly with the number of teams it supports; distributing ownership through embedded or platform-style models is the standard structural fix once one central team becomes the ceiling on everyone else's velocity.
Q8 (Module 8 — SRE, Other Frameworks & the Future). An organization running a traditional ITIL-style Change Advisory Board (CAB) — where every change, regardless of size or risk, must be approved at a fixed weekly meeting — wants to start adopting SRE practice. What's the key shift in how change risk gets governed?
- A. Nothing changes; SRE and ITIL rely on exactly the same change-advisory mechanism
- B. Replace the fixed calendar/committee gate with an error-budget-driven model, where the pace and risk tolerance of change is governed by remaining error budget and automated checks rather than one weekly meeting reviewing every change alike
- C. SRE requires abolishing all change review, for changes of every size and risk
- D. SRE and ITIL are mutually exclusive and can never inform practice at the same organization
Check the answer
B. SRE doesn't remove change governance — it replaces a fixed, calendar-gated committee with a continuous, risk-proportionate one driven by the remaining error budget, which is also the direction ITIL 4 itself moved in.
Block 2 — Q9–16 (minutes 12–24)
Q9 (Module 5 — SRE Tools & Automation). A critical page goes unacknowledged for 15 minutes because the primary on-call's phone was silenced, and there was no fallback configured. Which tool feature would have prevented the silent miss?
- A. A properly configured escalation policy — in PagerDuty, Opsgenie, or similar — that automatically re-routes an unacknowledged page to a secondary responder after a set timeout
- B. Sending the page by email only, since email is inherently more reliable than a push notification
- C. Removing the on-call rotation entirely, so everyone is nominally responsible for every page
- D. Configuring the page to keep retrying the exact same phone number every five seconds forever
Check the answer
A. An escalation policy with a timeout is specifically the guardrail against exactly this failure mode — a single unreachable responder silently swallowing a page. Retrying the same unreachable phone (D) does nothing to route around the actual failure.
Q10 (Module 7 — Organizational Impact of SRE). During a severe, company-wide outage touching five different teams, the response turns chaotic: three engineers independently try to roll back three different services with no shared picture of what's already been tried. What's missing?
- A. More Slack channels — one dedicated channel per engineer involved
- B. A formally designated Incident Commander, coordinating the response, assigning clear workstreams, and keeping one shared timeline and status — exactly what incident-command structure exists to provide
- C. Nothing — during a major incident, more people independently taking parallel action is always strictly better
- D. A vote among the five affected teams to decide who is most senior
Check the answer
B. Uncoordinated parallel action during a multi-team incident is precisely the failure mode incident command exists to prevent — one person directing the response and maintaining a shared picture, rather than everyone independently guessing.
Q11 (Module 2 — SLOs & Error Budgets). A monitoring system pages the on-call engineer because, at the current failure rate, the service is on pace to burn its entire 30-day error budget within the next two hours. What does this fast-burn alert mean, and what's the right response?
- A. It signals a slow, low-priority data-quality issue that can wait for the next business day
- B. It signals that the service is failing badly enough, right now, that the whole month's error budget will be gone in hours rather than weeks — this warrants an immediate page and active incident response, not a backlog ticket
- C. It signals the SLO itself was set too strictly and should be loosened immediately
- D. It's a false positive by definition, since a burn-rate alert can't mathematically fire faster than the SLO's own measurement window
Check the answer
B. A fast-burn alert is exactly what a high burn-rate multiplier is designed to catch — severe enough right now that waiting for the slow-burn signal to fire would mean discovering the exhausted budget only after it's already gone.
Q12 (Module 8 — SRE, Other Frameworks & the Future). A company already has a platform engineering team building a self-service internal developer platform. Leadership asks whether the organization still needs SRE as a separate discipline. What's the accurate answer?
- A. No — platform engineering and SRE are identical disciplines, so staffing both is always redundant
- B. Yes, and the existing platform engineering team should be disbanded in favor of SRE
- C. They're complementary, not identical: platform engineering focuses on the self-service developer experience for building and shipping software, while SRE focuses on defining and enforcing reliability targets and operating what's running in production — many organizations run both, with SRE practice often sitting on top of what the platform team builds
- D. SRE is strictly the older discipline, and platform engineering is a passing trend with no lasting distinction from it
Check the answer
C. The two disciplines answer different questions — how easy is it to build and ship, versus how reliable is what's running — and most mature organizations staff both rather than treating them as substitutes.
Q13 (Module 1 — SRE Principles & Practices). An on-call log shows a team spent 65% of the last quarter on paging, tickets, and manual restarts, and only 35% on engineering work. Per Google's toil-ceiling guidance, what's the correct next step?
- A. Praise the team for being highly responsive to customers
- B. Nothing — the guideline only applies once a team exceeds 90% operational work
- C. Immediately place the whole team on a performance improvement plan for spending too much time on operations
- D. Treat the sustained overage as a staffing and automation signal — push the toil back to the owning team, freeze new toil-generating launches, or add headcount, rather than simply asking the team to work harder
Check the answer
D. 65% sustained above the 50% ceiling is treated as a staffing and automation-investment signal, not a performance issue and not something to tolerate silently until it reaches 90%.
Q14 (Module 6 — Anti-Fragility & Learning from Failure). A team runs a "game day" simulating a full regional cloud outage. Two hours in, the exercise reveals that the documented failover runbook is badly out of date and would have caused a second outage if followed literally. What's the correct reaction?
- A. Cancel the game day and decide never to run another one, since this one clearly "went badly"
- B. Open a disciplinary process against whoever originally wrote the outdated runbook
- C. Treat the finding as the entire point of the exercise: fix the runbook, and count the discovery as a successful investment against a future real incident rather than a failure of the exercise
- D. Conclude that failover isn't actually needed, since the documented runbook obviously doesn't work
Check the answer
C. Finding a broken runbook in a drill, at no real cost, is the exercise working exactly as intended — the alternative is discovering it live, during an actual regional outage.
Q15 (Module 4 — Monitoring & SLIs). A dashboard shows CPU and memory both comfortably under 50%, yet users report the service feels slow. Which category of signal is most likely being missed?
- A. Saturation is under-measured — raw CPU and memory percentages don't capture queueing or backlog against effective capacity, such as thread-pool queue depth or connection-pool exhaustion, which is a distinct dimension from utilization
- B. Traffic — the dashboard just needs a plain requests-per-second graph
- C. This is impossible; if CPU and memory are both low, the service cannot be slow
- D. Latency, traffic, errors, and saturation are effectively the same measurement, so any one substitutes for the others
Check the answer
A. Low raw utilization can coexist with a saturated queue or exhausted connection pool — the four golden signals are distinct dimensions precisely because a system can look healthy on three of them and be starving on the fourth.
Q16 (Module 3 — Reducing Toil). A team wants to reduce the recurring toil around certificate rotation. What's the correct order of priority?
- A. Hire a contractor to perform the manual rotation faster
- B. Write a very detailed runbook so the manual rotation is at least performed consistently, and stop there
- C. First measure the actual toil — frequency and time cost — then automate the repeatable steps end-to-end, for example with cert-manager and ACME automation, so the task disappears rather than being merely documented for a human to follow
- D. Reduce how often certificates are rotated to once a year, cutting the toil roughly in half
Check the answer
C. A runbook (B) is a real rung on the automation ladder, but it's an early one — it makes manual execution consistent, it doesn't remove the human from the loop, which is what full automation actually does.
Block 3 — Q17–24 (minutes 24–36)
Q17 (Module 6 — Anti-Fragility & Learning from Failure). A postmortem's "five whys" chain ends at "...because the on-call engineer typed the wrong command." The facilitator says this isn't a complete root cause yet. What should the team ask next?
- A. Nothing further — a specific human action is always an acceptable place to end a five-whys chain
- B. Why did the system allow that command to run without a safeguard — a confirmation prompt, a dry-run default, or tighter permission scoping — pushing the chain past the human action to the system or process gap that let a typo become an outage
- C. Whose fault it ultimately was, so the finding can be attached to that person's next performance review
- D. Whether the engineer had been at the company long enough that they really should have known better
Check the answer
B. Stopping at a human action is the single most common way a five-whys chain falls short — the systemic gap that let a routine typo reach production is the fixable cause, and it's what's still unasked.
Q18 (Module 1 — SRE Principles & Practices). A postmortem for a database outage ends with the line: "Priya applied an untested migration under time pressure and should be more careful in the future." What's wrong with this postmortem from a blameless standpoint?
- A. Nothing — identifying who made the mistake is the entire point of a postmortem
- B. It should also have recommended a specific disciplinary consequence for Priya
- C. It stops at an individual's action instead of asking what allowed an untested migration to reach production under time pressure in the first place — the systemic, fixable cause is what a blameless postmortem is supposed to surface
- D. Nothing — as long as Priya's name is used only once and the tone stays calm
Check the answer
C. A calm tone or a single mention doesn't make a postmortem blameless — what matters is whether it stops at an individual's action or keeps asking what system-level gap let that action reach production.
Q19 (Module 8 — SRE, Other Frameworks & the Future). A team starts using an ML-based anomaly-detection tool that automatically correlates alerts across services and proposes a likely root cause during an incident. What's the SRE-appropriate way to treat its suggestion?
- A. Treat the tool's output as authoritative and resolve the incident based on it alone, with no further verification
- B. Treat it as a fast, useful hypothesis worth investigating — a starting point that speeds up triage — while still verifying against real evidence before acting, since a correlation-based suggestion can be wrong
- C. Ignore the tool's output entirely, since automated suggestions are fundamentally incompatible with SRE principles
- D. Disband the human on-call rotation and let the tool run incident response unattended
Check the answer
B. AIOps-style correlation tools speed up triage without replacing verification — a correlated suggestion is a hypothesis, not a diagnosis, until it's checked against real evidence.
Q20 (Module 3 — Reducing Toil). A team reports that toil has crept up because every new microservice they onboard requires the same six manual steps to wire up monitoring, logging, and alerting. What's the most durable fix?
- A. Assign a rotating "onboarding buddy" who performs the same six steps by hand for every new service, indefinitely
- B. Build a self-service golden-path template — a scaffolding tool or reusable module — that provisions monitoring, logging, and alerting automatically for every new service, eliminating the six manual steps rather than just redistributing who does them
- C. Require every new microservice proposal to pass through a six-week approval committee instead
- D. Declare monitoring optional for new services so the six steps are no longer needed
Check the answer
B. Rotating who performs a manual task (A) doesn't reduce the toil, it just spreads it — a self-service template removes the six steps from the critical path entirely.
Q21 (Module 7 — Organizational Impact of SRE). A team is asked to push a service from 99.9% to 99.999% availability ("five nines"). What's the correct framing to bring into that conversation?
- A. Every additional "nine" costs roughly the same to achieve as the last one, so the request can be approved without further analysis
- B. Reliability should always be maximized regardless of cost, since more reliability is categorically better
- C. Each additional nine typically costs substantially more than the one before it — more redundant infrastructure, more sophisticated failover, more engineering time — so the decision should weigh that marginal cost against what the business and its users will actually notice and need
- D. "Five nines" is a meaningless marketing phrase and shouldn't factor into the conversation at all
Check the answer
C. Reliability has diminishing returns and rising marginal cost — the correct conversation weighs that cost against real user and business need, not an assumption that more is unconditionally better.
Q22 (Module 2 — SLOs & Error Budgets). A customer complains the product is "always down," despite the internal dashboard showing the service comfortably inside its 99.9% quarterly SLO. Which distinction most likely explains the mismatch?
- A. The customer is simply mistaken, and there's nothing further worth checking
- B. The SLI being measured may not reflect what the customer actually experiences — for example, a server-side success rate that excludes client-side or network errors, or that doesn't cover the specific journey the customer is complaining about — so the SLI itself is worth re-examining
- C. The customer must be describing the SLA, which by convention is always stricter than the internal SLO
- D. SLOs are strictly internal and can never have any bearing on what a customer experiences
Check the answer
B. A comfortable SLO with an unhappy customer is usually an SLI-coverage problem — the measurement isn't capturing the specific journey or failure class the customer is actually hitting.
Q23 (Module 5 — SRE Tools & Automation). A team applies the same Terraform configuration twice in a row with nothing changed in between. What should happen the second time, and why does that property matter operationally?
- A. Terraform's plan should show no changes — applying it is idempotent, since the desired state already matches reality, and that idempotency is what makes it safe to reapply the same automation repeatedly, for example from a CI pipeline, without unintended side effects
- B. Terraform should create a second, duplicate copy of every resource
- C. Terraform should fail outright and demand that the state file be deleted and reset
- D. Terraform should silently skip the second apply and log nothing at all
Check the answer
A. Idempotent applies are what let infrastructure-as-code run unattended and repeatedly — a no-op plan on the second run is the correct, expected behavior, not an edge case.
Q24 (Module 4 — Monitoring & SLIs). A request that touches eight microservices is intermittently slow, but every individual service's own dashboard looks healthy in isolation. What's the correct tool for finding where the time is actually going?
- A. Set every one of the eight services' logs to DEBUG verbosity and grep through them by hand
- B. Assume whichever service has the highest CPU graph is the culprit and stop looking further
- C. Distributed tracing — following one request's spans across all eight services via a propagated trace context, for example with OpenTelemetry — shows exactly which hop absorbed the latency
- D. There's no reliable way to diagnose this; just add more replicas everywhere
Check the answer
C. When every service looks healthy alone but the end-to-end request is slow, the problem is almost always in the gaps between services — exactly what a per-request trace, not a per-service dashboard, is built to reveal.
Block 4 — Q25–32 (minutes 36–48)
Q25 (Module 2 — SLOs & Error Budgets). A product VP insists error budgets should be set unilaterally by engineering, with no product or business input, "because reliability is purely an engineering concern." What's the correct objection?
- A. This is correct — product should never have input on where an SLO target is set
- B. SLO targets, and the error budget they produce, represent a genuine business tradeoff between reliability investment and feature velocity, so they should be negotiated jointly between engineering and the product/business owners who bear the cost of both under- and over-investing in reliability
- C. Error budgets are a purely technical implementation detail with no organizational dimension worth discussing
- D. Only the CEO is authorized to approve an SLO target, regardless of company size
Check the answer
B. An SLO target sets the tradeoff between shipping speed and reliability spend — a decision with real business consequences on both sides, which is exactly why it's negotiated jointly rather than set unilaterally by either side.
Q26 (Module 4 — Monitoring & SLIs). An engineer adds a unique user_id label to a latency histogram "for better debugging," and within a day the metrics backend runs out of memory and starts dropping data. What happened?
- A. The metrics backend is simply undersized, and nothing about the metric's design needs to change
- B. Histograms are structurally incapable of carrying any label at all, which is why it broke
- C. High-cardinality labels — like a unique value per user — multiply the number of distinct time series combinatorially, overwhelming the metrics store; that kind of per-user debugging detail belongs in logs or traces, not in a metric label
- D. This proves a pull-based metrics system can never be used to measure latency
Check the answer
C. A label with unbounded unique values multiplies the time-series count, not the label count — this is the textbook cardinality-explosion failure, and the fix is moving that detail to logs or traces instead of resizing the metrics store.
Q27 (Module 6 — Anti-Fragility & Learning from Failure). A team has never run a failure-injection experiment before and is worried about causing a real outage on day one. What's the recommended way to begin?
- A. Start immediately in production with an unbounded, company-wide failure injection, on the theory that bigger experiments produce more useful data
- B. Skip experiment design entirely and go straight to disabling the primary production database to see what happens
- C. Start small and staged — run early experiments in a non-production or canary environment, or with a tightly bounded blast radius in production — and progressively widen the scope as tooling and confidence mature
- D. Wait for a real outage to happen on its own and learn from that instead, since deliberate experimentation isn't necessary
Check the answer
C. Chaos maturity is built progressively — small, bounded experiments first, wider blast radius only once confidence and tooling justify it — not by starting at company-wide scope on day one.
Q28 (Module 1 — SRE Principles & Practices). Leadership wants engineers to "just be more careful" as the primary reliability strategy for the coming year, with no change to review process, testing, or automation. What's the core SRE objection?
- A. It's the correct approach — reliability is fundamentally a matter of individual diligence
- B. Hoping people simply won't make mistakes isn't a strategy; SRE treats recurring operational failure modes as engineering problems to be solved with process, automation, and system design, not with willpower alone
- C. It's fine as long as the engineers involved are compensated with overtime pay
- D. It's fine as long as it's formally logged as a goal in the team's project tracker
Check the answer
B. "Be more careful" with no supporting process or automation change asks humans to substitute for a system fix — SRE's core move is to treat the recurring failure as an engineering problem instead.
Q29 (Module 3 — Reducing Toil). Which of the following is the best example of toil actually worth tolerating in the short term rather than automating right away?
- A. A one-off manual step needed exactly once, for a migration that will never repeat, where the cost of building automation for a single use clearly exceeds the cost of just doing it by hand
- B. A restart script that must be run by hand every single day for a chronically flaky job
- C. A capacity check performed manually before every weekly release, for a service that will keep releasing weekly indefinitely
- D. Manually rotating a shared credential every 90 days, indefinitely, for as long as the service exists
Check the answer
A. The "repetitive" gate is what disqualifies toil-worth-tolerating from toil-worth-automating — a genuinely one-off task never repeats, so the automation investment never pays itself back, unlike B, C, or D, all of which recur indefinitely.
Q30 (Module 8 — SRE, Other Frameworks & the Future). A team's systems have grown more distributed — more microservices, more managed cloud dependencies, more third-party APIs — and their metrics-only monitoring stack increasingly can't answer "why" a specific request failed. Which broader shift does this describe, and what's the recommended direction?
- A. Reduce the number of services back down to a single monolith, since distributed systems are unmonitorable as a matter of principle
- B. Nothing needs to change; more dashboards built from the same existing metrics will always resolve the problem
- C. The shift from narrow "monitoring known failure modes" toward broader "observability" — instrumenting systems, often via open standards like OpenTelemetry, so that novel, previously unanticipated failure modes in complex distributed systems can still be explored and explained after the fact
- D. This is a permanent, unsolvable limitation of distributed systems with no recommended direction at all
Check the answer
C. This is the monitoring-to-observability shift in a nutshell: a fixed set of pre-chosen metrics can only answer questions you thought to ask in advance, while observability is built to answer the question you didn't anticipate.
Q31 (Module 7 — Organizational Impact of SRE). During an active incident, the on-call SRE realizes the "outage" is actually being caused by an ongoing credential-stuffing attack overwhelming the auth service. What's the correct organizational response?
- A. Escalate to and coordinate with the security/incident-response team immediately — reliability and security incident response converge here, and scaling capacity alone would just absorb the attack rather than address its cause
- B. Treat it purely as a capacity incident and scale the auth service up to absorb the extra attack traffic
- C. Handle it alone without informing the security team, to avoid causing unnecessary alarm
- D. Close the incident as soon as the error rate on the dashboard returns to normal, regardless of the underlying cause
Check the answer
A. Scaling capacity treats the symptom of an attack, not the attack itself — once the cause is identified as adversarial, the incident belongs jointly to reliability and security response, not to reliability alone.
Q32 (Module 5 — SRE Tools & Automation). A recurring incident type — "disk fills up on node X, evict pods, resize the volume" — has a well-tested, low-risk runbook that's been executed by hand successfully more than 20 times. What's the correct next step for that runbook?
- A. Keep running it by hand indefinitely, since it already works and shouldn't be touched
- B. Convert the proven, low-risk, repetitive runbook into automated remediation — a script or an auto-remediation controller — so it executes without paging a human at all, reserving human judgment for genuinely novel situations
- C. Make the runbook longer and more detailed so new hires can follow each step more slowly
- D. Restrict who's allowed to execute it to only the most senior engineer, to reduce risk
Check the answer
B. Twenty clean executions of a low-risk, well-understood runbook is exactly the evidence base that justifies moving it up the automation ladder to unattended remediation.
Block 5 — Q33–40 (minutes 48–58)
Q33 (Module 8 — SRE, Other Frameworks & the Future). A company's cloud bill has grown faster than its traffic, and finance asks the SRE team why maintaining "five nines everywhere" behaves like a black-box cost center. Which emerging cross-discipline practice is most directly relevant here, and what's its core idea?
- A. FinOps — bringing shared, real-time cost visibility and accountability directly into engineering decisions, including reliability and redundancy tradeoffs, so cost becomes a first-class input alongside the reliability target rather than a surprise discovered after the fact
- B. DevSecOps — shifting security scanning earlier into the pipeline, which has no direct connection to a cloud cost question
- C. ITIL 4 — a change-management framework with no built-in concept of cost accountability at all
- D. Chaos engineering — deliberately injecting failure, which is unrelated to a billing conversation
Check the answer
A. FinOps is specifically the practice of bringing cost accountability into engineering decision-making in real time — exactly the missing piece when a reliability target is set without anyone pricing what it costs to hold.
Q34 (Module 3 — Reducing Toil). After building three automations this quarter, a team's toil-ticket count drops from 200 to 40. Two engineers report they're now spending that freed-up time entirely on ad hoc, low-priority feature requests instead of reliability project work. What's the correct read of this situation?
- A. This is fine — once toil is reduced, any use of the freed time is equally valuable to the organization
- B. The freed-up time should be reinvested in the engineering half of the roughly 50/50 operations-to-engineering split — further automation, capacity work, reliability projects — rather than absorbed by unrelated ad hoc work; left unmanaged, that ad hoc work can itself become a new source of toil
- C. The team should be assigned more toil immediately, since they clearly have spare capacity now
- D. Nothing should be measured here; toil-ticket counts have no bearing on how freed time gets used
Check the answer
B. The entire point of cutting toil is to grow the engineering half of the split — letting the freed capacity drift into unrelated, unmanaged ad hoc work quietly defeats that purpose and can seed a new toil source.
Q35 (Module 5 — SRE Tools & Automation). A platform team's Grafana dashboards are updated by hand-editing them directly in the production UI, with someone "supposed to" export the JSON back into the team's Git repository afterward — a step that routinely gets forgotten. What's the correct fix?
- A. Manage dashboards (and alerting rules) as code in the same version-controlled repository as everything else — provisioning them into Grafana via files or its API through CI — so a UI edit is never the source of truth and can't silently drift from what's actually deployed
- B. Ban all UI edits to Grafana outright, with no alternative workflow provided
- C. Accept the drift as unavoidable, since dashboards are considered too low-stakes to version-control
- D. Switch the team off Grafana entirely, since the UI is the source of the problem
Check the answer
A. Dashboards-as-code removes the manual "remember to export" step by making the repository, not the live UI, the source of truth — the same discipline applied to every other piece of production configuration.
Q36 (Module 2 — SLOs & Error Budgets). A checkout flow depends on three internal services — cart, payments, and inventory. The team defines one end-to-end SLO for "checkout succeeds" instead of three separate SLOs for each service. What's the main advantage of this composite approach?
- A. It hides which internal service is actually failing, which is desirable so users can't tell
- B. It reflects what the user actually experiences: a single failing dependency still breaks checkout for the user, so a composite (or dependency-aware) SLO won't overstate reliability the way three services' individual numbers, viewed in isolation, can
- C. A composite SLO is always mathematically identical to summing the three individual services' SLOs
- D. It removes any need to keep monitoring the three underlying services individually
Check the answer
B. Three individually healthy-looking services can still combine into a broken user journey — a composite SLO measures the thing the user actually experiences instead of three separate numbers that can each look fine in isolation.
Q37 (Module 4 — Monitoring & SLIs). A team standardizes on the RED method (Rate, Errors, Duration) for every service's dashboards and the USE method (Utilization, Saturation, Errors) for the resources underneath. During an incident, RED looks fine on the affected service but a downstream queue is badly saturated. Why does keeping both views matter here?
- A. RED and USE measure exactly the same thing from the exact same vantage point, so keeping both is pure redundancy
- B. RED describes what a service is doing from the request/consumer's point of view; USE describes the health of the resources underneath it — a service can look fine on RED while a resource it depends on is saturated, and vice versa, so both views are needed to fully triage an incident like this one
- C. The USE method only applies to Kubernetes clusters and has no meaning for a queue
- D. RED is a method for analyzing logs; USE is a method for analyzing traces
Check the answer
B. This exact split — clean RED, saturated resource underneath — is the textbook case for why service-level and resource-level views are tracked separately rather than assumed to move together.
Q38 (Module 7 — Organizational Impact of SRE). A manager says, "We already do blameless postmortems — we hold the meeting after every incident." A reviewing SRE points out this alone doesn't prove the culture is actually blameless. What additional evidence would actually demonstrate it?
- A. Whether the postmortem meeting has a fixed 30-minute slot reserved on the shared calendar every time
- B. Whether the resulting postmortem document is stored in a shared drive everyone can access
- C. Whether engineers actually volunteer uncomfortable details about their own actions without fear that it will affect a performance review or promotion — genuine blamelessness shows up in psychological safety and honest self-reporting, not merely in a meeting existing on the calendar
- D. Whether more than five people attend the scheduled meeting
Check the answer
C. A meeting existing on the calendar proves a process was followed, not that it's safe to be honest inside it — the real test is whether engineers volunteer the uncomfortable details without fear of career consequences.
Q39 (Module 6 — Anti-Fragility & Learning from Failure). Six months after a major outage, the postmortem's tracked action items are still open — they keep losing priority to feature work. What does this reveal, and what's the correct organizational fix?
- A. Nothing meaningful; open action items after six months are completely normal and require no response
- B. This is "postmortem theater" — the process looks right on paper, but no fix actually lands and the system remains exactly as fragile as the incident found it; the fix is giving postmortem action items real prioritization weight, for example tracked against the error budget or a protected reliability-capacity allocation, rather than letting them compete unprotected against the feature backlog
- C. The postmortem document should be deleted, since the team clearly isn't going to act on it
- D. The action items should be reassigned to a single engineer to complete alone, outside of sprint planning
Check the answer
B. A correct root cause with no completed fix is "postmortem theater" — the paperwork is right, the system is exactly as vulnerable as before, and the durable fix is protected prioritization, not more documentation.
Q40 (Module 1 — SRE Principles & Practices). A candidate paraphrases Ben Treynor Sloss's definition of SRE as "what happens when you ask a software engineer to design an operations function." A team lead asks what that framing actually implies day to day. What's the correct operational implication?
- A. It's a memorable slogan with no real operational implication for how a team behaves day to day
- B. It implies operators should be replaced entirely by engineers with no operations experience at all
- C. It implies that recurring manual operational work becomes something to eliminate with reviewed, tested software, rather than something to repeat by hand — the same standard applied to any other production code
- D. It implies SRE only applies to companies the size of Google
Check the answer
C. The operational payoff of "software engineer designs the operations function" is that recurring manual work gets treated as a bug to fix in code — reviewed, tested, version-controlled — not repeated by hand indefinitely.
Score yourself
☺ Like you're 10: Count how many letters you got right out of forty, turn it into a percentage against 65%, then look at which module — and which kind of mistake — actually cost you the points.
Total your correct answers out of 40 for your raw percentage against the 65% pass mark (26 correct or more). Then total each module separately out of its five questions, exactly as you did for Set 1 — a 30/40 spread evenly across eight modules and a 30/40 built from seven near-perfect modules plus a zero on one are very different results on the real exam, even though they score identically here.
| Module | Questions in this paper | Your score | If under 4/5, go here |
|---|---|---|---|
| 1 · SRE Principles & Practices | Q1, Q13, Q18, Q28, Q40 | /5 | SRE Principles & Practices |
| 2 · SLOs & Error Budgets | Q2, Q11, Q22, Q25, Q36 | /5 | Service Level Objectives & Error Budgets |
| 3 · Reducing Toil | Q3, Q16, Q20, Q29, Q34 | /5 | Reducing Toil |
| 4 · Monitoring & SLIs | Q4, Q15, Q24, Q26, Q37 | /5 | Monitoring & Service Level Indicators |
| 5 · SRE Tools & Automation | Q5, Q9, Q23, Q32, Q35 | /5 | SRE Tools & Automation |
| 6 · Anti-Fragility & Learning from Failure | Q6, Q14, Q17, Q27, Q39 | /5 | Anti-Fragility & Learning from Failure |
| 7 · Organizational Impact of SRE | Q7, Q10, Q21, Q31, Q38 | /5 | Organizational Impact of SRE |
| 8 · SRE, Other Frameworks & the Future | Q8, Q12, Q19, Q30, Q33 | /5 | SRE, Other Frameworks & the Future |
| Total | 40 questions | /40 | 65% (26/40) to pass |
A scenario paper produces a third kind of miss that Set 1 mostly couldn't. Sort what you got wrong into three piles. A genuine content gap means you didn't know the underlying mechanism at all — reread the linked module. The "sounds confident" trap means you knew the material but picked the option that read as decisive, cautious, or authoritative rather than the one that actually matched the mechanism the scenario was testing — that's not a knowledge gap, it's a habit of trusting tone over substance, and the fix is simply sitting more scenario papers and asking "does this option address the actual cause, or does it just sound like the kind of thing a careful person would say?" before you commit. A misread means the scenario contained a specific detail — how much budget was left, who'd already been notified, what role was already assigned — that you skimmed past; the fix there is the "read the whole stem first" habit from earlier on this page, not more content review.
For deeper, module-specific reps beyond a full 40-question sitting, this course's practice banks split the same ground into focused sets: Practice · SRE Principles, SLOs & Toil, Practice · Monitoring & SRE Tools, and Practice · Anti-Fragility & Organizational Impact — plus the full SREF Practice Questions bank and Answer Triage for the general skill of ruling out an option that merely sounds right, which is exactly the trap this paper spends most of its forty questions on.
Timmy the Turtle: Score?
Foxy: 34. But six of my misses cluster in exactly one place — Organizational Impact.
Timmy the Turtle: Which one, specifically?
Foxy: The team-topology bottleneck question. I picked "hire more into the same centralized team," because it sounded like the safe, cautious answer.
Timmy the Turtle: That's the trap on this whole paper. A scenario question always has one option that sounds careful and conservative and is actually just more of the thing that's already broken.
Rocky the Raccoon: Meanwhile I got the chaos-engineering one wrong — I picked "no planning needed, that's the whole premise." In my defense, that's basically my entire personality.
Timmy the Turtle: Which is exactly why blast radius exists, Rocky. Even you plan the boundary before you touch anything.
Sol the Sloth: The pattern's the same in both cases, though — you both answered fast, on vibes, instead of checking whether the option actually matched the mechanism.
Professor Owl: Which is the whole difference between Set 1 and Set 2. Set 1 asked what a word means. Set 2 asks whether you'd reach for the right lever under a running clock — and confidence isn't a substitute for checking.
1. What's the one thing that actually changed between Set 1 and Set 2, given that the module count, questions per module, timing, and pass mark all stayed identical? 2. Name one of the three common ways a candidate gets a scenario question wrong, beyond simply "never learned the term." 3. Why does a scenario stem typically deserve a slower first read than a straight definitional one, even though the clock still averages 90 seconds a question? 4. If your module tally shows one module noticeably weaker than the rest, what should you do before sitting Set 3 or Set 4?
Check your answers
- The structure stayed identical to Set 1 — eight modules, five questions each, one 60-minute sitting, 65% pass mark. What changed is the question's shape: almost every stem now describes a short situation and asks what a correctly-run SRE practice does next, instead of asking you to define or classify a term directly.
- Any of: a genuine content gap (never learned the underlying mechanism); the "sounds confident" trap (picking the option that reads as decisive or cautious rather than the one that actually matches the mechanism); or a misread (missing a specific stated detail in the scenario, such as time remaining, who's already been notified, or a role that's already assigned).
- Because the trap in a scenario question is usually hidden in one specific detail inside the situation, not in the wording of the options — skimming the setup and jumping straight to pattern-matching an answer is exactly how a candidate who knows the material still loses the point.
- Re-read that module's linked blueprint page, work the matching practice bank, and re-sit just that module's five questions cold in a day or two before moving on — fixing one soft module now is cheaper than carrying it forward into Set 3 or Set 4.
That's the paper. Set the timer once, keep every scenario's full detail in view before you touch the options, and let the forty questions above show you whether Set 1's vocabulary has actually turned into judgment. Score it by module, fix only what's soft, and when you're ready, Set 3 is a completely different kind of paper — the same four confusable term-pairs, over and over, until they stop being a coin flip.