SREF Mock Exam · Set 5
This is the fifth and final practice paper in this course's SREF set — forty questions, sat once, cold, under a single 60-minute timer that starts the moment you begin and doesn't stop for anything. If you've already worked through Set 1 through Set 4, this one won't ease you in: it's built to the exact difficulty, pacing, and single-best-answer format of the real SRE Foundation (SREF) exam, with questions spread evenly across all eight syllabus modules exactly the way Module 1 through Module 8 taught them. Nothing here reruns an earlier set's wording. Treat a strong score as your signal to go book the real exam this week — and treat anything else as a precise map of the one or two modules that still need one more pass before you do.
Imagine you've had four practice driving lessons with an instructor sitting next to you, ready to grab the wheel the moment you drift. This is the fifth lesson — except the instructor stays completely silent for the whole hour, and you don't find out how you did until after you've parked the car. That silence is the entire point. It's the closest thing to test day you can get without actually being on test day, and it's your last honest chance to find out whether you're really ready, or just used to having help nearby.
Where Set 5 fits
☺ Like you're 10: Five practice tests, each one harder to lean on than the last — this is the one you sit with nothing left to lean on at all.
Sitting Set 5 cold only tells you something useful if the four stages before it actually happened first. Skip straight here from the blueprint modules and a low score just tells you what re-reading the modules would have told you anyway, with none of the pacing information a timed paper is actually for.
| Stage | What you do | What it tells you |
|---|---|---|
| 1 · Learn | Work through all eight blueprint modules, Module 1 through Module 8. | Whether you hold the vocabulary and the arithmetic at all. |
| 2 · Drill | SREF Practice Questions, plus the three themed banks — SLOs & toil, monitoring & tools, resilience & culture. | Whether you can apply it to a described scenario, untimed. |
| 3 · Sit Sets 1–4 | Set 1 through Set 4, spaced a few days apart, timed. | Where the specific gaps are, with time left to fix them before it counts. |
| 4 · Sit Set 5 (this page) | Cold. Once. Exactly 60 minutes. No notes, no re-reads mid-sitting. | Your honest, final readiness reading. |
| 5 · Book it | The SREF Exam — logistics, format, and what to expect on the day. | Turns a readiness reading into a scheduled date. |
If Set 5 comes back below your target, the fix is not another cold sitting of this same paper — a second attempt mostly measures how well you remember this specific paper, not your actual readiness. Go back to the one or two modules where your misses clustered, re-drill them properly, and only then consider a fresh cold sitting before you book.
Sit it exactly like the real thing
☺ Like you're 10: One timer, no notes, no tabs, and an answer for all forty — even the ones you're guessing on.
A mock paper only measures what it claims to measure if you sit it under the real constraints. Loosen any one of these and you're no longer measuring exam readiness — you're measuring something more comfortable and considerably less useful.
- One 60-minute timer, started once. You don't pause it for a question you want to think about later, a phone call, or a coffee. If you lose four minutes to something unrelated, that's data about how the real hour will actually feel, not an excuse to reset the clock.
- Genuinely closed-book. No notes, no other tabs on this site, no search engine, and no AI assistant — the real SREF exam gives you nothing but the question and four options, and that's exactly the condition this sitting needs to reproduce. If a term feels fuzzy mid-question, that fuzziness is the result; don't resolve it by looking anything up.
- Single-best-answer, no partial credit. Every question below has exactly one correct option. Choosing a defensible-but-second-best answer scores the same as choosing an obviously wrong one — there's no partial credit to chase.
- Answer every question. Nothing published indicates the SREF exam deducts marks for a wrong answer — treat that as an assumption worth confirming officially rather than a guaranteed fact — but the advice holds regardless: a blank answer is a guaranteed zero, and an elimination-based guess between two plausible options costs you nothing and sometimes pays off.
- Don't open a single answer key until all 40 are answered. Each question below folds its explanation behind a summary, exactly like the sample questions on the blueprint pages — opening one mid-sitting converts a readiness test into a reading exercise, and you won't get an honest number back.
- Read "which is NOT" and "BEST represents" stems twice. A few questions on this paper invert the usual shape or ask you to pick the closest fit among several plausible options. Skimming those is the single most common way candidates lose points they actually knew.
This page uses 40 questions, 60 minutes, and a 65% pass mark throughout, matching what's stated across this course's own SREF blueprint pages. Those figures, along with price, retake policy, and eligibility windows, are set by the DevOps Institute and can be revised — verify the current specifics on DevOps Institute's own SRE Foundation page before you register, rather than trusting any third-party page, this one included.
How the 40 questions are spread across the eight modules
☺ Like you're 10: Eight topics, five questions each — nobody publishes an official split, so this paper uses the fairest one there is: equal.
DevOps Institute doesn't publish a confirmed per-module breakdown of the SREF exam's 40 questions the way some certification bodies publish domain weights — there's no legitimate "Module 6 is worth 17% of the exam" figure to build a study plan around, whatever number you might see quoted secondhand. This paper's five-questions-per-module split isn't a claim about the real exam's actual distribution; it's a deliberate study convention that mirrors the same advice given on every blueprint page — budget your attention roughly evenly across all eight modules, because any of them is fair game.
| # | Module | Questions | Blueprint page |
|---|---|---|---|
| 1 | SRE Principles & Practices | 5 | Module 1 |
| 2 | Service Level Objectives & Error Budgets | 5 | Module 2 |
| 3 | Reducing Toil | 5 | Module 3 |
| 4 | Monitoring & Service Level Indicators | 5 | Module 4 |
| 5 | SRE Tools & Automation | 5 | Module 5 |
| 6 | Anti-Fragility & Learning from Failure | 5 | Module 6 |
| 7 | Organizational Impact of SRE | 5 | Module 7 |
| 8 | SRE, Other Frameworks & the Future | 5 | Module 8 |
| Total | 40 | All eight modules, in syllabus order | |
The order on the page below is deliberately interleaved rather than grouped — Module 1, then 2, then 3, all the way through 8, repeated five times — so no two neighboring questions ever share a module. That mirrors a real risk on exam day: nothing guarantees the actual paper groups its questions by topic, so a Module 2 error-budget calculation can sit directly next to a Module 7 org-design question with no warning. Training that switch is part of what a cold, timed sitting is actually for.
The paper — 40 questions, cold
☺ Like you're 10: Forty questions, one shot — pick an answer for every single one before you fold back a single summary to check.
Work straight through in order. Each question names the module it's drawn from so you can tally your misses by module afterward, exactly the way the scoring table further down expects — that label carries no meaning during the sitting itself, since the real exam gives you no such hint. Choose an option, write it down (or hold it in your head), and move on. Only after you've committed to all 40 should you start folding back the "Check the answer" summaries below.
Q1 · Module 1, SRE Principles & Practices. Ben Treynor Sloss founded Google's SRE function around 2003, describing the resulting discipline in a single sentence the SREF exam likes to quote or paraphrase as a distractor. Which of the following is closest to how he actually described it, and which came first chronologically — Google's SRE practice or the term "DevOps"?
- A. "SRE is what happens when you ask a software engineer to design an operations function" — and SRE predates the term "DevOps" by roughly six years.
- B. "SRE is what happens when you ask an operations engineer to think like a software architect" — and DevOps predates SRE by roughly six years.
- C. "SRE is what happens when you ask a software engineer to design an operations function" — and DevOps predates SRE by roughly three years.
- D. "SRE is what happens when you automate every operational task a human used to do" — and the two terms were coined in the same year.
Check the answer
A. That's Treynor Sloss's own phrase, and SRE (c. 2003) predates the term "DevOps" (coined 2009, devopsdays Ghent) by roughly six years. B reverses both the quote and the chronology. C gets the quote right but the chronology backwards. D is a plausible-sounding paraphrase, not the actual quote, and its "same year" claim is false.
Q2 · Module 2, Service Level Objectives & Error Budgets. A checkout service's engineering team instruments "% of checkout requests returning a 2xx status within 300ms" and tracks it continuously. Separately, engineering and product agree internally to hold that measurement above 99.9% over a rolling 28-day window. Which term correctly names the internally agreed target, and who is primarily responsible for setting it?
- A. SLI — set unilaterally by product management
- B. SLO — set jointly by engineering and product
- C. SLA — set unilaterally by the SRE team
- D. Error budget — set jointly by legal and business
Check the answer
B. An internal target value for an SLI, over a defined window, set jointly by engineering and product, is exactly the SLO. The measurement itself is the SLI, not the target. An SLA is external and contractual; nothing in this scenario is external. The error budget is derived arithmetically from the SLO, not "set" as a separate agreement.
Q3 · Module 3, Reducing Toil. Every production release at a company requires a senior engineer to manually read the diff and sign off before deploy. This happens on every single release, by hand, and has for two years. Using the six-property toil test, is this toil?
- A. Yes — it's manual and repetitive, which is enough to qualify
- B. No — it fails the repetitive gate, because each diff is different
- C. Yes — it's toil because it's tedious and engineers dislike doing it
- D. No — it fails the automatable gate, because it requires genuine engineering judgment a script can't replicate
Check the answer
D. Manual and repetitive alone aren't sufficient — a task must clear all six gates. A genuine judgment call fails the automatable gate even on every release; "tedious" isn't a gate in the definition at all. Each diff being textually different doesn't fail "repetitive" either — the recurring task (review-and-sign-off) is what repeats, not the literal content.
Q4 · Module 4, Monitoring & Service Level Indicators. Of Google's four golden signals — latency, traffic, errors, and saturation — which one is essential to monitor and alert an engineer on, but should almost never be the direct answer to "name a good SLI," and why?
- A. Saturation, because users don't directly experience your CPU or queue depth — they experience the latency and error consequences a moment later
- B. Errors, because most errors are transient and self-correct without user impact
- C. Traffic, because demand alone says nothing about whether users are having a good experience
- D. Latency, because it can't be expressed as a ratio of good to valid events
Check the answer
A. Saturation is a resource-side signal, not a user-facing one — nobody feels your CPU utilization directly, only its downstream effect on latency and error rate a moment later. It stays essential as a dashboard and alerting signal, just not as the SLI itself; latency, in fact, is one of the most common SLIs (p95/p99 under a threshold).
Q5 · Module 5, SRE Tools & Automation. An on-call engineer types /restart checkout-prod in Slack, which triggers a webhook that runs a tested restart script. What rung of the automation maturity ladder does this describe?
- A. Rung 1 — plain runbook, no automation
- B. Rung 2 — runbook automation: a human still decides when to run it, but the mechanical steps are compiled into one action
- C. Rung 3 — human-approved remediation: the system detected the anomaly and proposed the fix
- D. Rung 4 — autonomous remediation: the system decided and acted without a human
Check the answer
B. A human is still noticing the problem and deciding to trigger the fix — the ChatOps command just compiles the mechanical steps into one tested action instead of manual typing. It's not rung 1, because a script genuinely runs; it's not rung 3 or 4, because nothing here detects the anomaly or proposes a fix on its own — a human did both.
Q6 · Module 6, Anti-Fragility & Learning from Failure. In Nassim Taleb's three-way fragile/robust/antifragile framework, which statement is correct?
- A. Antifragile is the opposite of fragile; robust is the neutral middle where a stressor does essentially nothing at all
- B. Robust is the opposite of fragile — both describe the two ends of the same spectrum
- C. Fragile and antifragile are synonyms for the same underlying property, viewed from different angles
- D. Robust and antifragile are identical states that only diverge under extreme, one-time shocks
Check the answer
A. This is the module's single most tested trap. Antifragile — which actively gains capability from a stressor — is the true opposite of fragile, which is actively harmed by one. Robust sits in the neutral middle. Robust and antifragile look identical after one event and only visibly diverge across many repeated stress events over time, not any single "extreme" shock.
Q7 · Module 7, Organizational Impact of SRE. Which of the following is NOT one of the five common triggers that lead an organization to formally adopt SRE?
- A. Hypergrowth outpacing the org's ability to hire traditional operators
- B. A high-profile outage that damages customer trust or revenue
- C. A mandate from an ISO 27001 auditor requiring a named reliability function
- D. Enterprise/SaaS contracts increasingly requiring a published, defensible uptime SLA
Check the answer
C. The five recognized triggers are hypergrowth, an architectural shift (monolith to microservices), competitive/contractual pressure, a wake-up-call outage, and retention (burnout in an overloaded rotation). An ISO 27001 audit mandate isn't one of them — a plausible-sounding distractor built from a real compliance concept that doesn't appear in the SREF adoption taxonomy.
Q8 · Module 8, SRE, Other Frameworks & the Future. ITIL 4's Service Level Management practice and SRE's SLIs/SLOs/error-budget mechanism both define a target for service quality. What is the real difference the SREF exam expects you to name?
- A. ITIL doesn't allow numeric targets at all, while SRE requires them
- B. SRE only applies to cloud-native services, while ITIL's Service Level Management only applies to on-premises infrastructure
- C. ITIL's targets are always stricter than SRE's, because ITIL predates cloud computing
- D. ITIL's SLAs are typically reviewed on a calendar cadence; SRE's SLOs are continuously measured and enforced by an automated policy the instant the budget hits zero
Check the answer
D. Both frameworks define a target — the real gap is enforcement timing. ITIL's SLAs are commonly negotiated documents reviewed quarterly or annually; SRE's SLOs are measured continuously against live telemetry, with an automated error-budget policy triggering the moment the budget is exhausted. The other three options are fabricated distinctions with no basis in either framework.
Q9 · Module 1, SRE Principles & Practices. In the "class SRE implements interface DevOps" framing, what does the "interface" represent, and what is one concrete mechanism the SRE "class" adds that DevOps as a philosophy never mandates?
- A. The interface represents SRE's tooling; the added mechanism is CALMS
- B. The interface represents a cultural contract (CALMS) with no mandated mechanism; the added mechanism is a formal, numeric error-budget policy
- C. The interface represents a specific vendor's product suite; the added mechanism is continuous deployment
- D. The interface represents SRE's org chart; the added mechanism is Agile sprint cadence
Check the answer
B. DevOps is the interface — a cultural contract (Culture, Automation, Lean, Measurement, Sharing) with no mandated mechanism. SRE is the class that must satisfy that contract but fills it in with concrete, testable mechanisms — a formal error-budget policy, a numeric toil ceiling, or mandatory blameless postmortems are all valid examples; DevOps mandates none of them.
Q10 · Module 2, Service Level Objectives & Error Budgets. A service holds a 99.95% availability SLO, measured over a rolling 30-day window (43,200 minutes). Approximately how much downtime does its error budget allow for that window?
- A. 4.32 minutes
- B. 43.2 minutes
- C. 21.6 minutes
- D. 432 minutes
Check the answer
C. (100% − 99.95%) × 43,200 = 0.05% × 43,200 = 21.6 minutes. A is the budget for 99.99%, B is the budget for 99.9%, and D is the budget for 99% — a full "one nine" off in either direction is this arithmetic's classic distractor set.
Q11 · Module 3, Reducing Toil. A team holds a recurring 30-minute weekly meeting to plan next quarter's headcount requests. It happens every week, is arguably tedious, and nobody particularly enjoys it. Is this toil?
- A. No — it's overhead, because it isn't operational work tied to running the live production service at all, so it never reaches the six-gate test
- B. Yes — it's manual, recurring, and adds no lasting value on its own
- C. Yes, but only because it fails the "automatable" gate
- D. No — it's engineering project work, because planning has enduring value
Check the answer
A. The prior question that sorts a task before you'd ever run the six gates is: is this about operating the production service, or not? Headcount planning would exist regardless of whether the service is in production — that's overhead, disqualified before the six-gate checklist is even relevant. It's not engineering project work either, since that category is about building things with enduring value to the service itself.
Q12 · Module 4, Monitoring & Service Level Indicators. A synthetic HTTP probe hits /checkout every 30 seconds from three external regions and gets no response, but every internal application metric — CPU, memory, error rate — looks completely healthy. What kind of monitoring caught this, and what's the likely cause it's uniquely positioned to catch?
- A. White-box monitoring; a slow database query
- B. White-box monitoring; a code path throwing an unhandled exception
- C. Black-box monitoring; a memory leak in the checkout service
- D. Black-box monitoring; an infrastructure-layer failure the app itself can't see, such as an expired TLS certificate or a broken load balancer
Check the answer
D. A black-box probe sees exactly what a real external user would see, catching problems white-box instrumentation inside your own code structurally can't — DNS misconfiguration, a broken load balancer, or an expired TLS certificate, all of which can fail a request before it reaches code that would log a clean internal metric. That mismatch — external failure, healthy internal metrics — is exactly the signature being tested.
Q13 · Module 5, SRE Tools & Automation. A team is about to promote an automated remediation system from rung 3 (human-approved) to rung 4 (fully autonomous, no approval gate). Which of the following is the strongest justification for requiring guardrails — blast-radius limits, an audit log, a kill switch — before making that promotion?
- A. Rung 4 systems are legally required to have them in most jurisdictions
- B. Guardrails are only needed if the system controls financial transactions specifically
- C. Removing the human from the loop also removes the human who would have caught a mistake before it caused real damage — the lesson widely drawn from Knight Capital's 2012 trading incident
- D. Rung 3 systems never make mistakes, so guardrails only matter once autonomy is added
Check the answer
C. The widely cited cautionary tale is Knight Capital's August 2012 incident, where a bad deployment left old test code live on one server; with no human reviewing trades in real time, the firm lost roughly $440–460 million in about 45 minutes. The lesson isn't "automation is dangerous" — it's that removing the human also removes the safety net that would have caught the mistake, exactly why guardrails aren't optional on a rung-4 system.
Q14 · Module 6, Anti-Fragility & Learning from Failure. Rocky the Raccoon kills one checkout-service instance in production. Traffic reroutes cleanly and nobody notices. What does this single clean result actually demonstrate?
- A. That the checkout service is now antifragile, because it survived a real production failure
- B. That the checkout service is robust to that specific fault at that specific blast radius, on that specific day — nothing about its future tolerance has changed yet
- C. That chaos engineering is unnecessary for this service going forward
- D. That the checkout service's error budget should be increased
Check the answer
B. A single clean pass only proves robustness to that specific fault, at that specific scope, on that specific day. It becomes evidence of antifragility only once a program shows a trend — failed hypotheses root-caused, fixed, and re-verified, so the fleet's tolerance for a whole class of failure measurably climbs over many repeated experiments. One passing result is the classic "robust vs. antifragile" trap.
Q15 · Module 7, Organizational Impact of SRE. A company's central SRE team spends six weeks embedded with the billing team, helping them define their first SLOs, build dashboards, and write runbooks — then formally hands ownership back to billing and moves on to the next team. Which adoption model does this describe?
- A. Embedded — because SREs were physically working inside the billing team
- B. Centralized/Platform — because a central SRE org was involved
- C. Hybrid — because it combines central expertise with local team work
- D. Consulting — the engagement has a defined start and end, and ownership is explicitly transferred back
Check the answer
D. The defining fact is permanence, not the presence of a central team or working closely with billing engineers. Consulting is explicitly temporary — anchored to an engagement, ending with a formal handback — which is exactly what's described. Embedded means permanent placement inside one team, which this isn't.
Q16 · Module 8, SRE, Other Frameworks & the Future. DORA's research originally identified four key metrics split across throughput and stability. In its 2021 report, DORA added a fifth metric. What is it, and how is it typically measured?
- A. Reliability — whether a team meets its own user-defined operational targets, measured through practices like SLOs and observability
- B. Security — measured by counting CVEs patched per quarter
- C. Cost efficiency — measured by cloud spend per deployment
- D. Developer satisfaction — measured by an internal engagement survey
Check the answer
A. DORA added Reliability in 2021, measured through exactly the mechanism this exam's Module 2 builds: SLIs, SLOs, and the error budget they create. It's a notable convergence point — two independent research traditions, DORA's delivery-performance statistics and Google's own SRE practice, arrived at the same underlying answer.
Q17 · Module 1, SRE Principles & Practices. Which of the following correctly matches a governing mechanism to the practice that uses it, per the SREF blueprint's three-way SRE/DevOps/traditional-ops comparison?
- A. Traditional ops: SLOs with an enforced error budget
- B. SRE: SLOs with an explicit, enforced error budget
- C. DevOps: a numeric toil ceiling of roughly 50%
- D. DevOps: uptime targets held informally, backed by change-control boards
Check the answer
B. SRE's governing mechanism is precisely SLOs with an explicit, enforced error budget. D actually describes traditional ops, mislabeled as DevOps; DevOps mandates no specific governing mechanism at all — CALMS is principle-level, with no numeric toil ceiling or formal error budget of its own.
Q18 · Module 2, Service Level Objectives & Error Budgets. A team's error budget is fully consumed twelve days before the window rolls forward. According to a properly implemented error-budget policy, what happens next?
- A. The SLO is lowered so the team is back in compliance immediately
- B. The on-call engineers involved in the incidents are reassigned as a corrective measure
- C. New feature releases and other budget-spending risk are frozen, and engineering effort is redirected to reliability work until the budget recovers or an exception process is invoked
- D. Nothing — the window will roll forward on its own regardless of current behavior
Check the answer
C. This is exactly what a pre-agreed error-budget policy exists to enforce. Lowering the SLO to match the failure defeats the entire purpose of having a target and is the single most tested wrong answer in this module; reassigning individuals is punitive and has no place in a blameless framework; ignoring the trigger is functionally the same as having no policy at all.
Q19 · Module 3, Reducing Toil. An SRE spends a week designing and building a new internal dashboard that will serve every team on the platform going forward. Is this toil?
- A. Yes, because it's a recurring type of work SREs are often asked to do
- B. Yes, because building dashboards is manual and repetitive by nature
- C. No, because dashboards are automated by definition
- D. No — it's engineering project work: it has enduring value, and a well-built dashboard doesn't scale linearly in effort with the size of the fleet it serves
Check the answer
D. This is the opposite of toil on purpose. Engineering project work has enduring value (the dashboard keeps paying off after it ships) and doesn't scale linearly with the service's growth (a well-built tool serves 10x the fleet without 10x the effort) — exactly the properties toil's six-gate test requires to be absent.
Q20 · Module 4, Monitoring & Service Level Indicators. A nightly batch ETL pipeline doesn't produce a clean stream of individually countable requests the way an HTTP API does. Which SLI aggregation shape is the better fit, and why?
- A. Windows-based — bucket time into fixed intervals and score each whole window good or bad against a threshold, since the workload isn't naturally made of discrete request events
- B. Request-based, because it's simpler to implement regardless of the workload shape
- C. Request-based, because batch pipelines always produce more requests than APIs
- D. Neither shape applies to batch pipelines; only synthetic black-box probes work here
Check the answer
A. Windows-based SLIs exist for exactly this case — systems that don't produce a clean stream of discrete, individually countable events, like a batch pipeline or a streaming job. Time gets bucketed into fixed intervals, and each window is scored good or bad as a whole, producing good windows ÷ total windows.
Q21 · Module 5, SRE Tools & Automation. A tool computes rolling-window SLI compliance against a target, tracks burn rate, and fires multi-window burn-rate alerts. Which toolchain category does this describe, and which pair of tools are representative examples?
- A. Metrics & monitoring — Prometheus and Grafana
- B. SLO / error-budget tracking — Nobl9 and Sloth
- C. Distributed tracing — Jaeger and Zipkin
- D. Chaos engineering — Gremlin and Litmus
Check the answer
B. Computing rolling-window SLI compliance and firing burn-rate alerts is the SLO/error-budget-tracking category's specific job — deliberately error-prone to hand-roll correctly, which is why dedicated tools exist. Prometheus and Grafana are metrics & monitoring (the raw telemetry these tools consume); Jaeger and Zipkin trace individual requests; Gremlin and Litmus inject failure deliberately.
Q22 · Module 6, Anti-Fragility & Learning from Failure. A chaos experiment finds that a service's retry logic has no backoff, which would make an instance loss much worse at scale. The finding is written up with a clear root cause, but no engineer is ever assigned to fix it and no deadline is set. What is this outcome called, and what does it leave the system as?
- A. A successful postmortem — the root cause was correctly identified, which is the main goal
- B. An antifragile improvement, since the gap is now documented for the future
- C. "Postmortem theater" — a document with a correct root cause but no actual fix, leaving the system exactly as fragile as the experiment found it
- D. Acceptable, as long as the finding is shared in the next all-hands meeting
Check the answer
C. Practitioners call this "postmortem theater" — good root-cause analysis with no tracked, owned, deadline-bound action item, leaving the system exactly as fragile as before. Antifragility requires the loop to close: the gap has to actually get fixed and the fix re-verified, not merely written down or announced.
Q23 · Module 7, Organizational Impact of SRE. What is the single defining fact that separates the centralized/platform adoption model from the consulting model, even when the same central team is running both?
- A. Centralized teams are larger than consulting teams
- B. Consulting teams never touch production systems directly
- C. Centralized teams report to engineering; consulting teams report to product
- D. Permanence — centralized is an ongoing service relationship; consulting is explicitly temporary and ends with ownership formally transferred back
Check the answer
D. Both models can involve a central SRE function, which is exactly why this pair is the exam's favorite point of confusion. The defining fact is permanence: centralized is ongoing with no planned end date; consulting has a defined start and end, with ownership explicitly handed back once the engagement closes.
Q24 · Module 8, SRE, Other Frameworks & the Future. A team holds a sprint retrospective every two weeks, regardless of whether anything went wrong, to discuss whatever the team wants to raise. Is this the same thing as a blameless postmortem?
- A. No — a retrospective is cadence-triggered with no formal root-cause requirement; a postmortem is incident-triggered and mandates a structured root-cause method plus tracked action items
- B. Yes — both are meetings where a team reflects on recent work
- C. Yes, as long as the retrospective happens after an incident
- D. No — retrospectives are exclusively about individual performance reviews
Check the answer
A. This is the exam's favorite Agile/SRE confusion. A retrospective happens on a fixed cadence whether or not anything went wrong, with no formal root-cause mandate. A postmortem is triggered by an incident crossing a defined severity threshold and requires a structured method (often five-whys) plus tracked, owned, deadlined action items.
Q25 · Module 1, SRE Principles & Practices. A company renames its operations team "Site Reliability Engineering" but changes nothing else — no SLOs, no error budget, no toil ceiling, no formal blameless postmortem process. According to the SREF blueprint, what is this team actually practicing?
- A. SRE, because the job title is the defining characteristic
- B. DevOps, because any team that works closely with developers counts as DevOps
- C. Traditional operations wearing a new job title — the mechanisms, not the label, are what make something SRE
- D. A hybrid adoption model, since the team has both an old function and a new name
Check the answer
C. Google's own SRE book, and the exam that follows its framing, is explicit: the job title without the mechanisms is traditional ops wearing a new name tag. Renaming a team doesn't retroactively install an SLO, an error budget, a toil ceiling, or a blameless postmortem process — the mechanisms are what make SRE falsifiable and auditable in the first place.
Q26 · Module 2, Service Level Objectives & Error Budgets. A team's internal SLO is 99.9% availability. Its external, contractual SLA with customers also promises 99.9%, with service credits for any shortfall. What's the problem with this setup?
- A. There's no problem — matching the two numbers exactly is best practice
- B. The SLA should be set looser than the SLO; with them identical, every ordinary SLO miss simultaneously becomes a contract breach, throwing away the team's early-warning window
- C. The SLO should be set looser than the SLA to avoid unnecessary internal pressure
- D. SLAs and SLOs are never allowed to share the same numeric value under any circumstances
Check the answer
B. A well-run SLA is set looser than the internal SLO on purpose, so ordinary operational noise doesn't trigger a customer-facing penalty every rough week. Setting them identical (or the SLA tighter) means every internal SLO miss is simultaneously a contract breach — the team has thrown away the buffer meant to let it self-correct before a customer-facing consequence kicks in.
Q27 · Module 3, Reducing Toil. An engineer spends four exhausting hours manually recovering from a completely novel failure mode nobody on the team has ever seen before. Is this toil?
- A. No — it fails the repetitive gate; "first time we've ever seen this" disqualifies it from toil immediately, however manual and tedious it was
- B. Yes — it was manual, tedious, and interrupt-driven
- C. Yes, because it took four hours, which exceeds the threshold for toil
- D. No, because it was resolved in under a business day
Check the answer
A. However manual, tedious, or tactical a task feels, a genuinely novel task fails the repetitive gate outright — toil requires the exact task to have recurred before and be expected to recur again. There's no duration threshold in the definition at all, which is exactly the fabricated criterion C and D are testing.
Q28 · Module 4, Monitoring & Service Level Indicators. A team's dashboards only ever answer questions someone anticipated when building them — if an on-call engineer needs to ask a brand-new question about an unfamiliar failure mode, the dashboards can't answer it without shipping new instrumentation first. What does this describe, and what property is missing?
- A. This describes strong observability; the missing property is alerting
- B. This describes white-box monitoring; the missing property is black-box coverage
- C. This describes a windows-based SLI; the missing property is a request-based SLI
- D. This describes monitoring without observability; the missing property is the ability to ask arbitrary new questions of the system's internal state using existing, high-cardinality telemetry — without deploying new code first
Check the answer
D. Monitoring watches known, predefined signals for known failure modes — exactly what's described. Observability is the further property of being able to ask new, previously unanticipated questions about a system's internal state from existing telemetry, without shipping new code to answer them.
Q29 · Module 5, SRE Tools & Automation. Two systems both restart a failing pod using the same underlying script. System X requires an engineer to click "approve" in Slack first; System Y detects, decides, and restarts entirely on its own, notifying the team afterward. What determines which rung each belongs to?
- A. The complexity of the script each system runs
- B. Whichever system runs faster is automatically the more mature rung
- C. Who decides to trigger the action — System X is rung 2 or 3 depending on who proposed the fix, System Y is rung 4, regardless of how sophisticated either script is
- D. Both are the same rung, because they use the same underlying script
Check the answer
C. The exam's favorite trap in this module is assuming "there's a script" settles the rung — it doesn't. The rung is set entirely by who decides to trigger the action: a human clicking approve is rung 2 or 3 no matter how sophisticated the script is; a system acting on its own analysis is rung 4 even if the resulting change is trivial.
Q30 · Module 6, Anti-Fragility & Learning from Failure. Service A loses an instance unpredictably, constantly, every day, because Chaos Monkey terminates instances at random. Service B loses an instance once a year, unplanned, and is simply rebuilt identically each time. Which service is more likely to be developing genuine antifragility, and why?
- A. Service B, because rare failures are more realistic than constant ones
- B. Service A — because instances vanish constantly, engineers can never get away with code that merely tolerates termination once; every latent assumption gets found and eliminated continuously, so fleet-wide tolerance for instance churn keeps climbing
- C. Neither — antifragility requires a formal chaos engineering certification program
- D. They're equally antifragile, since both eventually recover
Check the answer
B. This is Netflix's original Chaos Monkey logic exactly. Because termination happens constantly rather than as a rare fire drill, latent assumptions about instance lifetime get found and eliminated continuously, and fleet tolerance for that class of failure trends measurably upward. Service B's annual recovery is real and valuable — but it's a robust recovery, not an antifragile improvement, since nothing about its future tolerance actually changed.
Q31 · Module 7, Organizational Impact of SRE. Matthew Skelton and Manuel Pais's Team Topologies vocabulary maps cleanly onto the SREF adoption-pattern taxonomy. Which pairing is correct?
- A. An embedded SRE functions as a platform team; a centralized SRE org functions as a stream-aligned team
- B. A consulting SRE engagement functions as a complicated-subsystem team
- C. All four SREF models map to the same Team Topologies type: the enabling team
- D. An embedded SRE functions as part of a stream-aligned team; a centralized SRE org functions as a platform team that other teams consume as a service
Check the answer
D. An embedded SRE, permanently placed inside one product team, is functionally part of that stream-aligned team. A centralized SRE org, consumed by other teams as an internal service, is a platform team. A consulting engagement — temporary, capability-transferring — maps to an enabling team instead, and the four SREF models don't all collapse into one Team Topologies type. See SRE Team Topologies for the full mapping.
Q32 · Module 8, SRE, Other Frameworks & the Future. A candidate assumes "ITIL is the old, slow, change-advisory-board-driven framework that SRE replaced." What's wrong with that assumption for the current version of ITIL?
- A. ITIL 4 (2019) was deliberately redesigned around a "four dimensions" model and reframed rigid "processes" as more flexible "practices," explicitly to absorb Agile and DevOps thinking — the caricature describes ITIL v3, not ITIL 4
- B. Nothing — it's an accurate description of ITIL at any version
- C. ITIL was discontinued entirely once SRE became popular
- D. ITIL 4 removed Service Level Management entirely, replacing it with SRE's error-budget policy
Check the answer
A. The "old, slow" characterization is a reasonably fair critique of ITIL v3 (2007, refreshed 2011), often document-heavy and change-advisory-board-driven. ITIL 4, released in 2019, was deliberately rebuilt to absorb Agile and DevOps thinking. ITIL wasn't discontinued, and Service Level Management is still core to ITIL 4 — SRE offers a different implementation of similar territory, not a replacement folded into ITIL itself.
Q33 · Module 1, SRE Principles & Practices. Google's SRE book maps five widely cited DevOps pillars to the concrete SRE mechanism that implements each one. Which pairing is correct?
- A. "Leverage tooling and automation" maps to mandatory blameless postmortems
- B. "Accept failure as normal" maps to a formal error budget — failure spent on purpose, not merely avoided
- C. "Measure everything" maps to canary and progressive rollouts
- D. "Reduce organizational silos" maps to toil treated as a bug
Check the answer
B. "Accept failure as normal" maps to the error budget — a formal allowance for failure, spent deliberately rather than merely avoided. "Leverage tooling and automation" actually maps to toil treated as a bug, automated and code-reviewed; "measure everything" maps to SLIs and SLOs mandatory before launch; "reduce organizational silos" maps to shared on-call and common tooling across Dev and SRE.
Q34 · Module 2, Service Level Objectives & Error Budgets. A checkout API serves 2,000,000 requests over a 30-day window and carries a 99.5% success-rate SLO. How many failed requests does its error budget allow for the entire window?
- A. 1,000
- B. 5,000
- C. 10,000
- D. 20,000
Check the answer
C. (100% − 99.5%) × 2,000,000 = 0.5% × 2,000,000 = 10,000 failed requests, for the entire 30-day window — not per day. A and B correspond to tighter SLOs applied to the same volume, and D is roughly double what a 99.5% SLO actually allows.
Q35 · Module 3, Reducing Toil. A one-time, one-hour task migrates a legacy config format to a new schema across the whole fleet. It won't need to be repeated once done, regardless of how many more servers the fleet grows to. Even though it's manual, is it toil?
- A. Yes, because migrations are inherently toil-like work
- B. Yes, because it touches every server in the fleet, which satisfies the O(n) gate
- C. No, but only because it takes less than a full day to complete
- D. No — it's a fixed, one-time cost that doesn't scale with fleet growth, and it also fails the repetitive gate since it isn't expected to recur
Check the answer
D. Genuine toil scales linearly (O(n)) with fleet size on an ongoing basis — this task is a fixed, one-time cost, and it independently fails the repetitive gate since it isn't expected to happen again. "Touches every server" describes the scope of one instance of the task, not a recurring linear cost — exactly the confusion B is testing. Duration is not a criterion in the definition at all.
Q36 · Module 4, Monitoring & Service Level Indicators. A payment queue's depth is climbing steadily toward its configured maximum capacity, though requests are still being processed successfully with acceptable latency for now. Which golden signal is this, and what question does it answer?
- A. Saturation — how close is the system to its limit?
- B. Traffic — how much demand is the system under?
- C. Errors — how often did a request fail?
- D. Latency — how long did the request take?
Check the answer
A. Queue depth relative to a configured maximum is a textbook saturation signal — it answers "how close is the system to its limit," independent of whether requests are currently succeeding. It's exactly the kind of leading indicator worth alerting an engineer on before it becomes a latency or error problem the user actually feels.
Q37 · Module 5, SRE Tools & Automation. A tool shows where in a multi-service call chain time actually went, as a tree of timed spans stitched together by a propagated request ID. Which toolchain category is this, and which pair of tools represents it?
- A. Metrics & monitoring — Datadog and InfluxDB
- B. Distributed tracing — OpenTelemetry and Jaeger
- C. On-call & paging — PagerDuty and Opsgenie
- D. Chaos engineering — Chaos Monkey and AWS FIS
Check the answer
B. A tree of timed spans stitched together by a propagated request ID, showing where time went across a multi-service call chain, is the defining description of distributed tracing — OpenTelemetry (the instrumentation standard) and Jaeger (a tracing backend) are representative. Metrics tools collect numeric time series rather than per-request call trees; paging tools route alerts to humans; chaos tools inject failure.
Q38 · Module 6, Anti-Fragility & Learning from Failure. A service sits behind a load balancer with redundant replicas and automatic failover. An instance dies; traffic reroutes instantly and the service is exactly as capable afterward as before. Does this redundancy, by itself, make the service antifragile?
- A. Yes — surviving any failure automatically qualifies as antifragility
- B. Yes, because the failover was automatic rather than manual
- C. No — redundancy and failover make the service robust against that known failure mode repeating; they don't, by themselves, generate new capability against a broader class of failure
- D. No, because failover only counts as antifragile if it happens more than once a year
Check the answer
C. This is the textbook robust example, not antifragile — the service returns to exactly its prior capability, the defining shape of a flat response to a stressor. Reliability patterns like redundancy, retries, circuit breakers, and load shedding make a system robust against a known failure mode repeating; they don't, by themselves, generate new capability against a whole class of failure the way a closed chaos-engineering feedback loop does.
Q39 · Module 7, Organizational Impact of SRE. An organization runs a hybrid SRE model: a central platform team owns shared paging infrastructure and org-wide policy, while product teams own their own services' SLOs locally. One day the paging system itself goes down. Nobody is quite sure whose incident this is. What does this expose?
- A. That the hybrid model is fundamentally broken and should never be used
- B. That the platform team should be immediately dissolved in favor of a purely embedded model
- C. That paging systems should never be centrally owned under any adoption model
- D. The specific risk hybrid carries: without an explicit ownership line, ambiguity creeps in over who's responsible for a given failure — especially failures of the shared platform itself, not just an individual product team's service
Check the answer
D. This is precisely hybrid's named risk: it requires the clearest ownership split of all four models, and without an explicit line — who owns the shared platform's own uptime vs. who owns each product's service SLOs — hybrid quietly turns into nobody being clearly responsible, especially for failures of the shared infrastructure itself.
Q40 · Module 8, SRE, Other Frameworks & the Future. A vendor's AIOps product flags a metric deviation before a static threshold would have fired, and separately groups a storm of forty related alerts into a single incident automatically. Which rungs of the automation maturity ladder do these two capabilities sit on, and what do mature, widely-deployed AIOps capabilities generally NOT yet do reliably?
- A. Rungs 1–3 (anomaly detection and alert correlation); mature AIOps generally doesn't yet replace the judgment-heavy decisions — like whether an SLO target is acceptable — that stay distinctly human
- B. Rungs 3–4; they reliably replace human judgment on root cause with no oversight
- C. Rung 4 only; AIOps is definitionally fully autonomous by design
- D. Rungs 1–2 only; AIOps cannot perform any form of anomaly detection
Check the answer
A. Anomaly detection ahead of a static threshold and alert correlation are exactly the mature, widely-deployed AIOps capabilities, sitting on the lower-to-middle rungs of the automation maturity ladder. What AIOps doesn't reliably replace is the judgment-heavy work — deciding what SLO target a business can live with, running blameless postmortem culture, deciding when to freeze launches — which stays distinctly human even as the repetitive setup toil around it keeps getting automated away.
Score yourself — and read the pattern, not just the number
☺ Like you're 10: One number tells you whether you passed. Eight smaller numbers tell you exactly what to study next.
The pass mark for the real SREF exam is 65% — 26 of 40 — with no partial credit, matching the figure stated throughout this course's blueprint. Grade this paper honestly against that bar, then do the more useful arithmetic: tally your misses by module, using the labels each question carried, and fill in the table below.
| Module | Questions here | Your score | If under 4/5, revise here |
|---|---|---|---|
| 1 · SRE Principles & Practices | 5 | Module 1 | |
| 2 · SLOs & Error Budgets | 5 | Module 2 · practice bank | |
| 3 · Reducing Toil | 5 | Module 3 · practice bank | |
| 4 · Monitoring & SLIs | 5 | Module 4 · practice bank | |
| 5 · SRE Tools & Automation | 5 | Module 5 · practice bank | |
| 6 · Anti-Fragility & Learning from Failure | 5 | Module 6 · practice bank | |
| 7 · Organizational Impact of SRE | 5 | Module 7 · practice bank | |
| 8 · SRE, Other Frameworks & the Future | 5 | Module 8 | |
| Total | 40 | Pass mark: 26/40 (65%) |
A 30/40 built from four strong modules and one near-zero module is a completely different result from a 30/40 built from a couple of wrong answers scattered evenly across all eight — even though the headline percentage is identical. The first pattern means you have a specific, fixable gap; the second usually means fatigue or a reading-speed problem late in the hour. Only the module-by-module table tells the two apart.
If you cleared 65% comfortably, with no single module below 3/5, that's a strong signal you're ready to book. If you cleared it narrowly, or one module came back at 0 or 1 out of 5, treat this exactly like Sets 1 through 4: fix the specific gap the table just showed you before you consider yourself done, not after you've forgotten which questions you missed.
What's left after Set 5
☺ Like you're 10: Fix only what this paper found broken, then go book the exam — don't restart studying everything from scratch.
Whatever you scored, this paper has done its job only if it produces a short, specific list — not a vague feeling of "I should study more." Write down every question you got wrong, and for each one, write why the option you picked was wrong, not why the correct one was right; that's a harder, more useful sentence, and it's the one that stops the same mistake recurring when the real exam's wording doesn't match this page's.
mkdir -p ~/sref-revision cat > ~/sref-revision/set5-wrong-answers.md <<'EOF' # SREF Mock Set 5 — wrong answers | Q | Module | What I picked | Why MY answer was wrong | Page to re-read | |---|--------|---------------|--------------------------|------------------| | | | | | | EOF open ~/sref-revision/set5-wrong-answers.md
Once the ledger is filled in, three more pages turn a diagnosis into exam-day confidence: Know It Cold — SREF collects the small set of facts and formulas that need to be automatic rather than reasoned-through under time pressure; Closed-Book Strategy — No Docs Map covers what to do when a term feels almost-but-not-quite familiar with nothing to look it up in; and Answer Triage — SREF covers the elimination technique for exactly the "which is NOT" and "BEST represents" stems this paper leaned on. When the ledger stops growing and a re-sit of your weakest module or two clears comfortably, The SREF Exam has the actual booking logistics, and The SREF Study Plan is worth one more skim to confirm there's nothing left on it unchecked.
Sol the Sloth: Thirty-four out of forty. Slowly, correctly, and I checked my error-budget arithmetic twice before writing anything down.
Remy the Rabbit: Thirty-four?! I answered mine in eleven minutes flat!
Timmy the Turtle: And what did you score, Remy?
Remy the Rabbit: ...twenty-two. I picked "robust" every time the word "resilient" showed up. Every single time.
Professor Owl: That's not a knowledge gap, Remy — you know the material. It's a reading habit, and it's the cheapest ten points you'll ever recover, because the fix is just reading the stem twice before you commit.
Timmy the Turtle: Sol, your thirty-four clears the sixty-five percent bar with a lot of room. Any module under four out of five?
Sol the Sloth: One. Module 7 — I mixed up centralized and consulting on two of the five.
Professor Owl: Then re-read that one module tonight, not the other seven. You already proved you don't need to touch the rest again.
Remy the Rabbit: Fine. Ledger, one habit to fix, re-sit nothing blind. Then I book it.
1. How long do you get for the real SREF exam, how many questions does it have, and what's the pass mark? 2. What does "closed-book" mean for this exam specifically — name two things that are off-limits that might not be obvious. 3. Why should you answer every single question on this paper, even the ones you're genuinely guessing on? 4. Why does this practice paper split its questions evenly across all eight modules instead of using a weighted split? 5. You score 30/40 (75%) with six wrong answers clustered in two modules. What's your very next move?
Check your answers
- 40 questions, 60 minutes, and a 65% pass mark (26/40) — the figures used throughout this course's SREF blueprint. Confirm the current numbers at DevOps Institute's own SRE Foundation page before you register, since format details can change.
- No notes, no other tabs on this or any site, no search engine, and no AI assistant — the real exam gives you nothing but the question and its four options.
- Nothing published indicates the SREF exam deducts marks for a wrong answer, and a blank answer is a guaranteed zero either way — an elimination-based guess costs nothing and sometimes turns out right.
- DevOps Institute doesn't publish an official per-module weighting for its 40 questions, so an even five-per-module split is the fairest defensible study convention available — not a claim about the real exam's actual distribution, which stays unpublished.
- Don't re-sit Set 5 cold again — a second attempt on the same paper mostly measures how well you remember it, not your actual readiness. Re-read the two modules where the misses clustered, re-drill them in the practice bank, and only then consider booking, or a fresh sitting of an earlier set if you want one more timed data point.