Common Preparation · The assumed floor

The Baseline You Need First

The SRE Foundation exam is unusual among the certifications on this Academy: it has no formal prerequisites at all. PeopleCert's own product page says so, training is not mandatory, and nothing on this site or anywhere else has to be passed first. That is a booking rule, and people read it as a study plan. It isn't one. The syllabus underneath the exam is written in the vocabulary of production operations — toil, saturation, blast radius, burn rate, self-healing, auto-scaling, immutable infrastructure — and every one of those words is used as though you already know what it points at. This page is the honest floor: the Linux, the distributed-systems behaviour, the percentage arithmetic, the monitoring concepts and the on-call practice that make the rest of this course land instead of merely scroll past. It ends with ten diagnostic questions and a two-tier self-test with a full answer key, so you find the gaps here, in an afternoon, rather than in the sixtieth minute of a sixty-minute paper.

☺ Explain it like I'm 10

Imagine a class called "How to decide how late the town's buses are allowed to be." It's a real subject with real rules — how you measure lateness, how much lateness the town will tolerate before someone has to fix the timetable, and who gets woken up when a bus doesn't turn up at all. What the class never covers is what a bus is, how traffic works, or how to read a clock. It just assumes all of that. If you turn up not knowing how to read a clock, every lesson is secretly two lessons, and you'll spend the whole term translating instead of learning. This page is the quick check that you can already read the clock.

🦊🦥Your hosts for this topic: Foxy & Sol the Sloth — Foxy asks the question you were hoping to skip ("do you actually know this, or have you read about it?"), and Sol does the part almost everyone underrates: the slow, correct arithmetic that turns a percentage into minutes and a burn rate into a decision. Between them they cover the two ways a baseline gap shows up — the concept you can't explain, and the number you can't work out.

"No prerequisites" is a booking rule, not a study plan

☺ Like you're 10: Anyone is allowed to sign up. That's not the same as everyone being ready.

Three official statements set the frame for this whole page, and they pull in slightly different directions.

The first is PeopleCert's, on its SRE Foundation product page: there are no prerequisites, and accredited training is recommended rather than required. The second is the DevOps Institute's, in the free SREF Exam Study Guide: although there are no formal prerequisites, it recommends candidates complete at least 16 contact hours of instruction and labs through an accredited Education Partner. The third is the same study guide's statement about difficulty: the exam is constructed at Bloom levels 1 and 2 only — knowledge and comprehension. There are no Bloom 3+ items. Nothing on the paper asks you to apply a technique to a novel situation, analyse a system, or design anything.

Put those together and you get an exam with a genuinely low ceiling and a genuinely real floor. The ceiling is low because Bloom 1 and 2 means recall and understanding, in a fixed vocabulary, with four options and a single best answer. Chris McLeod, who sat and passed the SREF in December 2023 and wrote it up on his personal blog, reached the same conclusion from the inside: he found it broad rather than deep, and said his DevOps background helped with general familiarity but was not essential. That is one person's one sitting, and it was the v1.1-era syllabus — but it corroborates the official Bloom statement exactly, which is the most useful thing a single account can do.

The floor is real because the syllabus never defines the world its concepts live in. The official concept and terminology list — the exam's own vocabulary inventory, printed in that same free PDF — includes Kubernetes, Amazon Web Services, cloud computing, auto-scaling, software-defined networking, immutable infrastructures, application performance management, canary testing, latency, response time and self-healing. Not one of those is taught by the syllabus. They are the furniture of the room the syllabus is describing.

What the official sources requireWhat the material actually assumes
No prerequisites — anyone may book and sit it (PeopleCert)That you have seen a production system misbehave, so "the service degraded" is a memory rather than a phrase
Training recommended, not required — 16 contact hours through an accredited partner (DevOps Institute)That someone will define the syllabus's own named models for you. If you self-study, you have to find them yourself
Bloom 1 and 2 only — knowledge and comprehensionThat the concepts have something to attach to. Comprehension of "toil" is cheap if you have done some; expensive if you haven't
Open book — the Official Training Materials onlyThat you either bought those materials or already hold the content in your head. There is no third option

Two different meanings of "I know this"

Everything on this page splits along one line, and it is worth being blunt about it, because the two halves have completely different costs to close.

The exam's standard is recognition: shown four sentences about error budgets, can you pick the one that matches the syllabus? That standard is reachable by a determined beginner in a few weeks of disciplined memorisation, and this course's flashcards, Know It Cold and concept reference exist to get you there efficiently.

The practitioner's standard is different: could you set an SLI for a service you own, defend the number to a product manager, and know what to do at 3am when the budget is gone? Nothing on the SREF paper measures that. Everything in the job does, and so does every serious interview.

◆ Key idea

The SREF is one of the few exams where you can legitimately pass without the baseline — and that is exactly why the baseline is worth checking anyway. A hands-on exam fails you loudly, at a specific task, and hands you a list of what to fix. A Bloom 1–2 multiple-choice paper does no such thing: it will let you through on recognition alone and hand you a badge that says something slightly different from what you can actually do. Decide deliberately which of the two standards you are aiming at. Both are legitimate. Confusing them is not.

The whole floor, at a glance

Five areas. Each gets its own section below, each says which of the eight examined topic areas leans on it, and each ends with the specific pages in this course that close the gap if you find one.

Baseline areaWhat the syllabus does with itWhere it bites hardest
Linux & the production OSUses it as the setting for every example of toil, automation and self-healingSRE-3 Reducing Toil · SRE-5 Tools & Automation
Distributed-systems behaviourAssumes partial failure, latency tails and cascading effects are already intuitiveSRE-1 Principles · SRE-6 Anti-fragility
Availability arithmeticAsks you to convert a percentage into a budget — the one calculation on the paperSRE-2 SLOs & Error Budgets
Monitoring conceptsBuilds the heaviest topic area directly on top of metrics, signals and telemetry vocabularySRE-4 Monitoring & SLIs — the heaviest area
On-call & process practiceDescribes one particular model of incident response and organizational culture as though it were the water you swim inSRE-7 Organizational Impact · SRE-8 Frameworks
NEVER EXAMINED Writing code · configuring a monitoring stack · cloud consoles · running a cluster There is no terminal and no console anywhere near this exam. EXAMINED — 40 questions across 8 topic areas Principles & practices · SLOs & error budgets · reducing toil · monitoring & SLIs tools & automation · anti-fragility · organizational impact · other frameworks Bloom levels 1 and 2 only — knowledge and comprehension, never application or analysis. ASSUMED — never taught, never tested, always present Linux & the production OS · distributed-systems behaviour · availability arithmetic monitoring concepts · on-call and incident practice This page is about this band, and only this band.
⚠ The most expensive mistake in SREF prep

Starting the blueprint modules while the floor is thin feels productive and is not. You read a paragraph about reducing toil, and half your attention silently goes to reconstructing what the example task actually is — so you absorb a fraction of what you should and remember it as a phrase instead of an idea. On this particular exam the damage is delayed rather than obvious: recognition-level memory will still carry you through many questions, right up until a scenario-flavoured item asks you to pick the best of four workable answers, which is precisely where having seen the thing happen decides it. Spend the week this page identifies before you start. The SREF study plan quietly assumes you have.

Where the questions actually are

☺ Like you're 10: Not all the topics are worth the same. Three of them are nearly half the paper.

The DevOps Institute publishes something most vendors don't: an explicit maximum number of questions per topic area, summing to the full 40. It appears in the free SREF Exam Study Guide, and it is the single most useful planning artefact for this exam.

Maximum questions per topic area (of 40) SRE-1 Principles & practices 4 SRE-2 SLOs & error budgets 6 SRE-3 Reducing toil 5 SRE-4 Monitoring & SLIs 7 SRE-5 Tools & automation 6 SRE-6 Anti-fragility 4 SRE-7 Organizational impact 4 SRE-8 Other frameworks & trends 4 Highlighted: the three areas that together can reach 19 of the 40 questions.

The arithmetic below is ours, derived from that official table — it is not a DevOps Institute claim. SRE-2 (6) plus SRE-4 (7) plus SRE-5 (6) is 19 of 40. The pass mark is 65%, which on a 40-question paper is 26 correct. So those three areas alone can carry roughly three-quarters of the marks you need, and two of the three — SLOs and monitoring — sit directly on top of two of this page's baseline areas: availability arithmetic and monitoring concepts. If you close only two gaps before you start, close those two.

◆ Key idea

The weighting is a maximum per area, not a guaranteed count, and it comes from the v1.1 study guide. Treat it as a strong steer rather than a contract — and note that it steers you toward exactly the areas that also matter most in the job. That alignment is not always true of certifications. Here it is, and it means the efficient study path and the useful study path are the same path.

⚠ The free study guides are two versions behind

PeopleCert's official training-material update notices record the SRE Foundation as v1.2, updated October 2024, adding coverage of generative AI, AIOps, platform engineering, value stream management, observability and progressive deployments. The freely downloadable exam study guides are v1.0 and v1.1, from 2020 and 2021. Their syllabus map, topic weighting and sample paper are still the best free assets in existence for this exam — but they predate that content, and their administration details are stale too (see the trap section below). Use them for the map; do not use them to decide what is out of scope.

Linux and the production operating system

☺ Like you're 10: Nobody will ask you to type a command. Every example quietly assumes you've watched a computer misbehave.

⚖ Where it bites — SRE-3 (Reducing Toil, up to 5 questions) and SRE-5 (Tools & Automation, up to 6). Between them, up to 11 of 40.

This is not a Linux exam. There is no terminal, no console, no cluster, and you will not be asked what a signal number means. What you will be asked is to recognise which of four described activities counts as toil, what kind of work automation should absorb, and what "self-healing" means in practice — and every single one of those questions is set in an operating system whose behaviour is taken as read.

The five places the syllabus quietly touches the OS

Read the concept list and the topic descriptions closely, and the OS shows up at exactly five points. Each one is a place where prior experience turns a memorisation task into an obvious answer.

WhereWhat is assumed
Toil examplesThat you recognise recurring manual chores when described: restarting a stuck process, rotating a certificate, clearing a full disk, re-running a failed job, applying the same fix on twenty hosts
Saturation (one of the four golden signals)That "how full the thing is" is a real, measurable property — CPU, memory, disk, file descriptors, connection pools, queue depth — and that saturation predicts trouble before errors appear
Self-healing & automated rollbackThat something is doing the healing: a supervisor restarting a process, an orchestrator replacing a container, a health check removing an instance from a pool
Auto-scalingThat there is a unit being added and removed, that adding one is not instant, and that scaling does not fix a bug
Immutable infrastructureThat the alternative — patching servers in place until no two are alike — is a thing that actually happened to people, and why it hurt

What "comfortable" looks like

You do not need to be a systems administrator. You need the following to be unremarkable rather than novel — if you read this list and none of it is familiar, that is a real gap, and it is the one most likely to make the toil and automation material feel abstract:

# where did it go wrong, and when
journalctl -u myservice --since "10 min ago"
tail -n 200 /var/log/myservice.log

# is the process even running, and what is it doing
systemctl status myservice
ps aux --sort=-%cpu | head

# what is full
df -h                 # disk
free -m               # memory
ulimit -n             # file descriptor ceiling

# was it killed rather than crashed
dmesg -T | grep -i "out of memory"

The point of that list is not the commands. It is the five questions behind them: did it run, did it stop, was it killed, is something full, and what did it say before it went? Those five questions are the shape of almost every real operational incident, and they are the reason "restart it" feels like a fix when it is only a delay.

🦫 Benny's-eye view

"My first automation was a cron job that restarted the service every night at three. It worked. Nobody was paged for four months and I was quietly proud of it. Then the memory leak got twice as fast and it started falling over at midnight instead, and I discovered that what I had actually built was a machine for hiding a bug from the people who could have fixed it. That's the lesson the exam is pointing at when it says automation should remove toil rather than absorb it — but I only understood the sentence after I'd built the wrong thing. If you've never built the wrong thing, read that sentence twice."

If this section found a gap

Toil & automation is the direct fix, and it defines toil precisely enough to apply as a test rather than a vibe. Capacity planning & performance covers saturation as a measurable quantity rather than a feeling. The SRE toolchain names the things that do the healing, and the toil audit drill makes you apply the definition to your own week, which is by far the fastest way to make it stick.

How distributed systems actually fail

☺ Like you're 10: Big systems don't break all at once. One slow part quietly ruins everything attached to it.

⚖ Where it bites — SRE-1 (Principles & Practices) and SRE-6 (Anti-fragility & learning from failure), up to 8 of 40 — plus it silently underwrites most of SRE-4.

The SREF's whole worldview rests on a single premise stated in Google's framing of the discipline: failure is normal. Not rare, not preventable, not a sign that someone did something wrong — normal, expected, and budgeted for. That premise only makes sense if you already know that a distributed system is almost never fully up or fully down. It is partially degraded, for some users, on some requests, some of the time. If your mental model of "down" is a light switch, error budgets will read as accounting rather than as physics.

The eight behaviours worth already knowing

None of these will be asked about by name in a Bloom 1–2 question. All of them make the concepts that are asked about obvious rather than memorised.

BehaviourWhy the syllabus needs you to have it
Partial failure — some requests fail while the service is "up"It is the reason an SLI is a ratio or a percentile rather than a yes/no
Tail latency — p50, p95, p99, and why the mean liesNearly every real latency SLI is a percentile. Without this, SRE-4 is pure memorisation
Timeouts — every call needs one, and it should be shorter than the caller'sThe absence of a timeout is the most common single cause of a cascading failure
Retry amplification — retries multiply load on a service that is already strugglingExplains why "just retry" is a failure mode and why backoff and jitter exist
Cascading failure — one slow dependency exhausts a shared resource and fails unrelated workThe mechanism behind blast radius, bulkheads and circuit breakers
The utilization knee — latency rises non-linearly as a resource approaches fullWhy headroom is bought deliberately, and why "we're only at 85%" is not reassuring
Idempotency — running it twice is the same as running it onceThe property that makes automated remediation safe. Without it, self-healing is self-harming
Graceful degradation & load shedding — dropping some work on purpose to protect the restThe engineering answer to an error budget that is burning
⚠ The average is not a user

Two arithmetic facts about percentiles that trip up experienced engineers, not just beginners. First, the mean hides the tail: a service where 99% of requests take 50ms and 1% take 10 seconds has a mean around 150ms, which looks fine and describes nobody's experience. Second, you cannot average percentiles. The p99 of two services is not the mean of their p99s, and the p99 over a day is not the mean of the hourly p99s — you need the underlying distribution (in practice, mergeable histograms) to recompute it. Any dashboard that averages a percentile is quietly lying to you, and that is worth knowing before you design an SLI on top of one.

🐢 Timmy's dependency trace · 20 min

Take one service you have access to — yours, or the one in this course's worked case study — and trace a single request out loud, hop by hop: client, load balancer, service, each downstream call, each datastore. At every hop, name three things without looking them up: the timeout, what happens on failure, and what the user sees. Then go back along the chain and find the first hop where you had to guess. That hop is your distributed-systems gap, located in twenty minutes and with a name attached. If you can do the whole trace without guessing, this section is closed for you.

If this section found a gap

Distributed-systems reliability fundamentals is the direct read. Reliability patterns covers timeouts, retries, circuit breakers, bulkheads and load shedding as a connected set rather than a list. Queueing theory for SRE is the honest answer to why the utilization knee exists at all, and dependency management and Hyrum's law covers what happens when the thing you depend on changes underneath you.

The arithmetic floor — nines, minutes and budgets

☺ Like you're 10: There is exactly one bit of maths on this exam, and it comes up more than once. It's worth being fast at it.

⚖ Where it bites — SRE-2 (SLOs & Error Budgets), up to 6 of 40 — and it is the only place the paper asks you to compute anything.

The free official study guide includes a full 40-question sample paper with its answer key, and one of the items on it does exactly this: it gives a team a monthly availability SLO of 99.9% and asks how much error budget that allocates. The published key's answer is 43 minutes. That single item tells you two things — that percentage-to-minutes conversion is genuinely examinable, and that the exam uses the simple time-based reading of a "month" (30 days, 43,200 minutes) rather than anything more careful.

The table worth holding in your head

Two columns, one calculation. A 30-day month is 43,200 minutes; a 365-day year is 525,600 minutes. Multiply by one minus the target.

Availability targetBudget per 30-day monthBudget per 365-day year
99% ("two nines")432 min — 7 h 12 m5,256 min — 3 d 15 h 36 m
99.5%216 min — 3 h 36 m2,628 min — 1 d 19 h 48 m
99.9% ("three nines")43.2 min525.6 min — 8 h 45 m
99.95%21.6 min262.8 min — 4 h 22 m
99.99% ("four nines")4.32 min52.6 min
99.999% ("five nines")~26 seconds5.26 min

The mental shortcut is worth more than the table: 1% of a 30-day month is about 7.2 hours, and every extra nine divides by ten. 7.2 hours, 43 minutes, 4.3 minutes, 26 seconds. If you can produce that ladder from memory in ten seconds you have covered every version of this question anyone can construct, and you never have to trust a memorised row again.

Budget, burn rate, and the two ways to count

Three formulas, all of which follow from the same idea — the budget is the unreliability the target permits:

Be aware that the two counting methods can genuinely disagree about whether a service met its SLO — a short outage during peak traffic costs far more request budget than time budget, and vice versa. The exam works in the simple time-based frame; production increasingly does not. Both are correct; they answer slightly different questions. SLO window types and composite SLOs covers the difference properly, including rolling versus calendar windows.

🦥 Sol's arithmetic drill · 15 min

Do these five in your head, slowly and correctly, before checking anything. (1) 99.95% monthly — how many minutes? (2) A service takes 40 million requests a month with a 99.9% success SLO — how many failures may it spend? (3) Halfway through that month it has spent 28,000 — what is its burn rate, and will it make it? (4) A quarterly SLO of 99.9% — roughly how many minutes over a 90-day quarter? (5) A service is at 99.7% against a 99.9% target — has it overspent, and by roughly how much?

Answers: (1) 21.6 minutes. (2) 40,000. (3) 28,000 against a sustainable 20,000 at the halfway point is a 1.4× burn rate, projecting to 56,000 against a 40,000 budget — it will not make it, and something needs to change now rather than at month end. (4) About 130 minutes (129,600 × 0.001). (5) Yes — it has spent three times its budget, because 0.3% is three times 0.1%. That last one is the conversion people get wrong under time pressure: the gap between the achieved number and the target is not the overspend, the ratio of the failure rates is.

⚠ We could not confirm whether a calculator is provided

We found no official statement from PeopleCert or the DevOps Institute about whether an on-screen calculator is available during an SREF sitting, and we are not going to guess — practice varies across PeopleCert's exam portfolio. Assume you will have to do it in your head, because that assumption costs you nothing and the reverse assumption could cost you a mark. If it matters to you, ask PeopleCert before you book. This is also a good general habit for this exam: several widely-repeated "facts" about it are wrong, and the vendor is the only source worth trusting on delivery mechanics.

If this section found a gap

SLIs, SLOs & error budgets is the core page and does the arithmetic slowly. Measuring & reporting reliability covers what you do with the numbers once you have them. Multi-window burn-rate alerting is the practical use of burn rate, and the SLO and error-budget drill makes you compute one for a service rather than read about it.

Monitoring — the heaviest area, and the highest assumed floor

☺ Like you're 10: The biggest chunk of the exam is about measuring things. It assumes you already know what a measurement is.

⚖ Where it bites — SRE-4 (Monitoring & Service Level Indicators), up to 7 of 40 — the single heaviest topic area on the paper.

This is where the gap between "no prerequisites" and "the syllabus assumes plenty" is widest. The topic area's own description talks about understanding SLIs and how they relate to SLOs, the monitoring landscape, observability, and setting measurable objectives. Every one of those phrases contains an unexploded assumption: that you know what a metric is, what a time series is, what it means to aggregate one, and what a percentile computed from a histogram actually represents.

The vocabulary the questions stand on

ConceptThe one thing to have straight
Counter, gauge, histogramA counter only goes up (requests served); a gauge goes up and down (queue depth); a histogram buckets observations — and percentiles come from histograms, not from gauges
Time seriesA named stream of numbers over time, identified by its labels. Everything on a dashboard is one
CardinalityThe number of distinct label combinations. High cardinality (user ID, request ID) is what makes a metrics system fall over — which is why those belong in traces or logs instead
Metrics vs logs vs tracesMetrics are cheap and aggregate; logs are detailed and expensive; traces follow one request across services. Different questions, different tools
Black-box vs white-boxBlack-box measures from outside as a user would; white-box measures internals the service reports about itself. Black-box tells you it's broken; white-box tells you why
Symptom vs cause alertingPage on what the user feels (errors, latency, budget burn), not on what a machine is doing (CPU is high). The second is a dashboard, not a page
Monitoring vs observabilityMonitoring answers questions you thought of in advance; observability is the property of being able to answer ones you didn't. The syllabus treats them as distinct and so should you

The named frames you should recognise on sight

Three of these are general SRE knowledge; the fourth is specific to this syllabus, and that distinction is worth making explicitly.

⚠ SLI, SLO, SLA — the distinction the distractors are built on

An SLI is a measurement: the proportion of requests served successfully, or the 95th-percentile latency. An SLO is a target for that measurement, set internally: "99.9% of requests succeed over 30 days." An SLA is a contract with a customer, with consequences — money, credits, penalties — attached to missing it. The one-line test: the SLA is the one with money attached, and it is almost always looser than the SLO, deliberately, so that you breach your internal target well before you breach your customer's contract. On a single-best-answer paper, a question that hinges on this distinction will offer you two options that are both true statements about reliability and only one that is a true statement about the specific term. Precision here is worth more than any other definition on the syllabus.

🐘 Ellie's dashboard audit · 20 min

Open any dashboard you have access to. Answer four questions about it, in writing. (1) Which of the four golden signals are actually on it, and which are missing? (2) For each panel, is it a counter, a gauge or a percentile — and if it's a percentile, do you know how it was computed? (3) Which panels would tell you the user is having a bad time, and which only tell you a machine is? (4) Which panel has an alert attached, and is that alert a symptom or a cause? Most dashboards fail question 1 on saturation and question 4 on symptoms. If yours does, you have just learned the two things this topic area is really about.

If this section found a gap

Monitoring & observability is the core page. Alert design and alert fatigue covers symptom-versus-cause properly and is the single best return on this area. The SRE-4 blueprint module maps it to the exam directly, and if the tooling vocabulary is the gap rather than the concepts, Prometheus, Grafana and OpenTelemetry cover the three names most likely to appear in a sentence you are asked to interpret.

On-call and incident practice

☺ Like you're 10: The exam describes one particular way of handling emergencies as though everyone does it that way. Not everyone does.

⚖ Where it bites — SRE-7 (Organizational Impact, including sustainable incident response and blameless post-mortems) and SRE-6 (Anti-fragility), up to 8 of 40.

The syllabus describes a specific, Google-derived model of incident response: a rotation with a primary and a secondary, an escalation policy, pages that mean something and tickets that don't, a structured response with named roles, and a blameless review afterwards that produces tracked action items. If you have carried a pager under something like that model, this whole area is nearly free marks. If you have carried a pager under a different model — or never carried one — it is vocabulary to learn deliberately rather than intuition to lean on.

The vocabulary that shows up undefined

TermWhat it means, briefly
Rotation, primary, secondaryWho holds the pager now, and who is the backstop if they don't answer
Escalation policyThe predefined chain of who gets paged next, and after how long, if nobody acknowledges
Page vs ticketA page means "a human must act now." A ticket means "someone will act during working hours." Confusing the two is the definition of alert fatigue
Severity levelsA shared scale for how bad this is, agreed in advance so nobody negotiates it at 3am
Incident command rolesAn incident commander who decides, a communications lead who tells people, an operations lead who does the hands-on work. Separating them is what stops one person doing all three badly
RunbookThe written procedure for a known failure, ideally the thing that gets automated next
Blameless postmortemA review that treats failure as a systems property rather than a personal one, on the premise that people act reasonably given what they knew at the time

The mean-time family, and why it is contested

The official concept list names three of these specifically — MTTD (mean time to detect), MTTR (mean time to repair/recover) and MTRS (mean time to restore service) — and listing MTRS separately is itself informative, because most practitioner writing collapses everything into MTTR.

Be aware that in the wider industry, MTTR is genuinely ambiguous: different organizations expand the R as repair, recover, restore or resolve, and those measure different intervals ending at different moments. That ambiguity is a real, documented source of confusion in reliability reporting, not a trick. On the exam, answer from the syllabus's definitions rather than your employer's. In a job interview, the better move is to say which definition you are using before you quote a number.

⚠ On-call practice varies enormously — do not assume one model

The syllabus describes one model. Reality is far more varied, and the variation is regional as well as organizational. Whether on-call is compensated, and how, differs by country and by employment law. Whether pages route to the developers who wrote the code or to a separate operations team differs by company and by decade. Whether a team runs follow-the-sun handoffs across time zones or simply wakes people up locally depends on how many time zones the company has staff in. Whether a postmortem is genuinely blameless depends on the organization's culture far more than on its template. This course covers the model the exam describes and says where practice diverges; treat any source that presents one arrangement as universal — including a training course — with appropriate scepticism.

🐦 Pip's pager audit · 15 min

If you are on a rotation, pull the last month of alerts and sort them into three piles: this needed a human immediately, this could have waited until Tuesday, and this needed nobody at all. Count the piles. If the first pile is not the largest, you have just measured alert fatigue in your own team, and you now understand SRE-4 and SRE-7 better than any amount of reading would have taught you. If you are not on a rotation, do the same exercise on the incident timeline in this course's worked case study — the reasoning transfers.

If this section found a gap

Incident management & on-call covers the model the syllabus assumes. Postmortems & blameless culture covers the review that follows, and incident command for large incidents covers the role separation. The incident tabletop drill is the cheapest way to acquire the experience if you have never been on a rotation — it is a rehearsal, and rehearsal is most of what on-call training actually is.

The change, delivery and process vocabulary

☺ Like you're 10: A last set of words the exam uses without stopping to explain — mostly about how software gets released, and how other frameworks talk about the same problems.

⚖ Where it bites — SRE-5 (Tools & Automation, up to 6) and SRE-8 (SRE, Other Frameworks, Trends, up to 4).

Delivery vocabulary

The official v1.1 concept list already names canary testing, automated rollback, immutable infrastructures, auto-scaling, functional and non-functional tests and the software delivery lifecycle. The v1.2 update, per PeopleCert's own training-material notices, added progressive deployments explicitly. You need the release vocabulary at recognition level: what a pipeline is, and the difference between a rolling update (replace instances gradually), blue/green (two full environments, cut traffic across, roll back by cutting back) and a canary (a small slice of real traffic, compared against the baseline before proceeding). The distinction that matters for a best-answer question: blue/green buys you a fast rollback; a canary buys you a comparison — an actual signal about whether the new version is worse, on real traffic, before most users see it. Release engineering & progressive delivery covers all of it.

You also need Kubernetes at reading level — the concept list names it directly. You do not need to operate a cluster, and nothing on the paper will ask you to. You need to be able to read a sentence containing "the orchestrator rescheduled the pod" without stalling. If that sentence stalls you, Kubernetes reliability patterns gives you enough, and this Academy's CKA page describes the proper route if you want more than enough.

Process and framework vocabulary

SRE-8 positions SRE against DevOps, Agile and ITSM. That means a handful of process terms are assumed:

◆ The cheapest study asset for this exam is free

The concept and terminology list in the free SREF exam study guide is, in effect, the exam's own vocabulary inventory, published by the people who wrote the exam. Roughly sixty terms, alphabetical, no explanations. Print it. Go down it and mark every term you could not define in one sentence to a colleague. That marked list is your study plan for the vocabulary half of this exam, and it took you fifteen minutes and nothing to produce. This course's glossary and flashcards then close most of it — with one honest caveat: the free list is the v1.1 list, so it will not include terms added by the v1.2 update.

If ITSM is the gap, this Academy's ITIL 4 Foundation page describes the neighbouring credential, and the SRE-8 blueprint module covers exactly what this exam wants from that world — which is considerably less than ITIL itself would want. The SRE-7 module covers the culture material, including Westrum.

Where your intuition will be marked wrong

☺ Like you're 10: Sometimes the exam wants the book's answer, not the true-in-your-job answer. Learn to spot when.

This is the section that separates a working SRE from a working SRE who passes comfortably. Experience helps enormously on this paper — right up until it doesn't, and there are four specific places where it doesn't.

1. When a question names its source, answer from that source

The free official sample paper contains a definition item that explicitly attributes the definition it wants to a named industry survey. Its published answer key selects the terse, technical wording over the more intuitive user-facing one that most practitioners would give. Whether or not you agree with that key, the lesson generalises and is worth internalising: when a question says "according to X," it is testing your recall of X, not your understanding of the concept. Reading past that clause and answering from experience is a reliable way to lose a mark you knew the answer to.

2. Reliability is targeted, not maximised

The same sample paper includes a true/false item on whether the error budget should be spent down over the period, and the published key treats a fully-spent budget as the intended outcome rather than a failure. That surprises people whose instinct is that less downtime is always better. It is, in fact, the core economic argument of the whole discipline: an unspent error budget means you bought reliability nobody asked for, at the cost of the features and velocity you could have bought instead. Reliability economics covers the argument properly, and Google's own error-budget policy is the primary source. Learn the reasoning, not just the answer — it is the thing most likely to come back at you in an interview.

3. "Open book" does not mean what you think it means

PeopleCert's product page states that the SREF is open book and that the only permitted material is the Official Training Materials, whether supplied directly or through an accredited training organization. That is a much narrower permission than it first sounds. Chris McLeod's write-up makes the practical consequence concrete: he did not have the official materials with him and therefore worked from memory throughout, which is the situation almost every self-studier will be in. If you did not buy an accredited course, plan as though it is closed book, because functionally it is. That is exactly the assumption this course's closed-book strategy page is built on, and it is the right one. Also note the arithmetic: 40 questions in 60 minutes is 90 seconds each. Even with materials open in front of you, there is no time to study from them.

4. Two different certifications share this name

⚠ Check whose exam you are reading about

The Global Skill Development Council (GSDC) issues a separate credential also called Site Reliability Engineering (SRE) Foundation, and its specifications differ from the DevOps Institute / PeopleCert exam this course covers — different duration, a different open-book policy and a different validity period. Someone searching "SRE Foundation exam duration" can very easily land on the wrong vendor's numbers and prepare against the wrong constraints. Beyond that, third-party blogs publish specifications for this exam that are simply wrong: we found sites asserting 90 minutes, 50 questions and a 70% pass mark, all three of which contradict both official sources. And the search results for this exam are heavily populated by "exam dump" vendors, whose material is both unreliable and a breach of the certification's terms. Confirm every number on PeopleCert's own SRE Foundation page before you plan around it. Our SREF exam guide is the full logistics briefing and carries the same warning.

Ten questions that find the gap

☺ Like you're 10: Cover the page. Answer out loud. The ones you can't answer are the ones worth a week.

Deliberately not multiple choice. Recognising the right option is exactly the skill that hides a baseline gap, and this course already gives you plenty of that in the exam simulator, the practice question bank and five mock exams. Answer each of these in a sentence or two, from memory, before opening the key. Track which baseline area each miss belongs to — that mapping is the output of the exercise, not the score.

  1. (Arithmetic) A team has a monthly availability SLO of 99.9%. How many minutes of error budget is that? Now say what changes if the window is a rolling 30 days rather than a calendar month.
  2. (Arithmetic) A service handles 20 million requests a month against a 99.9% success SLO. How many failed requests may it spend? Exactly halfway through the window it has spent 14,000. What is its burn rate, and is it on track?
  3. (Distributed systems) One downstream dependency gets slow. Your service starts failing requests that never touch that dependency at all. Name the mechanism, and give two mitigations.
  4. (Distributed systems) Mean latency is unchanged; p99 has doubled. What can that mean? And why can you not compute a service's overall p99 by averaging the p99s of its instances?
  5. (Monitoring) Define a metric, an SLI, an SLO and an SLA in one sentence each. Which one has money attached, and which of the SLO and SLA is normally the tighter number?
  6. (Monitoring) Name the four golden signals. Which of them is measured at the resource rather than at the request, and why is it the one most often missing from a dashboard?
  7. (Linux / toil) A cron job restarts a service every night at three and nobody remembers why. Apply the definition of toil and give a verdict — then say what the right fix is and why the cron job is not it.
  8. (On-call) What do MTTD, MTTR and MTRS each measure? Give one concrete reason a team's MTTR can fall while users are having a worse time.
  9. (Delivery) Blue/green versus canary: what does each one actually reduce, and which of the two gives you evidence rather than just an escape route?
  10. (Process) A team's incident reviews name an individual in the summary line. Which of Westrum's organizational types does that indicate, and what does it cost the organization in practical terms?
Check your answers
  1. 43.2 minutes — a 30-day month is 43,200 minutes, and 0.1% of that is 43.2. (The official sample paper's key rounds this to 43 minutes.) A rolling 30-day window changes the behaviour rather than the size: the budget never resets on a fixed date, so a bad day keeps counting against you until it ages out 30 days later, and there is no month-end amnesty to wait for. Calendar windows are easier to report; rolling windows are harder to game and better reflect what a user experienced recently.
  2. 20,000 failed requests (20,000,000 × 0.001). At the halfway point the sustainable spend is 10,000, so 14,000 is a burn rate of 1.4×, projecting to 28,000 against a 20,000 budget. It is not on track — and the useful part of that answer is that you know it now rather than at month end, which is the entire point of tracking burn rate instead of just tracking the final number.
  3. Cascading failure through a shared, exhausted resource — typically a thread pool, connection pool or worker queue. Requests to the slow dependency stop returning, occupy every worker, and requests that need none of it then queue behind them and time out. Mitigations, any two: a timeout short enough that a call releases its worker; a circuit breaker that fails fast once the dependency's failure rate crosses a threshold; a bulkhead that gives each dependency its own bounded pool so one cannot consume all of them; load shedding that rejects excess work early; and graceful degradation that serves the request without the optional dependency. Reliability patterns covers all five.
  4. It means a subset of requests got much slower while the bulk did not — one bad shard or instance, one slow customer's data, a cold cache for some keys, garbage-collection pauses, or a retry path only some requests take. The mean is dominated by the common case and is nearly blind to it. And you cannot average percentiles because a percentile is a property of a distribution, not a quantity that composes linearly: to get the true overall p99 you have to merge the underlying distributions (in practice, mergeable histograms) and recompute. The same applies over time — the p99 of a day is not the mean of 24 hourly p99s.
  5. A metric is any measured quantity about the system. An SLI is a metric deliberately chosen to represent the user's experience, usually as a ratio of good events to valid events, or a latency percentile. An SLO is an internal target for an SLI over a window. An SLA is an external contract with consequences. The SLA has the money attached, and the SLO is normally the tighter of the two — deliberately, so that you notice and react well before you owe anyone credits.
  6. Latency, traffic, errors, saturation. Saturation is the resource-level one — how full the thing is, relative to what it can hold. It is most often missing because the other three fall out of request logs and a load balancer with no extra work, while saturation requires you to know what the binding resource actually is for this service, and to be honest about the fact that it might be the connection pool rather than the CPU. It is also the most predictive of the four: saturation degrades before errors appear.
  7. Yes, it is toil — it is manual in origin, repetitive, automatable-away rather than merely automated, tactical rather than enduring, and it produces no lasting improvement; it scales linearly with the number of services you do it to. The cron job is not the fix; it is toil that has been hidden rather than removed, and it makes the underlying defect invisible to the people who could repair it. The right fix is to find why the service needs restarting — a leak, an unbounded cache, a connection that is never released — and repair that, with the restart kept only as a temporary guard with an alert attached so that somebody still knows it is firing. Toil & automation covers the full definition.
  8. MTTD is the mean time to detect — from the failure starting to somebody or something noticing. MTTR is the mean time to repair or recover, and MTRS the mean time to restore service; the official concept list names all three separately, and in the wider industry the "R" in MTTR is expanded inconsistently, so always say which definition you are using. MTTR can fall while users suffer more when: a flood of short, auto-resolved, low-impact incidents drags the mean down while the few long ones that actually hurt are unchanged; the clock starts at detection, so getting slower at detecting failures paradoxically improves the number; or incidents are marked resolved when the symptom clears rather than when the fix is verified. A mean over a heavily skewed distribution is a poor summary — which is the same lesson as question 4, arriving from the other direction.
  9. Blue/green reduces rollback time: two complete environments, traffic cut from one to the other, and reversal is another cut rather than a redeployment. Its weakness is that the switch is all-or-nothing, so the first evidence of a bad release is every user getting it. Canary reduces blast radius and, more importantly, produces evidence: a small slice of real traffic on the new version, its error rate and latency compared against the baseline, and promotion only if the comparison holds. Canary is the one that gives you a signal; blue/green is the one that gives you an exit. Many teams run both. Release engineering & progressive delivery covers the combination.
  10. Pathological, in Westrum's typology — power-oriented, where information that reflects badly on someone is suppressed and messengers are punished. (Bureaucratic is rule-oriented; generative is mission-oriented.) The practical cost is not moral, it is informational: people stop volunteering the details that make root causes findable, near-misses go unreported entirely, and the organization loses its cheapest source of reliability data. That is precisely why "blameless" is treated in the syllabus as a mechanism rather than a courtesy — see postmortems & blameless culture.
🦉 Professor Owl's free diagnostic · 60 min

The DevOps Institute's free SREF Exam Study Guide PDF contains a complete 40-question sample paper with an answer key, plus the topic weighting table and the concept list this page has quoted throughout. It is vendor-authored rather than a leaked dump, it is free, and it is the closest thing to a real exam form that exists in public. Sit it cold, before any study, under a 60-minute clock. Then read your result the useful way: not as a score, but as a map of which baseline area each miss belonged to.

Two caveats to hold while you do it. It is the v1.0/v1.1 sample, so it cannot contain the material added by the v1.2 update in October 2024. And its administration table is out of date — see the next callout.

⚠ "Supervised: No" in that PDF is stale — the exam is proctored

The free study guide's exam-format table lists Supervised: No. That reflects the pre-PeopleCert delivery platform and no longer describes the exam. DevOps Institute certifications migrated to the PeopleCert platform in late 2023, and PeopleCert's SRE Foundation is delivered online with live human proctoring. Chris McLeod's December 2023 account describes exactly that: a real proctor, a webcam room scan including over and under the desk, and a requirement to reposition his desk and camera so that he, the room's door and a cupboard door were all visible at once. He also reports a provisional score at the end of the session with official confirmation following a couple of days later, rather than within hours. Use the PDF for the syllabus map; use PeopleCert's own online-proctoring documentation, and this course's exam guide, for what exam day actually involves.

Score yourself — the baseline self-test

☺ Like you're 10: Tick only what you could genuinely do right now. The empty boxes are the useful part.

Two tiers, for the two standards from the top of this page. Tier 1 is the vocabulary floor — the standard the exam actually applies. If you cannot tick these, blueprint study will be uphill, because you will be translating rather than learning. Tier 2 is the production floor — each item asks whether you have done a thing, not whether you could describe it. Be strict with yourself on Tier 2 in particular; it is the tier that decides the last two options on a best-answer question, and the one nobody can bluff in an interview. Your ticks are saved on this device.

Tier 1 — the vocabulary floor

Tier 2 — the production floor

What your score means

Count the unticked boxes in each tier separately. The gaps are the signal; a combined total tells you nothing useful, because the two tiers fail in completely different ways and are closed by completely different work.

UntickedTier 1 — vocabularyTier 2 — production
0–2Ready. Start the SREF study plan at week one, todayReady for full-length timed practice — go to Mock Exam Set 1 and sit it under a real clock
3–6One focused week. Re-read the matching sections above, then work the flashcards and the glossary entries you stumbled on, and re-take this testTwo to four weeks. Work the SLO drill, the toil audit and the incident tabletop — they exist precisely to convert these boxes into ticks without needing a production outage
7+Do the foundations properly first: What is SRE?, then SLIs, SLOs & error budgets, monitoring and incident management. Blueprint study on this foundation will not stickYou can still pass — this exam is genuinely open to people without production experience, and that is not a criticism of it. But be honest with yourself about what the badge will and won't say about you. The reliability capstone is the closest substitute this course can offer
◆ Key idea

The two tiers can legitimately diverge, and the direction tells you what to do next. Strong Tier 1, weak Tier 2 is the well-read profile: you will pass, comfortably, and you should — but the exam will have measured almost none of what a hiring manager will ask you about three weeks later, so pair the badge with something you actually built. Strong Tier 2, weak Tier 1 is the practitioner profile, and it is the cheaper gap of the two by a wide margin: you already have the judgement, and what you are missing is the syllabus's specific names for things you do every week — VALET, Westrum, MTRS, the Three Ways. That is a fortnight of vocabulary, not a career of experience. Do not let it convince you the exam is beneath you; the terminology list is short, free, and the difference between 24 and 28 correct.

Closing the gaps you found

☺ Like you're 10: You now have a specific list. Fix the things on it — don't start the whole course over.

Resist the urge to re-read everything. You have a named list of gaps, and named lists are closed by specific work rather than by general study.

For vocabulary gaps

Say the item out loud, from memory, in one paragraph — then check it against the section above and mark what came out thin. That retrieval loop is the entire method, and it is what the closed-book strategy page builds a schedule around. The flashcards and the glossary are for the terms you keep dropping; Know It Cold is the shortlist of definitions that have to be exact rather than approximately right; and the concept reference is the lookup when a term needs more than a card. Print the official concept list alongside them and use it as your checklist of record — it is the exam's own inventory, and nothing this course writes outranks it.

For production gaps

Experience is only ever built one way, but it can be simulated more cheaply than people expect. The toil audit makes you apply a definition to your own week. The SLO and error-budget drill makes you compute a budget for something real. The incident tabletop is a rehearsal you can run at a desk with colleagues, which is most of what on-call training actually is anywhere. The alert-design drill and the postmortem-writing drill close the two most common Tier 2 gaps on this page. And if you want the whole arc rather than a piece of it, the reliability capstone runs from defining SLOs through building monitoring to an incident and its postmortem, in order.

Then, and only then, start the blueprint

Once this page's boxes are ticked, everything downstream gets cheaper — because "error budget" is now arithmetic you can do rather than a phrase you recognise, "saturation" is a thing you have watched fill up, "toil" is a task you can name from last Tuesday, and "blameless" is a mechanism with a reason rather than a nice sentiment. That is the whole return on this week: not new knowledge, but the disappearance of the quiet translation cost that makes foundation-level study feel harder than it is. Go to SRE-1 and work forward, or straight to SRE-2 and SRE-4 if you want the heaviest areas first.

⚠ Verify the numbers before you book

Every exam figure on this page comes from an official source and is stated as true at the time of writing. PeopleCert's SRE Foundation product page is the authority for the commercial and delivery details: 40 questions, 60 minutes, a minimum passing score of 65%, open book restricted to the Official Training Materials, no prerequisites, online proctored delivery, and renewal every three years under its Continuing Professional Development programme. The DevOps Institute's free SREF Exam Study Guide is the authority for the syllabus map: the eight topic areas and their maximum question counts, the Bloom 1–2 difficulty ceiling, the recommended 16 contact hours, the concept and terminology list and the sample paper.

Two honest caveats. First, that study guide's administration table (notably "Supervised: No") predates the PeopleCert migration and no longer describes delivery. Second, the renewal period is genuinely inconsistent across sources: the DevOps Institute's certification page states three years in its table while its own overview prose elsewhere on the page says two, and some training partners have advertised these certifications as never expiring. Three different claims. We are not going to pick one for you — confirm the renewal terms with PeopleCert directly before you assume anything about how long the credential lasts. Price is regional and moves; confirm that too. Our SREF exam guide carries the full logistics briefing.

🎬 At the Reliability Watch
🦫

Benny the Beaver: Booked the SREF for a fortnight's time. No prerequisites, it says. So there's nothing I need to do first, right?

🦊

Foxy: Quick one before you celebrate. Monthly SLO of 99.95%. How many minutes?

🦫

Benny the Beaver: Er. It's less than the 43 one. Twenty-something? Let me open a calculator —

🦥

Sol the Sloth: Twenty-one point six. One percent of a thirty-day month is about seven hours, and every extra nine divides by ten. Seven hours, forty-three minutes, four minutes, twenty-six seconds. That's the whole ladder, and it's the only sum on the paper.

🐘

Ellie the Elephant: Which is the heaviest area, Benny — the one worth up to seven of the forty?

🦫

Benny the Beaver: ...toil? I'm good at toil. I've made loads of it.

🐘

Ellie the Elephant: Monitoring and service level indicators. Toil is five. And it's published — the study guide prints the maximum per topic area, free, for anyone who opens it.

🐢

Timmy the Turtle: Note what Benny does have, though. He's watched a service fall over at midnight. He knows what a restart hides. That's the expensive half, and he already owns it.

🦉

Professor Owl: Which is the useful shape of this. Benny's gap is a fortnight of vocabulary. Someone arriving with no production experience at all can still pass this exam — it's built at knowledge and comprehension level and it says so — but they'd be carrying the harder gap, and it's worth knowing which one you have.

🐦

Pip the Hummingbird: Also — open book, Benny. Official training materials only. Did you buy the course?

🦫

Benny the Beaver: ...no.

🐰

Remy the Rabbit: Then it's closed book. Ninety seconds a question. Flashcards. Now. Go.

🐢 Timmy's checkpoint

1. The SREF has no formal prerequisites — so why does this page exist, and what exactly does the exam assume that it never teaches? 2. What Bloom levels is the exam built at, and what does that rule out appearing on the paper? 3. Which topic area is the heaviest, how many questions can it reach, and which two baseline areas sit underneath it? 4. Convert 99.9% and 99.99% monthly SLOs into minutes, from memory, and state the shortcut you used. 5. What does "open book" actually permit on this exam, and what does that mean for a self-studier? 6. Name two ways a working SRE's intuition can be marked wrong on this paper. 7. What is stale about the free official study guide, and what is still the best free thing in it? 8. What is the practical difference between a Tier 1 gap and a Tier 2 gap on this page's self-test?

Check your answers
  1. "No prerequisites" is a booking rule, not a readiness assessment — it means anyone may sit the exam, not that the material assumes nothing. The syllabus is written in the vocabulary of production operations and its own official concept list names Kubernetes, AWS, cloud computing, auto-scaling, software-defined networking, immutable infrastructure, APM, canary testing and self-healing without ever teaching any of them. The five assumed areas are Linux and the production OS, distributed-systems behaviour, availability arithmetic, monitoring concepts, and on-call and process practice.
  2. Bloom 1 and 2 only — knowledge and comprehension. That rules out Bloom 3 and above: nothing on the paper asks you to apply a technique to a genuinely novel situation, analyse a system, evaluate a design or create anything. It is why the exam is broad rather than deep, and it is stated by the DevOps Institute in its own study guide rather than being an inference.
  3. SRE-4, Monitoring and Service Level Indicators, at up to 7 of 40 — the heaviest single area. Underneath it sit monitoring concepts (metrics, time series, percentiles, signals, telemetry vocabulary) and availability arithmetic, since the SLI/SLO relationship it examines is the numeric one.
  4. 99.9% → 43.2 minutes; 99.99% → 4.32 minutes. The shortcut: a 30-day month is 43,200 minutes, so 1% of it is about 7.2 hours, and every additional nine divides the budget by ten — 7.2 hours, 43 minutes, 4.3 minutes, 26 seconds. Produce the ladder rather than memorising rows and every version of the question is covered.
  5. It permits the Official Training Materials only — supplied by PeopleCert or through an accredited training organization — and nothing else. For a self-studier who has not bought an accredited course, that means there is nothing permitted to bring, so the exam is functionally closed book; Chris McLeod's write-up describes exactly that situation. And at 40 questions in 60 minutes, 90 seconds each, there is no time to study from materials even if you have them.
  6. Any two of: when a question names its source ("according to X"), it is testing recall of that source rather than your understanding, and answering from experience loses the mark; reliability is targeted rather than maximised, so an unspent error budget is a signal of over-investment rather than a success; MTTR is expanded inconsistently across the industry, so answer from the syllabus's definitions; and the syllabus carries named models you may never have met at work — VALET, Westrum's typology, the Three Ways — which are vocabulary to learn rather than intuition to derive.
  7. Stale: it is the v1.0/v1.1 guide, so it predates the v1.2 update of October 2024 that added generative AI, AIOps, platform engineering, value stream management, observability and progressive deployments — and its administration table says "Supervised: No", which no longer describes a live-proctored PeopleCert sitting. Still excellent: the topic weighting table, the Bloom 1–2 statement, the concept and terminology list, and a complete 40-question sample paper with answer key — all free, all vendor-authored.
  8. A Tier 1 gap is vocabulary — it costs marks directly, it is closed by retrieval practice against flashcards, the glossary and the official concept list, and a fortnight is usually enough. A Tier 2 gap is experience — it costs you very little on this particular exam, because Bloom 1–2 recognition will carry you, but it costs you on the scenario-flavoured best-answer items, in interviews, and in the job. The two are closed by different work, which is why the score table reads them separately rather than adding them up.