How to Study for the Exam
Most people who fail the SRE Foundation (SREF) do not fail because the material was beyond them — this is an entry-level exam, and its own syllabus caps the difficulty on purpose. They fail because they studied the way they studied at school: reading front to back, highlighting, re-reading the highlights. That method is almost perfectly mismatched to what this exam measures. SREF is a knowledge exam, delivered by PeopleCert on behalf of the DevOps Institute: multiple choice, no terminal, no cluster, no dashboard. Nobody watches your hands. What gets scored is whether you can hold four plausible sentences side by side and pick the one that is true by the syllabus's own definitions — roughly one every ninety seconds, for an hour. This page is the method, not the material. The eight blueprint pages are the syllabus, the SREF study plan is the calendar, and this is how to work through them so any of it is still there when the clock starts.
There are two kinds of test. In the first, someone hands you a bike and says "ride to the end of the street and back, you have four minutes." In the second, someone hands you a sheet of paper with four sentences about bikes and asks which one is the correct explanation of why a bike stays upright. You would not train for those the same way. For the first you get on a bike, over and over, until your legs stop thinking. For the second you have to be able to explain — because all four sentences will sound sensible, and three of them are slightly wrong in a way you only notice if you know the real explanation cold. This exam is the second kind. Reading the bike book again does not help. Being asked, out loud, with the book closed, does.
What kind of exam this is — and why that changes everything
☺ Like you're 10: One kind of test watches your hands. This one reads your mind. You practice hands by doing; you practice knowing by explaining, and by arguing with wrong answers.
Everything on this page follows from a single distinction, so get it straight before you plan an hour. A performance-based exam grades the state you leave behind. Nobody cares how you got there — if the service ends up healthy, you scored, whether you typed it, scripted it, or copied it from a runbook. A knowledge-based exam grades which option you selected. Nobody cares whether you could have built the thing; they care whether you can tell one correct statement from three that are also plausible.
SREF is squarely the second kind, and its published format leaves no ambiguity about that. Both official sources — the PeopleCert product page and the DevOps Institute certification page — describe the same paper: 40 multiple-choice questions, 60 minutes, 65% to pass, web-based delivery, no formal prerequisites. There is no lab component and nothing to build. (Those figures were current when this page was written. Confirm them on the vendor pages before you book — exam formats get revised, and this one has already changed platforms once.)
This distinction matters more on this site than on most, because the Certifications page in this course also covers exams of the other kind, and the performance-based exams in the sister Platform Engineering course are graded entirely on cluster state. If you are working toward both, do not run one study regime for both. They reward opposite skills, and a regime tuned for one shortchanges the other.
The syllabus tells you how hard to study — literally
Most exams make you guess at depth. SREF does not. The official SREF Exam Study Guide states that the certification is built on the Bloom taxonomy and that the paper contains Bloom level 1 questions (knowledge of SRE concepts and vocabulary) and Bloom level 2 questions (comprehension of those concepts in context). Those are the only two levels it names.
Read that as an instruction, because it is one. Level 1 and level 2 mean the exam asks you to define and to recognize in context. It does not ask you to design an SLO regime for a described organization, weigh two architectures against each other, or critique a postmortem process — those are levels 3 and up, and they belong to higher-level credentials such as SRE Practitioner, and to job interviews, not here. So the study target is not "be a better site reliability engineer." It is: be able to state, precisely and quickly, what every term on the syllabus means, and recognize it when it is dressed up in a one-paragraph scenario.
A Bloom 1–2 ceiling is the most useful sentence in the whole official study guide. It tells you that breadth beats depth here: eight topic areas at a definitional level will pass you comfortably, and three topic areas understood at architect level will not. If you catch yourself two hours deep in multi-window burn-rate alerting math for this exam, you are studying above the ceiling — that is a deep-dive page, not an SREF question.
The arithmetic that drives every tactic on this page
Three numbers, and everything downstream of them. All three are derived from the published format above, so re-derive them yourself if the vendor revises the paper.
| Derived number | Working | What it changes about how you study |
|---|---|---|
| 90 seconds per question | 60 minutes ÷ 40 questions | There is no look-up window. Whatever is not already in your head is not going to arrive in time, however the reference rules are worded |
| 26 correct to pass | 65% of 40 | Your target is a raw count, not a feeling. Score every practice paper out of 40 and compare against 26 |
| 14 wrong is still a pass | 40 − 26 | One weak topic area does not sink you — but two do. This is why coverage beats depth here |
That third row is worth sitting with, because it cuts both ways. Fourteen is a generous allowance by exam standards, and it is why a candidate who has honestly covered all eight topic areas at a definitional level rarely fails. It is also exactly why skipping an area is fatal: the heaviest area is worth seven questions on its own, so writing one off spends half your entire margin before you have answered anything.
"Open book" is not the safety net it sounds like
Both official sources list this exam as open book — and PeopleCert's product page qualifies that in a way most third-party summaries drop: the permitted material is the Official Training Materials, whether supplied directly by PeopleCert or through an Accredited Training Organization. Not your notes. Not this site. Not a search engine. If you self-study without buying those materials, you have nothing admissible to bring, and you are sitting a closed-book exam with an open-book label on it.
Even if you do have them, the arithmetic above closes the gap: ninety seconds a question is enough to check a term you nearly know, and nowhere near enough to learn one you do not. Chris McLeod, who sat and passed the exam on 9 December 2023 and wrote it up on his own blog the following week, makes precisely this point from the other side: open book only helps if you have the official training materials to hand, and he did not, so he answered from memory. His account describes the exam as broad rather than deep — which is the Bloom 1–2 ceiling showing up in practice, from a candidate who had never read that table.
Other pages here — notably Closed-Book Strategy — train you as if no reference existed at all. That is deliberate, and for most readers it is the accurate posture: no official materials in hand, and no time to use them if you had. Treat "open book" as a technicality that changes nothing about your preparation, and confirm the current reference policy on PeopleCert's own page before you book — it is one of the details third-party sites most often get wrong, and one candidates most often plan around incorrectly.
Retrieval, not re-reading
☺ Like you're 10: Reading a page again feels like learning because the words look friendly. Shutting the page and saying it out loud feels awful — and that is the one that works.
Here is the single most important sentence on this page: the act of pulling something out of your head with the page shut is what moves it from "recognized" to "owned." Everything else here is a wrapper around that move.
Re-reading fails for a specific and slightly cruel reason. The second pass over a page is fluent — the sentences arrive smoothly, nothing surprises you — and your brain reads that fluency as evidence of knowing. It is not. It is evidence that the page is nearby. Remove the page, and the fluency goes with it, which is exactly the condition the exam creates. Retrieval feels worse and works better: it is effortful, you fumble, and the fumbling is the part that does the work.
For SREF, retrieval has two units, and you need both.
Unit one — the definition
A single sentence you can produce cold, in your own words, that would survive a pedant. "An SLI is the measurement itself — a metric of some aspect of the service, taken at the boundary the user actually crosses." Not "SLIs are about measuring things." The difference between those two sentences is roughly the difference between passing and failing, because the second one cannot reject a distractor and the first one can.
The discipline: pick a term from the glossary or from the official study guide's concept list, say your sentence out loud before you look, then compare. Where you were vague, you have found a real gap. Where you were confidently wrong, you have found a much more valuable one.
Unit two — the discrimination
A knowledge exam is not really testing whether you know the right answer. It is testing whether you can reject three wrong ones that were written to be attractive. That is a separate skill, and it is trained separately: given four options, say why each of the other three is wrong, in words, before you commit. Not "this one feels right" — an actual reason.
This is where the self-check and the SREF practice bank earn their place. Three rules make them training rather than entertainment:
- Before you pick, argue down the other three out loud. If you cannot say why an option is wrong, you do not yet know the concept it is testing — you are pattern-matching on familiar words.
- When you get one right, read the explanation anyway. Right answers arrived at for wrong reasons are invisible unless you check, and they are the ones that fail you on a differently-worded question.
- When you get one wrong, classify it before moving on. Did you not know the fact, did you confuse it with its neighbor, or did you misread the question? Those three need completely different fixes, and the mistake log below is where the classification lives.
How to tell "I recognize this" from "I can do this"
Recognition is the counterfeit currency of exam prep. It is produced by every pleasant study activity — reading, watching, listening, highlighting — and it is spent by exactly none of the exam. Four tests separate them, and all four are cheap:
| Test | How to run it | What a fail looks like |
|---|---|---|
| The blank page | Close everything. Write the topic area's key terms and one-line definitions from memory | You produce four of eleven terms and stall. You had "read" all eleven an hour ago |
| The five-second rule | On a flashcard, start answering within five seconds or mark it unknown | You take twenty seconds, get there, and tell yourself you knew it. At 90 seconds a question, you did not |
| The teach test | Explain the concept to someone who does not work in infrastructure, with no jargon | You can only say it in the syllabus's own phrasing. You have memorized wording, not meaning |
| The neighbor test | State the difference between the term and its nearest neighbor (SLO vs SLA, MTTR vs MTRS) | You can define each separately but blur when asked to contrast. This is the exact seam distractors are built on |
Run the neighbor test most often. On an exam capped at Bloom 2, near-neighbor confusion is the dominant failure mode — not ignorance. There is a whole table of the pairs worth drilling further down this page.
Feeling familiar is a property of the page. Being able to produce it is a property of you. Every hour you spend making pages feel familiar is an hour that will not show up in your score. Study in the direction of production: shut the tab, say the sentence, then check.
Read the syllabus as a checklist, not as prose
☺ Like you're 10: The exam publishes a list of everything it is allowed to ask, and even says how many questions each part gets. Turn that into a to-do list with ticky boxes instead of reading it once like a story.
The single most wasted asset in SREF prep is the official study guide, because almost everyone reads it once, nods at the topic list, and goes off to watch videos. It is not an introduction. It is the specification of the question pool: nothing outside it can be asked, and everything inside it can.
The eight topic areas, and exactly what each is worth
Here is the part most third-party study advice gets wrong by omission. The DevOps Institute's certification web page publishes no percentage split — but the official Exam Study Guide PDF carries a table headed Exam Topic Areas and Question Weighting, with a maximum question count for each of the eight areas. That table is reproduced below, alongside where each area lives in this course.
| # | Topic area | Max questions | Where it is taught here |
|---|---|---|---|
| SRE-1 | SRE Principles and Practices | 4 | Blueprint 2.1 · What is SRE? |
| SRE-2 | Service Level Objectives and Error Budgets | 6 | Blueprint 2.2 · SLIs, SLOs & error budgets |
| SRE-3 | Reducing Toil | 5 | Blueprint 2.3 · Toil & automation |
| SRE-4 | Monitoring and Service Level Indicators | 7 | Blueprint 2.4 · Monitoring & observability |
| SRE-5 | SRE Tools and Automation | 6 | Blueprint 2.5 · The SRE toolchain |
| SRE-6 | Anti-fragility and Learning from Failure | 4 | Blueprint 2.6 · Chaos engineering |
| SRE-7 | Organizational Impact of SRE | 4 | Blueprint 2.7 · Postmortems & blameless culture |
| SRE-8 | SRE, Other Frameworks, Trends | 4 | Blueprint 2.8 · SRE team topologies |
Two observations, and both are ours, derived from that table rather than claims the vendor makes:
- The maxima sum to exactly 40. 4 + 6 + 5 + 7 + 6 + 4 + 4 + 4 = 40. A set of "maximums" that adds up to precisely the paper length has very little room to vary — every area has to sit at its maximum for the arithmetic to work. Corroborating this, the 40-question sample paper inside the same official document is laid out in exactly that pattern: its own answer key tags questions 1–4 to area 1, 5–10 to area 2, 11–15 to area 3, 16–22 to area 4, and so on down to 37–40 for area 8.
- Three areas carry 19 of the 40 marks. Monitoring & SLIs (7), SLOs & error budgets (6) and Tools & automation (6) are, between them, worth nearly half the paper — and just seven short of the 26 you need to pass.
What this does not license is writing off the four-question areas. Four questions is 10% of the paper each, and the four smallest areas together are 16 of 40 — two more than the 14 wrong answers you are allowed. Write all four off and you have failed before the paper starts, however solid the heavy three are. The weighting decides order and hours, not what you can skip. Study everything; study the heavy three first and revisit them last.
The weighting above comes from the freely published v1.1 study guide (September 2021). The syllabus has moved since — see the next section — and PeopleCert has not published a comparable free table for the current version. Treat the counts as the best available public map, not as a guarantee about the paper you will sit, and check the vendor pages for anything newer before you commit study hours on the strength of a single row.
Turn each topic area into claims you can defend
Syllabus lines are written as nouns: "Reducing Toil." Nouns cannot be ticked off, because there is no moment at which you have finished a noun. Rewrite each one as a sentence you can either say or not say:
| Published topic (noun) | Rewritten as a claim you can defend |
|---|---|
| Understanding toil and why it is bad | "Toil is manual, repetitive, automatable, tactical, devoid of enduring value, and it scales O(n) with the service" — and I can name a task that fails exactly one of those tests and therefore is not toil |
| Service level objectives and error budgets | "A 99.9% monthly SLO leaves roughly 43 minutes of error budget" — and I can produce that number in under thirty seconds without a calculator |
| Error budget policies | "The policy is the pre-agreed rule for what changes when the budget is exhausted" — and I can state why a policy agreed after the budget burns is not a policy |
| Understanding SLIs and how they relate to SLOs | "The SLI is the measurement, the SLO is the target set against it, the SLA is the contractual floor" — in that fixed order, and I can say why the SLA is looser |
| Anti-fragility defined | "Resilient systems survive stress unchanged; anti-fragile systems get stronger because of it" — and I can put chaos engineering on the correct side of that line |
| Patterns for SRE adoption | "Here are two named adoption patterns and the organizational trade-off each one makes" — from team topologies, stated in one sentence each |
Notice the shape. Every rewrite is a sentence you could say out loud, and most carry a second clause that proves you did not just memorize the first. If your rewrite has neither, you have written a mood, not a study item.
Build a confidence ledger
Now make it queryable. A plain text file beats any app, because you want to sort and count it without leaving the terminal you already live in. One line per topic area, carrying the published question count so the file can do the hour arithmetic for you:
mkdir -p ~/sref-prep && cd ~/sref-prep printf 'area|title|maxq|conf|last\n' > ledger.psv cat >> ledger.psv <<'EOF' SRE-1|SRE Principles and Practices|4|0|never SRE-2|Service Level Objectives and Error Budgets|6|0|never SRE-3|Reducing Toil|5|0|never SRE-4|Monitoring and Service Level Indicators|7|0|never SRE-5|SRE Tools and Automation|6|0|never SRE-6|Anti-fragility and Learning from Failure|4|0|never SRE-7|Organizational Impact of SRE|4|0|never SRE-8|SRE, Other Frameworks, Trends|4|0|never EOF
Score each line 0–3 on a scale that has nothing to do with feelings. 0 = never studied. 1 = read the blueprint page once. 2 = wrote the area's terms out from a blank page, with gaps. 3 = wrote them out cleanly and scored 80%+ on that area's practice questions, cold, on two different days. Only a 3 counts as done — anything looser lets "I read that on Tuesday" masquerade as knowledge, and that is exactly the item that collapses at question 31.
Then the file answers the only question that matters at the start of a session: what should I work on today?
cd ~/sref-prep
awk -F'|' 'NR>1 {printf "%-6s %-46s %sq conf=%s exposure=%s\n", $1, $2, $3, $4, $3*(3-$4)}' ledger.psv | sort -t= -k3 -rn
awk -F'|' 'NR>1 {t+=$3*(3-$4)} END {printf "total exposure: %d (0 = ready)\n", t}' ledger.psvThat exposure column is the whole point: questions at stake multiplied by how far you are from owning the area. SRE-4 at confidence 1 scores 14; SRE-8 at confidence 1 scores 8. Both need work, and the ledger has just told you which one to open first — without any appeal to which topic you happen to find interesting. Total exposure starts at 120 and should be walking toward zero. If it has not moved in a week, you have been reading, not retrieving.
Let the counts allocate your hours
This course budgets roughly 20–25 hours in total for SREF prep, which is proportionate to a 60-minute foundation paper — see the study plan for the calendar version. Split a 24-hour budget by published question count and it lands like this:
| Topic area | Max questions | Share of paper | Hours out of 24 |
|---|---|---|---|
| SRE-4 Monitoring & SLIs | 7 | 17.5% | ~4 |
| SRE-2 SLOs & error budgets | 6 | 15% | ~3.5 |
| SRE-5 Tools & automation | 6 | 15% | ~3.5 |
| SRE-3 Reducing toil | 5 | 12.5% | ~3 |
| SRE-1 Principles & practices | 4 | 10% | ~2.5 |
| SRE-6 Anti-fragility | 4 | 10% | ~2.5 |
| SRE-7 Organizational impact | 4 | 10% | ~2.5 |
| SRE-8 Frameworks & trends | 4 | 10% | ~2.5 |
Allocating purely by question count assumes you start from zero everywhere, and nobody does. Multiply each row by your gap — that is what the ledger's exposure column computes. An on-call veteran may be able to take SRE-4 to confidence 3 in a single session and should spend the saved time on SRE-8, which is the area experience does not hand you. The counts break ties; the ledger sets priority.
The syllabus has moved since the free material was written
☺ Like you're 10: The free study guide you can download is from a few years ago. The real exam got updated since. Most of it is the same — but some new topics were added, and you would never know from the free file.
This is the trap that catches careful, thorough candidates, so it gets its own section. The free study guides that everybody finds — and that this page points you at, because they are genuinely excellent — are v1.0 (2020) and v1.1 (2021). PeopleCert's official training-material update notices list the current SRE Foundation syllabus as v1.2, updated October 2024, and describe that update as adding insights on modern DevOps practices including generative AI, AIOps, platform engineering, value stream management, observability and progressive deployments, alongside refreshed references.
Two consequences, and neither is optional:
- Study the v1.1 guide for structure, not for scope. Its eight topic areas, its Bloom ceiling and its sample paper are all still the best free map of how the exam thinks. Its topic list is one revision behind.
- Its administration details are stale. The v1.1 format table says Supervised: No — a leftover from the pre-PeopleCert delivery platform. The exam is now delivered through PeopleCert, and the online-proctored route — the one you get if you book the exam yourself rather than sitting it at a training centre — involves a live human proctor, an ID check and a webcam room scan; McLeod's December 2023 account describes exactly that, including panning the camera over and under his desk. Do not plan your exam-day environment from a 2021 PDF.
The good news for anyone studying here: the v1.2 additions map onto pages this course already has, so covering them costs you an evening rather than a rewrite.
| Added in v1.2 (Oct 2024) | Read here | Why it is likely to be tested at Bloom 1–2 |
|---|---|---|
| Observability | Monitoring & observability | The monitoring-vs-observability distinction is a definitional question waiting to happen, and SRE-4 is the heaviest area |
| Progressive deployments | Release engineering & progressive delivery | Canary, blue-green and automated rollback are already named concepts in the syllabus's own terminology list |
| AIOps | The SRE toolchain | Sits naturally in SRE-5, where the syllabus asks for a tooling-landscape overview rather than tool expertise |
| Platform engineering | The Platform Engineering course | An SRE-8 "other frameworks" item: know what a platform team is and how it differs from an SRE team |
| Value stream management | Measuring & reporting reliability | Another SRE-8 neighbor concept — flow and value measurement next to reliability measurement |
| Generative AI | Best practices & the SRE operating model | Expect vocabulary and framing, not implementation. Bloom 1 means "what is it and where does it fit" |
Three passes over one syllabus
☺ Like you're 10: Go over everything three times — fast, then slow, then from memory. Not once, slowly, forever.
The instinct is to start at topic area one and go deep until it is perfect, then move on. That produces a candidate who is superb on error budgets and has never read the word "anti-fragility," because time ran out. Three passes over the whole syllabus beats one perfect pass over a third of it — especially on a paper where every area is guaranteed to appear and a zero anywhere costs the same as a zero everywhere.
Pass 1 — survey (about 15% of your time, so 3–4 hours)
Skim all eight blueprint pages and the official study guide's topic table. Take no notes. The only artifact you produce is the ledger, scored honestly. This pass exists to kill surprises: by the end there should be nothing in the syllabus you have never heard of. It will feel unproductive. It is buying you the ability to plan, and it is the pass people skip and then regret.
Pass 2 — read and produce (about 60%, so 14–15 hours)
The long middle, worked in exposure order. For each area: read the blueprint page, close it, write the area's terms and definitions out from a blank page, then run that area's practice set — principles, SLOs & toil, monitoring & tools, or anti-fragility & organizational impact. Finish every session by updating the ledger. An area reaches confidence 2 when the blank page comes out mostly right; it reaches 3 later, in pass three, when it survives a timed paper.
Pass 3 — compress (about 25%, so 6 hours)
Nothing new is allowed in pass three. You are doing three things only: retrieving from a blank page, drilling the lines your ledger and mistake log agree on, and sitting timed papers from the five SREF mock exams. If you find yourself learning a brand-new concept in pass three, you have mis-scheduled — bank it as a known gap, accept the four-question risk, and protect the revision time instead.
Save at least two of the five mock sets for pass three, unseen. A paper you have already worked through measures your memory of that paper, not your readiness. The single most common self-deception in the last week is a 92% on a set you sat a fortnight ago.
Spacing — review just before you forget
☺ Like you're 10: Things you learn leak out of your head. Grab them again just as they are about to leak and they stick much longer — and you have to do it less and less often.
Memory for anything you are not using daily decays, and it decays fastest right after you learn it. Reviewing something just as it starts to fade resets the clock and flattens the curve, so each successive review can sit further out than the last. That is the whole idea: review at increasing intervals, timed to just-before-you-forget. It is not a productivity fad; it is the most evidence-backed study technique there is, and it costs nothing.
It matters more here than it would on a longer, deeper exam, because SREF prep is short. Twenty-odd hours spread over four to six weeks means most of what you learn in week one has to survive three or four weeks of not being touched. Without spacing, it will not.
A schedule you will actually keep
Ignore the elaborate algorithms. A fixed 1 / 3 / 7 / 21 ladder captures nearly all the benefit and needs no software: learn something, revisit it tomorrow, then three days later, then a week later, then three weeks later. Anything you fumble on a review drops back to the start of the ladder. Anything you nail at +21 is done — trust it and stop spending time on it.
The cheapest implementation is a dated file. Add a line when you learn something; each morning, grep for today:
cd ~/sref-prep
item='SRE-2: 99.9% monthly SLO -> ~43 min error budget, from memory'
for d in 1 3 7 21; do
printf '%s|%s\n' "$(date -v+${d}d +%F)" "$item" >> reviews.psv
done
grep "^$(date +%F)|" reviews.psv || echo "nothing due today — go take a fresh area instead"That date -v+1d form is BSD/macOS; on Linux it is date -d '+1 day' +%F. Either way, the point is that scheduling a review must cost you one line, or you will stop doing it by Thursday.
A 1/3/7/21 ladder started on day one lands its final review on day twenty-two. If your exam is four weeks out, everything you learn after the first week gets a truncated ladder — 1/3/7 at best. That is a real argument for front-loading the heavy areas (SRE-4, SRE-2, SRE-5) into week one, where they get the full ladder, and letting the four-question areas take the shorter one. It is also an argument against booking the exam for next Saturday.
The three drills on this site, and the order to run them in
☺ Like you're 10: Cards to find out what you have forgotten, questions to practice picking the right one, and a checklist to prove you can actually apply it. Different jobs — do not use one and skip the others.
This course ships three general-purpose drills plus a whole SREF-specific shelf. They are not interchangeable, and most people misuse the first one by treating it as training rather than as a diagnostic.
Flashcards — a diagnostic you have to schedule yourself
The flashcards here are 38 cards drawn from every module, laid out as a searchable grid you flip by clicking. Be clear about what that is and is not: it is a deck, not a scheduling engine. Nothing tracks which cards you fumbled or brings them back tomorrow — that bookkeeping is yours, which is exactly why the reviews.psv file above exists. Search the deck by topic, work the cards for the area you studied today, and log the misses into the ladder.
Two rules make the deck work, and both are about honesty. Answer out loud before you flip. Flipping first and thinking "yes, I knew that" trains recognition, which is the currency this exam does not accept. Use the five-second rule. Anything you cannot start answering within five seconds is not known, however familiar it feels — at 90 seconds a question you cannot afford a term that takes twenty seconds to surface.
Self-check — where discrimination is actually trained
The self-check is the closest thing here to the exam's own mechanic: pick a domain, take ten questions cold, commit, and get the explanation immediately. Options are reshuffled every session, so you are recalling the answer rather than remembering that it was the third one — and anything you miss is remembered on your device, so ↻ Past misses always has your weak spots queued up. That misses queue is the most valuable feature on the page and the one people never open.
Run it with the three rules from the retrieval section: argue down the other three options before committing, read the explanation even when you were right, and classify every miss. Ten questions done that way is worth an hour of questions done quickly.
The checklist — the one drill that is not about the exam
The on-call readiness checklist is not an exam-readiness checklist, and it would be dishonest to sell it as one. It is a phase-by-phase pass over SLO and measurement readiness, on-call and incident readiness, reliability engineering readiness, and reporting — with progress saved in your browser.
Use it anyway, for a specific reason. An exam capped at Bloom 2 still asks you to recognize concepts in context, and the fastest way to make a definition concrete is to try to apply it to a real service. Walking the checklist against something you actually operate turns "error budget policy" from a phrase into a thing your team either has or does not have — and concepts you have argued about are dramatically more resistant to a well-written distractor than concepts you have only read.
The SREF shelf, in the order it becomes useful
| When | Page | What it is for |
|---|---|---|
| Before you plan | The SREF Exam | Format, registration paths and logistics. Read once, then stop re-reading it — logistics are not study |
| Planning | The SREF Study Plan | The four-week calendar. This page is the method; that page is the schedule |
| Pass 2 | The SREF Concept Reference | Tightening definitions that feel "close enough" until a distractor proves they were not |
| Pass 2 | The three focused sets: SLOs & toil, monitoring & tools, anti-fragility & culture | Area-by-area discrimination practice while the material is fresh |
| Pass 2 → 3 | The practice bank | The full pool, mixed. Mixing is what stops you answering from context rather than knowledge |
| Pass 3 | Know It Cold — SREF | The compression sheet: the handful of facts and numbers that must be instant |
| Pass 3 | Mock exams 1–5 | Full 40-question papers under the clock. Keep two unseen for the final week |
| After every paper | Answer Triage — SREF | Why the distractor was tempting. This is where a wrong answer becomes a lesson |
| Any time | Exam Simulator | 20 questions in 20 minutes across all three certifications on the Certifications page — broader than SREF alone, useful for pacing practice |
Flashcards and question banks are diagnostics, not training. Their real output is a list of things you do not know. When a card keeps coming back, the fix is never to flip it again — it is to re-read the blueprint page it came from, write the concept out from a blank page, and put it back on the +1 rung of the ladder.
Definitions are the exam — so drill the near neighbors
☺ Like you're 10: Most wrong answers on this test are not things you never heard of. They are two things you know that live right next to each other, and you grabbed the wrong one.
On a paper capped at Bloom 1–2, the question writer's main tool is the near neighbor: an option that is a true statement about a different concept, or a true statement about the right concept in the wrong role. You defeat those by drilling pairs rather than terms. Cover the right-hand column, produce the distinction out loud, then check.
| The pair | The distinction, in one line |
|---|---|
| SLI vs SLO vs SLA | The measurement, the internal target set against it, the external contractual floor — in that fixed order, with the SLA deliberately looser than the SLO |
| Error budget vs error budget policy | The budget is the arithmetic remainder (100% minus the SLO); the policy is the pre-agreed rule for what the organization does when it runs out |
| MTTD vs MTTR vs MTRS | Time to detect the defect, time to repair/recover, time to restore service. All three appear on the official concept list, which is a strong hint they appear on the paper |
| Toil vs ordinary operational work | Toil needs all the properties — manual, repetitive, automatable, tactical, no enduring value, O(n) with growth. Hard work that produces something lasting is not toil |
| Resilience vs anti-fragility vs stability | Stability resists change, resilience absorbs stress and returns to normal, anti-fragility improves because of the stress |
| Monitoring vs observability vs telemetry | Monitoring watches for known failure modes, observability is the property that lets you ask new questions of a system, telemetry is the signal data both are built on |
| Availability vs reliability | Availability is the fraction of time the service is usable; reliability is the probability it performs correctly for a given period. A service can be highly available and unreliable |
| Symptom-based vs cause-based alerting | Alert on what the user is experiencing, not on the internal condition you suspect causes it — the classic alert-design distinction |
| Canary vs blue-green vs automated rollback | Two release strategies and one response. Canary exposes a slice, blue-green swaps environments wholesale, rollback is what either does when the signals go bad |
| SRE vs DevOps | The syllabus's own framing: SRE is one concrete implementation of DevOps principles, with prescribed practices where DevOps offers a philosophy |
| NRE vs CRE vs DBRE vs HRE | Network, Customer, Database and Heritage Reliability Engineering — all four are on the official terminology list, and the last one catches people who have never met the term at work |
| Pathological vs bureaucratic vs generative culture | Westrum's three organizational culture types; the first two are named explicitly on the official concept list, and the distinction turns on how each type handles information and bad news |
| Simian Army vs Chaos Monkey | Chaos Monkey terminates instances; the Simian Army is the wider family of failure-injection tools around it. Both are syllabus vocabulary, taught here in chaos engineering |
The arithmetic has to be reflexive
One item type deserves separate drilling because it is the only one where being slow costs you as much as being wrong: error-budget arithmetic. The official sample paper's question 7 asks for the error budget of a 99.9% monthly SLO, and the keyed answer is 43 minutes. That is not a hard calculation — 0.1% of a 30-day month — but doing it from first principles under a clock, in a room with a proctor watching, is where people lose ninety seconds they needed elsewhere.
So do not calculate it on the day. Know it:
30-day window = 43,200 minutes 99.9% SLO -> 0.1% -> ~43 min of allowed downtime per 30 days 99.95% SLO -> 0.05% -> ~21.6 min 99.99% SLO -> 0.01% -> ~4.3 min Each extra "nine" divides the allowance by roughly ten.
Sol's arithmetic drills in SLIs, SLOs & error budgets are where to build that reflex, and Know It Cold is where the finished numbers live.
Slowly, and out loud. Take these six SLO targets — 99%, 99.5%, 99.9%, 99.95%, 99.99%, 99.999% — and produce the allowed downtime for a 30-day window for each, without a calculator and without writing anything down until you have said it. Then check all six. Whichever ones you fumbled go on the +1 rung of the ladder tonight. Do it again on day three. On this exam, the arithmetic is worth more than the theory around it, because the theory has four plausible options and the arithmetic has exactly one.
Your production experience is an asset — and a specific liability
☺ Like you're 10: Knowing the job really helps. But your team probably uses some words slightly differently from the exam, and the exam only marks its own version right.
This exam assumes a working baseline — Linux, the basics of distributed systems, some monitoring exposure, and enough on-call experience to know what a page feels like. If you have that, most of the syllabus will read as things you already do, and you will be tempted to skip straight to the mocks. Two reasons that backfires.
First, the exam tests the syllabus's definitions, not your industry's. The official sample paper contains a question asking for the definition of latency according to the Catchpoint survey the course cites, and the keyed answer is the narrow one — the delay incurred in communicating a message — not the broader "total time from request to response" that many teams use day to day. Both are defensible in the wild. Only one is marked correct. That is not a flaw in the exam; it is what a vendor-neutral foundation certification is: a shared vocabulary, defined by the syllabus that issues it.
Second, the syllabus contains named artifacts your workplace probably does not use. The SLO VALET model is a good example — the official sample paper asks what its "T" stands for, and the keyed answer is Tickets. If you have never met the model, no amount of production experience will get you there, and it is a one-line fact to learn.
It is not "I did not know how reliability works." It is "I answered from how we do it." Whenever a practice question's keyed answer clashes with your shop's convention, do not argue with it and do not quietly override it — log it as source-drift in the mistake log and learn the syllabus's version alongside your own. You are being examined on a vocabulary, not on your platform.
The upside of experience is real, though, and worth using deliberately: attach every abstract term to something you have actually seen. Which service in your estate has a genuine SLO rather than a dashboard? What was the last piece of toil you automated, and which of the six properties made it toil? When did you last see an alert that was cause-based and should have been symptom-based? Concepts anchored to a memory survive distractors far better than concepts anchored to a paragraph.
Keep a mistake log
☺ Like you're 10: Write down every mistake in one place. After two weeks you will notice you keep making the same three — and those three are your whole revision list.
Individual mistakes are noise. Patterns of mistakes are the highest-value information you will produce during the entire plan, and you cannot see patterns without writing things down. One file, appended to the moment something goes wrong, never edited afterwards.
What an entry contains
Five fields, and the second-to-last is the one everybody skips and the one that does the work:
cat >> ~/sref-prep/mistakes.md <<'EOF' ## 2026-08-24 · SRE-4 Monitoring & SLIs · question 17 - what was asked: which signal is the SLI in the described scenario - what I picked: the internal queue depth metric - what was really wrong: I do not have "measured at the user-facing boundary" as part of my definition of an SLI — only "a metric about the service" - class: term-collision - redrill: 2026-08-27 EOF
"What was really wrong" is a different sentence from "what I got wrong," and forcing yourself to write it is what converts a mistake into a lesson. "I picked C" is not actionable. "My definition of SLI is missing a clause" tells you exactly what to fix and exactly which card to make.
Classify it — the fix depends entirely on the class
| Class | What it looks like | The actual fix |
|---|---|---|
knowledge-gap | You had never met the term or the model at all | Re-read the blueprint page, write it out from a blank page, put it on the ladder. The cheapest class to fix |
term-collision | You confused a concept with its nearest neighbor — SLO for SLA, MTTR for MTRS | Drill the pairs table, not the individual terms. Fixing one side of a pair does not fix the pair |
near-miss | You picked an option that was true, but was not what the question asked for | Argue down all three distractors out loud before committing. Work Answer Triage |
misread | You missed a qualifier — NOT, BEST, monthly rather than quarterly | A reading ritual: restate the question in your own words before you look at the options. This is a habit fix, not a study fix |
arithmetic-slip | Right method, wrong number — usually an error-budget conversion | Memorise the table rather than deriving it. Sol's drill above, twice, on different days |
source-drift | You answered from your workplace's convention rather than the syllabus's definition | Learn both, and label which is which. Common in experienced candidates and invisible without the log |
recall-lag | You got it right, but it took forty seconds | Speed work: flashcards at five seconds a card, and Know It Cold. This is the class that runs the clock out |
Once entries are classified, one command each Sunday tells you what kind of candidate you currently are:
grep -h '^- class:' ~/sref-prep/mistakes.md | sort | uniq -c | sort -rn grep -h '^- redrill:' ~/sref-prep/mistakes.md | awk -v today="$(date +%F)" -F': ' '$2 <= today'
A log dominated by knowledge-gap means you are early — keep reading and producing. Dominated by term-collision means you know the material and need the pairs table, not another pass over the blueprint. Dominated by misread means you need a ritual, not more study — and that is a thirty-second fix you would never have found by studying harder. The prescription is completely different in each case, which is why the classification is not busywork.
The traps
☺ Like you're 10: Some study activities feel productive and teach nothing. And some of the exam facts you will find on the internet are simply wrong, including for a different exam with the same name.
Trap one — re-reading the study guide
The official study guide is good, free, and pleasant to read, which is exactly what makes it dangerous. Reading it a third time produces fluency and no retrieval. It has two legitimate roles: a source for pass one's survey, and a reference you open with a specific question already in your head, after which you close it. The discipline that keeps both honest: never open a document without a question. If you cannot say what you are looking for, you are browsing, and browsing is entertainment.
One specific piece of that guide is worth knowing about: its "Value Added Resources" section — a long list of talks, articles and videos — is explicitly labeled by DevOps Institute as not examinable. It is enrichment. It is also the single easiest way to lose six hours that belonged to retrieval practice. Read it after you pass.
Trap two — third-party specs that are simply wrong
Search for this exam and you will land, overwhelmingly, on training-partner pages, practice-test vendors and content farms. Several publish exam specs that contradict both official sources. Verified examples worth inoculating yourself against:
| What you will read somewhere | What the official sources say |
|---|---|
| "90 minutes, 50 questions, 70% to pass" | 60 minutes, 40 questions, 65% — on both the PeopleCert and DevOps Institute pages. All three numbers in that claim are wrong |
| "Closed book" | Officially open book, restricted to the Official Training Materials. Plan as though closed-book for the reasons above — but do not repeat the wrong fact |
| "No waiting period between attempts" | We could not confirm this from PeopleCert or DevOps Institute. Treat retake terms as unverified and check PeopleCert's policy directly |
| "Certification is valid for two years" / "three years" / "evergreen" | All three appear in the wild — and DevOps Institute's own page contains both the two-year and the three-year claim, the latter attributed to PeopleCert's CPD programme after the acquisition. Renewal terms are genuinely inconsistent across sources; confirm with PeopleCert before assuming |
The Global Skill Development Council (GSDC) issues a separate credential also named Site Reliability Engineering (SRE) Foundation, with different specifications — 40 questions but 90 minutes, not open book, and a five-year validity. Google "SRE Foundation exam duration" and you can easily end up preparing your timing strategy against the wrong exam. Before you trust any spec you find, check which body issued it. This course prepares you for the DevOps Institute / PeopleCert exam.
Trap three — banking on the open-book label
This earns its own entry because of how badly it changes behavior. A candidate who believes they can look things up studies for recognition and stops at "I would know it if I saw it." Ninety seconds a question converts that into a fail. Study as though there is nothing in the room but you.
Trap four — dumps
Exam-dump sites dominate the search results for this exam more thoroughly than for almost any other on this site. They are worthless and they are a risk: memorized answers do not transfer to a re-worded question pool, and certification bodies generally treat their use as malpractice — read the candidate agreement you accept when you book rather than taking anyone's word for the consequences. There is a legitimate alternative, and those dump sites bury it — the official study guide contains a full 40-question sample paper with an answer key, written by the people who write the exam, laid out in the real topic distribution. That is the closest thing to a real form that exists in public, it is free, and it is allowed.
Trap five — expecting a community to study with
Worth saying plainly, because it surprises people. We went looking for first-hand accounts of sitting this exam and found one — McLeod's, cited above. There is no rich forum record here of the kind you find around the hands-on cloud-native exams, and the search results that look like experience reports are mostly vendor marketing. (We could not survey Reddit directly, so treat that as "we found none," not "none exists.") Practically: do not wait to find a study group or a curated video course. The official study guide, the eight blueprint pages here, and the practice bank are the material. That is not a gap in your research — it is the actual state of the public record.
Build a weekly cadence
☺ Like you're 10: Small amounts, most days, always at the same time. That beats one giant panic weekend, every time.
Study plans fail on scheduling, not on content. Three and a half focused hours spread across a week beats seven hours on a Sunday, because spacing is doing half the work and because a long block is where fatigue quietly turns study into transcription. Build the week around one session length you can actually protect — 45 minutes is the right unit for this exam.
The shape of a study week
| Slot | Length | What |
|---|---|---|
| Mon | 45 min | New topic area — read the blueprint page, close it, write the area's terms out from a blank page |
| Tue | 30 min | Yesterday's area explained aloud, then that area's flashcards at five seconds a card |
| Wed | 45 min | Second half of the same area, plus its focused practice set, distractors argued down out loud |
| Thu | 30 min | Reviews due from the 1/3/7/21 ladder, plus mistake-log redrills |
| Fri | — | Off. Deliberately. Spacing needs gaps, and you need to still be doing this in five weeks |
| Sat | 60 min | Whole-area pass: the deeper course lesson behind the blueprint page, then 20 mixed questions from the bank, timed at 90 seconds each |
| Sun | 20 min | Bookkeeping: update the ledger, re-sort by exposure, mine the mistake log, pick next week's area |
That is about three and a half hours a week — roughly six weeks to the 20–25 hours this course budgets, or four weeks if you take the near-daily version in the study plan. The Sunday twenty minutes is the slot people cut and the one that makes the plan self-correcting; without it you will keep studying whichever area you enjoyed most last week.
A 45-minute session template
| Minutes | What | Why |
|---|---|---|
| 0–5 | Retrieval warm-up: last session's area, from a blank page | Starts with the hard move while you are fresh, and doubles as a spaced review |
| 5–13 | Read today's blueprint page | Input, bounded — the timer is what stops reading from eating the session |
| 13–30 | Close it and produce: terms, definitions, one worked example each | The only part that actually moves a ledger score |
| 30–40 | Ten practice questions on it, distractors argued down aloud | Converts fragile knowledge into knowledge that survives a plausible wrong option |
| 40–45 | Update the ledger, log misses with a class, schedule the +1 review | Five minutes that make every future session better targeted |
When you miss a week
You will. The rule is: do not restart, and do not try to make it up. Come back with a single 30-minute retrieval session on whatever the ledger says has the highest exposure, take the score hit honestly, and resume the normal week. Candidates who try to repay lost hours in one weekend usually abandon the plan within a fortnight. A study plan you resume is worth ten study plans you abandon.
Do this before you study anything else. Open the eight topic areas next to a blank file and build the ledger from the code block above. For each area, write your confidence 0–3 using the strict definition — a 3 means you wrote the area out cleanly from a blank page and scored 80%+ on its questions, cold, on two separate days. Do not look anything up while scoring; the entire value is in the honesty. Then run the exposure command and read the top line. That area is Monday's session. You have just built a study plan in twenty-five minutes, and it is better targeted than any generic one you could buy — because it is the only one that knows what you do not know.
How to know you are actually ready
☺ Like you're 10: "I feel ready" is not evidence. "I scored 34 out of 40 on a paper I had never seen, twice, on different days" is evidence. Book on evidence.
Confidence is not a readiness signal — it is a mood, and it tracks how recently you read something rather than whether you can produce it. This exam has an objective bar, so measure yourself against an objective test.
| Signal | The bar | Why this bar |
|---|---|---|
| Unseen timed papers | Two full 40-question mock sets at 85%+ (34/40), on different days, unseen, inside 60 minutes | The real cut is 26/40, but our wording is not the vendor's. A 34 here is the margin that makes 26 there a non-event |
| No area dragging | Every topic area at or above roughly 70% across your last two papers | You can afford 14 wrong; you cannot afford them clustered in one area that is worth 7 |
| Blank-page coverage | All eight areas written out from memory — terms and one-line definitions — including the four-question ones | The 10% areas are where under-prepared candidates lose the margin they were counting on |
| Arithmetic reflex | Any common SLO converted to an error budget in under 30 seconds, no calculator | 90 seconds a question means a slow calculation costs you a different question, not just this one |
| Distractor argument | You can cover the correct answer and argue down the other three out loud | This is the actual exam skill. Picking right without being able to say why is luck you cannot repeat |
| Quiet mistake log | What remains is recall-lag, not knowledge-gap or term-collision | The class of your remaining mistakes tells you whether you need speed or study — and only one of those is fixable this week |
Signals that mean nothing
| Feels like readiness | Why it is not |
|---|---|
| "I've read every page on this site" | Input is not output. Reading measures your patience, not your recall |
| "I've been an SRE for six years" | Helpful, and not the same thing. The exam tests a published vocabulary, some of which your workplace does not use — and one term of it is "Heritage Reliability Engineer" |
| "I scored 92% on a mock" (that you had already sat) | A repeated paper measures your memory of that paper. Keep two sets unseen for the final week |
| "It all feels familiar now" | Familiarity is recognition. The exam scores discrimination, which is a different thing entirely |
| "It's only a foundation exam" | True, and it is still 26 out of 40 with no partial credit. Foundation means shallow, not automatic |
| "I finished the accredited course" | Completion is the course's metric, not the exam's. The official recommendation is at least 16 contact hours of instruction and labs — a recommendation, not a pass |
When the objective signals are green, book it. Waiting for the feeling is how people spend six months "almost ready" on an exam whose whole syllabus is twenty-odd hours of work. The registration walkthrough is on The SREF Exam.
The last 72 hours
☺ Like you're 10: The three days before the test are for polishing what you already have, not for learning anything new. And sleep counts as studying.
Whatever you do not know 72 hours out, you will not know on the day — so stop trying. The final stretch has one job: arrive rested, fast and calm, with the logistics rehearsed so none of your attention is spent on them.
| Window | Do | Don't |
|---|---|---|
| T-72 → T-48 | Your final unseen 40-question paper, under real conditions: right time of day, no phone, no pausing. Mark it, then drill its two worst findings | Open a topic area you have never touched. If the paper finds something big and unlearnable in two days, that is real information about rescheduling — better now than an hour before |
| T-48 → T-24 | Read your own artifacts: the mistake log end to end, the ledger's remaining 1s and 2s, and Know It Cold. One pass through the deck at five seconds a card | Sit another full paper — you cannot act on the result in time, and a bad score here costs you sleep you need more |
| T-24 → T-8 | Rehearse the logistics: test the machine, the camera and the connection, check your photo ID matches your registration name exactly, clear the desk and the room, and know how early you must be online. Then stop | Cram. It costs more in fatigue than it returns in recall — and PeopleCert's proctored sitting has a real, published set of environment rules that a tired candidate gets wrong |
| T-8 → 0 | Sleep, eat, log in early, breathe | Any studying at all |
This sitting is proctored live. Published PeopleCert guidance and McLeod's first-hand account agree on the shape: a system check beforehand, ID verification, and a webcam pan around the room — in his case over and under the desk, with the desk and camera re-angled so that he, the room's door and a cupboard door were all visible at once. None of that is hard; all of it is slow if you meet it for the first time at the start of your 60 minutes. Read the current rules on PeopleCert's online proctoring page the day before, not the day of — and note that the free v1.1 study guide's "Supervised: No" line predates this process entirely.
The last 72 hours cannot add knowledge, but they can absolutely subtract points — through fatigue, a failed system check, or a panic-induced topic switch. Treat the final stretch as protecting the score you already have.
The SREF Study Plan
This method, laid out as a four-week schedule with a day count for each of the eight topic areas.
✓ · tick it offReadiness checklist
SLO, on-call, engineering and reporting readiness — the fastest way to make abstract terms concrete against a real service.
🃏 · recallFlashcards
38 cards across every module. Answer out loud before you flip, and use the five-second rule.
✎ · discriminationSelf-Check
Ten questions at a time, reshuffled options, an explanation on every answer, and a past-misses queue that remembers your weak spots.
🎯 · the bankSREF practice questions
The full mixed pool. Mixing is what stops you answering from context instead of knowledge.
⏱ · 40 in 60Mock exams 1–5
Full papers under the clock. Keep two of the five unseen for the final week.
🧠 · compressionKnow It Cold — SREF
The facts and numbers that have to be instant, because 90 seconds a question leaves no room to derive them.
📋 · logisticsThe SREF Exam
Format, registration paths and exam-day mechanics. Read it once, then get back to retrieving.
7 · heaviest areaMonitoring & SLIs
The single largest slice of the paper. If one blueprint page gets an extra hour, it is this one.
Foxy: I've read all eight blueprint pages twice. Book me in for Friday.
Professor Owl: Splendid. Close the laptop and give me the difference between an SLI, an SLO and an SLA. In that order, out loud, now.
Foxy: …The SLI is the measurement. The SLO is the target. And the SLA is — the same thing but for customers?
Professor Owl: That is recognition, not recall. You've read it. You don't own it. And "the same thing but for customers" is precisely the option a question writer puts in slot C.
Sol the Sloth: …It's also worth six questions. Area two. Plus seven for monitoring and six for tooling — nineteen of forty, before you've reached anything else.
Remy the Rabbit: Ninety seconds a question. If it takes you that long to say it in a quiet room, it isn't happening in a proctored one.
Nutty the Squirrel: Every card you fumble goes on the list, and the list comes back tomorrow, then in three days, then in a week. That's the whole system.
Timmy the Turtle: And before anyone books anything — check the format on the vendor's own page. Two different bodies publish an exam with this name, and their numbers don't match.
Sol the Sloth: …Sixty-five percent. Twenty-six of forty. I checked. Twice.
Professor Owl: Measure, then study. In that order. Foxy — the ledger, please. Twenty-five minutes.
One exam, one method. Now point it at something: open the eight topic areas next to a blank file, build the ledger, and let exposure pick tomorrow's session. When the objective signals turn green — two unseen papers clear, no area dragging, the mistake log gone quiet — work through the registration steps on The SREF Exam, confirm every number on the vendor's own page, and book it. Not before.
1. Name the skill a performance-based exam trains that SREF does not — and the skill SREF trains instead. 2. Why is re-reading a blueprint page a poor use of study time, and what should replace it? 3. What does a confidence score of 3 mean in the ledger, and why is the definition deliberately strict? 4. The syllabus caps difficulty at Bloom levels 1 and 2. What does that change about how deep you study? 5. Three topic areas carry 19 of the 40 questions — so why can you not skip the four-question areas? 6. Your mistake log is 60% term-collision. What should you change, and what should you not change? 7. SREF is officially open book. Give two reasons to prepare as though it were not. 8. Give three objective readiness signals and one signal that feels convincing but means nothing.
Check your answers
- A performance exam trains production under time — building a correct end state without hunting. SREF trains discrimination — telling one true statement from three plausible neighbors, at roughly 90 seconds each. Neither transfers automatically, which is why one regime for both this exam and a hands-on one shortchanges both.
- Re-reading is fluent, and that fluency gets mistaken for learning — but it produces recognition, and the exam scores discrimination. Replace it with retrieval: close the page, produce the definition out loud or on a blank page, then check.
3means you wrote the area's terms out cleanly from a blank page and scored 80%+ on its questions, cold, on two different days. It is strict because anything looser lets "I read that on Tuesday" pass as done — and that is exactly the item that collapses at question 31.- Bloom 1–2 means the paper asks you to define and to recognize in context, not to design or evaluate. So breadth beats depth: eight areas at a definitional level passes comfortably, three areas at architect level does not. If you are deep in burn-rate alerting math, you are studying above the ceiling.
- Because the four-question areas are 10% of the paper each, and the four smallest together are 16 of 40 — two more than the 14 wrong answers the pass mark allows you. Writing all four off is an automatic fail no matter how strong the heavy three are. The published counts decide order and hours, not what you can skip.
term-collisionmeans you know the material and are losing on near neighbors, so drill the pairs — SLO vs SLA, MTTR vs MTRS, resilience vs anti-fragility — as contrasts rather than as separate terms. Do not re-read the blueprint pages: that fixesknowledge-gap, which is not your problem, and it will eat the time the actual fix needs.- Any two of: the permitted material is only the Official Training Materials, which you probably do not have if you self-studied; 90 seconds a question is enough to check a term you nearly know and nowhere near enough to learn one you do not; and the one published first-hand account of sitting this exam reports answering from memory for exactly that reason. Confirm the current policy on PeopleCert's page regardless.
- Objective signals: two unseen full papers at 85%+ on different days, inside the hour; no topic area dragging below roughly 70%; all eight areas writable from a blank page; error-budget arithmetic in under 30 seconds; a mistake log whose remaining entries are
recall-lag. Meaningless: "it all feels familiar now" — that is recognition. Also meaningless on its own: years in the job, since the exam tests a published vocabulary rather than your platform's conventions.