SRE Interview Questions
An SRE interview is a debugging interview with a measurement interview attached. The coding rounds are real but moderate; what differentiates candidates is whether they can define reliability in numbers that drive decisions, and whether they can narrow a live failure from evidence rather than from intuition.
The most common weakness is a candidate who can recite the SRE vocabulary — SLI, SLO, error budget, toil — and cannot use it. Asked to define an SLO for a specific service, they produce "99.9% uptime" with no definition of what "up" means, no measurement point, and no consequence attached to breaching it. That answer is the gap the interview is designed to find.
Written and reviewed by Sahil Srivastav
What the bar actually is
You should be able to define an SLI precisely enough to implement: what event counts, what makes it good, where it is measured. "Proportion of requests to the checkout endpoint returning non-5xx within 300ms, measured at the load balancer" is an SLI. "Availability" is not. The measurement point matters because client-side and server-side numbers differ, and the user experiences the client-side one.
You are also expected to treat the error budget as a decision mechanism rather than a report. The useful answer connects a breach to a specific action — freeze feature work, prioritise the top failure source, revisit the target because it was never justified. An SLO with no consequence attached is a dashboard, and interviewers will say so.
How the rounds are structured
Debugging from evidence
A symptom and partial telemetry. The round that decides most SRE loops: can you form one narrow hypothesis and name the measurement that would confirm it?
SLO and reliability design
Define SLIs and SLOs for a described service, justify the target, and say what happens when the budget is exhausted.
Systems and Linux depth
File descriptors, memory and the OOM killer, cgroup limits, networking basics, what a process is actually doing when it hangs.
Coding, usually automation-flavoured
Parse logs, reconcile state, write a safe operational script. Correctness and idempotency matter more than elegance — the script may run during an incident.
What this interview bar tests
SLIs you could implement tomorrow
A countable event, a precise goodness criterion, and a named measurement point. Vagueness here invalidates everything built on top.
Error budget as a lever
Breaching it must change what the team does. If nothing changes, the number was decoration and the interviewer will test whether you know that.
Narrowing, not listing
Given a failure, the weak answer enumerates everything that could be wrong. The strong one picks the cheapest measurement that eliminates the most hypotheses.
Toil identified and removed
Concrete instances of manual, repetitive, automatable work you eliminated, with the time recovered. This is the SRE impact narrative.
A preparation plan that works
- 1Write SLIs and SLOs for three real services you know, to an implementable level of precision, including the measurement point and the consequence of a breach. Most candidates have never done this in writing and it shows immediately.
- 2Practise incident narration: given a symptom, say your hypothesis, the measurement that would confirm it, and what you would do if it came back negative. Narrating the narrowing is the skill being assessed.
- 3Get genuinely comfortable with the diagnostic commands — what is holding file descriptors, what the OOM killer logged, what a hung process is blocked on, whether a query is slow or waiting. Fluency here is visible instantly.
- 4Prepare two toil-removal stories with numbers: what was manual, what you automated, how much time it returned and what it freed the team to do.
Questions you should expect
Define an SLO for a checkout service.
Name the SLI precisely — proportion of checkout requests returning non-5xx within 300ms, measured at the load balancer — then a target justified by user impact rather than by convention, then the consequence: at budget exhaustion, feature work pauses until the top failure source is addressed. A target with no justification and no consequence is the answer that fails.
Error rate just went from 0.1% to 4%. What do you do first?
Determine blast radius and whether it correlates with a change, in that order. Which endpoints, which regions, which customers, and did it start at a deploy, a config push, or a dependency event. Mitigate before diagnosing — roll back if a change correlates — because restoring service and finding root cause are separate activities with different urgency.
Why might your availability number look fine while users complain?
Measurement point and aggregation. Server-side metrics miss failures that never reached you — DNS, TLS, client network — and averaging across endpoints or customers hides a total outage for one tenant inside a healthy global number. Both are reasons to measure closer to the user and to slice by tenant.
A service hangs but CPU is near zero. What is happening?
It is waiting, not working. Candidates: a lock cycle, a connection pool with no free slots, an outbound call with no timeout, a thread pool whose workers are all blocked on each other. Zero CPU with no progress is diagnostic — it rules out most "too slow" explanations immediately.
What gets candidates rejected
- Defining an SLO as "99.9% uptime" with no SLI, no measurement point and no consequence
- Treating the error budget as reporting rather than as a decision rule
- Listing every possible cause in the debugging round instead of narrowing
- Diagnosing before mitigating during an incident
- Reciting SRE vocabulary without being able to apply it to a specific service
- Measuring availability only server-side and missing the failures users actually see
- No concrete toil-removal story with numbers attached
What to practise, in order
Database connection pool exhaustion
The archetypal zero-CPU hang: slots consumed by leaks and sessions left in transaction. Diagnose from pool and database evidence.
Database overload under a traffic spike
Compounding causes where partial fixes leave the invariant broken. Tests assert database work per request rather than wall-clock time.
Unbounded cache heap exhaustion
Memory growth from retention, with the evidence an on-call engineer would actually have. Raising the limit does not pass.
Thread pool starvation with nested tasks
Throughput collapsing at low load with everything apparently healthy. The incident that looks impossible until you read the thread dump.
FAQ
How much coding is in an SRE interview?
Real but moderate, and usually automation-shaped: parse this, reconcile that, write a script that is safe to re-run. Idempotency matters more than algorithmic sophistication, because operational scripts get run twice during incidents.
Do I need Kubernetes depth?
For most SRE roles now, yes — but the useful depth is failure-shaped. Why a pod is OOMKilled, what a readiness probe does to traffic, what happens to in-flight requests on a rolling update. Not API surface recall.
What if I have never been on call?
It is a genuine gap for an SRE role and worth naming rather than hiding. Compensate by demonstrating evidence-driven reasoning on debugging problems; what you cannot credibly claim is incident judgement you have not exercised.
Are SRE and DevOps interviews the same?
They overlap and emphasise differently. SRE loops weight reliability measurement and incident reasoning; DevOps loops weight delivery pipelines, infrastructure as code and rollout mechanics. Candidates who prepare one and interview for the other usually notice in the first round.