DevOps Interview Questions
DevOps interviews are reversibility interviews. Almost every substantive question reduces to the same thing: when this change is wrong, how quickly and safely can you undo it? Pipelines, rollout strategies, infrastructure as code and secret management are all mechanisms serving that one property.
The common weakness is tool fluency without failure reasoning. A candidate lists the pipeline stages they built and cannot say what happens to a request that was in flight when the old pod received SIGTERM, or why a database migration that cannot be rolled back makes the whole deploy irreversible regardless of how good the pipeline is. Interviewers go straight there.
Written and reviewed by Sahil Srivastav
What the bar actually is
You should be able to describe a deploy as a sequence of independently safe steps, and identify which step is the point of no return. The schema change is usually the honest answer, which is why the expected approach is additive migration, then code tolerant of both shapes, then backfill, then narrow — each step reversible on its own.
On infrastructure as code, you are expected to treat state as the hard part rather than the syntax. Where state lives, who can write it, what happens when two pipelines apply concurrently, and what a plan showing an unexpected destroy means. Candidates who have only ever applied from a laptop tend to be exposed by these questions.
How the rounds are structured
Pipeline and release design
Design CI/CD for a described service. The scoring is mostly about gates, reversibility and what runs where, not about which tool you name.
Rollout and failure scenarios
What a rolling update does to in-flight requests, how a readiness probe gates traffic, what a canary actually proves and what it cannot.
Infrastructure as code
State management, drift, module structure, and how you review a plan that proposes a destroy you did not expect.
Linux and container troubleshooting
Why a container is OOMKilled, why a build works locally and fails in CI, image layer and architecture mismatches.
What this interview bar tests
The irreversible step
Every deploy has one. Usually the schema change. Identifying it and making it reversible is the core DevOps judgement being tested.
Graceful shutdown that actually works
SIGTERM reaching the right PID, the process stopping new work, draining in-flight requests, then exiting. Broken in a surprising proportion of real deployments.
IaC state as the real risk
Remote state with locking, least-privilege write access, and reviewing plans for unexpected replacements. Syntax is the easy half.
Secrets never in the image
Injected at runtime, rotatable without a rebuild, and absent from build logs and layer history. A baked secret is a rebuild-and-redeploy to rotate.
A preparation plan that works
- 1Take a service you know and write its deploy as numbered steps, marking the point of no return and how you would reverse each step. This exercise produces better answers than any amount of tool reading.
- 2Verify graceful shutdown on something real: send SIGTERM mid-request and check whether the response completes. Most people discover it does not, which is the lesson.
- 3Practise reading a plan that proposes a replacement rather than an update, and explain what would cause it. This is a daily IaC skill and a frequent interview question.
- 4Reproduce a container OOMKill with a JVM or Node process whose heap exceeds the container limit, and fix it by sizing against the cgroup rather than the host.
Questions you should expect
How do you make a deploy with a schema change reversible?
Split it so no single step is a cliff: additive migration first, then code that tolerates both old and new shapes, then backfill, then remove the old column in a later release. Mention lock behaviour too — DDL waiting on a long transaction queues every reader behind it, so migrations need a short lock timeout and a retry rather than an unbounded wait.
What happens to in-flight requests during a rolling update?
They are dropped unless the application handles SIGTERM and the platform waits. The correct sequence is: fail readiness so traffic stops arriving, finish in-flight work, close connections, exit — within the termination grace period. If SIGTERM goes to a shell wrapper rather than the process, none of this happens.
Your container is OOMKilled but the application logged no error. Why?
Because the kernel killed the process rather than the runtime raising an error — exit code 137, no application-level trace. Common with a JVM or Node heap sized against host memory instead of the cgroup limit. Set the heap as a percentage of the container limit and verify what the runtime actually believes the limit is.
Two pipelines apply the same infrastructure concurrently. What happens?
Without remote state and locking, they interleave and the state file ends up describing neither reality — the failure that takes a day to untangle. With locking, the second waits or fails cleanly. This is why laptop-local state is a production hazard rather than a style preference.
What gets candidates rejected
- Listing tools fluently with no account of what happens when a change is wrong
- Treating a migration as just another pipeline step rather than the irreversible one
- Assuming the platform handles in-flight requests without application cooperation
- Baking secrets into images, so rotation requires a rebuild
- IaC state on a laptop or without locking
- Sizing a runtime's memory against the host rather than the container limit
- Claiming canary deployments without being able to say what signal would stop the rollout
What to practise, in order
Zero-downtime database migration
The irreversible step, made executable. Tests run traffic throughout and assert no failed request and no lost write.
Database connection pool exhaustion
What a restart appears to fix and does not. Useful for reasoning about why "redeploy and it goes away" is not a diagnosis.
Linux log parsing
The operational scripting that DevOps rounds actually include — correct, re-runnable, and safe on malformed input.
Bash script hardening
Why `set -e` does not do what you think in a pipeline, and why unquoted variables corrupt scripts silently.
FAQ
Does the specific tool matter — Jenkins, GitHub Actions, GitLab?
Much less than candidates expect. Interviewers care about gates, reversibility, artefact promotion and what runs where. Tool-specific syntax is learnable in a week; release judgement is not.
How much Kubernetes depth is needed?
Failure-shaped depth. Probe semantics and their effect on traffic, OOMKill and exit codes, what a rolling update does to connections, why a pod is evicted. API recall is rarely the differentiator.
Is DevOps the same interview as SRE?
Overlapping but differently weighted. DevOps leans toward delivery, pipelines and infrastructure; SRE leans toward reliability measurement and incident reasoning. Prepare for the one you are actually interviewing for.
Will I be asked to write code?
Usually scripting rather than application code: parse, reconcile, automate. The quality bar is idempotency and safe failure, because these scripts get run under pressure and often twice.