DevOps CI/CD interview questions

Reviewed by Mark Dickie · Last updated

DevOps CI/CD is the practice of automating the build, test, and deployment stages of software delivery so that changes move from a developer's commit to production with minimal manual intervention. For interviews, you should understand pipeline stages (build, test, stage, deploy), common deployment strategies (blue-green, canary, rolling), and the trade-offs between speed and safety in release automation. Expect questions on tooling choices (Jenkins, GitHub Actions, GitLab CI), pipeline-as-code concepts, artifact management, and how to handle failures and rollbacks.

TopicWhat to Study
Pipeline stagesSource, build, test, stage, deploy — and what happens at each
Deployment strategiesBlue-green, canary, rolling, and when to pick each
Pipeline-as-codeYAML/declarative pipeline definitions, templating, reuse
Artifacts & imagesDocker image tagging, registry storage, immutable artifacts
Rollback & failureAutomated rollback triggers, health checks, circuit breakers
Security in CI/CDSecrets management, scanning in pipeline, least-privilege runners

What does a DevOps CI/CD interview test?

Interviewers want to see that you can design a pipeline end to end and reason about failure modes, not just list tools. You will likely be asked to diagram a pipeline for a sample application, explain where tests fit, and describe what happens when a stage fails. They also probe your understanding of environment promotion — how a build artifact moves from dev through staging to production — and what controls prevent a bad release from reaching users.

How should I prepare for CI/CD pipeline questions?

  1. Draw a full pipeline for a web application from commit to production, labeling every stage and the tool responsible for it.
  2. Compare at least two CI/CD tools (e.g., GitHub Actions vs. Jenkins) and be ready to explain why you would choose one over the other for a given team size or project.
  3. Practice explaining a canary deployment and a blue-green deployment step by step, including how traffic is shifted and how rollback works in each case.
  4. Study how secrets and credentials are injected into pipelines without being exposed in logs or source control.
  5. Review common pipeline anti-patterns — long-running pipelines, flaky tests, manual approval bottlenecks — and how to address them.

What are the most common CI/CD interview mistakes?

Candidates often confuse continuous delivery with continuous deployment. Continuous delivery means every change is ready to deploy but release is a manual gate; continuous deployment pushes every passing build straight to production with no human approval. Getting this distinction wrong signals a surface-level understanding. Another frequent error is describing a pipeline without mentioning testing — if you cannot say where unit, integration, and security tests run and what happens when they fail, the interviewer will assume you have not operated a real pipeline.

Key facts

  • Tarmac has 31 DevOps interview questions on this topic, 25 of them on this page, at difficulty 1–5 of 5.
  • Tarmac tracked 2,674 job postings asking for DevOps in August 2026.
  • Roles asking for DevOps advertise a median base salary of £80,000, across 596 job postings as of August 2026.
  • Tarmac last reviewed these DevOps interview questions on 31 August 2026.

At a glance

Questions25 shown · 31 in the bank
Difficulty1–5 of 5
FormatsMultiple answer, Multiple choice, Ordering, True / false, Fill in the blank, Short answer, Code output, Design exercise, Flashcard

What you'll review

  1. build artifacts
  2. pipeline as code
  3. deployment gates
  4. pipeline stages
  5. continuous delivery vs deployment

Practice questions

Try one before you open the answer. Pick an option and press Check; it's marked on the spot.

In a CI/CD pipeline, which of the following are generally accepted characteristics of a well-designed build artifact?#

Options

Pick every one that applies.

Show answer

A well-designed build artifact is immutable once published, produced from a specific reproducible source commit, and promoted unchanged across environments. It should NOT bundle environment-specific secrets or config; those are injected at runtime so the same artifact can be safely deployed everywhere.

Why:

A well-designed build artifact is immutable (it never changes once published), built from a reproducible source state, and promoted as-is across environments so that what was tested is what is deployed. Baking environment-specific secrets into the artifact is an anti-pattern — secrets and environment config should be injected at runtime, not embedded in the artifact.

Which of the following are commonly used as build artifacts in CI/CD pipelines?#

Options

Pick every one that applies.

Show answer

Common build-artifact formats in CI/CD pipelines include Docker container images, compiled JAR or WAR files, and ZIP archives of binaries and assets. Git commit messages are source-control metadata, not build artifacts.

Why:

Docker images, compiled JAR/WAR files, and ZIP archives of binaries and assets are all standard build-artifact formats that CI pipelines produce, version, and store in an artifact registry for downstream deployment. Git commit messages are metadata about source changes, not build artifacts.

Which of the following is a defining characteristic of pipeline-as-code?#

Options

Show answer

Pipeline-as-code means defining CI/CD pipeline configuration in version-controlled text files such as YAML or Groovy, rather than configuring it manually through a web UI. This lets teams review, branch, and reproduce their build and deployment workflows exactly like application code.

Why:

Pipeline-as-code means the CI/CD pipeline definition lives in version-controlled text files—commonly YAML (e.g., GitHub Actions, GitLab CI) or code in a DSL like Groovy (Jenkinsfile). The other options each describe the opposite: UI-only configuration, compiled binaries, or manual stage management all defeat the purpose of treating the pipeline itself as reviewable, reproducible code.

Order these steps in a Docker-based CI/CD deployment pipeline, from first to last.#

Put these in order

Show answer

In a Docker CI/CD deployment, the steps proceed in strict dependency order: first clone the source repository, then build the image from the Dockerfile, push it to a container registry, pull it on the deployment target, and finally start a container from it. Each step requires the previous step's output — you cannot build without source, push without a build, pull without a push, or run without a local image.

Why:

Each step depends on the completion of the previous one. The CI/CD runner must clone the repository before it can build an image from the Dockerfile; the image must be built before it can be pushed to a registry; the image must exist in the registry before the deployment target can pull it; and the image must be present locally before a container can be started from it. This gives a strict total order: clone → build → push → pull → run.

Order the steps in the evaluation cycle of a single automated deployment gate when determining whether to allow a pipeline to proceed after a stage completes.#

Put these in order

Show answer

The evaluation cycle of a single automated deployment gate is: the preceding stage completes and triggers the gate; the gate collects data from configured data sources; the gate evaluates the collected data against defined criteria; the gate produces a pass or fail result; and if all criteria pass, the pipeline proceeds to the next stage. Each step strictly depends on the previous one — you must collect data before evaluating it, and evaluate before reporting a result.

Why:

A deployment gate's evaluation cycle begins when the preceding stage completes and the gate is triggered. The gate then collects data from its configured data sources — for example, querying monitoring APIs or checking work item status. Next, it evaluates the collected data against the defined quality criteria (such as error rate thresholds or alert rules). Based on that evaluation, the gate produces a pass or fail result. Only when all gate criteria pass does the pipeline proceed to the next stage; a failed gate blocks promotion and may retry after a configured interval. Note that this ordering describes the automated gate evaluation cycle only — manual approvals, where used, are a separate, independently configured mechanism on pre-deployment or post-deployment conditions.

In a CI/CD pipeline, build artifacts (such as compiled binaries, JAR files, or Docker images) should be committed to the Git repository alongside source code so that they are version-controlled and available to every pipeline stage.#

Options

Show answer

False. Build artifacts like compiled binaries, JAR files, and Docker images should be stored in a dedicated artifact repository (e.g., Artifactory, Nexus, or a container registry), not committed to Git. Committing artifacts bloats the repository with large binary files, creates unnecessary merge conflicts, and duplicates data that can always be reproduced by rebuilding from source.

Why:

Build artifacts are derived from source code and should be stored in an artifact repository (e.g., Nexus, Artifactory, a container registry), not committed to Git. Committing artifacts bloats the repository, creates merge conflicts, and duplicates information already recoverable by rebuilding from source. Git is designed for source code and text, not large binary files.

In a mature CI/CD practice, the same immutable build artifact that passed tests in the staging environment should be promoted to production — rather than rebuilding a new artifact from source for the production deployment.#

Options

Show answer

True. The same immutable build artifact that passed tests in staging should be promoted directly to production. Rebuilding from source for production risks producing a subtly different artifact due to changes in build environment, dependency versions, or transient state, which would break the guarantee that the tested software is exactly the shipped software.

Why:

Promoting the exact same artifact that was tested ensures that what was verified is what gets deployed. Rebuilding from source introduces risk because differences in build environment, dependency versions, or transient states can produce a subtly different artifact, breaking the guarantee that the tested software is the shipped software.

A _____ gate is a deployment checkpoint that requires explicit human approval before a build can be promoted to the next pipeline stage, adding a manual decision point to an otherwise automated release flow.#

Show answer

A manual approval gate is a deployment checkpoint that requires explicit human approval before a build can be promoted to the next pipeline stage, adding a manual decision point to an otherwise automated release flow.

Why:

A manual approval gate (also called an approval gate) pauses the pipeline at a defined stage boundary and waits for a designated person to click approve or reject. Unlike automated quality gates that evaluate metrics programmatically, a manual approval gate inserts a human judgement call into the promotion path.

In CI/CD pipelines, a _____ gate automatically blocks deployment when a measurable signal — such as test pass rate, code coverage, or static-analysis score — falls below the configured _____.#

Show answer

In CI/CD pipelines, a quality gate automatically blocks deployment when a measurable signal — such as test pass rate, code coverage, or static-analysis score — falls below the configured threshold.

Why:

A quality gate is an automated checkpoint that evaluates one or more pipeline metrics against predefined thresholds. If a metric like test pass rate or code coverage does not meet the configured minimum threshold value, the gate fails and the pipeline halts, preventing a non-compliant build from being deployed.

In a canary deployment, the progressive rollout is governed by automated health gates that continuously monitor signals such as error rate and latency; if those metrics breach their thresholds, the pipeline triggers an automatic _____ to the previous stable version.#

Show answer

In a canary deployment, the progressive rollout is governed by automated health gates that continuously monitor signals such as error rate and latency; if those metrics breach their thresholds, the pipeline triggers an automatic rollback to the previous stable version.

Why:

Canary deployments incrementally route traffic to a new version while health gates watch real-time metrics. When error rates or latency exceed acceptable bounds, the automated safety mechanism performs a rollback — reverting traffic to the last known-good version — to minimize user impact without requiring manual intervention.

Which of the following is a primary benefit of defining a CI/CD pipeline as code in a version-controlled file (e.g., Jenkinsfile, .gitlab-ci.yml, or a GitHub Actions workflow YAML) rather than configuring it through a web UI?#

Options

Show answer

Pipeline-as-code lets you version, review, and branch pipeline changes alongside application code. Because the pipeline definition lives in a committed file (e.g., Jenkinsfile or .gitlab-ci.yml), every modification goes through the same code-review and history workflow as source code, giving you traceability, reproducibility, and team collaboration that a web-UI configuration cannot match.

Why:

Pipeline-as-code stores the pipeline definition as a text file in the same version-control repository as the application source. This means every change to the pipeline—new stages, modified build steps, environment updates—goes through the same commit, branch, review, and history workflow as application code. The claim that the pipeline executes faster because the YAML definition is compiled to native machine code before running is wrong because YAML is parsed and interpreted, not compiled to native code. The claim that the pipeline becomes immune to infrastructure outages because the definition is stored as text is wrong because storing the definition as text has no bearing on whether the CI/CD infrastructure or target environments are available. The claim that pipeline-as-code eliminates the need for build tools, test frameworks, or deployment dependencies is wrong because the pipeline still orchestrates real build tools, test runners, and deployment tooling—pipeline-as-code does not remove those dependencies.

A well-structured CI/CD pipeline for a web application typically moves through a defined sequence of stages. Arrange the following stages in the order they should execute, from earliest to latest.#

Put these in order

Show answer

The canonical CI/CD pipeline stage order is Source → Build → Test → Deploy. The pipeline first checks out code from version control, then compiles and packages the artifact, then runs automated tests against that artifact, and finally releases the validated artifact to the target environment. Each stage depends on the output of the prior stage.

Why:

A standard CI/CD pipeline begins with Source (pulling code from version control on a commit), then Build (compiling and packaging the artifact), then Test (running automated test suites against the built artifact), and finally Deploy (releasing the validated artifact to an environment). Each stage depends on the output of the one before it, so the order Source → Build → Test → Deploy is the canonical sequence.

What is the precise difference between continuous delivery and continuous deployment?#

Options

Show answer

Continuous delivery and continuous deployment both keep the codebase continuously releasable — every change is automatically built, tested, and proven deployable. The one difference is the production gate. Continuous delivery makes each change release-ready but a human approves the final push to production. Continuous deployment removes that gate: any change that passes the automated pipeline ships to production with no manual step. Continuous integration is the upstream practice both build on — frequent merges with automated testing on each commit.

Why:

Both practices keep the codebase in a continuously releasable state — every change is automatically built, tested, and proven deployable. The single distinction is the production gate. Continuous delivery stops just short of production: the artifact is release-ready and could ship at the click of a button, but a human decides when. Continuous deployment removes that button — any change that passes the full automated pipeline is released to production with no manual step. The option that says continuous deployment requires a manual approval step inverts the definitions. The option that says continuous delivery means merging to main frequently and continuous deployment means running tests on every commit confuses these with continuous integration (frequent merges plus automated testing on each commit), which is the upstream practice both build on. The option that says continuous delivery ships to staging only and that the two terms are the same thing renamed for marketing is wrong: continuous delivery is about being able to release to prod at any time, not deploying only to staging, and the two terms are genuinely different, not a rename. The interview tell is naming the manual-approval gate as the one differentiator.

In a CI/CD pipeline, what is the practice of compiling/packaging a build artifact exactly once and then promoting that same immutable artifact across environments (e.g., dev → staging → prod) commonly called? Name the practice and briefly state one reason it is preferred over rebuilding from source for each environment.#

Show answer

The practice is called "build once, promote everywhere" (also referred to as promoting immutable artifacts or single-source artifact promotion). One key reason it is preferred is that it guarantees the exact binary tested in staging is the one deployed to production, eliminating drift caused by non-deterministic builds (e.g., differing dependency versions, toolchain updates, or environment-specific build flags) that could introduce defects between environments.

Why:

The 'build once, promote everywhere' principle holds that a CI pipeline produces a single, versioned, immutable artifact (e.g., a Docker image, JAR, or tarball) that is then promoted through environments without recompilation. This eliminates non-deterministic build drift — the risk that a rebuild picks up different dependency versions, compiler patches, or environment variables, producing a binary that differs from the one that passed earlier test gates. The keyword 'immutable' captures the core constraint: once built, the artifact is never modified, only re-tagged or copied as it moves downstream.

In a CI/CD pipeline, a _____ gate samples live telemetry — such as error rate and p99 latency — for a fixed observation window after routing a small percentage of production traffic to the new release. The pipeline proceeds to full rollout only if every metric threshold stays within its defined limit throughout that window; otherwise the release is automatically rolled back.#

Show answer

In a CI/CD pipeline, a canary gate samples live telemetry — such as error rate and p99 latency — for a fixed observation window after routing a small percentage of production traffic to the new release. The pipeline proceeds to full rollout only if every metric threshold stays within its defined limit throughout that window; otherwise the release is automatically rolled back.

Why:

A canary gate (or canary release gate) is the standard deployment-gate pattern in which a new version receives a small slice of real traffic while automated checks observe health metrics for a set duration. If any threshold is breached during the observation window, the gate fails and the pipeline triggers a rollback; only a clean window allows the rollout to expand to full traffic. The name comes from the 'canary in a coal mine' metaphor, where the small group serves as an early-warning signal for the broader deployment.

A team stores their Jenkins pipeline definition in a Jenkinsfile committed to the application repository. The Jenkinsfile loads a shared pipeline library from a separate Git repository via library 'shared-pipeline-lib' (no version or commit pinned). Which statement about this pipeline-as-code setup is correct?#

Options

Show answer

A change to the shared library's default branch can modify every consuming pipeline with no commit in any application repository. Because library 'shared-pipeline-lib' specifies no pinned version, Jenkins resolves the library from its default branch at each build, so the library and the Jenkinsfile are not coupled to the same commit and the library's changes are reviewed in its own repo, not the application's.

Why:

Because library 'shared-pipeline-lib' is declared without a version or commit qualifier, Jenkins resolves it from the library's default branch at build time. A merge to that default branch therefore changes the pipeline behavior for every application that loads it — even though none of those application repositories received a commit. The claim that every change to the shared library is reviewed through the application's pull-request process is wrong because the library lives in its own repository, so its changes are reviewed through that repository's PR process, not the application's. The claim that the pipeline is fully reproducible from the application repository's commit is wrong because the library is not pinned to the application repo's commit; only the Jenkinsfile is. The claim that moving pipeline logic into the shared library eliminates the need for a Jenkinsfile in each consuming application repository is wrong because a Jenkinsfile is still required in each application repo to declare the library load and any app-specific configuration.

In a CI/CD pipeline for a compiled application, arrange the following artifact-lifecycle stages in their correct sequential execution order — from pipeline trigger to the artifact being available in a registry for deployment. Each stage depends on the output of the previous one.#

Put these in order

Show answer

The correct order is Source → Build → Package → Publish. A commit triggers the pipeline; source is compiled and linked into a binary; the binary is bundled with its runtime dependencies into a deployable artifact; and the packaged artifact is pushed to a container registry. Each stage depends on the direct output of the one before it, forming a strict prerequisite chain.

Why:

Each stage here has a hard technical dependency on the output of the previous one. A commit to version control (Source) must exist before anything can be built. Building (Build) compiles and links source into an executable binary — regardless of whether the toolchain performs compilation and linking as one invocation or as separate steps, the binary is the output of this stage. Packaging (Package) bundles that binary with its runtime dependencies into a deployable artifact; you cannot package a binary that has not been built. Finally, Publishing (Publish) uploads the finished artifact to a container registry so it can be pulled and deployed; you cannot publish what has not been packaged. These stages form a strict prerequisite chain: each step consumes the direct output of the one before it.

A pipeline builds an immutable artifact tagged with the commit SHA and promotes the same artifact through environments (it never rebuilds per env). After deploying v=git-9f3a1c to prod, a regression appears. The rollback script runs the commands below. What does the rollback do, and is it safe?#

PREV=$(deploy-history prod --nth 2 --field artifact)   # -> app:git-2b7e08
echo "rolling back to $PREV"
deploy prod --artifact "$PREV"   # re-deploys the exact bytes that ran before

Options

Show answer
It re-deploys the exact previously-running artifact (app:git-2b7e08) with no rebuild, so the rollback is fast and reproducible — you get back the bytes that were known-good in prod
Why:

The whole value of building once and promoting an immutable, SHA-tagged artifact shows up at rollback time. The script looks up the artifact that was running two deploys ago (app:git-2b7e08) and re-deploys those exact bytes — no recompilation, no dependency resolution, no fresh image build. That makes the rollback fast and, crucially, reproducible: you are restoring the artifact that was already proven in prod, not a hopefully-equivalent rebuild. The claim that it rebuilds the previous commit from source describes the rebuild-per-environment anti-pattern, which the prompt explicitly rules out — rebuilding from an old commit can silently pull newer transitive dependencies or a newer base image and yield a different binary, defeating the point of rolling back. The claim that it cannot roll back because immutable artifacts can never be re-deployed once superseded is the opposite of true: immutability is what enables clean rollback by tag. The claim that it reverts the prod database to the previous schema automatically as part of re-deploying the artifact is the important caveat stated as a wrong answer — re-deploying an artifact rolls back code only; database migrations are not reversed by this and must be handled separately (which is why backward-compatible, expand-then-contract migrations matter).

You are the platform lead for a large e-commerce platform running 40+ microservices on Kubernetes. Each service team owns its own CI/CD pipeline and can deploy independently. The organization wants to introduce deployment gates so that no service reaches production without passing automated checks, but without re-centralizing release approval or creating a single bottleneck team.#

Show answer

Gate architecture: Each pipeline has five stages with both team-configurable and org-mandated gates.

  1. Build stage: Org-mandated gate — container image build + SAST scan + dependency vulnerability scan (fail on Critical CVEs). Team gate — unit tests with ≥80% coverage delta threshold.
  2. Integration stage: Org-mandated gate — consumer-driven contract tests against the service's published API spec (Pact or equivalent). The gate fails if any registered consumer's contract is broken. Team gate — integration tests against a shared ephemeral environment.
  3. Staging stage: Org-mandated gate — load test reaching 2× expected peak QPS; p99 latency must stay within SLO. Team gate — smoke tests and synthetic transaction checks.
  4. Pre-prod (canary launch): The service is deployed to 5% of production traffic. Canary gate evaluates error rate (must not exceed baseline + 0.5 percentage points), p99 latency (must not exceed baseline × 1.2), and business KPI delta (e.g., checkout conversion rate must not drop more than 1% relative to control). If any metric breaches its threshold, the gate triggers an automatic rollback to the previous version. Metrics that are informative but not safety-critical (e.g., cache hit rate, memory usage trending up but within limits) trigger a hold-and-notify — deployment pauses at 5%, alerts the team, and waits for manual promotion or rollback.
  5. Production (progressive rollout): After canary passes, traffic is ramped 25% → 50% → 100% with the same canary gate re-evaluated at each step. A failure at any step rolls back to the last known-good percentage.

Cross-service dependency gating: Every service publishes an OpenAPI/Protobuf spec in a central schema registry. Before a service can deploy to production, the gate runs backward-compatibility checks — if the new spec removes a required field or changes a type in a breaking way, the gate fails unless all downstream consumers have already deployed a compatible version. The dependency graph is maintained by a service catalog; the gate queries it to identify consumers. Consumer-driven contract tests run at the integration stage and must all pass; the gate blocks if any registered consumer's contract is broken by the new version.

Decentralization with enforced baseline: Gates are defined as code in a shared gate-library repository. Each team's pipeline imports a 'base-gate-profile' that contains the org-mandated gates (SAST, CVE scan, contract tests, load test, canary SLO gate). Teams can add team-specific gates on top but the base profile is enforced by a policy controller (OPA/Conftest) that validates every pipeline definition at commit time — if the base gates are missing or disabled, the pipeline config is rejected by the CI system and cannot run. This gives teams autonomy to extend while guaranteeing the safety floor.

Why:

This design exercise tests the candidate's ability to architect a multi-layered deployment-gate system that balances safety with team autonomy. A strong answer covers stage-specific gates with concrete pass/fail criteria, canary rollback signals with numeric thresholds, cross-service dependency validation with an explicit block condition, and a policy-enforced baseline that prevents teams from disabling safety gates while still allowing extension. The rubric was corrected so that c1 (stage taxonomy) and c3 (cross-service dependency gating) no longer both credit consumer-driven contract tests: c1's integration-stage example now references integration tests against an ephemeral environment, while c3 exclusively owns contract tests, schema-compatibility checks, and dependency-graph mechanisms, and additionally requires the candidate to specify the exact block condition — making it discriminating against shallow answers that merely name a technique.

You are designing the deployment pipeline for a financial-services application subject to SOX and PCI-DSS compliance. The business wants to move from monthly releases to weekly production deployments, but compliance requires that every production deployment has an auditable approval trail, separation of duties (the person who builds cannot be the person who approves), and evidence that specific quality gates were met.#

Show answer

Gate taxonomy and automation boundary:

Automated gates (no human action required, evidence captured automatically):

  • Build gate: compilation + unit tests + coverage threshold (≥80%). Pass/fail is deterministic.
  • Security gate: SAST, DAST, dependency CVE scan, container image scan. Fail on Critical/High.
  • IaC gate: Terraform plan validation + policy-as-code check (no public S3 buckets, encryption required, etc.).
  • Integration gate: contract tests + integration tests against a staging environment mirroring production config.
  • Compliance gate: automated check that the deployment package matches the approved change ticket (e.g., Jira issue linked in the commit), and that all required automated gates have passed with their evidence stored.

Human-approval gate (the only manual step):

  • Production release authorization: a designated approver (not the builder) reviews the compiled evidence package — test results, scan reports, change-ticket linkage, artifact hash — and explicitly approves promotion to production. This gate exists because SOX requires management authorization for changes to production financial systems, and a human must attest that the change was reviewed.

Separation of duties enforcement: The CI/CD platform enforces two disjoint RBAC roles: 'Deployer' (can trigger pipeline runs, merge code, push artifacts through staging) and 'Approver' (can authorize production promotion). The pipeline's production-promotion step checks the identity of the approver against the identity of the user who triggered the current pipeline run. If they match, the step is blocked and returns an error. Additionally, pull requests require at least one review from a member of the Approver group before merge, and the PR system prevents self-review. These are enforced by the platform (e.g., pipeline engine identity checks + SCM branch-protection rules), not by verbal policy.

Compliance evidence capture: Every gate — automated and manual — writes a structured evidence record to an immutable, append-only audit store (e.g., a dedicated evidence repository or a tamper-evident log service). Each record contains: gate name, gate version/config hash, input artifact hash, pass/fail result, detailed output (test report, scan findings), timestamp (UTC, from a trusted time source), and the identity that triggered or approved it. Records are correlated by a deployment ID that links every gate result to the specific artifact hash and change ticket. An auditor can query by deployment ID or date range and receive a complete, ordered timeline of every gate executed, its result, and who acted. The audit store uses hash-chaining (each record references the hash of the previous record) to make tampering detectable.

Bottleneck reduction: (1) The human approver never runs tests or scans manually — by the time approval is requested, all automated gates have passed and the evidence package is pre-assembled. The approver reviews a summary dashboard with green/red indicators and drill-down links, taking 2–5 minutes. (2) Production rollout uses canary delivery: the approver authorizes a 5% canary, automated health gates evaluate metrics for 15 minutes, and if they pass, the rollout auto-promotes to 100% without a second human approval. The human is in the loop once, not at every step. (3) For low-risk changes (documentation, non-functional config via feature flags), an expedited path auto-approves if the change classifier determines the diff touches only flag-config files — the compliance gate still records all evidence and the auto-approval decision with its rationale.

Why:

This exercise tests the candidate's ability to design deployment gates that satisfy regulatory compliance (SOX/PCI-DSS) without reintroducing slow manual release processes. A strong answer distinguishes automated from human gates with rationale, enforces separation of duties through technical controls, captures tamper-evident audit evidence, and uses techniques like pre-assembled evidence packages and canary delivery to minimize human bottleneck.

You run a continuous delivery pipeline for a polyglot microservices platform where several services share a common PostgreSQL database (not ideal, but it's the current state). Database schema migrations are a frequent source of deployment failures and production incidents — breaking changes, long-running ALTER TABLE locks, and migrations that succeed in staging but fail in production due to data volume differences.#

Show answer

Breaking-change detection and expand-contract enforcement:

The gate includes a migration analyzer that parses each migration script (or diffs the before/after schema) and classifies every DDL operation:

  • Breaking: DROP COLUMN, DROP TABLE, RENAME COLUMN, type narrowing (VARCHAR(255) → VARCHAR(50)), adding NOT NULL constraint without a default, removing a default.
  • Non-breaking: ADD COLUMN (nullable or with default), ADD INDEX (using CONCURRENTLY), widening types, adding a table.

When the gate detects a breaking change, it checks the migration history and the currently-deployed application code version. For a DROP COLUMN, the gate requires that: (a) a prior migration added the replacement column, (b) the currently-deployed code version no longer references the old column (verified by checking that the deployed artifact's schema-config or ORM mapping doesn't include it), and (c) at least one full deployment cycle has passed since the expand migration. If any condition is unmet, the gate blocks the migration with a message like 'Contract phase blocked: column X still referenced by deployed code version Y.' This enforces expand-then-contract by making the gate stateful — it tracks which expand migrations have been deployed and only permits the corresponding contract migration after code deployment is verified.

Production-scale performance validation:

The gate does not trust staging data volume. Instead, it runs the migration against a production-sized data clone — a logical replica of production created from the latest PITR snapshot, restored to a disposable compute instance. The migration is executed against this clone with instrumentation: the gate captures wall-clock time, lock-wait time, and whether any statement holds an AccessExclusiveLock for more than a configurable threshold (e.g., 5 seconds). If the migration exceeds the time threshold or acquires a blocking lock, the gate fails and suggests an alternative (e.g., use CREATE INDEX CONCURRENTLY, or batch the UPDATE). For very large tables, the gate also runs EXPLAIN ANALYZE on destructive statements to estimate row-scan costs. The clone is sized to match production's largest table row counts, closing the staging-vs-production gap.

Migration-application deployment ordering:

The gate enforces ordering through a deployment manifest that each service submits with its pipeline. The manifest declares the migration type (expand, contract, or none) and the code compatibility level (backward-compatible with old schema, requires new schema). The gate enforces:

  1. If the migration is 'expand' (additive): migration runs first, then new code is deployed. Old code remains compatible because the migration is additive.
  2. If the migration is 'contract' (removing old schema): the gate verifies the currently-deployed code is the backward-compatible version from the expand phase, then runs the contract migration, then (optionally) deploys code that assumes the old schema is gone.
  3. If no migration: standard code deployment. The pipeline engine refuses to proceed if the manifest's declared ordering doesn't match the gate's classification — e.g., if a team labels a DROP COLUMN migration as 'expand,' the analyzer catches the mismatch and blocks.

Rollback safety:

The gate classifies each migration as reversible or irreversible at analysis time:

  • Reversible: ADD COLUMN (nullable), ADD INDEX, CREATE TABLE. These can be rolled back with DROP, and no committed data is lost (the column was new).
  • Irreversible: DROP COLUMN (data in that column is destroyed), DROP TABLE, TRUNCATE, type narrowing (data may be truncated).

Policy: When a deployment fails after a migration has been applied, the system checks the migration's reversibility classification. If reversible, the migration is rolled back and the previous code version is redeployed. If irreversible, the system rolls back the code to the previous version but leaves the schema change in place — rolling back the migration would cause data loss. The incident is flagged for a forward-fix: the team must fix and redeploy rather than revert the schema. The gate also requires that all contract-phase migrations are preceded by a verified backup or that the data being removed has been archived — the gate checks for an archive artifact (e.g., a dump of the column's data stored in object storage) before permitting an irreversible migration to proceed.

Why:

This exercise tests the candidate's ability to design a database-migration gate that prevents schema-related deployment failures. A strong answer includes static analysis for breaking changes with expand-contract enforcement, production-scale performance validation that closes the staging-vs-production data gap, enforced migration-code deployment ordering, and a rollback policy that distinguishes reversible from irreversible migrations to prevent data loss.

A monorepo team has a linear CI/CD pipeline: Build → Unit Tests → Integration Tests → Security Scan (SAST) → Build Docker Image → Deploy to Staging → E2E Tests → Deploy to Production. Pipeline wall-clock time is 25 minutes and developers frequently wait for the full run before getting feedback. The team wants to reduce feedback time without weakening deployment confidence. Which redesign is the soundest approach?#

Options

Show answer

Build the deployment artifact once, run fast feedback gates (lint, unit tests) first, then parallelize independent slower checks (integration tests, SAST) after those pass, and promote the same image through staging and production. This follows fail-fast sequencing and build-once-promote-everywhere: you get early signal on cheap failures, avoid redundant rebuilds, and guarantee the tested artifact is the one deployed.

Why:

Building the Docker image once in the Build stage, running fast gates (lint, unit tests) first, then parallelizing independent slower checks (integration tests, SAST) after fast gates pass, and promoting that same image through staging and production deployments applies two well-established CI/CD principles. First, fail-fast: cheap, high-signal checks (lint, unit tests) run before expensive ones (integration tests, SAST), so a trivial failure is caught in seconds rather than after waiting for slower parallel jobs to finish — which is why running Unit Tests, Integration Tests, Security Scan, and E2E Tests all in parallel immediately after Build wastes compute and delays the most common feedback. Second, build-once-promote-everywhere: the artifact tested in later stages is the exact artifact deployed to production, which rebuilding the Docker image inside each deployment stage violates by rebuilding per environment (breaking traceability and the guarantee that the tested image is the deployed image). Merging Integration Tests, Security Scan, and E2E Tests into a single 'Quality Gate' stage collapses distinct gates into one, losing granular feedback on which check failed and removing the ability to short-circuit early on a fast failure. Building the Docker image once in the Build stage, running fast gates (lint, unit tests) first, then parallelizing independent slower checks (integration tests, SAST) after fast gates pass, and promoting that same image through staging and production deployments preserves the same artifact across stages while sequencing fast gates before slow ones and parallelizing only the independent slow checks that remain.

You are the platform lead for a large-scale microservices platform serving 50M+ daily active users across multiple regions. Each service deploys independently 5-20 times per day. Leadership wants to move from manual pre-production approvals to an automated progressive delivery system with canary analysis, but the platform must satisfy two hard constraints: (1) any deployment that causes a regression in user-facing SLOs must auto-rollback within 90 seconds of the regression becoming statistically detectable, and (2) the system must support per-service customization of gate logic while enforcing platform-wide safety invariants (e.g., no service can skip the canary stage).#

Show answer

The deployment pipeline progresses through these gates: (1) Commit gate — CI builds the artifact, runs unit tests and SAST; (2) Integration gate — deploys to a staging cluster and runs integration and contract tests; (3) Canary gate — deploys to 1% of traffic in a single AZ in the primary region with a 5-minute soak; (4) Progressive rollout gates — traffic shifts through 5%, 25%, 50%, and 100% in the primary region with a 10-minute soak and gate evaluation at each stage; (5) Multi-region gates — repeat stages 3-4 region-by-region in a defined order (primary → secondary → edge regions).

Complete gate pipeline from commit to production (c1): Every change starts at the commit gate, where CI builds the container image, runs unit tests, executes SAST/DAST scans, and publishes a signed artifact to the registry. Only a passing commit gate triggers the integration gate, which deploys the artifact to a staging cluster that mirrors production topology (minus scale) and runs integration tests, consumer-driven contract tests, and schema-compatibility checks. A failing integration gate blocks the artifact from ever reaching production traffic. Only after both pre-deployment gates pass does the system proceed to the canary gate, ensuring no code reaches live traffic without automated pre-deployment validation.

Canary SLI selection and baseline comparison (c2): Each service declares its SLIs via a policy file: error rate (HTTP 5xx / total requests), p99 latency, and saturation (CPU and connection pool utilization). The baseline is the previous deployed version running concurrently in the same environment. The platform collects metrics for both the canary and baseline cohorts using the same observability stack (e.g., Prometheus with canary/baseline labels injected by the service mesh). The comparison is relative: canary error rate must not exceed baseline by more than the service-defined tolerance.

Statistical rollback decision boundary (c3): The platform uses sequential hypothesis testing (specifically, a sequential probability ratio test, SPRT) on the error-rate delta between canary and baseline. Parameters: minimum sample size of 10,000 requests per cohort before evaluation begins; significance threshold of p < 0.01 for rollback (null hypothesis: canary error rate ≤ baseline); evaluation window of 60 seconds sampled every 10 seconds. If at any sample the SPRT rejects the null hypothesis (canary is worse), the system triggers automatic rollback. If after the full soak period the SPRT has not rejected the null and the canary is not worse, the gate promotes to the next traffic stage. A secondary check on p99 latency uses a non-inferiority test: canary p99 must not exceed baseline p99 by more than 10ms with 95% confidence. If either check fails, rollback fires. The 10-second sampling interval combined with the SPRT's sequential nature means that once statistical power is reached (minimum sample size met), a regression is detected and rollback initiated within the 90-second window.

Progressive traffic-shifting stages with soak gates (c4): Traffic shifts via the service mesh load balancer: 1% (5-min soak) → 5% (10-min soak) → 25% (10-min) → 50% (10-min) → 100% (10-min). At each stage boundary, the gate evaluates all SLIs using the same SPRT. If any gate fails, traffic is immediately shifted back to 0% for the canary version (full rollback to previous version) and the deployment is marked failed with an alert to the owning team. The soak times are minimums; the gate cannot advance until both the soak time has elapsed and the SPRT has not triggered rollback.

Blast radius containment (c5): Two mechanisms limit impact. First, sticky-session routing: users assigned to the canary cohort are pinned via a cookie or consistent hash so that the canary population is stable and does not rotate new users in during evaluation — this means at most 1% of users are exposed at the first canary stage. Second, single-AZ canary: the initial canary deploys to one AZ in one region, so a canary failure affects at most one AZ's worth of traffic in one region, and the mesh health checks drain the canary instances within seconds. On rollback, the service mesh removes canary endpoints from the load balancer pool; existing canary sessions are redirected to baseline instances. For stateful services, session draining waits for in-flight requests to complete (up to a 30-second drain timeout) before terminating canary instances.

Per-service customization with enforced platform invariants (c6): Services define their gate policy in a YAML file versioned alongside their code: custom SLIs (e.g., a payment service might add a 'payment success rate' SLI), custom thresholds (tolerance for error-rate delta, latency budget), and pre-deploy checks (e.g., a database migration dry-run). The platform loads this policy but applies a platform-level override layer that enforces non-negotiable invariants: the canary stage cannot be skipped or disabled; minimum soak times cannot be set below platform-defined floors (5 min for canary, 10 min for progressive stages); the SPRT significance threshold cannot be weakened below p < 0.01; and the rollback action is always platform-controlled (services cannot override the rollback trigger). If a service's policy conflicts with an invariant, the platform rejects the policy at the integration gate and the deployment cannot proceed. This model lets services customize their sensitivity while the platform guarantees that no service can weaken safety below the floor.

Multi-region rollout failure handling (c7): During multi-region rollout, each region runs its own independent canary and progressive stages. If region B's canary fails, region B rolls back immediately and independently — its traffic returns to the previous version without waiting for any cross-region coordination. The platform then pauses the rollout to all remaining un-deployed regions (region C, D, etc.) to prevent propagating a potentially broken build. Region A (already promoted to 100%) is not automatically rolled back. Instead, the platform raises a high-severity alert and surfaces the region B failure analysis (which SLIs failed, magnitude of regression, error signatures) to the owning team, who must explicitly decide whether to rollback region A. This alert-and-halt approach is chosen over automatic multi-region rollback because the failure signal in region B may be region-specific (e.g., a regional dependency outage, a data residency edge case) and would not necessarily reproduce in region A. An unnecessary multi-region rollback of healthy regions carries its own blast radius — it disrupts users currently served by the new (and apparently healthy) version in region A and consumes deployment capacity. The platform provides a one-click rollback for region A and, if the team identifies the failure as a global code defect, a single action to rollback all regions simultaneously.

Why:

This is a staff-level design exercise requiring the candidate to integrate progressive delivery, statistical canary analysis, blast-radius engineering, and policy-as-code extensibility into a coherent system with hard performance constraints (90-second rollback) and organizational tensions (per-service customization vs. platform safety invariants). The rubric now covers all seven areas the prompt explicitly requires: the full commit-to-production gate sequence, SLI/baseline selection, statistical decision boundaries, progressive traffic shifting, blast radius containment, customization with invariants, and multi-region failure handling.

In a distributed, multi-node CI build system using remote build caching (e.g. Bazel remote execution or Buildkit's cache exporter), what is the key design principle that makes build artifact cache hits safe across heterogeneous worker nodes with different OS kernels and CPU architectures, and how is it enforced?#

Show answer

Use a content-addressable storage (CAS) model keyed on the hash of inputs (source + toolchain + env vars + transitive deps), not on branch/commit alone. Bazel and Buck2 exemplify this: the build graph is hermetically sealed (all inputs explicitly declared), outputs are stored by their content hash, and remote execution nodes share a CAS-backed cache. A cache hit on node A is therefore deterministic on node B because the hash captures the full input set—toolchain version, compiler flags, env var values, and dependency hashes—making locality irrelevant.

Why:

This tests staff-level understanding of hermetic, content-addressable build systems—the foundation of safe remote caching. The key insight is that correctness of cross-node cache hits depends on hashing the entire input closure (including toolchain, flags, env, transitive deps), not just source. Without hermeticity, a cache hit on one node may reflect a different effective toolchain or environment, producing non-reproducible artifacts. Systems like Bazel enforce this through declared inputs and sandboxed execution.

Distinguish between 'build reproducibility' and 'build hermeticity' in the context of producing trustworthy, cacheable build artifacts. Why is hermeticity the stronger property, and what specific ambient state must be eliminated to achieve it?#

Show answer

Reproducibility means the build is a pure function of declared inputs: the same inputs produce byte-identical output regardless of time, worker, or environment. Reproducibility is the prerequisite—without it, cache keys derived from inputs are unreliable. Hermeticity is the enforcement mechanism: all inputs are explicitly declared and ambient state (wall clock, network, env vars, filesystem layout) is sandboxed or forbidden. Together, hermeticity guarantees that the input hash captures everything, which makes reproducibility provable and cache keys trustworthy.

Why:

At staff level, the distinction matters because many teams conflate reproducibility (same inputs → same output) with hermeticity (no undeclared inputs at all). Reproducibility can exist accidentally; hermeticity makes it structural by banning ambient state—wall-clock time, random number generators, network access, environment variables, absolute paths, and host filesystem assumptions. Only hermetic builds produce cache keys that are safe to share across workers.

Related interview questions

Job market

See devops salaries and hiring demand from live job postings.

The other 6 questions

This page shows 25 and marks what you pick. That's as far as a page can go. A free account opens the other 6 and keeps every answer. What you miss comes back until it's right: after a day, then at longer gaps.

Start with this topic

Free · the whole bank · 100 marked answers per 30 days · written feedback on the paid plan

What moved, monthly

One email a month when the bulletin comes out: what moved in the markets we track, and the new question topics we published. Confirm your address to join. Unsubscribe any time.