39 practice questions for Domain 3 of the AWS Certified DevOps Engineer - Professional (DOP-C02) exam, which makes up 15% of its scored content. Your answers count towards one score and one timer for the whole exam.
Domain 3: Resilient Cloud Solutions
88. A stateful service must survive the loss of an Availability Zone with its data intact and no manual intervention. Which solution meets these requirements?
Answer and explanation
Answer: D. A Multi-AZ managed database maintains a synchronous standby and fails over automatically, and multi-zone compute survives the same event. Snapshots require a manual restore. Instance storage does not survive instance loss. Manual replica promotion is not automatic.
89. Single points of failure must be identified across many existing workloads rather than discovered during an incident. Which solution meets these requirements?
Answer and explanation
Answer: B. Resilience Hub assesses a workload against a defined resiliency policy and reports where it cannot meet its recovery objectives, which is systematic discovery. Annual diagram review depends on the diagram being current. Waiting for an incident is discovery by failure. Trusted Advisor checks specific resource conditions rather than assessing a workload against recovery objectives.
90. An Amazon S3 bucket's contents must remain available if the primary Region becomes unavailable. Which solution meets these requirements?
Answer and explanation
Answer: A. Cross-Region Replication maintains a copy in another Region that readers can be directed to during a Regional event. Versioning protects against overwrite within one Region. Intelligent-Tiering and lifecycle transitions manage cost rather than Regional availability.
91. An Application Load Balancer must distribute traffic evenly across targets in three Availability Zones where target counts differ per zone. Which configuration meets these requirements?
Answer and explanation
Answer: A. Cross-zone load balancing distributes requests across every registered target rather than dividing traffic equally between zones first. Equalising target counts is a workaround that constrains capacity planning. Sticky sessions bind clients to targets. Health check timing governs failure detection.
92. A business requirement states a service level of 99.95 percent availability for a workload currently deployed in one Availability Zone. Which step should be taken first?
Answer and explanation
Answer: B. Translating an availability target into design changes starts with identifying what can fail and take the service down. A larger instance is still a single instance in a single zone. A support plan does not change the architecture's availability. More frequent backups shorten recovery rather than preventing the outage.
93. An Auto Scaling group scales on average CPU, but the workload is bound by queue depth rather than CPU. Which solution meets these requirements?
Answer and explanation
Answer: C. Scaling should follow the metric that actually constrains the workload, which here is queue backlog per instance rather than CPU. A lower CPU target scales on the wrong signal. A larger maximum permits more capacity without triggering it. A shorter cooldown reacts faster to the wrong metric.
94. An Amazon ECS service must scale its tasks with demand while the underlying capacity scales to match. Which solution meets these requirements?
Answer and explanation
Answer: C. Task-level scaling paired with a capacity provider scales both layers together so tasks are not blocked waiting for capacity. A fixed-size group caps task scaling. A fixed task count does not scale with demand. Fargate removes capacity management but a fixed task count still does not scale.
95. A read-heavy workload's database is saturated by repeated identical queries. Which solution most reduces load on the database?
Answer and explanation
Answer: B. A cache removes the repeated work entirely, which is the largest reduction available when the same queries recur. Read replicas move the work to another instance that still executes it. A larger instance and higher IOPS both make the same queries faster at higher cost.
96. A serverless application must handle a predictable morning traffic spike without cold start latency. Which solution meets these requirements?
Answer and explanation
Answer: A. Provisioned concurrency keeps environments initialized and can be scheduled ahead of a predictable spike. Reserved concurrency caps and guarantees concurrency without pre-initializing. More memory shortens rather than removes initialization. Buffering converts a synchronous call to asynchronous.
97. A workload must serve users on three continents with low latency and must tolerate the loss of a Region. Which solution meets these requirements?
Answer and explanation
Answer: C. Multi-Region deployment with health-checked latency routing gives both proximity and Regional failure tolerance. A content delivery network helps static content but leaves dynamic requests crossing continents. Larger instances do not shorten distance. Read replicas serve reads locally while writes and failover remain single-Region.
98. A workload must recover from Regional failure with a recovery time objective of fifteen minutes and a recovery point objective of five minutes. Which strategy meets these requirements?
Answer and explanation
Answer: D. Warm standby keeps a running scaled-down environment with continuous replication, so failover is scaling up and shifting traffic within the stated window. Pilot light must provision and start the application tier, which typically exceeds fifteen minutes. Backup and restore takes hours. Active-active meets the targets but exceeds what the requirement calls for.
99. Failover testing must run on a schedule and produce evidence, rather than depending on an engineer initiating it. Which solution meets these requirements?
Answer and explanation
Answer: B. Automating the exercise makes it repeatable and produces evidence without depending on anyone starting it. A calendar reminder still relies on a person. A runbook review validates understanding rather than behaviour. Replication lag confirms a precondition rather than that failover works.
100. A backup policy must be applied uniformly across every account in an organization, including accounts created later, without per-account configuration. Which solution meets these requirements?
Answer and explanation
Answer: D. An Organizations backup policy is inherited by accounts in the organizational unit, including those that join later, and is managed centrally. Per-account plans drift and miss new accounts. StackSets deploy the plan but each account can then modify it locally. A custom function recreates an inherited-policy capability with more failure modes.
101. A stateful service must fail over between Availability Zones without losing in-flight data. Which approach is appropriate?
Answer and explanation
Answer: A. Synchronous replication acknowledges a write only once the standby has it, so committed data survives the failover. Asynchronous replication and periodic snapshots both lose whatever occurred after the last replication point. Rebuilding from source extends the outage.
102. A cross-Region solution must keep an Amazon DynamoDB table readable and writable in two Regions. Which solution meets these requirements?
Answer and explanation
Answer: C. A DynamoDB global table replicates between Regions with both accepting writes. DynamoDB does not offer read replicas in that form. A nightly export is stale and read-only. A point-in-time restore takes time and is a recovery action rather than an active second Region.
103. A load balancer must continue serving when one Availability Zone's targets all become unhealthy. Which configuration matters most?
Answer and explanation
Answer: A. Surviving the loss of a zone requires healthy targets in another zone. A shorter interval detects the failure faster without providing an alternative. Deregistration delay governs draining. Sticky sessions bind clients to targets that may be the failed ones.
104. An Auto Scaling group launches instances that take eight minutes to become ready, and scaling lags demand. Which improvement is appropriate?
Answer and explanation
Answer: D. Eight minutes of startup is the constraint, addressed by shortening it and by adding capacity before demand arrives. A shorter cooldown launches more instances that are all still slow. A lower threshold starts earlier but still waits eight minutes. A higher maximum does not make instances ready sooner.
105. A container platform must scale its worker nodes only when pods cannot be scheduled. Which approach is appropriate?
Answer and explanation
Answer: B. Unschedulable pods are the direct signal that capacity is insufficient, which node utilization can miss when resources are reserved but idle. A fixed peak-sized group pays for idle nodes. A schedule does not respond to actual demand.
106. A serverless application's downstream database is overwhelmed when the function scales out. Which solution meets these requirements?
Answer and explanation
Answer: C. Capping concurrency bounds the load the function can generate and a connection proxy pools connections so each invocation does not open its own. More memory makes invocations faster and therefore more numerous. A larger database raises the ceiling without bounding the caller. A longer timeout holds connections longer.
107. A workload must be deployed to several Regions with one definition and one deployment action. Which approach is appropriate?
Answer and explanation
Answer: A. StackSets deploy one definition across Regions and accounts from a single action. Manual per-Region deployment drifts. A replication script recreates StackSets with more failure modes. Single-Region deployment with global routing does not meet a multi-Region requirement.
108. A recovery procedure must be validated to meet a stated recovery time objective. Which approach is appropriate?
Answer and explanation
Answer: A. Measuring an actual execution is the only way to know whether the objective is met. An estimate from documentation omits what goes wrong. Backup duration is not recovery duration. Resource existence is a precondition.
109. An automated recovery must restore a database to a point immediately before a bad deployment. Which capability is required?
Answer and explanation
Answer: A. Point-in-time recovery restores to a chosen moment within the retention window, which is what recovering to just before a specific event requires, and it must be enabled beforehand. Daily snapshots restore to the snapshot time. A replica carries the same bad change. Cross-Region copies address Regional loss.
110. An Aurora cluster must fail over to a replica with minimal disruption when the writer instance fails. Which configuration is appropriate?
Answer and explanation
Answer: A. A reader with failover priority is promoted automatically and the cluster endpoint follows the new writer. A direct instance endpoint points at the failed instance. Snapshots require a restore. A larger instance is still a single writer.
111. A workload must survive the loss of an entire Region. Which design is required?
Answer and explanation
Answer: A. Regional survival requires presence in another Region with data and traffic direction. Multiple zones protect within a Region only. Snapshot copies enable recovery but not continuity without an environment. More instances in the failed Region do not help.
112. An Elastic Load Balancer must stop routing to a target whose dependencies have failed even though the target process is running. Which configuration is appropriate?
Answer and explanation
Answer: D. A deep health check verifying dependencies reports the real state. A TCP check confirms the port is open. Interval and threshold govern detection timing rather than what is checked.
113. An Auto Scaling group must replace instances gradually when a new launch template version is published. Which capability is appropriate?
Answer and explanation
Answer: D. An instance refresh replaces instances in controlled batches honouring a health threshold. Terminating everything causes an outage. Capacity manipulation is a crude substitute. Manual termination does not scale.
114. A scaling policy must add capacity proportionally to how far a metric exceeds its threshold. Which policy type is appropriate?
Answer and explanation
Answer: D. Step scaling defines different adjustments for different breach magnitudes. Target tracking maintains a target without explicit steps. Simple scaling applies one adjustment. Scheduled scaling is time-based.
115. A predictable weekly traffic pattern must be met without waiting for reactive scaling. Which capability is appropriate?
Answer and explanation
Answer: D. Predictive scaling forecasts recurring patterns and provisions ahead of demand. Target tracking and step scaling both react after the metric moves. A larger instance does not anticipate demand.
116. A recovery process must restore an application and its data in a second Region with the least manual work. Which approach is appropriate?
Answer and explanation
Answer: B. Infrastructure as code with continuous replication makes recovery a deployment rather than a rebuild. Runbooks, manual rebuilds, and spreadsheets all depend on human execution during an incident.
117. An AWS Backup plan must copy recovery points to a second Region and retain them for a defined period. Which configuration is appropriate?
Answer and explanation
Answer: C. A copy action in the backup rule replicates recovery points with independent retention. A separate plan in the second Region backs up different resources. Manual copying is unreliable. Storage replication is not the same as backup copies.
118. An Aurora cluster's reader endpoint must distribute read traffic across replicas. Which configuration is required?
Answer and explanation
Answer: B. The reader endpoint balances across replicas and adapts as they change. The writer endpoint sends reads to the primary, an instance endpoint targets one replica, and a load balancer is not how Aurora distributes reads.
119. A multi-Region active-passive workload must detect that the primary Region is unhealthy from outside that Region. Which approach is appropriate?
Answer and explanation
Answer: A. Detection must not depend on the failing Region. In-Region checks and alarms may be unavailable during a Regional event, and user reports are slow.
120. A service must remain available when a dependency in another Availability Zone becomes slow rather than failing outright. Which pattern applies?
Answer and explanation
Answer: D. A bounded timeout with a fallback prevents a slow dependency exhausting resources. Unbounded retries and longer timeouts hold resources, and a larger pool delays exhaustion.
121. An ElastiCache cluster must survive the loss of a node without losing cached data. Which configuration is appropriate?
Answer and explanation
Answer: D. A replication group with Multi-AZ and automatic failover promotes a replica. Snapshots require a restore, more memory does not add redundancy, and ElastiCache clusters are Regional.
122. An Auto Scaling group must avoid scaling on a metric spike caused by instances starting up. Which configuration is appropriate?
Answer and explanation
Answer: C. An instance warm-up excludes starting instances from the metric so their startup load does not trigger further scaling. Cooldown governs the interval between actions, a higher threshold delays all scaling, and maximum size caps capacity.
123. A serverless application's downstream API allows a fixed number of concurrent connections. Which configuration bounds the load?
Answer and explanation
Answer: B. Reserved concurrency caps parallel executions and therefore downstream connections. Timeout and memory affect execution characteristics, and retries add load.
124. An ECS service must scale on a metric published by the application rather than on infrastructure utilization. Which configuration is appropriate?
Answer and explanation
Answer: D. Target tracking supports custom metrics, which lets scaling follow the application's own signal. CPU may not correlate, a fixed count does not scale, and schedules ignore actual demand.
125. A recovery process must be validated without an engineer initiating it each time. Which approach is appropriate?
Answer and explanation
Answer: C. An automated scheduled exercise runs without depending on a person and records evidence. Calendar entries, post-incident exercises, and document reviews all rely on human initiation or are retrospective.
126. An automated recovery must not trigger during a planned maintenance window. Which approach is appropriate?
Answer and explanation
Answer: A. Checking within the automation makes the decision at the moment of action with no state to restore. Disabling and re-enabling alarms is error-prone and leaves genuine failures undetected, a longer evaluation period delays all alerts, and letting it fire causes unintended change.