22 practice questions for Domain 5 of the AWS Certified DevOps Engineer - Professional (DOP-C02) exam, which makes up 14% of its scored content. Your answers count towards one score and one timer for the whole exam.
Domain 5: Incident and Event Response
188. An AWS Health event indicating a scheduled instance retirement must automatically create a remediation task rather than only notifying the team. Which solution meets these requirements?
Answer and explanation
Answer: A. Invoking a runbook from the Health event turns the notice into an automated replacement with an auditable execution history. An email leaves the work to a person. A daily dashboard review may miss a short notice window. A status check alarm fires once the instance is already impaired rather than acting on the advance notice.
189. A high volume of events must be processed by several independent consumers, each receiving every event. Which solution meets these requirements?
Answer and explanation
Answer: C. The fan-out pattern publishes once to SNS, which delivers a copy to each subscribed queue, giving every consumer an independent buffer and retry behaviour. A shared queue delivers each message to only one consumer. A FIFO queue with one group serialises processing. Sequential invocation couples the consumers.
190. A specific EventBridge event must trigger a predefined sequence of remediation steps whose execution is auditable. Which solution meets these requirements?
Answer and explanation
Answer: B. Automation runbooks encode remediation declaratively, execute with a defined role, and record every step in execution history. Emailing steps requires a human to perform them. SSH access from Lambda reintroduces key management and produces no structured audit. A dashboard annotation records a note without acting.
191. A fleet's configuration must be corrected automatically when it drifts from the desired state. Which solution meets these requirements?
Answer and explanation
Answer: D. State Manager reapplies an association on a schedule, restoring the desired state whenever it drifts. A Config rule reports without correcting. Run Command executed on report is manual unless automated separately. Recreating instances is disproportionate to a configuration drift.
192. A remediation must not run twice for the same event even when the event is delivered more than once. Which design meets these requirements?
Answer and explanation
Answer: D. EventBridge delivers at least once, so the remediation must recognise an event it has already handled. A longer retry interval does not prevent duplicate delivery. A FIFO queue orders messages without deduplicating arbitrary redeliveries. Manual approval removes the automation the design requires.
193. A CodePipeline execution failed at the deployment stage, and the team must establish why. Which approach is appropriate?
Answer and explanation
Answer: C. Deployment lifecycle event logs record which hook failed on which instance with its output, which is the direct evidence. Execution duration indicates that something took longer without explaining why. Billing metrics report cost. Commit history identifies what changed without showing how it failed.
194. An Amazon ECS service fails to reach its desired task count, and tasks stop shortly after starting. Which source should be examined first?
Answer and explanation
Answer: B. The stopped task reason and the container's exit code and logs state directly why the task ended. Cluster reservation metrics indicate capacity pressure rather than the stop reason. Load balancer request counts describe traffic. Revision history shows what changed without explaining the failure.
195. An event processing workflow must preserve the order of events that relate to the same entity. Which approach is appropriate?
Answer and explanation
Answer: B. Partitioning by entity preserves order within each entity while retaining parallelism across entities. Serial processing preserves order at the cost of throughput. More consumers without partitioning breaks ordering. Sorting after processing cannot undo out-of-order side effects.
196. An event source occasionally delivers the same event twice. Which approach is appropriate?
Answer and explanation
Answer: B. At-least-once delivery makes duplicates inevitable, so the processor must recognise an event it has already handled. A longer retry interval does not prevent duplicate delivery. Disabling retries loses events that genuinely failed. Accepting duplicate side effects is the defect.
197. A remediation must modify a resource's configuration and record what changed for audit. Which approach is appropriate?
Answer and explanation
Answer: A. An Automation runbook records each step with its inputs and outputs, which is a complete audit of the action. A logged summary records what the author chose to log. Manual execution with a ticket note is inconsistent. Config records the resulting state without recording what performed the change or why.
198. A fleet configuration change must be applied progressively so a faulty change does not affect every instance. Which approach is appropriate?
Answer and explanation
Answer: B. A staged rollout with verification bounds the blast radius of a faulty change. Simultaneous application exposes the whole fleet. One surviving instance is a weak signal to proceed on. Monitoring afterwards means the change is already everywhere.
199. A CloudFormation stack update fails and rolls back, and the team must find the cause. Which source should be examined first?
Answer and explanation
Answer: A. The earliest resource failure states the cause, and subsequent events describe the rollback it triggered. The final event reports the rollback outcome. A template diff shows what changed without saying why it failed. CloudTrail records that the update was requested.
200. An Auto Scaling group repeatedly launches and terminates instances. Which cause should be investigated first?
Answer and explanation
Answer: B. A launch and terminate cycle usually means instances are marked unhealthy before they finish starting, which the health check grace period governs. A low maximum caps capacity without cycling. An unsupported instance type would fail to launch at all. Zone count affects distribution rather than cycling.
201. An event-driven workflow must process events in the order they were produced for a given entity. Which mechanism is appropriate?
Answer and explanation
Answer: C. A FIFO queue preserves order within a message group. A standard queue does not guarantee order regardless of consumer count. SNS fan-out does not preserve ordering across subscribers.
202. An event source must deliver to several targets with different processing requirements and independent failure handling. Which pattern is appropriate?
Answer and explanation
Answer: A. Topic to queue fan-out gives each target an independent buffer and retry. A shared queue delivers each message once. Direct invocation couples the producer. A polled database adds latency and schema coupling.
203. A configuration change applied by automation must be reversible if it proves incorrect. Which approach is appropriate?
Answer and explanation
Answer: D. Capturing the prior state makes reversal possible. Waiting for the next deployment leaves the incorrect state in place. A maintenance window limits exposure without enabling reversal. Prior testing reduces risk but does not provide a rollback.
204. An automated response must apply a change to many resources without exceeding an API rate limit. Which approach is appropriate?
Answer and explanation
Answer: D. Controlled batching with backoff respects the limit while completing efficiently. Unbounded parallelism triggers throttling. Sequential processing without rate control may still throttle and is slow. Repeated limit increase requests are not an operational solution.
205. A Lambda function invoked by EventBridge is not executing, and no invocation appears in its metrics. Which cause should be investigated first?
Answer and explanation
Answer: B. No invocation at all points to the rule not matching or lacking invoke permission. Memory, code errors, and log retention all presuppose the function was invoked.
206. An ECS task fails to start with an error indicating the image could not be pulled. Which cause should be investigated first?
Answer and explanation
Answer: C. A pull failure points to registry permissions or network reachability. CPU allocation, application bugs, and desired count all presuppose the image was pulled.
207. An API Gateway endpoint returns 502 errors while the backing Lambda function reports successful invocations. Which cause should be investigated first?
Answer and explanation
Answer: C. A 502 with successful invocations indicates the response format does not match the integration's expectation. Timeouts produce different errors, throttling produces 429, and concurrency limits also produce throttling rather than 502.
208. A configuration change applied in response to an event must be recorded so its effect can be reviewed later. Which approach is appropriate?
Answer and explanation
Answer: A. Automation execution history records inputs, steps, and outcome together. A ticket note is manual, Config records state without the action's context, and a summary records what the author chose to log.
209. An event-driven configuration change must be applied to resources across several Regions. Which approach is appropriate?
Answer and explanation
Answer: D. A multi-Region automation applies the change consistently from one invocation. Per-Region rules multiply maintenance, replication is not a configuration mechanism, and manual application does not scale.