Monitoring and Logging

Domain 4: Monitoring and Logging

61 practice questions for Domain 4 of the AWS Certified DevOps Engineer - Professional (DOP-C02) exam, which makes up 15% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 4: Monitoring and Logging

15% of scored content · 61 practice questions

127. Application log lines containing a specific error pattern must become a metric that an alarm can evaluate. Which solution meets these requirements?

Answer and explanation

Answer: B. A metric filter converts matching log events into a numeric metric an alarm can evaluate. A subscription filter streams events without raising an alert. A saved query must be run by a person. EventBridge does not match on log event content in a log group.

128. CloudWatch metrics must be delivered continuously to a third-party analytics platform in near real time. Which solution meets these requirements?

Answer and explanation

Answer: A. Metric streams deliver metrics continuously through Firehose to a destination, which is the purpose-built path for near real time export. Scheduled polling adds latency and API cost. A daily export is batch. A dashboard is a visualization rather than a data feed.

129. Memory and disk utilization must be collected from EC2 instances into CloudWatch. Which solution meets these requirements?

Answer and explanation

Answer: C. Memory and disk are guest operating system metrics the hypervisor cannot observe, so the agent must run inside the instance. Detailed monitoring raises the frequency of existing EC2 metrics without adding new types. Enhanced networking improves network performance. A dashboard displays metrics that already exist.

130. Log data must be retained for ninety days in CloudWatch Logs and for seven years at lower cost. Which solution meets these requirements?

Answer and explanation

Answer: C. CloudWatch Logs is priced for active querying, so a short retention there with long-term copies in S3 under a lifecycle policy meets both requirements economically. Seven years in CloudWatch Logs is expensive. Accepting loss fails the retention requirement. Manual deletion is unreliable and still pays CloudWatch rates.

131. Log events must be processed in near real time by a Lambda function as they arrive. Which solution meets these requirements?

Answer and explanation

Answer: B. A subscription filter streams matching log events to a destination as they arrive. Scheduled querying adds latency and repeats work. Hourly export is batch. An alarm fires on a threshold rather than delivering individual events.

132. Log data written to CloudWatch Logs must be encrypted with a key the security team controls. Which solution meets these requirements?

Answer and explanation

Answer: C. Associating a customer managed key with the log group encrypts the data with a key whose policy the security team authors, and the service principal must be permitted in that policy. Log groups are encrypted by default with a service key rather than a customer key. Exporting and deleting removes the ability to query recent logs. Access restriction is not encryption.

133. An engineer must find every occurrence of a correlation identifier across several log groups over a two-day window. Which solution meets these requirements?

Answer and explanation

Answer: B. Logs Insights queries across multiple log groups and time ranges and returns matching events in seconds. Downloading and searching locally is slow and not repeatable. A metric filter counts occurrences without returning the events. Manual searching of exported objects is slower still.

134. An application must publish a custom business metric at high volume without a synchronous API call per event. Which solution meets these requirements?

Answer and explanation

Answer: A. Embedded metric format writes structured log entries that CloudWatch parses into metrics asynchronously, avoiding a synchronous call and its latency and throttling. Per-event PutMetricData adds latency and cost. An approximate built-in metric measures something different. Publishing through SNS adds a component without removing the synchronous call.

135. An on-call rotation receives many alerts that require no action, and the team must reduce them without losing genuine signals. Which approach is appropriate?

Answer and explanation

Answer: D. Triaging actioned against non-actioned alerts and adjusting each individually preserves the signals that matter while removing the noise. A shared mailbox hides every alert including the urgent ones. Deleting the most frequent alarms may remove the most important. A blanket period increase delays genuine alerts as well.

136. Alarms with consistent thresholds and actions must be created automatically for every new service deployed by any team. Which solution meets these requirements?

Answer and explanation

Answer: A. Defining alarms alongside the service in the same template means monitoring is created, versioned, and destroyed with the resource it watches. A runbook depends on teams remembering. A weekly sweep leaves new services unmonitored for up to a week. Manual creation on request scales poorly and drifts.

137. A request traversing API Gateway, Lambda, and DynamoDB must be traced end to end. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, D. Active tracing produces segments for the managed services and instrumenting the SDK captures subsegments for the DynamoDB calls, which together complete the trace. Flow Logs record network metadata and cannot see inside managed service calls. Retention governs log persistence. CloudTrail data events audit API activity rather than tracing a request path.

138. Security findings from several AWS services must be aggregated and evaluated against a standard across an organization. Which solution meets these requirements?

Answer and explanation

Answer: B. Security Hub ingests findings in a common format, runs automated standards checks, and aggregates across accounts. Detective investigates a single finding. A Config aggregator collects configuration compliance rather than security findings. Exporting to S3 requires building the normalization and standards logic.

139. Configuration changes across the estate must be evaluated continuously against defined rules, with non-compliant resources reported. Which solution meets these requirements?

Answer and explanation

Answer: A. Config records resource configuration over time and evaluates it against rules, reporting compliance per resource. CloudTrail records the API calls that caused changes rather than evaluating the resulting state. Inspector assesses software vulnerabilities. Flow Logs capture network traffic.

140. A specific error pattern in application logs must automatically restart a service across hundreds of instances. Which solution meets these requirements?

Answer and explanation

Answer: D. A metric filter converts matching log lines into a metric, an alarm fires on it, and Systems Manager executes the remediation across the fleet without inbound access. Connecting to each instance does not scale. An S3 notification requires the logs to land first and adds latency. A blind hourly restart disrupts healthy instances.

141. An EC2 instance failing its system status check must be recovered automatically while preserving its instance identity and attached volumes. Which solution meets these requirements?

Answer and explanation

Answer: D. The recover action migrates the instance to new underlying hardware while preserving its instance ID, private address, and attached volumes. A notification requires a person to act. An Auto Scaling group replaces the instance with a new one, losing identity. A weekly reboot is unrelated to failure response.

142. An event processing workflow must handle bursts far larger than its downstream system can absorb, without losing events. Which solution meets these requirements?

Answer and explanation

Answer: C. A queue decouples arrival rate from processing rate, absorbing the burst durably while the consumer works at a sustainable pace. Raising concurrency passes the burst straight through to the downstream system. Retrying keeps pressure on a system already overwhelmed. Sizing the downstream for the largest burst pays for that capacity continuously.

143. An AWS Config rule detects a security group allowing unrestricted inbound access, and the rule must revert the change automatically. Which solution meets these requirements?

Answer and explanation

Answer: D. An attached remediation action revokes the offending rule automatically when the resource is found non-compliant. A notification requires a person to act. Manual review leaves the exposure open overnight. An SCP denying modification would block legitimate security group changes as well.

144. An Application Load Balancer must stop sending traffic to targets whose application has crashed while the operating system remains reachable. Which solution meets these requirements?

Answer and explanation

Answer: D. A health check against an application endpoint detects a crashed process that a reachability check would not. System status checks observe instance and hypervisor health. A shorter interval checks the same inadequate signal more often. Deregistration delay governs connection draining.

145. An object landing in Amazon S3 must trigger a processing workflow without polling. Which solution meets these requirements?

Answer and explanation

Answer: B. Event notifications and EventBridge rules invoke a target as objects arrive, which is the event-driven pattern required. A listing function polls by definition. A scheduled association introduces delay and repeated listing. The object count metric is reported at daily granularity.

146. An Amazon DynamoDB table's provisioned capacity must follow demand without manual adjustment. Which solution meets these requirements?

Answer and explanation

Answer: C. Auto scaling adjusts provisioned capacity toward a target utilization as demand changes. Peak provisioning pays for idle capacity. Average provisioning throttles at peak. Manual adjustment on an alarm depends on someone responding.

147. Log data from many accounts must be centralized for correlation while each account retains local access. Which approach is appropriate?

Answer and explanation

Answer: D. A subscription or delivery destination streams logs centrally in near real time while local groups remain for account-level use. Weekly export is stale. Cross-account read access requires querying each account separately rather than correlating centrally. Bucket replication carries raw objects without the streaming path.

148. Metrics must be published from an application at a resolution finer than one minute. Which configuration is required?

Answer and explanation

Answer: D. High-resolution custom metrics support sub-minute granularity, which standard and detailed monitoring do not. Detailed monitoring provides one-minute resolution for EC2 metrics. A longer collection interval is coarser. Metric streams export metrics without changing their resolution.

149. An architect must control the cost of log storage without losing the ability to investigate incidents from the past year. Which approach is appropriate?

Answer and explanation

Answer: D. A short queryable retention with long-term copies in cheaper storage preserves investigation capability at a fraction of the cost. A year in the queryable store is expensive. Reduced verbosity loses detail incidents need. Metrics do not substitute for log detail.

150. An application's logs must be searchable by a structured field rather than by free-text matching. Which approach is appropriate?

Answer and explanation

Answer: A. Structured JSON allows queries to filter and aggregate on named fields rather than matching text. Pattern matching on formatted text is brittle. A consistent prefix helps filtering without providing fields. Retention governs how much data exists rather than how it is queried.

151. Logs must be encrypted with a key the security team controls and remain queryable by the operations team. Which approach is appropriate?

Answer and explanation

Answer: C. A customer managed key gives the security team an author-controlled policy while readers with decrypt permission still query normally. Client-side encryption makes the content unqueryable. Unencrypted storage fails the requirement. The service default key exposes no key policy for the security team to control.

152. An alarm must fire when a metric that is normally reported stops being reported at all. Which configuration is required?

Answer and explanation

Answer: B. Treating missing data as breaching is what makes the absence of a metric raise the alarm. Treating it as not breaching leaves the alarm silent. Evaluation period and threshold govern how reported values are assessed rather than their absence.

153. An investigation must establish the sequence of events across several services during an incident. Which approach is appropriate?

Answer and explanation

Answer: B. Correlating on a shared request identifier across aggregated logs reconstructs the sequence. Per-service dashboards show metrics rather than ordered events. CloudTrail records API calls rather than application behaviour. Scaling history covers one subsystem.

154. An architect must detect a gradual increase in application latency that stays within the existing alarm threshold. Which approach is appropriate?

Answer and explanation

Answer: D. An anomaly band detects drift away from the established pattern even when the absolute value remains within a static threshold. Lowering the threshold converts a gradual trend into frequent alerts at the new level. A longer period averages the trend away. Daily review depends on someone noticing.

155. A trace must show which downstream call caused a request's latency to exceed its target. Which approach is appropriate?

Answer and explanation

Answer: D. Subsegments attribute time to individual downstream calls, which is what identifies the slow one. Total duration shows the request was slow. Call count describes volume. A higher sampling rate captures more traces without adding detail within each.

156. An automated remediation must not run when the condition is already being handled by an in-progress deployment. Which approach is appropriate?

Answer and explanation

Answer: A. Checking deployment state inside the remediation makes the decision at the moment of action and needs no external coordination. Disabling and re-enabling alarms is error-prone and leaves genuine failures undetected. A longer evaluation period delays all alerts. Allowing a conflict risks two processes changing the same resources.

157. An EventBridge rule must invoke a target in a different account. Which configuration is required?

Answer and explanation

Answer: C. Cross-account delivery requires an event bus in the target account whose resource policy permits the source, with the rule targeting that bus. A target ARN alone lacks the permission. IAM user credentials are not how event routing authorises. S3 replication is not an event path.

158. Alarms must be created automatically for every new resource a team deploys, without the team remembering to add them. Which approach is appropriate?

Answer and explanation

Answer: C. Defining alarms alongside the resource means monitoring is created, versioned, and destroyed with what it watches. A weekly sweep leaves new resources unmonitored in the interim. A standard relies on compliance. Central creation on request depends on the request being made.

159. A health check must detect that an application is serving stale data even though it responds successfully. Which approach is appropriate?

Answer and explanation

Answer: C. A health endpoint that asserts a freshness condition reports a state a simple availability check cannot see. A shorter interval checks the same inadequate signal more often. A 200 response and a running process both indicate availability rather than correctness.

160. A CloudWatch agent must collect logs from a fleet with a consistent configuration. Which approach is appropriate?

Answer and explanation

Answer: A. A configuration in Parameter Store is fetched by the agent, so one change reaches the fleet. Per-instance editing does not scale. Baking into the image requires a rebuild per change. Console configuration is per instance.

161. Logs must be delivered to a destination in another account in near real time. Which mechanism is appropriate?

Answer and explanation

Answer: C. A subscription filter with a cross-account destination streams events as they arrive. Export tasks are batch. Polling adds latency. Manual transfer is not near real time.

162. A metric must be available at one-second resolution for a short-lived spike. Which configuration is required?

Answer and explanation

Answer: B. High-resolution custom metrics support one-second granularity. Detailed monitoring provides one-minute resolution. A one-minute interval is coarser. A metric filter produces standard-resolution metrics.

163. Trace data must be retained beyond the service's default retention for later analysis. Which approach is appropriate?

Answer and explanation

Answer: A. Exporting to durable storage preserves traces beyond the service's retention. Sampling affects how many are captured rather than how long they persist. An alarm monitors. Fewer instrumented services reduce coverage.

164. An alarm must evaluate a percentage derived from two metrics. Which capability is appropriate?

Answer and explanation

Answer: C. Metric math computes a derived value an alarm can evaluate. Separate alarms evaluate each metric independently. A composite alarm combines alarm states rather than computing a ratio. A dashboard displays without alarming.

165. A dashboard must show the same view of a workload across several accounts. Which capability is appropriate?

Answer and explanation

Answer: B. Cross-account observability presents source account telemetry in a monitoring account. Per-account dashboards fragment the view. Spreadsheet export is manual and stale. Broad account access is disproportionate.

166. An investigation must determine whether a configuration change preceded an incident. Which source is appropriate?

Answer and explanation

Answer: B. Config records the configuration state over time and shows what changed and when. Error logs, scaling history, and access logs describe behaviour rather than configuration changes.

167. An EventBridge rule must invoke a target only for events matching a specific field value. Which configuration is appropriate?

Answer and explanation

Answer: B. An event pattern filters before invocation, which avoids unnecessary invocations and cost. Filtering inside the target pays for every invocation. A separate bus requires the producer to route. Reducing publishing loses events.

168. A scheduled automation must run at a precise time with timezone awareness and no drift. Which capability is appropriate?

Answer and explanation

Answer: C. EventBridge Scheduler supports timezone-aware schedules without managing infrastructure. An instance cron job requires an instance. A sleeping function wastes duration and is unreliable. Alarms evaluate metrics rather than schedule.

169. An operational event must create a tracked item that responders work through to resolution. Which capability is appropriate?

Answer and explanation

Answer: D. Incident Manager creates a tracked incident with engagement and a timeline. A notification alerts without tracking. An alarm state records a condition. A log entry records without workflow.

170. An automated remediation must be prevented from running repeatedly for the same ongoing condition. Which approach is appropriate?

Answer and explanation

Answer: D. An in-progress marker prevents concurrent remediation of the same resource. A shorter evaluation period fires more often. A longer timeout extends each run. Disabling the alarm loses detection of future occurrences.

171. A CloudWatch agent must collect a custom application metric alongside system metrics. Which approach is appropriate?

Answer and explanation

Answer: A. The agent collects StatsD and collectd input alongside system metrics with one configuration. A separate API call adds latency, a metric filter works but adds a processing step, and a second agent duplicates the deployment.

172. Log data must be retained in a queryable form for compliance while controlling cost. Which approach is appropriate?

Answer and explanation

Answer: B. A short queryable retention with long-term copies in S3 balances investigation against cost. Full retention in CloudWatch Logs is expensive, S3 only loses fast investigation, and reduced verbosity loses detail.

173. An organization must collect logs from on-premises servers into CloudWatch Logs. Which approach is appropriate?

Answer and explanation

Answer: D. The agent with role-based credentials delivers logs continuously from on-premises servers. Nightly copies are batch, email is not a log pipeline, and remote querying requires access and is not centralized.

174. Metrics must be retained beyond CloudWatch's standard retention for long-term trend analysis. Which approach is appropriate?

Answer and explanation

Answer: A. Metric streams deliver to durable storage for indefinite retention. CloudWatch metric retention is fixed by resolution, dashboards display within retention, and manual export does not scale.

175. An organization must ensure log data cannot be modified after delivery. Which approach is appropriate?

Answer and explanation

Answer: B. Compliance mode prevents modification and deletion by any principal. Encryption protects confidentiality, write restrictions can be changed, and versioning preserves history without preventing changes.

176. Trace sampling must capture more traces for a specific high-value request path. Which approach is appropriate?

Answer and explanation

Answer: B. A targeted sampling rule raises coverage where it matters without the cost of tracing everything. A higher default and no sampling both increase cost broadly, and retention governs how long traces persist.

177. An organization must aggregate metrics from several accounts into one monitoring view. Which approach is appropriate?

Answer and explanation

Answer: B. Cross-account observability presents source account telemetry in a monitoring account. Per-account dashboards fragment the view, spreadsheets are manual, and console access requires switching accounts.

178. An alarm must be suppressed while a related upstream alarm is already firing, to avoid duplicate paging. Which approach is appropriate?

Answer and explanation

Answer: B. A composite alarm expresses the dependency and suppresses the derived alert. Deleting loses detection, a shared topic still pages twice, and a longer period delays rather than suppresses.

179. An operator must find which requests contributed to a latency increase across a distributed application. Which approach is appropriate?

Answer and explanation

Answer: C. Filtering traces by latency and examining segments identifies both the requests and the slow stage. Aggregate metrics show the increase without attribution, log volume describes activity, and error rate covers failures rather than slow successes.

180. An audit must determine whether a configuration was compliant at a specific date in the past. Which source is appropriate?

Answer and explanation

Answer: C. Config records state over time and supports point-in-time queries. Current configuration reflects now, CloudTrail records the calls that caused changes, and tags are metadata.

181. An organization must detect when a security finding remains unresolved beyond an agreed period. Which approach is appropriate?

Answer and explanation

Answer: B. Tracking age against a threshold surfaces findings that have stalled. Alerting on every new finding does not measure resolution time, weekly review is manual, and archiving hides the problem.

182. An automated response must run in a different account from the one where the event occurred. Which approach is appropriate?

Answer and explanation

Answer: D. Cross-account event forwarding with a role for the action keeps the automation centralized. Direct invocation reverses the trust direction, per-account replication multiplies maintenance, and polling adds latency.

183. An operational event must trigger different actions depending on the affected resource's environment tag. Which approach is appropriate?

Answer and explanation

Answer: D. A branching workflow selects the action from the resource's attributes. Identical rules per environment cannot distinguish by tag alone, a uniform action ignores the difference, and operator decisions remove the automation.

184. An automated remediation must be tested before it is enabled in production. Which approach is appropriate?

Answer and explanation

Answer: A. Testing against a non-production resource verifies behaviour safely. Enabling in production, with or without a quiet period, experiments on live resources, and document review does not execute the logic.

185. A monitoring system must detect that a scheduled job did not run at all. Which approach is appropriate?

Answer and explanation

Answer: A. A heartbeat with a missing-data alarm detects absence, which an error-based alarm cannot. Manual review is slow, and a timeout governs a running job.

186. An organization must ensure an automated remediation does not act on a resource an operator is already investigating. Which approach is appropriate?

Answer and explanation

Answer: B. A suppression tag checked by the automation scopes the exclusion to the specific resource. Disabling the remediation removes it for everything, off-hours execution is arbitrary, and post-hoc notification is too late.

187. An operational dashboard must show the health of a service rather than the health of its individual components. Which approach is appropriate?

Answer and explanation

Answer: B. A service-level indicator built from user-facing metrics represents service health. Component metrics, instance counts, and alarm counts describe the system rather than the experience.