Monitoring, Logging, Analysis, Remediation, and Performance Optimization

Domain 1: Monitoring, Logging, Analysis, Remediation, and Performance Optimization

49 practice questions for Domain 1 of the AWS Certified CloudOps Engineer - Associate (SOA-C03) exam, which makes up 22% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 1: Monitoring, Logging, Analysis, Remediation, and Performance Optimization

22% of scored content · 49 practice questions

1. An operations team must be paged only when a high error rate and elevated latency occur together, because isolated spikes in either metric are common. Which solution meets these requirements?

Answer and explanation

Answer: A. A composite alarm evaluates a rule expression over other alarms, so an AND across the two fires only when both conditions hold. Two independent alarms on one topic page on either condition, which is the noise being eliminated. A longer evaluation period on one metric delays the alert without adding the second condition. A dashboard requires someone to be watching it.

2. Memory utilization for Amazon EC2 instances does not appear in Amazon CloudWatch. Which solution meets these requirements?

Answer and explanation

Answer: C. Memory is a guest operating system metric the hypervisor cannot observe, so the agent must run inside the instance and publish it. Detailed monitoring raises the reporting frequency of existing EC2 metrics but adds no new metric types. Enhanced networking improves network performance. A dashboard visualises metrics that already exist and cannot create them.

3. An application writes structured JSON log entries, and the team must alert when the errorCount field exceeds 10 within five minutes. Which solution meets these requirements?

Answer and explanation

Answer: C. A metric filter parses matching log events into a numeric metric that an alarm can evaluate against a threshold over a period. A subscription filter streams events to a destination but raises no alert. EventBridge does not match on the content of log events in a log group. A saved query must be run by a person and cannot alarm.

4. A metric's normal range varies predictably by time of day, and static thresholds produce false alarms overnight. Which solution meets these requirements?

Answer and explanation

Answer: A. Anomaly detection models the metric's expected range including daily seasonality and alarms on deviation from that band. A lower threshold breaches more often overnight, not less. A longer evaluation period averages away detail without accounting for the pattern. Enabling and disabling alarms on a schedule is brittle and misses holidays and traffic shifts.

5. An organization must query metrics, logs, and traces from 40 member accounts in one place without copying the data. Which solution meets these requirements?

Answer and explanation

Answer: B. Cross-account observability lets a monitoring account query metrics, logs, and traces in linked source accounts natively with no duplication. Copying metrics duplicates data and adds lag and failure modes. Per-account dashboards defeat the single-pane requirement. Exporting logs to a bucket loses the live metric and alarm experience.

6. CloudWatch Logs storage costs have risen steadily because every log group retains data indefinitely. Which solution meets these requirements?

Answer and explanation

Answer: B. Log groups never expire by default, so setting retention is the direct fix and can be enforced across accounts. Deleting whole log groups destroys history and recreates the problem as they refill. Reducing verbosity helps future volume but does nothing about data already retained forever. Exporting and deleting loses the ability to query recent logs in place.

7. A CloudWatch alarm on a metric that reports only when errors occur sits in INSUFFICIENT_DATA during quiet periods and never alarms. Which solution meets these requirements?

Answer and explanation

Answer: B. A metric with no datapoints leaves the alarm in insufficient data, so either the missing data treatment must be set explicitly or the application must publish a zero. A shorter period does not create datapoints. A different threshold cannot evaluate absent data. The namespace does not affect reporting frequency.

8. An operations team must derive an error rate by dividing an error count metric by a request count metric and alarm on the result. Which solution meets these requirements?

Answer and explanation

Answer: B. Metric math combines existing metrics with arithmetic expressions and can be alarmed on directly. A metric filter extracts a metric from log text but does not combine two existing metrics. Config evaluates resource configuration. Athena queries stored data rather than live metrics and cannot alarm.

9. Application logs in CloudWatch Logs must be streamed to Amazon OpenSearch Service in near real time. Which solution meets these requirements?

Answer and explanation

Answer: D. A subscription filter streams matching log events to a destination such as OpenSearch as they arrive. Scheduled exports to S3 are batch and introduce delay. A metric filter produces metrics rather than forwarding events. An alarm notifies rather than streams log content.

10. A team must identify which client IP addresses generate the most throttled requests from log data, without creating a metric for each address. Which solution meets these requirements?

Answer and explanation

Answer: C. Contributor Insights analyses log events and ranks top contributors by a chosen field without creating a metric per value. A dimension per address would create unbounded metrics. Anomaly detection models one metric's expected range. Sampling rules control which requests are traced.

11. An alarm must fire only when a threshold is breached in at least three of the last five evaluation periods, so that brief spikes are tolerated. Which solution meets these requirements?

Answer and explanation

Answer: C. The datapoints to alarm setting implements M out of N evaluation, which tolerates isolated breaches while catching sustained ones. A single long period averages away the detail. Combining an alarm with itself adds no information. A wider anomaly band changes sensitivity rather than the counting rule.

12. The CloudWatch agent must collect system metrics and application logs from a fleet of EC2 instances with one configuration that can be updated centrally. Which solution meets these requirements?

Answer and explanation

Answer: A. Holding the configuration in Parameter Store and applying it through Systems Manager keeps the fleet consistent and lets it change without rebuilding instances. Editing each instance does not scale. A baked configuration cannot change without a new image. User data runs only at launch and leaves later drift uncorrected.

13. An operations team requires a single view of an application's metrics, alarm states, and recent log queries across two Regions. Which solution meets these requirements?

Answer and explanation

Answer: B. CloudWatch dashboards support cross-Region metric widgets alongside alarm and Logs Insights widgets, which produces the single view required. Two dashboards defeat the single-view requirement. A Grafana workspace is a valid tool but rebuilding the widgets is unnecessary work when CloudWatch already spans Regions. Exporting to S3 produces a report rather than a live view.

14. An Amazon EC2 instance fails its instance status check while its system status check passes. Which conclusion should the operations team draw?

Answer and explanation

Answer: D. An instance status check failure points to a software or configuration problem inside the instance, which is the customer's responsibility. A system status check failure indicates underlying AWS infrastructure. Security group and DNS problems do not surface as status check failures.

15. A team runs Amazon EKS workloads already instrumented for Prometheus and must retain and query those metrics without operating a Prometheus server. Which solution meets these requirements?

Answer and explanation

Answer: C. Managed Service for Prometheus ingests Prometheus metrics and supports PromQL with no server to operate, preserving the existing instrumentation. Custom CloudWatch metrics would require reworking the instrumentation and lose PromQL. Logs Insights queries log data rather than time series metrics. A self-managed server reintroduces the operational burden being avoided.

16. A team must build dashboards that combine metrics from Amazon Managed Service for Prometheus and Amazon CloudWatch in Grafana. Which solution meets these requirements with the LEAST operational overhead?

Answer and explanation

Answer: A. Managed Grafana provides a hosted workspace with native data sources for both services and no server to patch or scale. A self-managed server carries that burden. Republishing Prometheus metrics duplicates data and loses Grafana. Exporting to S3 introduces a batch pipeline and a different tool.

17. An application must publish a custom metric from a high-volume workload without a synchronous PutMetricData call for every event. Which solution meets these requirements?

Answer and explanation

Answer: D. Embedded metric format writes structured log entries that CloudWatch parses into metrics asynchronously, avoiding a synchronous call and its latency and throttling. Per-event PutMetricData adds latency and cost even when batched. An approximate built-in metric measures something different. Publishing through SNS adds a component without removing the synchronous call.

18. A specific error pattern in application logs must automatically trigger a service restart across hundreds of EC2 instances. Which solution meets these requirements?

Answer and explanation

Answer: C. A metric filter converts matching log lines into a metric, an alarm fires on it, and Systems Manager executes the remediation across the fleet without inbound access. Connecting to each instance does not scale and is not automatic. An S3 notification requires the logs to land in the bucket first and adds latency. A blind hourly restart disrupts healthy instances and masks the problem.

19. An EC2 instance that fails its system status check must be recovered automatically without human action, preserving its instance ID and attached volumes. Which solution meets these requirements?

Answer and explanation

Answer: D. The recover action migrates the instance to new underlying hardware while preserving its instance ID, private address, and attached volumes. An SNS notification requires a person to act. An Auto Scaling group replaces the instance with a new one, losing identity and instance store data. A weekly reboot is unrelated to failure response.

20. An Amazon EventBridge rule must forward EC2 state-change events to a target that requires more instance detail than the event carries. Which solution meets these requirements?

Answer and explanation

Answer: C. An input transformer reshapes the event before delivery and an enrichment step can add detail the source event lacks. A retry policy governs delivery failures rather than payload content. A duplicate rule delivers the same insufficient event twice. The choice of event bus does not change what the event contains.

21. An operations team must run a predefined sequence of remediation steps against a managed instance, with each step and its outcome recorded for audit. Which solution meets these requirements?

Answer and explanation

Answer: C. Automation runbooks express a multi-step procedure declaratively and record each step's input, output, and status. Run Command executes a single document rather than an orchestrated sequence with per-step history. Interactive steps depend on the operator. A Lambda function would need that orchestration and audit trail built by hand.

22. An EventBridge rule that invokes an AWS Lambda function occasionally loses events when the function is throttled. Which solution meets these requirements?

Answer and explanation

Answer: B. A dead-letter queue and retry policy on the target capture events that cannot be delivered after retries, so nothing is silently lost. Memory and timeout affect execution rather than delivery when the function is throttled. A duplicate rule doubles invocations and worsens throttling.

23. A security finding must trigger an automated containment action within minutes, and the action must be recorded for later review. Which solution meets these requirements?

Answer and explanation

Answer: D. An EventBridge rule matched to the finding invokes a runbook that performs containment within seconds and records every step in execution history. An SNS notification requires human latency. An hourly query is far too slow. A dashboard records nothing and acts on nothing.

24. An Amazon EBS gp2 volume is hitting its baseline throughput, and performance must increase without increasing volume size. Which solution meets these requirements?

Answer and explanation

Answer: C. gp3 decouples IOPS and throughput from capacity, so performance rises without paying for unneeded storage. Increasing gp2 size raises baseline IOPS but is what the requirement excludes. st1 is a throughput-optimised HDD unsuited to random access. Striping adds complexity and still pays for extra capacity.

25. Large objects uploaded to Amazon S3 from an on-premises network frequently fail partway and restart from the beginning. Which solution meets these requirements?

Answer and explanation

Answer: A. Multipart upload splits the object so a failed part is retried alone rather than restarting the whole transfer. A longer timeout does not prevent a mid-transfer failure. Versioning retains successfully written objects rather than protecting an in-flight upload. Transfer Acceleration improves throughput but a failure still restarts the affected part.

26. A shared file system must serve Windows workloads over SMB with Active Directory integration. Which solution meets these requirements?

Answer and explanation

Answer: D. FSx for Windows File Server provides native SMB with Active Directory integration. EFS is an NFS file system and is not the native protocol for Windows clients. Presenting S3 as a drive does not give true SMB semantics. EBS volumes attach to a single instance and are not shared.

27. An Amazon RDS instance exhausts its connection limit under bursty application load while CPU remains moderate. Which solution meets these requirements?

Answer and explanation

Answer: D. RDS Proxy pools and multiplexes connections so bursts do not each open a new one, which addresses exhaustion at moderate CPU. A larger instance raises the ceiling at higher cost without changing the pattern. A read replica cannot accept writes. Storage changes address I/O rather than connection count.

28. Amazon RDS Performance Insights shows that a small number of SQL statements dominate database load. Which action should the team take first?

Answer and explanation

Answer: C. Performance Insights attributes load to individual statements, so addressing the dominant ones removes the cause rather than paying to absorb it. A larger instance masks inefficient queries at ongoing cost. A Multi-AZ standby serves no traffic. Backup retention affects recovery rather than load.

29. EC2 instances tagged as a batch workload are consistently under-utilized, and right-sizing candidates must be identified across the fleet. Which solution meets these requirements?

Answer and explanation

Answer: C. Compute Optimizer analyses utilization history and recommends specific configurations, and tag filtering scopes the review to the workload. Per-instance metrics provide raw data without recommendations. Cost Explorer reports spend rather than utilization. The named Trusted Advisor checks address different resource types.

30. A tightly coupled cluster of EC2 instances requires the lowest possible network latency between nodes. Which solution meets these requirements?

Answer and explanation

Answer: C. A cluster placement group packs instances onto closely connected hardware within one Availability Zone. A spread placement group deliberately separates instances to reduce correlated failure. A partition placement group isolates groups for large distributed workloads. Distributing across zones adds latency.

31. An Amazon EFS file system backing a media rendering workload cannot sustain the throughput the application requires, although storage volume is small. Which solution meets these requirements?

Answer and explanation

Answer: D. Bursting throughput scales with stored volume, so a small file system cannot sustain high throughput; Elastic or provisioned throughput decouples performance from size. Storing extra data to raise the baseline pays for capacity that is not needed. The One Zone class changes durability rather than throughput. Lifecycle transitions reduce cost and can reduce throughput for transitioned files.

32. A data transfer of 40 TB from an on-premises NFS share to Amazon S3 must preserve file metadata and verify integrity, over an adequate network link. Which solution meets these requirements?

Answer and explanation

Answer: A. DataSync transfers between NFS and S3 with metadata preservation and built-in integrity verification. The CLI copies objects without preserving NFS metadata or verifying at task level. Mounting a bucket does not provide true file semantics or metadata fidelity. Transfer Acceleration speeds uploads without preserving metadata or verifying integrity.

33. A CloudWatch alarm on a metric that reports only when an error occurs stays in INSUFFICIENT_DATA between errors. Which configuration is appropriate?

Answer and explanation

Answer: B. Treating missing data as notBreaching resolves the gap for a sparse metric, though publishing zeros is also valid where the application can do it. Evaluation period and datapoints govern how breaches are counted rather than how gaps are treated.

34. An alarm must fire only when a condition persists for fifteen minutes rather than on a single spike. Which configuration is appropriate?

Answer and explanation

Answer: B. Requiring multiple breaching datapoints over consecutive periods filters transient spikes. A higher threshold changes what counts as a breach. Maximum makes the alarm more sensitive to spikes. Time-based action enabling is unrelated.

35. An operator must be alerted when an EC2 instance's memory utilization exceeds a threshold. Which prerequisite is required?

Answer and explanation

Answer: D. Memory is a guest operating system metric requiring the agent. Detailed monitoring increases the frequency of existing EC2 metrics. Enhanced networking improves network performance. A metric filter operates on log content.

36. A metric filter must count occurrences of a specific error code in JSON-formatted application logs. Which filter pattern form is required?

Answer and explanation

Answer: D. JSON logs require a JSON filter pattern selecting on the field. Space-delimited patterns apply to positional text. A literal string match works but is brittle and does not select on the field. Metric filter patterns are not general regular expressions.

37. An EC2 instance fails its instance status check while the underlying hardware is healthy. Which cause should be investigated first?

Answer and explanation

Answer: C. An instance status check failure points inside the instance, such as a corrupted file system or exhausted memory. A system status check failure would indicate host or infrastructure problems. Security groups and zone status produce different symptoms.

38. An Auto Scaling group is not launching instances although the desired capacity was increased. Which cause should be investigated first?

Answer and explanation

Answer: A. The scaling activity history records why a launch failed. Application logs, target health, and alarm configuration all presuppose instances launched.

39. An automated remediation must run when an alarm enters the ALARM state and must not run on recovery. Which configuration is appropriate?

Answer and explanation

Answer: B. Alarm actions are configured per state, so attaching only to ALARM runs the remediation on breach. Attaching to OK runs it on recovery. INSUFFICIENT_DATA fires on missing data. A schedule is unrelated to the alarm.

40. Logs from an EC2 instance are not appearing in CloudWatch Logs. Which cause should be investigated first?

Answer and explanation

Answer: C. Missing logs usually trace to agent permissions or configuration. Retention affects how long logs persist once delivered. Instance type does not govern logging. Encryption does not prevent delivery when the agent has key permission.

41. A CloudWatch dashboard shows no data for a metric that exists in the console. Which cause should be investigated first?

Answer and explanation

Answer: D. A Region or dimension mismatch is the common cause of an empty widget for a metric that exists. Widget count, namespace reservation, and creation date do not cause this.

42. An application's Amazon S3 read throughput is limited although the bucket is not near any documented limit. Which cause should be investigated first?

Answer and explanation

Answer: B. S3 throughput scales with parallel requests, so a serial client limits itself. Storage class affects retrieval characteristics for archival classes. Versioning adds stored objects. Account ownership does not throttle throughput.

43. An EC2 instance's disk throughput is lower than expected immediately after a volume is created from a snapshot. Which cause explains this?

Answer and explanation

Answer: D. A volume restored from a snapshot retrieves blocks lazily on first access, which slows early reads. Volume type would cap throughput permanently. A cross-zone volume cannot attach. An unformatted volume would not mount.

44. A Lambda function used in an operational workflow runs slower than its allocated timeout suggests it should. Which adjustment should be evaluated first?

Answer and explanation

Answer: C. Lambda allocates CPU in proportion to memory, so CPU-bound work runs faster with more memory and may cost less overall. A longer timeout permits slowness. Less memory slows it further. Concurrency governs parallel executions.

45. An EBS volume's performance is below its provisioned IOPS, and the attached instance's network is saturated. Which cause explains this?

Answer and explanation

Answer: B. An instance's EBS bandwidth ceiling can cap throughput below the volume's provisioning. Volume type would prevent provisioning IOPS at all. A cross-zone volume cannot attach. Snapshot initialization affects first-access latency rather than sustained throughput.

46. An application's read performance against Amazon S3 must be improved for a workload issuing many parallel GET requests to one prefix. Which approach is appropriate?

Answer and explanation

Answer: D. Spreading requests across prefixes increases achievable request rate. Storage class affects retrieval characteristics rather than request rate. Versioning adds stored objects. Fewer parallel requests reduces throughput.

47. An RDS database's query performance has degraded, and Performance Insights shows waits on a specific query. Which action is appropriate?

Answer and explanation

Answer: D. Performance Insights identifies the query, and its plan and indexes explain the waits. A larger instance runs an inefficient query faster at higher cost. Multi-AZ provides availability. Backup retention is unrelated.

48. An operator must reduce the time an Auto Scaling group takes to bring new instances into service. Which approach is appropriate?

Answer and explanation

Answer: D. Startup time is dominated by what happens after launch, so pre-baking removes it. Maximum size governs capacity. Cooldown governs the interval between scaling actions. Health check type governs what is evaluated.

49. An alarm must notify a team only during the hours they are on duty. Which approach is appropriate?

Answer and explanation

Answer: A. An on-call schedule in the notification path routes to whoever is on duty. Per-shift alarms multiply configuration, disabling loses detection, and evaluation period does not control routing.