Reliability and Business Continuity

Domain 2: Reliability and Business Continuity

32 practice questions for Domain 2 of the AWS Certified CloudOps Engineer - Associate (SOA-C03) exam, which makes up 22% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 2: Reliability and Business Continuity

22% of scored content · 32 practice questions

50. An Amazon DynamoDB table configured with provisioned capacity and auto scaling is still throttling during short, sharp traffic spikes. Which solution meets these requirements?

Answer and explanation

Answer: D. Provisioned auto scaling reacts over minutes, so a spike shorter than the adjustment interval throttles regardless of the policy; on-demand mode removes that lag. A higher target utilization delays scaling further. A lower minimum reduces baseline capacity and worsens the spike. An index adds write cost and does not increase table throughput.

51. An Auto Scaling group must maintain capacity across three Availability Zones even when one zone lacks capacity for the requested instance type. Which solution meets these requirements?

Answer and explanation

Answer: A. A mixed instances policy lets the group launch any of several compatible types, so a shortfall in one type does not prevent scaling. A larger maximum does not help when the single type is unavailable. Health check frequency and scheduled scaling do not address capacity availability.

52. A read-heavy application repeatedly issues identical queries against Amazon RDS, and database CPU is high. Which solution most reduces load on the database?

Answer and explanation

Answer: B. Caching removes the repeated work entirely, which is the largest available reduction when the same queries recur. A read replica moves the work to another instance that still executes it. A Multi-AZ standby serves no read traffic. A larger instance executes the same queries faster at higher cost.

53. An Amazon DynamoDB table serves traffic that swings unpredictably between near zero and very high with no notice. Which solution meets these requirements?

Answer and explanation

Answer: C. On-demand billing charges per request and absorbs sudden swings without capacity planning. Provisioning for the peak pays for idle capacity most of the time. Provisioning for the average throttles during spikes. Provisioned capacity without auto scaling cannot respond to change at all.

54. A globally distributed audience downloads large static assets from a single Region, and origin load must be reduced. Which solution meets these requirements?

Answer and explanation

Answer: B. CloudFront caches assets at edge locations near viewers, which cuts latency and absorbs requests that would otherwise reach the origin. Larger or additional instances address compute capacity rather than origin request volume or distance. Enhanced networking improves throughput within the VPC.

55. An Amazon Aurora cluster's read load varies unpredictably between the demand of one replica and eight. Which solution meets these requirements at the lowest cost?

Answer and explanation

Answer: C. Aurora Auto Scaling adds and removes replicas against a target such as average CPU, matching capacity to unpredictable read load. Eight fixed replicas pay for peak continuously. One large replica cannot absorb eightfold variation. Sending reads to the writer removes the separation that protects write performance.

56. An Auto Scaling group configured with ELB health checks replaces instances before the application finishes starting, causing a replacement loop. Which solution meets these requirements?

Answer and explanation

Answer: B. The grace period tells Auto Scaling to ignore health checks while a newly launched instance boots and the application starts. A longer check interval delays detection without covering the startup window. Switching to EC2 health checks stops the group noticing genuine application failures later. More capacity adds instances that hit the same race.

57. An Amazon Route 53 failover record is not failing over even though the primary application is returning HTTP 500 responses. Which cause should the operations team investigate first?

Answer and explanation

Answer: A. A TCP-level health check confirms the port is open and passes while the application returns errors, so the check must inspect the HTTP status or a specific response body. Time to live affects how quickly resolvers observe a change rather than whether the health check evaluates. Weights apply to weighted routing, not failover. Both records being in the hosted zone is a prerequisite that would prevent resolution entirely if wrong.

58. Traffic must be sent to a secondary Region only when the primary endpoint stops responding. Which solution meets these requirements?

Answer and explanation

Answer: A. Failover routing pairs a primary and secondary record with a health check and resolves to the secondary only when the primary is unhealthy. Weighted routing sends a share of traffic to the secondary at all times. Latency routing chooses by proximity rather than health. Simple routing performs no health evaluation.

59. An Amazon RDS database must remain writable when an Availability Zone fails, with minimal data loss and no operator action. Which solution meets these requirements?

Answer and explanation

Answer: A. Multi-AZ maintains a synchronous standby and redirects the endpoint automatically. Promoting a read replica is manual and loses asynchronously replicated writes. Point-in-time restore and hourly snapshots are recovery paths measured in minutes to hours with operator involvement.

60. Users report intermittent failures reaching an application behind an Application Load Balancer, and only some requests fail. Which cause should the team investigate first?

Answer and explanation

Answer: D. Partial failure across requests points to a subset of targets being unhealthy, because the load balancer distributes requests and only those routed to a failing target break. A deleted hosted zone would cause total DNS failure. A missing internet gateway or a security group blocking all inbound traffic would fail every request.

61. An operations team must be notified of AWS-side events affecting its own resources, such as scheduled EC2 maintenance or a Regional service disruption. Which solution meets these requirements?

Answer and explanation

Answer: B. AWS Health publishes account-specific events, including scheduled maintenance affecting the account's own resources, to EventBridge for routing. A metric alarm reacts to symptoms without identifying an AWS-side cause. A public status page carries no account-specific detail. Config records configuration changes rather than AWS operational events.

62. Instances in an Auto Scaling group are terminated during scale-in while still writing log files that must be preserved. Which solution meets these requirements?

Answer and explanation

Answer: A. A terminating lifecycle hook holds the instance in a wait state so a document can flush and upload the logs before termination completes. Termination protection does not apply to Auto Scaling scale-in. A larger minimum reduces how often scale-in happens without handling it safely. Hourly uploads still lose up to an hour of logs at termination.

63. A recovery plan across two Regions must be tested and its readiness continuously verified. Which solution meets these requirements?

Answer and explanation

Answer: A. Application Recovery Controller provides readiness checks and routing controls that continuously verify the standby is prepared to take traffic. AWS Backup copies recovery points without verifying readiness to serve. A canary tests an endpoint from outside without confirming capacity or quota readiness. Config confirms resources exist rather than that they are ready.

64. Daily backups of Amazon EBS volumes, Amazon RDS databases, and Amazon EFS file systems must be governed by one policy with cross-Region copies. Which solution meets these requirements?

Answer and explanation

Answer: A. AWS Backup applies a single policy across many supported services and its copy rules replicate recovery points to another Region. Data Lifecycle Manager covers EBS snapshots and AMIs but not RDS or EFS. A custom function recreates that capability with more failure modes. Monthly manual snapshots cannot produce daily backups.

65. Backups must be protected so that no principal, including an account administrator, can delete them before their retention period expires. Which solution meets these requirements?

Answer and explanation

Answer: C. Vault Lock in compliance mode prevents deletion before retention expires and cannot be disabled once locked. Governance mode can be bypassed by principals holding the bypass permission. An IAM policy can be modified by an administrator. A separate account improves isolation but its own administrator can still delete.

66. A quarterly restore test must verify that an Amazon RDS backup is usable, without disrupting production. Which solution meets these requirements?

Answer and explanation

Answer: D. Restoring to a new instance exercises the whole recovery path and lets the data be validated with no effect on production. Failing over tests Multi-AZ availability but never reads the backup. Restoring over production is destructive. Confirming a snapshot exists proves it was created, not that it can be restored.

67. An Amazon RDS major version upgrade must be tested against production-like data with the ability to switch over quickly and roll back. Which solution meets these requirements?

Answer and explanation

Answer: B. Blue/Green Deployments create a synchronised staging environment that can be switched over quickly with the previous environment retained for rollback. Promoting a replica is one-way. A restored snapshot is a point-in-time copy without continuous synchronisation. Upgrading in place with snapshot rollback means a long outage if it fails.

68. An Amazon DynamoDB table with point-in-time recovery enabled must be rolled back to its state before an accidental bulk delete two hours ago. Which solution meets these requirements?

Answer and explanation

Answer: D. A point-in-time restore always creates a new table rather than overwriting the source, so the application must be repointed or the tables swapped as part of the recovery. There is no in-place restore. Streams retain records for 24 hours and replaying deletions in reverse is not a supported recovery path. Backing up the current damaged state preserves the damage.

69. An Amazon EFS file system must be replicated to a second Region for disaster recovery with a recovery point objective of minutes. Which solution meets these requirements?

Answer and explanation

Answer: A. EFS replication continuously replicates a file system to another Region with a recovery point objective measured in minutes. A daily DataSync task gives a recovery point up to a day old. Copying to S3 changes the storage model and adds a conversion step. Weekly backup copies give a recovery point up to a week old.

70. Snapshots of production EBS volumes must be copied to a second Region and encrypted with a key managed in that Region. Which solution meets these requirements?

Answer and explanation

Answer: A. Copying a snapshot to another Region allows a destination KMS key to be specified for re-encryption. EBS volumes are not replicated across Regions directly. A lifecycle policy without a copy rule creates snapshots only in the source Region. EBS snapshots are not customer-visible S3 objects.

71. A disaster recovery plan must be validated rather than assumed to work. Which practice meets these requirements?

Answer and explanation

Answer: C. Only exercising the failover proves that the standby can serve traffic and that the procedure works end to end. Reviewing documentation, confirming backup completion, and checking resource existence all verify preconditions rather than the outcome.

72. An Auto Scaling group's instances are terminated and relaunched repeatedly shortly after launch. Which configuration should be adjusted?

Answer and explanation

Answer: D. A grace period that is too short marks starting instances unhealthy. Maximum size and desired capacity govern how many instances run. Switching to EC2 health checks would mask an application problem rather than fix the timing.

73. An Auto Scaling group must terminate the oldest instances first during a scale-in event. Which configuration is appropriate?

Answer and explanation

Answer: B. A termination policy determines which instances are removed. Scale-in protection can achieve a similar effect but requires managing protection per instance. Cooldown governs timing. Manual capacity does not select instances.

74. A DynamoDB table with auto scaling still throttles during sharp traffic spikes. Which cause explains this?

Answer and explanation

Answer: A. Provisioned auto scaling is reactive and lags a sharp spike, which is why on-demand mode suits unpredictable traffic. Auto scaling applies to both read and write capacity, is for provisioned mode, and adjusts far more often than daily.

75. An Amazon RDS Multi-AZ deployment has failed over, and the application cannot reconnect. Which cause should be investigated first?

Answer and explanation

Answer: A. An application caching the resolved address rather than re-resolving after failover cannot reconnect. Multi-AZ standbys are within the Region. Backups and storage capacity are separate concerns.

76. An Application Load Balancer reports targets as unhealthy although the application responds correctly when tested directly. Which cause should be investigated first?

Answer and explanation

Answer: D. A security group blocking the load balancer's health check traffic is the common cause when direct testing succeeds. Scheme, protocol version, and zone count produce different symptoms, although protocol version mismatch is also worth checking.

77. A Route 53 health check reports a healthy endpoint as unhealthy. Which cause should be investigated first?

Answer and explanation

Answer: C. Health checkers originate from published address ranges that must be permitted. Zone type, time to live, and record type affect resolution rather than health checking.

78. An AWS Backup restore of an EBS volume must be attached to an instance in a different Availability Zone. Which step is required?

Answer and explanation

Answer: D. EBS volumes are zonal, so the restore must target the zone where the instance runs. Cross-zone attachment is not possible, a volume's zone cannot be modified, and a snapshot is not attachable.

79. A point-in-time restore of an RDS database creates a new instance rather than modifying the existing one. Which consequence must be planned for?

Answer and explanation

Answer: B. A point-in-time restore creates a new instance with a new endpoint, so cutover requires renaming or repointing. The original is retained, is not modified, and keeps its own endpoint.

80. An AWS Backup vault must be prevented from having its recovery points deleted before their retention expires. Which configuration is appropriate?

Answer and explanation

Answer: C. Vault Lock enforces retention against deletion. An IAM policy can be changed by an administrator. Encryption protects confidentiality. A second copy adds redundancy without preventing deletion of either.

81. An AWS Backup plan must protect resources across several accounts in an organization. Which configuration is required?

Answer and explanation

Answer: A. An organization backup policy applies the plan across member accounts including new ones. Per-account plans drift and miss new accounts. Manual copying does not scale. Sharing a vault does not cause other accounts to be backed up.