Design Resilient Architectures

Domain 2: Design Resilient Architectures

46 practice questions for Domain 2 of the AWS Certified Solutions Architect - Associate (SAA-C03) exam, which makes up 26% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 2: Design Resilient Architectures

26% of scored content · 46 practice questions

54. A payment service must not lose messages when its consumer fails, and failed messages must be reprocessed later without blocking new work. Which solution meets these requirements?

Answer and explanation

Answer: C. A redrive policy moves a repeatedly failing message aside into a dead-letter queue where it can be inspected and replayed, which unblocks the main queue. Raising the visibility timeout only delays redelivery of the same poison message and never removes it. SNS delivers notifications without retaining messages for later reprocessing. A one-hour Kinesis retention risks permanent loss during a longer outage.

55. Messages in an Amazon SQS queue are occasionally processed twice because a consumer sometimes takes longer than expected to finish. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, D. A message becomes visible again when the visibility timeout expires before deletion, so extending it past the longest realistic processing time stops the redelivery, and an idempotent consumer makes any remaining duplicate harmless. Decreasing the timeout causes more duplicates. Retention governs how long unprocessed messages survive. A higher receive count allows more redeliveries.

56. An application must continue accepting orders when a downstream fulfilment service becomes unavailable, and must process the accepted orders once that service recovers. Which solution meets these requirements?

Answer and explanation

Answer: C. A queue between the two decouples their availability entirely, so orders are durably accepted while the consumer is down and drained when it recovers. A synchronous call with a short timeout still fails when the service is fully down. Retrying until success blocks the caller and exhausts connections. Sharing an Auto Scaling group couples their failure domains more tightly.

57. An AWS Lambda function invoked asynchronously by Amazon S3 events must not discard events that fail after all retry attempts. Which solution meets these requirements?

Answer and explanation

Answer: B. Asynchronous invocation discards the event once built-in retries are exhausted unless an on-failure destination captures it. A longer timeout addresses duration rather than downstream failure. Reserved concurrency limits how many executions run and does not preserve a failed event. More memory addresses resource exhaustion, not an event that ultimately failed.

58. An AWS Step Functions workflow calls an external API that is intermittently unavailable. Transient failures must not fail the whole execution. Which solution meets these requirements?

Answer and explanation

Answer: D. State-level Retry with backoff handles transient failures inside the execution, and Catch defines behaviour once retries are exhausted. A longer overall timeout lets a stuck execution persist without retrying. Invoking the API twice in parallel duplicates side effects. Restarting the execution repeats work that already completed.

59. Orders are processed through an Amazon SQS queue. During a prolonged downstream outage, messages older than the retention period were lost. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: A, B. Messages expire at the retention period, so extending it to cover the longest plausible outage prevents loss, and alarming on the age of the oldest message gives warning before that point is reached. A shorter visibility timeout redelivers sooner without preventing expiry. FIFO changes ordering and deduplication rather than retention. More consumers cannot process messages while the downstream is unavailable.

60. A queue-backed worker fleet must finish processing the message in flight before Auto Scaling terminates an instance during scale-in. Which solution meets these requirements?

Answer and explanation

Answer: D. A lifecycle hook holds the instance in a terminating wait state so the worker can complete its current message and drain cleanly. Termination protection does not apply to Auto Scaling scale-in. A larger minimum reduces how often scale-in happens without handling it safely. A shorter visibility timeout redelivers the message but the work already done is lost.

61. An application runs on a single Amazon EC2 instance that writes to one Amazon EBS volume. The business requires the application to survive the loss of an Availability Zone with a recovery time objective of minutes. Which solution meets these requirements?

Answer and explanation

Answer: D. Surviving the loss of an Availability Zone requires capacity in another one, and moving shared state off instance-local storage lets a replacement instance reach the same data immediately. An Auto Scaling group confined to one Availability Zone still fails with that zone. A larger instance leaves the zone as a single point of failure. Snapshots in the same zone depend on that zone and would need a manual restore that exceeds an RTO of minutes.

62. An Amazon RDS for PostgreSQL database must fail over automatically to another Availability Zone with the least possible data loss. Which solution meets these requirements?

Answer and explanation

Answer: C. Multi-AZ maintains a synchronous standby and redirects the endpoint automatically, which is the only option here that is both automatic and free of asynchronous replication lag. Promoting a read replica is manual and loses writes not yet replicated. Point-in-time restore is a recovery path measured in hours. Instance class and storage type affect performance rather than failover.

63. Multiple Amazon EC2 instances across several Availability Zones must mount the same file system with read and write access at the same time. Which solution meets these requirements?

Answer and explanation

Answer: A. EFS is a managed NFS file system reachable concurrently from many instances across Availability Zones with standard read and write semantics. An EBS volume attaches within a single Availability Zone. Multi-Attach also works only within one Availability Zone and requires a cluster-aware file system. Presenting S3 as a drive does not provide true shared file semantics.

64. A workload must survive the loss of an entire Region with a recovery point objective near zero and a recovery time objective of a few minutes, at the lowest cost that still meets those targets. Which solution meets these requirements?

Answer and explanation

Answer: C. A warm standby keeps a functioning scaled-down environment with continuous replication, so failover is scaling up and shifting traffic, which fits minutes with near-zero data loss. A full-capacity active-active deployment meets the targets at the highest cost, which the question excludes. Provisioning the application tier at failover time typically exceeds a few minutes. Nightly snapshots give a recovery point up to a day old.

65. A stateless web tier behind an Application Load Balancer must lose no serving capacity when one Availability Zone fails. Peak load requires 8 instances. Which solution meets these requirements at the lowest instance count?

Answer and explanation

Answer: B. Twelve instances across three zones leave eight serving when one zone is lost, which is exactly peak capacity at the lowest total. Eight across three zones drops to roughly five. Sixteen across two zones also survives with eight but requires four more instances. Eight across two zones drops to four, half of what is required.

66. An application writes to an Amazon Aurora cluster. The business requires that the loss of a Region be survivable with a recovery point objective under one minute. Which solution meets these requirements?

Answer and explanation

Answer: C. An Aurora global database replicates to a secondary Region with typical lag well under a second, which meets a sub-minute recovery point objective against Region loss. Multi-AZ and replicas protect against zone and instance failure but all reside in one Region. Backtrack rewinds a cluster in place and does not survive the loss of the Region hosting it. Nightly snapshots give a recovery point up to a day old.

67. An Auto Scaling group repeatedly launches and terminates instances in a short cycle as load fluctuates around the scaling threshold. Which solution meets these requirements?

Answer and explanation

Answer: C. Rapid launch and terminate cycling is thrashing, and a cooldown or warm-up value prevents a further scaling action until the previous one has taken effect. Lowering the threshold makes scaling trigger sooner and worsens the oscillation. The minimum size sets a floor and does not damp the cycling. Load balancer type has no bearing on scaling behaviour.

68. A critical batch job runs nightly on one Amazon EC2 instance. If the instance fails partway through, the job must resume rather than restart from the beginning. Which solution meets these requirements?

Answer and explanation

Answer: B. Recording progress externally lets a replacement instance read the last checkpoint and continue, which is what resumability requires. Termination protection prevents accidental API termination and does not survive a hardware fault. A larger instance shortens the exposure window without making a failure resumable. Restoring a pre-run snapshot returns to the starting state, which is a restart.

69. A static website hosted in Amazon S3 must continue serving read traffic if the primary Region becomes unavailable. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, C. Replication keeps a current copy of the site in a second Region, and failover records with a health check resolve to the secondary when the primary is unhealthy. Versioning protects against overwrite within one bucket in one Region. A load balancer operates within a single Region and cannot span two. EBS snapshots are irrelevant to an S3-hosted static site.

70. Read traffic against an Amazon Aurora cluster varies unpredictably between the load of one replica and twelve. Which solution meets these requirements at the lowest cost?

Answer and explanation

Answer: D. Aurora Auto Scaling adds and removes replicas against a target such as average CPU, matching capacity to unpredictable read load. Twelve fixed replicas pay for peak continuously. One larger replica cannot absorb twelvefold variation. Converting the writer to Serverless v2 does not scale read replica count and sends reads to the writer.

71. An Amazon DynamoDB table must be restorable to any second within the previous 35 days after an accidental bulk delete. Which solution meets these requirements?

Answer and explanation

Answer: B. Point-in-time recovery maintains continuous backups allowing restore to any second within the retention window. Weekly on-demand backups restore only to each backup instant. Streams retain change records for 24 hours and are not a restore mechanism. A global table replicates the deletion to the other Region.

72. An application must survive the failure of an entire Region with a recovery time objective under five minutes and a recovery point objective near zero. Cost is not the primary constraint. Which solution meets these requirements?

Answer and explanation

Answer: D. An active-active deployment already serves traffic from both Regions, so recovery is removing the failed Region from rotation, which fits under five minutes. A scaled-down standby must scale up first. Provisioning the application tier at failover time takes longer still. Restoring from snapshots takes hours.

73. An Auto Scaling group replaces instances that fail EC2 status checks, but leaves instances running when the application process crashes while the operating system remains reachable. Which solution meets these requirements?

Answer and explanation

Answer: C. EC2 status checks observe only instance and system reachability, so the group must consult the load balancer's application-level target health to notice a crashed process. A shorter grace period would terminate instances before they finish booting. A shorter load balancer interval speeds detection at the load balancer but Auto Scaling still acts on EC2 checks. More capacity adds instances without removing unhealthy ones.

74. On-premises users require low-latency access to a shared file set whose authoritative copy must reside in Amazon S3 and survive an on-premises outage. Which solution meets these requirements?

Answer and explanation

Answer: C. File Gateway presents an NFS or SMB share with a local cache while objects are stored durably in S3, so an on-premises failure does not lose data. Hourly DataSync copies leave up to an hour of data only on premises. A file server on EBS keeps the authoritative copy on a single instance. Mounting EFS over the internet is not supported.

75. An application deployed in two Regions must send users to the Region with the lowest latency and must stop sending users to a Region that fails its health check. Which solution meets these requirements?

Answer and explanation

Answer: C. Latency records select the Region with the lowest latency for each user, and attached health checks remove an unhealthy endpoint from consideration. Latency records without health checks continue resolving to a failed Region. Weighted records distribute by proportion rather than proximity. Simple records return values with no health evaluation.

76. An Amazon EFS file system holds data where recent files are accessed constantly and older files are rarely read. All data must remain immediately available. Which solution meets these requirements at the lowest cost?

Answer and explanation

Answer: D. Lifecycle management moves files not accessed within a configured period to a cheaper class while keeping them immediately available in the same file system. One Zone reduces cost by reducing durability across Availability Zones rather than by access pattern. Deleting files breaks immediate availability. Throughput settings affect performance rather than storage cost.

77. A stateless application must lose no capacity when one of two Availability Zones fails. Peak load requires 10 instances. Which solution meets these requirements?

Answer and explanation

Answer: C. With only two zones, surviving the loss of one at full capacity requires each zone to carry the entire peak, so 20 instances are needed. Fifteen leaves roughly seven after a failure and ten leaves five, both short of peak. Keeping an image ready means an outage while capacity is rebuilt.

78. A critical object in Amazon S3 must not be permanently deletable by a compromised credential, while routine updates to the object must continue. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, D. Versioning means a delete leaves prior versions intact, and MFA Delete requires a second factor before a version can be removed permanently. Denying PutObject blocks the routine updates the requirement preserves. Object Lock in compliance mode would also block those updates. Replication copies objects but a delete marker can propagate.

79. An application's web tier must not fail when its processing tier is unavailable. Which solution meets these requirements?

Answer and explanation

Answer: B. A queue decouples their availability so the web tier accepts work while the processing tier is down. A synchronous call fails when the tier is unavailable. A shared group couples their failure domains. More capacity reduces but does not remove the coupling.

80. A workload must scale its compute capacity based on the number of messages waiting in a queue. Which solution meets these requirements?

Answer and explanation

Answer: B. Backlog per instance follows the actual work waiting. CPU may not correlate with queue depth. Scheduled scaling ignores actual demand. A fixed count pays for the peak continuously.

81. Several microservices must react to the same business event without the publisher knowing about them. Which solution meets these requirements?

Answer and explanation

Answer: B. EventBridge decouples publisher from subscribers, each adding a rule without publisher changes. Direct calls couple the publisher to every service. A polled database adds latency and coupling to a schema. A single queue delivers each message to one consumer.

82. A workload processes long-running jobs that must survive an instance failure mid-job. Which solution meets these requirements?

Answer and explanation

Answer: B. External state with a visibility timeout returns an unfinished job to the queue after a failure. A larger instance does not remove failure. Local disk state is lost with the instance. Synchronous processing ties the job to a request.

83. An application must run a series of steps with branching and error handling, and each step's status must be visible. Which solution meets these requirements?

Answer and explanation

Answer: D. Step Functions expresses branching, retries, and error handling with per-step execution history. Chained functions are fragile and opaque. A single function hides the steps. Independent schedules cannot express dependencies.

84. A relational database must remain available if its Availability Zone fails, with automatic failover. Which solution meets these requirements?

Answer and explanation

Answer: C. A Multi-AZ deployment maintains a synchronous standby and fails over automatically. A read replica is asynchronous and requires manual promotion. Snapshots require a restore. A larger instance is still in one zone.

85. A stateless web tier must remain available during an Availability Zone failure. Which solution meets these requirements?

Answer and explanation

Answer: B. A multi-zone Auto Scaling group behind a load balancer keeps serving when a zone fails. A standby AMI, a larger instance, and snapshots all leave a single zone dependency.

86. An application's file storage must be accessible from instances in several Availability Zones simultaneously. Which solution meets these requirements?

Answer and explanation

Answer: C. EFS is a shared file system accessible from multiple zones. EBS is zonal and single-attach in the general case. Instance store is local and ephemeral. S3 lacks full file system semantics.

87. A solution must direct users to a secondary Region if the primary Region's application becomes unhealthy. Which solution meets these requirements?

Answer and explanation

Answer: D. Failover records with a health check direct traffic to the secondary when the primary fails. Weighted routing distributes without health awareness in this form. Simple records return values without evaluating health. A shorter time to live speeds propagation without detecting failure.

88. An application must tolerate the loss of an Availability Zone without losing queued work. Which characteristic of Amazon SQS supports this?

Answer and explanation

Answer: A. SQS stores messages redundantly across zones within the Region. It does not use single instances, does not replicate cross-Region automatically, and does not hold messages in producer memory.

89. An application's components must communicate without either knowing the other's location. Which solution meets these requirements?

Answer and explanation

Answer: C. Queues and topics decouple producers from consumers entirely. Direct addressing, co-location, and DNS all require the caller to know the callee.

90. A workload must process messages exactly once where duplicates would cause incorrect results. Which solution meets these requirements?

Answer and explanation

Answer: C. A FIFO queue with deduplication plus an idempotent consumer addresses duplicates from both delivery and retry. Standard queues deliver at least once, visibility timeout governs redelivery timing, and SNS fan-out delivers to every subscriber.

91. An application must invoke a long-running process without the caller waiting for it to complete. Which solution meets these requirements?

Answer and explanation

Answer: C. Queueing the work and returning an identifier decouples the caller from the duration. Longer timeouts and larger instances keep the caller waiting, and returning an error loses the work.

92. A microservice must be able to scale independently of the services that call it. Which characteristic is required?

Answer and explanation

Answer: C. Statelessness or external state allows instances to be added and removed freely. Shared databases, co-location, and session affinity all couple the service to its callers or its instances.

93. An application must react to changes in a DynamoDB table without polling it. Which solution meets these requirements?

Answer and explanation

Answer: D. Streams deliver change records to a consumer as they occur. Scheduled querying and scanning are polling, and read capacity does not provide change notification.

94. A workload's components must be able to fail independently without the failure cascading. Which pattern applies?

Answer and explanation

Answer: D. Bulkhead isolation prevents one component exhausting another's resources. A shared pool, a single process, and synchronous retries all propagate failure.

95. An application must distribute work across consumers so each message is processed once. Which solution meets these requirements?

Answer and explanation

Answer: C. An SQS queue delivers each message to a single consumer. SNS and multiple EventBridge rules deliver to every subscriber, and listing S3 requires coordination to avoid duplicate processing.

96. An application must buffer requests during a traffic spike so the backend processes them at a sustainable rate. Which pattern applies?

Answer and explanation

Answer: B. Queue-based load levelling decouples arrival rate from processing rate. Peak scaling pays for idle capacity, rejection loses work, and longer timeouts hold connections.

97. A workflow's steps must run in order with error handling visible per step. Which solution meets these requirements?

Answer and explanation

Answer: C. Step Functions provides ordering, per-state error handling, and execution history. Chained functions are opaque, a single function hides the steps, and independent schedules cannot express ordering.

98. An Aurora cluster must continue serving reads if its writer instance fails. Which configuration is required?

Answer and explanation

Answer: C. A reader in another zone is promoted on writer failure. Instance size, backups, and Performance Insights do not provide failover.

99. An application must continue accepting requests while a downstream component is being deployed. Which solution meets these requirements?

Answer and explanation

Answer: D. A queue buffers requests so the front end continues accepting them. Synchronous retries fail while the component is down, a maintenance window is an outage, and extra capacity does not help during replacement.