Troubleshooting and Optimization

Domain 4: Troubleshooting and Optimization

46 practice questions for Domain 4 of the AWS Certified Developer - Associate (DVA-C02) exam, which makes up 18% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 4: Troubleshooting and Optimization

18% of scored content · 46 practice questions

154. An AWS Lambda function returns HTTP 502 through Amazon API Gateway. Which cause is most likely?

Answer and explanation

Answer: B. With Lambda proxy integration API Gateway expects a specific response shape, and a payload that does not conform produces a 502 Bad Gateway. Insufficient memory surfaces as a timeout or an out-of-memory error. An undeployed stage returns a missing resource error. Throttling returns 429.

155. An application receives ThrottlingException responses from an AWS service during traffic spikes. Which solution meets these requirements?

Answer and explanation

Answer: D. Exponential backoff spaces retries out and jitter prevents many clients retrying in lockstep, which lets the service recover. A tight retry loop amplifies the overload. More concurrent connections increase the request rate and worsen throttling. Failing immediately discards requests that would succeed a moment later.

156. An AWS Lambda function shows duration close to its limit while memory usage stays well below the configured allocation. Which step should the developer take first?

Answer and explanation

Answer: B. Lambda allocates CPU in proportion to memory, so when memory is underused and duration is high the constraint is often waiting on an external call, and raising memory buys CPU the function is not using. Adding memory is the reflex fix but is wasted spend here. A shorter timeout causes failures instead of speeding work. Provisioned concurrency removes cold starts, not steady-state duration.

157. An application's Amazon DynamoDB writes intermittently fail with a validation error stating that the item size exceeds the limit. Which solution meets these requirements?

Answer and explanation

Answer: D. DynamoDB items are capped at 400 KB, so oversized payloads belong in S3 with the object key held in the item. Write capacity governs throughput and has no bearing on the item size limit. BatchWriteItem sends several separate items and cannot split one oversized item. Streams emit a change record and do not change item size constraints.

158. A developer must search several gigabytes of AWS Lambda logs for one correlation ID across a two-day window. Which solution meets these requirements?

Answer and explanation

Answer: A. Logs Insights runs an indexed query across log groups and time ranges and returns matching events in seconds. Downloading gigabytes to search locally is slow and expensive. Retention controls how long logs persist rather than how they are searched. X-Ray cannot be enabled retroactively for traffic that has already occurred.

159. An Amazon DynamoDB table shows occasional throttling on one partition while overall provisioned capacity is not exhausted. Which solution provides the durable fix?

Answer and explanation

Answer: D. Adaptive capacity shifts throughput toward busy partitions automatically but cannot overcome sustained skew, so the key design must distribute traffic. More capacity does not change how items map to partitions. A local secondary index shares the base partition key and inherits the skew. Auto scaling adjusts total capacity without addressing distribution.

160. An application writes unstructured text to logs, making incident investigation slow. Which solution meets these requirements?

Answer and explanation

Answer: B. Structured logs let queries filter and aggregate on named fields rather than parsing free text, which is what makes incident search fast. Retention governs how long logs persist. Reducing the level discards information needed during an incident. Changing destination does not add structure.

161. An AWS Lambda function intermittently receives server errors from a downstream service. Which client behaviour should the developer implement?

Answer and explanation

Answer: B. Bounded retries with backoff and jitter handle transient server errors without amplifying load, and a clear failure after exhaustion lets callers respond appropriately. Indefinite retries can exhaust the function timeout and the downstream service. Returning success on failure hides the problem and corrupts downstream behaviour. Failing immediately discards recoverable transients.

162. An Amazon API Gateway endpoint returns HTTP 429 to callers during peak traffic. Which cause is most likely?

Answer and explanation

Answer: C. A 429 indicates too many requests, which API Gateway returns when a throttling limit or usage plan quota is exceeded. A malformed response produces 502. An undeployed stage produces a missing resource error. An integration timeout produces 504.

163. A developer must identify which AWS Lambda invocations exceeded three seconds over the past week across millions of log events. Which solution meets these requirements?

Answer and explanation

Answer: C. Logs Insights parses the duration field from Lambda report lines and filters across log groups and time ranges in seconds. Exporting to S3 and querying with Athena works but requires building the export and table definition for a query Logs Insights answers directly. Retention controls persistence rather than search. X-Ray cannot be enabled retroactively.

164. A developer must find which downstream call causes intermittent latency in a request that traverses Amazon API Gateway, two AWS Lambda functions, and Amazon DynamoDB. Which solution meets these requirements?

Answer and explanation

Answer: D. X-Ray records a trace across every instrumented hop and attributes latency to individual segments and subsegments, which localises the slow call. Per-service metrics show that something is slow without tying it to one request path. Flow Logs capture network metadata and cannot see inside managed service calls. CloudTrail records API activity for audit rather than per-request timing.

165. A developer must record a value in an AWS X-Ray trace so that traces can later be filtered on it. Which solution meets these requirements?

Answer and explanation

Answer: C. Annotations are indexed key-value pairs that can be used in filter expressions. Metadata is recorded with the trace but is not indexed for searching. A subsegment name identifies a unit of work rather than storing arbitrary searchable values. Sampling rules control which requests are traced.

166. An AWS Lambda function must publish a business metric without adding a synchronous API call to each invocation. Which solution meets these requirements?

Answer and explanation

Answer: A. Embedded metric format writes structured log entries that CloudWatch parses into metrics asynchronously, avoiding a synchronous call and its latency and throttling. PutMetricData per invocation adds latency and cost. An alarm on invocation count measures something different. X-Ray metadata is not indexed as a metric.

167. An operations team must be notified when an application approaches its AWS Lambda concurrency quota. Which solution meets these requirements?

Answer and explanation

Answer: B. An alarm on concurrent executions set below the quota warns before invocations start being throttled. An alarm on errors fires only once throttling has already caused failures. A weekly console review may miss rapid growth. Raising the quota postpones the problem without providing warning.

168. A distributed application must record which customer tier each traced request belongs to so that traces can be filtered by tier. Which solution meets these requirements?

Answer and explanation

Answer: D. Annotations are indexed key-value pairs usable in X-Ray filter expressions. Metadata is stored with the trace but is not indexed for filtering. Putting a variable value in the segment name fragments the service map. A metric dimension aggregates counts rather than letting individual traces be retrieved.

169. An application writes log entries that must record the request identifier, the operation, and the outcome in a form that can be queried without regular expressions. Which solution meets these requirements?

Answer and explanation

Answer: B. Structured JSON with consistent field names lets CloudWatch Logs Insights query on named fields directly. A formatted sentence requires parsing to extract values. Splitting one event across several entries loses the correlation between them. Raising the log level increases volume without adding structure.

170. An AWS Lambda function connecting to Amazon RDS exhausts the database connection limit under load. Which solution meets these requirements?

Answer and explanation

Answer: B. RDS Proxy maintains a shared connection pool so many concurrent executions do not each open a database connection. More memory does not reduce connection count. Opening a connection per invocation is the behaviour causing exhaustion. A larger instance raises the ceiling at higher cost without addressing the pattern.

171. An application makes many small sequential calls to Amazon DynamoDB, and latency is dominated by connection setup. Which solution meets these requirements?

Answer and explanation

Answer: D. Reusing connections through keep-alive avoids repeating the TLS handshake on every call, which dominates latency for many small requests. Capacity settings affect throttling rather than connection setup. A shorter timeout causes premature failures. Creating a client per request forces a new connection each time.

172. A team must determine the AWS Lambda memory setting that minimizes cost per invocation for a CPU-bound function. Which solution meets these requirements?

Answer and explanation

Answer: B. Because CPU scales with memory, a higher setting can reduce duration enough to lower total cost, so the optimum is found by measuring across settings. The minimum can be more expensive if duration rises sharply. The maximum wastes allocation for functions that do not use it. Package size is unrelated to the memory a function needs at run time.

173. An application's Amazon DynamoDB reads intermittently fail with ProvisionedThroughputExceededException during brief spikes. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, D. Exponential backoff with jitter spreads retries so clients do not re-converge on the same instant, and auto scaling or on-demand capacity raises throughput to absorb spikes. A tight retry loop amplifies the overload it is reacting to. Larger items consume more read capacity units and worsen throttling. Disabling SDK retries removes the resilience that handles transient throttling gracefully.

174. A CloudWatch Logs subscription filter forwards every log event to a processing function, and most events are not relevant. Which solution meets these requirements?

Answer and explanation

Answer: A. A filter pattern on the subscription evaluates events before delivery, so irrelevant events never invoke the function and are never billed. More memory processes the same irrelevant volume faster. Retention governs stored history rather than what is delivered. A second subscription doubles delivery rather than reducing it.

175. An Amazon CloudFront distribution serves responses that differ depending on a request header the application sets, and users are receiving responses intended for other clients. Which solution meets these requirements?

Answer and explanation

Answer: B. Including the header in the cache key causes CloudFront to store and serve a separate object per header value, which is what prevents cross-client responses. Disabling caching removes the problem but also the benefit. A shorter TTL narrows the window without fixing the collision. An origin request policy forwards the header without differentiating cached objects.

176. A request traverses API Gateway, Lambda, and DynamoDB, and the team must identify which segment is slow. Which solution meets these requirements?

Answer and explanation

Answer: C. X-Ray attributes latency to individual segments across the request path. Function logs show what the developer chose to log. Access logs show total latency without a breakdown. Capacity metrics describe one service.

177. A Lambda function is failing intermittently, and the team must find the error message for the failed invocations. Which source should be examined?

Answer and explanation

Answer: C. Function logs contain the error output from failed invocations. Configuration, creation records, and concurrency settings do not carry runtime error detail.

178. An application must emit a custom metric without a synchronous API call on every event. Which solution meets these requirements?

Answer and explanation

Answer: C. Embedded metric format writes structured log entries CloudWatch parses into metrics without a synchronous call. Per-event PutMetricData adds latency and throttling risk. Alarms evaluate metrics rather than creating them. Retention governs log persistence.

179. A developer must search application logs across several Lambda functions for a correlation identifier. Which solution meets these requirements?

Answer and explanation

Answer: C. Logs Insights queries across multiple log groups and returns matching events. Downloading is slow and not repeatable. A metric filter counts occurrences without returning events. Retention governs how much data exists.

180. An application's logs must be machine-parseable so fields can be queried individually. Which approach is appropriate?

Answer and explanation

Answer: A. Structured JSON allows queries to filter and aggregate on named fields. Prose, fixed columns, and error-only logging all limit what can be queried.

181. A developer must add custom segments to an X-Ray trace to measure a specific block of application code. Which approach is appropriate?

Answer and explanation

Answer: C. A subsegment measures a specific code block within the trace. Log statements record timing without appearing in the trace. Sampling rate governs how many requests are traced. Active tracing produces the parent segment.

182. An application must report a business event count that can be alarmed on. Which approach is appropriate?

Answer and explanation

Answer: B. A custom metric can be alarmed on directly. Logs require a metric filter to become alarmable. A table and an API response are not metric sources for alarms.

183. An API's responses for a frequently requested resource rarely change. Which solution reduces backend load?

Answer and explanation

Answer: B. API caching serves repeated requests without reaching the backend. A larger backend serves the same requests at higher cost. A higher throttling limit permits more load through. Request validation rejects malformed requests.

184. A Lambda function is slower than expected, and profiling shows most time in CPU-bound work. Which adjustment is appropriate?

Answer and explanation

Answer: B. Lambda allocates CPU proportionally to memory, so more memory speeds CPU-bound work and may reduce total cost. A longer timeout permits slowness. Less memory slows it further. Concurrency governs parallel executions.

185. An application makes many small writes to DynamoDB that could be grouped. Which approach improves efficiency?

Answer and explanation

Answer: B. BatchWriteItem groups writes into fewer round trips. A loop of individual calls is the inefficiency. Transactions add coordination overhead where atomicity is not required. More capacity serves the same inefficient pattern.

186. A Lambda function is throttled because downstream requests exceed a third-party API's rate limit. Which approach is appropriate?

Answer and explanation

Answer: B. Reserved concurrency caps how many invocations run in parallel, bounding the downstream request rate. Timeout and memory do not limit concurrency. Immediate retries amplify the overload.

187. A developer must identify which downstream call in a request caused an error. Which approach is appropriate?

Answer and explanation

Answer: B. Subsegments attribute the error to a specific call. Aggregate error rate, duration, and invocation counts do not identify which call failed.

188. An intermittent failure must be correlated across several services handling the same request. Which approach is appropriate?

Answer and explanation

Answer: A. A propagated identifier ties entries together. Timestamps are ambiguous under concurrency, higher verbosity adds volume without linkage, and independent review cannot connect the entries.

189. A developer must determine whether a failure originated in the application or in an AWS service call. Which approach is appropriate?

Answer and explanation

Answer: D. Trace segments distinguish where the fault occurred. Error counts, quotas, and memory configuration do not attribute the fault.

190. An error occurs only under concurrent load and cannot be reproduced with a single request. Which approach is appropriate?

Answer and explanation

Answer: B. Reproducing the concurrency under a load test with tracing captures the condition. Retrying a single request does not create concurrency, timeouts address duration, and code review may miss a race.

191. An application must emit a metric with dimensions allowing it to be filtered by customer. Which approach is appropriate?

Answer and explanation

Answer: A. A dimension supports filtering while keeping one metric name, provided cardinality stays reasonable. A metric per customer explodes the namespace, a log message is not a metric, and undimensioned metrics cannot be filtered.

192. An application's traces must include a business identifier so specific transactions can be located. Which approach is appropriate?

Answer and explanation

Answer: B. Annotations are indexed and filterable; metadata is stored but not indexed for filtering. Logs and response payloads are outside the trace.

193. An application must record structured context with every log entry without repeating it in each statement. Which approach is appropriate?

Answer and explanation

Answer: C. Logger-level context is attached automatically and consistently. Manual concatenation is error-prone, a single context entry requires correlation, and a database lookup is disproportionate.

194. An application must expose a health signal distinguishing a running process from a functioning service. Which approach is appropriate?

Answer and explanation

Answer: B. A dependency-aware health endpoint reports whether the service can actually function. Process state, an open port, and a schedule all report availability rather than function.

195. A developer must measure how long a specific block of application code takes within a traced request. Which approach is appropriate?

Answer and explanation

Answer: C. A custom subsegment records the block's duration within the trace. Log timestamps require manual correlation, total duration does not isolate the block, and sampling rate governs how many requests are traced.

196. An application must record the outcome of each business transaction in a form that supports alerting. Which approach is appropriate?

Answer and explanation

Answer: D. A dimensioned metric supports alarms directly. Logs require a metric filter, a database requires querying, and a response reaches only the caller.

197. A DynamoDB query retrieves items and then discards most of them in the application. Which improvement is appropriate?

Answer and explanation

Answer: B. Refining the key condition reduces items read and capacity consumed; a filter reduces returned data but the read is still charged, so the key condition is the stronger lever. More capacity serves the same waste, caching helps repeats, and memory is unrelated.

198. An application makes the same read request to a downstream service many times within a short period. Which improvement is appropriate?

Answer and explanation

Answer: B. Caching removes the repeated calls entirely. More downstream capacity serves the same waste, more concurrency increases it, and backoff addresses throttling.

199. An application must record the duration of each external call for later analysis. Which approach is appropriate?

Answer and explanation

Answer: C. A dimensioned metric supports aggregation and alarming per service. Timestamps require post-processing, total duration does not isolate the call, and counts describe volume.