42 practice questions for Domain 5 of the AWS Certified Generative AI Developer - Professional (AIP-C01) exam, which makes up 11% of its scored content. Your answers count towards one score and one timer for the whole exam.
Domain 5: Testing, Validation, and Troubleshooting
253. Two candidate model configurations must be compared on quality and cost before one is adopted. Which solution meets these requirements?
Answer and explanation
Answer: D. A model evaluation over a representative dataset measures quality on the actual task, and pairing it with cost and latency gives the complete comparison. Published benchmarks describe different workloads. Complaint volumes are a lagging and noisy signal. Parameter count and context window are specifications rather than outcomes.
254. A quality gate must prevent a GenAI deployment when output quality falls below the current production standard. Which solution meets these requirements?
Answer and explanation
Answer: C. A blocking evaluation against a baseline prevents a regression reaching production at all. Deploying and comparing afterwards exposes users. Manual sampling is inconsistent and slow. A checklist entry relies on someone applying it.
255. An agent's performance must be measured on whether it completes assigned tasks correctly rather than on individual response quality. Which solution meets these requirements?
Answer and explanation
Answer: A. Task completion rate and tool usage effectiveness measure whether the agent achieved the goal, which is the unit of value for an agent. Fluency assesses individual text. Token consumption and latency are cost and speed measures.
256. Retrieval quality must be measured separately from generation quality in a RAG application. Which approach is appropriate?
Answer and explanation
Answer: A. Labelled relevance judgements isolate retrieval, and evaluating generation with correct context supplied isolates the generator. Proportional attribution of end-to-end failures is guesswork. Latency measures speed. Passage counts describe volume.
257. User feedback must be collected continuously and fed back into evaluation of a GenAI application. Which solution meets these requirements?
Answer and explanation
Answer: A. In-product ratings tied to the interaction identifier produce continuous labelled feedback that can calibrate automated scoring. A quarterly survey is infrequent and detached from specific responses. Email relies on users making the effort. Follow-up questions are ambiguous, since they may indicate either dissatisfaction or engagement.
258. An application intermittently fails with an error indicating the input exceeded the model's context window. Which solution meets these requirements?
Answer and explanation
Answer: D. Context window overflow requires the input to be reduced deliberately, with the least relevant content trimmed or summarized. Output tokens are a separate budget and increasing them reduces the space available for input. Retrying the same input fails identically. Temperature affects sampling rather than context length.
259. A RAG application returns an irrelevant answer, and the team must determine whether retrieval or generation failed. Which step should be taken first?
Answer and explanation
Answer: B. The retrieved context shows immediately whether the right passage was supplied, which determines which half of the system to fix. A larger model cannot compensate for context that was never retrieved. Temperature affects sampling. Re-embedding is expensive guesswork before the cause is known.
260. Responses have become inconsistent after a prompt template change, and the team must establish which revision caused it. Which solution meets these requirements?
Answer and explanation
Answer: C. Version comparison against fixed inputs isolates the revision that changed behaviour, which is what diagnosis requires. Reverting resolves the symptom without identifying the cause or preserving the improvement. Higher temperature increases inconsistency. Rewriting discards the information the version history holds.
261. Retrieval quality degraded after the embedding model was upgraded, although the corpus is unchanged. Which cause is most likely?
Answer and explanation
Answer: D. Embeddings from different models occupy different spaces, so a partially re-embedded index produces meaningless similarity scores for the mixed portion. The question states the corpus is unchanged. A threshold change would affect all results uniformly rather than degrading quality selectively. Exceeding storage produces write errors rather than poor relevance.
262. An evaluation framework must assess factual accuracy, relevance, and fluency separately rather than as one score. Which rationale applies?
Answer and explanation
Answer: A. The dimensions fail independently, and a combined score cannot say which one did, which is what makes the failure actionable. Computational ease and tooling requirements are not the reason. A combined score can be calculated but tells you less.
263. A canary evaluation must validate a model update against real traffic before full rollout. Which design is appropriate?
Answer and explanation
Answer: A. Comparing quality and business metrics between versions over the same period isolates the change as the variable. Confirming responses return tests availability rather than quality. Routing all traffic removes the comparison. Offline scores predict but do not establish live behaviour.
264. A user feedback mechanism must produce data useful for improving a GenAI application. Which design is appropriate?
Answer and explanation
Answer: D. Linking feedback to the specific interaction is what makes it diagnosable and usable for calibrating automated evaluation. An unlinked rating cannot be traced to a cause. A separate channel loses the connection. A monthly survey measures sentiment without identifying which responses caused it.
265. A quality gate must prevent a deployment whose hallucination rate exceeds an agreed level. Which approach is appropriate?
Answer and explanation
Answer: B. A blocking measurement against a labelled set prevents the regression reaching users. Measuring in production means users encounter it first. Manual sampling is inconsistent. Recording a number in release notes relies on someone acting on it.
266. An agent evaluation must establish whether multi-step reasoning is sound rather than only whether the final answer is correct. Which approach is appropriate?
Answer and explanation
Answer: B. A correct answer reached by faulty reasoning will fail on the next similar task, so the trace must be examined. Final-answer accuracy misses that. Step count and duration describe efficiency rather than soundness.
267. A stakeholder report must communicate GenAI evaluation results to a non-technical audience. Which content is most appropriate?
Answer and explanation
Answer: B. A non-technical audience needs capability, failure modes, and business consequence. Metric tables, dataset composition, and configuration detail are the evidence behind that conclusion rather than the conclusion itself.
268. An application intermittently returns truncated answers, and the token usage shows responses hitting the configured limit. Which action is appropriate?
Answer and explanation
Answer: A. Responses hitting the configured ceiling indicate the limit is too low for that class of request, or that the answer should be shorter by design. Temperature affects sampling. More context lengthens the input rather than permitting a longer output. A timeout governs time rather than tokens.
269. A model integration fails intermittently with errors that name a validation problem in the request. Which step should be taken first?
Answer and explanation
Answer: D. A validation error names a malformed request, so the actual payload must be compared against the schema. Backoff addresses throttling rather than validation. A different model does not fix a malformed request. Timeouts address latency.
270. Retrieval quality has degraded, and the team must determine whether the cause is the embeddings or the chunking. Which approach is appropriate?
Answer and explanation
Answer: A. Isolating one variable at a time attributes the change to its cause. Changing both makes the contributions indistinguishable. More candidates masks the problem. Re-embedding with the same model reproduces the same vectors.
271. A prompt performs well in testing but poorly in production, and the inputs differ in ways the test set did not represent. Which action is appropriate?
Answer and explanation
Answer: A. A test set that does not represent production traffic cannot predict production behaviour, so it must be extended with real examples. More of the same kind of example does not broaden coverage. Reverting discards the improvement without addressing the gap. Temperature does not compensate for unrepresentative testing.
272. An evaluation must measure relevance, factual accuracy, consistency, and fluency of GenAI outputs. Which approach is appropriate?
Answer and explanation
Answer: A. Separate per-dimension metrics reveal which quality fails. A single score conceals it. Measuring one dimension omits the others.
273. A multi-model evaluation must compare cost-performance across candidates. Which metric is appropriate?
Answer and explanation
Answer: A. Quality per cost captures the trade-off. Quality alone ignores cost. Price alone ignores quality. Parameter count is neither.
274. An annotation workflow must produce reliable quality labels for GenAI outputs. Which design is appropriate?
Answer and explanation
Answer: C. Multiple annotators with a rubric and agreement checks produce reliable labels. Single annotators and no rubric are unreliable. Automated-only lacks human ground truth.
275. A continuous evaluation workflow must run automatically as the application changes. Which trigger is appropriate?
Answer and explanation
Answer: B. Any behaviour-affecting change warrants re-evaluation. Model-only triggers miss prompt and retrieval changes. Monthly and complaint-driven triggers leave gaps.
276. A RAG evaluation must measure whether the retrieved context was actually used in the answer. Which metric applies?
Answer and explanation
Answer: A. Faithfulness measures grounding in context. Latency, count, and length do not.
277. A retrieval quality test must measure whether relevant passages are ranked highly, not just retrieved. Which metric applies?
Answer and explanation
Answer: C. Rank-aware metrics measure position. Total recall ignores rank. Index size and latency are unrelated.
278. Deployment validation must detect semantic drift in a GenAI application after an update. Which approach is appropriate?
Answer and explanation
Answer: A. Synthetic workflows with baseline comparison detect meaning change. Completion, response, and error rate do not.
279. An evaluation report for stakeholders must support a decision between two model configurations. Which presentation is appropriate?
Answer and explanation
Answer: A. A side-by-side with recommendation supports the decision. Raw tables and datasets require interpretation. Preference is not evidence.
280. A GenAI application returns truncated responses only for certain inputs. Which cause should be investigated first?
Answer and explanation
Answer: C. Input-dependent truncation points to context overflow leaving no response room. Overload, network, and key issues are not input-dependent.
281. A prompt template produces inconsistent output formats despite instructions. Which troubleshooting step is appropriate?
Answer and explanation
Answer: B. Systematic testing identifies the failure pattern and schema constraint enforces format. Higher temperature worsens consistency. Removing instructions loses guidance. Switching without testing is guesswork.
282. Embedding quality is suspected of degrading retrieval. Which diagnostic is appropriate?
Answer and explanation
Answer: D. Labelled recall comparison isolates embedding quality. Visual inspection is not measurement. More results masks the problem. Same-model rebuild reproduces the same vectors.
283. A prompt maintenance process must detect when a model update changes how the prompt is interpreted. Which approach is appropriate?
Answer and explanation
Answer: D. Behaviour tests on each model update detect interpretation change. Assumption and complaint monitoring are reactive. Rewriting every time is unnecessary.
284. A retrieval issue must be traced to whether the vectorization step or the chunking step is responsible. Which approach is appropriate?
Answer and explanation
Answer: D. Inspecting chunks then embeddings isolates the stage. Blind re-embedding, chunk size changes, and store replacement do not diagnose.
285. An evaluation dataset must represent the application's real usage. Which approach is appropriate?
Answer and explanation
Answer: D. Sampled production traffic represents actual usage. Team-written examples, public benchmarks, and prompt examples all reflect something other than real requests.
286. An automated evaluator's scores must be trusted before being relied upon. Which approach is appropriate?
Answer and explanation
Answer: B. Measured agreement with human ratings establishes whether the evaluator can be trusted. Capability does not imply calibration, error rate is unrelated, and selective use does not validate.
287. An evaluation must detect that a change improved one request class while degrading another. Which approach is appropriate?
Answer and explanation
Answer: D. Per-class reporting reveals offsetting effects an aggregate hides. Aggregates, single-class reporting, and volume counts do not.
288. An evaluation must establish whether a retrieval change improved the final answer. Which approach is appropriate?
Answer and explanation
Answer: B. Measuring both establishes whether the retrieval gain reached the answer. Either metric alone leaves the link unproven, and latency is a different dimension.
289. An evaluation harness must produce comparable results across runs. Which approach is appropriate?
Answer and explanation
Answer: C. Fixing everything but the variable under test makes runs comparable. New samples, changing metrics, and current traffic all introduce confounds.
290. A human review process must scale to a large volume of generated output. Which approach is appropriate?
Answer and explanation
Answer: C. Risk-weighted sampling concentrates review where errors matter most. Full review does not scale, uniform sampling ignores consequence, and complaint-driven review is reactive.
291. An evaluation must assess whether an agent completed a multi-step task correctly. Which approach is appropriate?
Answer and explanation
Answer: A. Verifying the end state and checkpoints establishes correctness. A returned response, step count, and duration do not.
292. Evaluation results must be communicated so a product owner can decide whether to ship. Which approach is appropriate?
Answer and explanation
Answer: B. Quality by class, cost, and residual risk in business terms support the decision. Metric tables and methodology are supporting evidence, and a recommendation without evidence is not a basis for the owner's decision.
293. An agent intermittently fails to complete tasks it previously handled correctly. Which approach is appropriate?
Answer and explanation
Answer: C. Comparing traces localises where behaviour changed. A higher step limit, a larger model, and restarts all act without diagnosis.
294. A GenAI application returns correct answers in testing but poor answers for a subset of production users. Which approach is appropriate?
Answer and explanation
Answer: B. Comparing failing requests against the test set identifies the unrepresented characteristic. Temperature, retraining, and more passages all act before the cause is known.