Operational Efficiency and Optimization for GenAI Applications

Domain 4: Operational Efficiency and Optimization for GenAI Applications

48 practice questions for Domain 4 of the AWS Certified Generative AI Developer - Professional (AIP-C01) exam, which makes up 12% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 4: Operational Efficiency and Optimization for GenAI Applications

12% of scored content · 48 practice questions

205. A team enables prompt caching but sees no cost reduction, and the prompt's shared instruction block sits after the variable user content. Which change is required?

Answer and explanation

Answer: C. Prompt caching reuses a matching prefix, so any variable content placed ahead of the static block makes every request's prefix unique and nothing can be reused. Retention governs how long an entry lives rather than whether it matches. A shorter block would still be preceded by variable content. More concurrency multiplies requests that all miss.

206. An assistant answers many semantically similar questions each day. Which solution most reduces inference cost?

Answer and explanation

Answer: C. A semantic cache eliminates inference entirely for repeated questions, which is the largest available saving when queries recur. More output tokens increase cost. A larger model maximises cost. Varying responses removes the repetition that makes caching effective.

207. Context supplied to the model has grown as retrieval was tuned, and cost has risen with it. Which approach is appropriate?

Answer and explanation

Answer: A. Context volume is the variable that grew, so measuring quality against it identifies how much is actually needed. Output limits address a different part of the cost. A larger window permits more tokens rather than fewer. Raising the threshold reduces passages without checking whether quality survives.

208. A time-sensitive feature must minimise the delay before the user sees the first word of a response. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, D. Streaming delivers the first tokens immediately and a latency-optimized model shortens generation, which together address time to first word. More output tokens and more context both increase the work before and during generation. Temperature does not affect when generation starts.

209. Retrieval latency dominates the end-to-end response time of a RAG application. Which approach is appropriate?

Answer and explanation

Answer: D. When retrieval dominates, the index configuration and query path are where the time is spent, and both must be tuned against measured latency and recall. More candidates increases retrieval work. A larger generation model addresses the wrong stage. A longer timeout tolerates the latency.

210. A workflow makes several independent model calls that currently run one after another. Which solution meets these requirements?

Answer and explanation

Answer: D. Running independent calls in parallel reduces total latency to the longest call rather than their sum. Combining unrelated tasks into one prompt degrades quality on each. More output tokens does not reduce the number of distinct tasks. A longer timeout tolerates the delay.

211. A team must detect a sudden change in token consumption that indicates a prompt or retrieval regression. Which solution meets these requirements?

Answer and explanation

Answer: D. Tokens per request isolates the change from volume, and an anomaly band detects a shift without a preset threshold. Monthly spend is retrospective and aggregates volume with efficiency. Invocation counts measure volume. An error rate does not move when a prompt simply becomes more verbose.

212. An investigation requires the full prompt and completion for a specific past interaction. Which solution meets these requirements?

Answer and explanation

Answer: D. Model invocation logging captures the full request and response payload, which is what reconstructing an interaction requires. CloudTrail records that an API was called without the payload. A request identifier alone omits the content. X-Ray records timing and the call path rather than payloads.

213. A vector store's retrieval quality must be monitored so degradation is detected before users report poor answers. Which solution meets these requirements?

Answer and explanation

Answer: C. A recurring known-answer set measures whether the right passages are still being found, which is retrieval quality. Storage size and ingestion counts describe volume. Query latency describes speed rather than relevance.

214. An application's context grows through a conversation until cost per request becomes unacceptable. Which approach is appropriate?

Answer and explanation

Answer: C. Summarizing older turns compresses the history while preserving what the conversation needs. Blind truncation may discard a fact the conversation still depends on. Starting over loses the conversation. A larger window permits more tokens and therefore more cost.

215. An application's provisioned throughput is underused outside business hours. Which approach is appropriate?

Answer and explanation

Answer: C. Committing only to the sustained period captures the discount without paying for idle hours, with the remainder served flexibly. Retaining capacity continuously is the cost being addressed. More capacity worsens it. Batching delays requests that users expect answered.

216. A caching strategy must avoid returning a stale answer when the underlying documents change. Which approach is appropriate?

Answer and explanation

Answer: A. Keying the cache so entries can be mapped to their sources allows precise invalidation when a document changes. A short uniform expiry discards valid entries and still serves stale ones within the window. Disabling caching forfeits the saving. A nightly refresh leaves staleness through the day.

217. An application must reduce perceived latency for a response that genuinely takes several seconds to generate. Which approach is appropriate?

Answer and explanation

Answer: D. Streaming reduces time to first token, which is what the user perceives, even when total generation time is unchanged. More output tokens lengthen generation. Fewer passages may degrade the answer. A longer timeout changes nothing the user experiences.

218. A team must decide between a lower temperature and a higher one for a factual question-answering application. Which consideration applies?

Answer and explanation

Answer: A. Lower temperature reduces sampling variation, which suits a task with one correct answer. Higher temperature increases variation and the likelihood of unsupported content. Retrieval supplies context but generation still samples. A single setting for all applications ignores the task.

219. A retrieval step is fast in isolation but slow within the application's request path. Which cause should be investigated first?

Answer and explanation

Answer: C. A component fast in isolation but slow in context points to the path between them rather than the component itself. Index size would slow it in isolation too. An embedding change would affect quality. The generation model is a separate stage.

220. An application must handle a burst of concurrent model invocations without exceeding its throughput allocation. Which approach is appropriate?

Answer and explanation

Answer: C. A concurrency limit with queuing keeps utilisation at the allocation without breaching it. Invoking everything and retrying amplifies the overload. A longer timeout holds resources while throttling continues. Fewer tokens per request reduces work per call rather than concurrent calls.

221. A dashboard must show whether a GenAI application is delivering business value rather than only operating correctly. Which metric belongs on it?

Answer and explanation

Answer: D. Task completion measures whether the application delivers its purpose. Invocation counts, latency, and error rate confirm the system is working without establishing that it is useful.

222. A team must detect when a model update changes an application's behaviour on inputs that previously worked. Which approach is appropriate?

Answer and explanation

Answer: A. A golden dataset with recorded baseline outputs detects behavioural change directly, including changes that produce no error. Error rates miss changes that are wrong but well-formed. User feedback is delayed and partial. Benchmarks describe the model rather than this application.

223. Vector store operational health must be monitored so degradation is detected before it affects answers. Which combination of signals is appropriate? (Select TWO.)

Answer and explanation

Answer: B, D. Latency at the target percentile and hit rate against known questions together cover whether the store is fast enough and still returning the right content. Vector count, cost, and job count describe scale and activity rather than health.

224. An agent's tool usage must be monitored to establish whether it is calling tools appropriately. Which signal is most informative?

Answer and explanation

Answer: D. Comparing the per-task distribution against a baseline reveals an agent calling the wrong tools or calling them too often. Daily totals aggregate across tasks. Tool latency is a performance measure. Registered tool count is a configuration fact.

225. A request's token count must be estimated before invocation to enforce a per-request budget. Which approach is appropriate?

Answer and explanation

Answer: C. The model's own tokenizer measures accurately. Character and word heuristics are unreliable across languages and content. Checking afterwards does not enforce.

226. A tiered model strategy must be validated before rollout. Which validation is appropriate?

Answer and explanation

Answer: C. Evaluation on the actual query class establishes quality holds. Assumptions and list prices do not measure quality. Complaints are reactive.

227. Batching must be applied to a workload of many independent short requests to improve throughput. Which consideration applies?

Answer and explanation

Answer: C. Batching improves throughput and cost efficiency but adds latency for individual requests. It does not reduce latency, does affect cost, and does not require related requests.

228. Predictable queries must be answered faster than the model can generate them. Which approach is appropriate?

Answer and explanation

Answer: C. Pre-computation serves known queries instantly. Throughput, length, and network do not eliminate generation time for known answers.

229. A hybrid search implementation must weight lexical and semantic results appropriately. Which approach is appropriate?

Answer and explanation

Answer: A. The right balance is corpus-dependent and must be measured. Fixed equal weights and single-method retrieval forgo the tuning.

230. Top-p and temperature parameters must be selected for a creative writing feature. Which configuration is appropriate?

Answer and explanation

Answer: A. Creative tasks benefit from sampling variation, tuned and validated. Zero temperature or top-p removes variation. Defaults without evaluation are unvalidated.

231. A GenAI workload's capacity must be planned in terms of token processing rather than request count. Which rationale applies?

Answer and explanation

Answer: A. Token throughput is what the model consumes and what varies. Request count is not proportional. Tokens are central to capacity. Per-user planning does not map to model capacity.

232. API call profiling reveals that most latency is in a post-processing step rather than the model. Which action is appropriate?

Answer and explanation

Answer: D. Profiling identified post-processing as the constraint. Model changes, prompt length, and throughput address the wrong stage.

233. A GenAI application's observability must cover business impact as well as operational health. Which metric belongs in the business view?

Answer and explanation

Answer: A. Conversion or resolution measures business outcome. Latency, errors, and tokens are operational.

234. Response drift must be detected in a deployed GenAI application. Which approach is appropriate?

Answer and explanation

Answer: A. Comparing fixed-input outputs over time detects drift. Volume, errors, and latency do not reveal changed output semantics.

235. Forensic traceability must allow reconstruction of exactly what a GenAI system did for a given request. Which data must be logged?

Answer and explanation

Answer: D. Full provenance requires every element correlated. Timestamp, response, or user alone cannot reconstruct the path.

236. Vector store index optimization must be automated to maintain performance as the corpus grows. Which approach is appropriate?

Answer and explanation

Answer: B. Scheduled maintenance triggered by metrics keeps the index healthy. Complaint-driven and never-optimize approaches degrade. Per-change rebuilds are wasteful.

237. A hallucination must be detected in production without a reference answer. Which technique applies?

Answer and explanation

Answer: A. Grounding checks and consistency analysis detect unsupported or unstable content. Length, grammar, and self-report do not detect hallucination.

238. A GenAI application's cost must be attributed to the features that drive it. Which approach is appropriate?

Answer and explanation

Answer: A. Per-invocation feature tagging attributes spend accurately. Equal division, totals, and frequency estimates all ignore that features differ in token consumption.

239. A GenAI application's cost must be forecast before a traffic increase. Which approach is appropriate?

Answer and explanation

Answer: D. Modelling invocations and tokens gives a defensible forecast. Scaling the bill assumes composition stays constant, list prices ignore actual token use, and user counts do not map to tokens.

240. A GenAI application must reduce cost for requests that repeat a long static preamble. Which approach is appropriate?

Answer and explanation

Answer: A. Prompt caching reuses a matching prefix, which requires the static content first. Shortening loses instruction, fewer requests reduce function, and a cheaper model may reduce quality.

241. A GenAI application's end-to-end latency must be reduced where retrieval dominates. Which approach is appropriate?

Answer and explanation

Answer: A. Optimizing the dominant stage is what reduces total latency. Generation changes, output limits, and concurrency address other stages or throughput.

242. A GenAI application must handle a burst without degrading latency for existing requests. Which approach is appropriate?

Answer and explanation

Answer: C. Bounding concurrency protects in-flight requests. Immediate processing of everything degrades all of them, longer timeouts hide the degradation, and shorter responses reduce quality for everyone.

243. A GenAI application's throughput must be increased without raising cost per request. Which approach is appropriate?

Answer and explanation

Answer: C. Batching raises throughput efficiently where latency allows. Provisioned throughput and more instances raise cost, and a larger model raises cost per request.

244. A GenAI application must reduce the time to first token for interactive use. Which approach is appropriate?

Answer and explanation

Answer: B. Streaming plus minimising pre-generation work reduces time to first token. Response length affects completion time, and more tokens and retrieval results add work.

245. A GenAI application's performance regression must be detected before users report it. Which approach is appropriate?

Answer and explanation

Answer: A. Continuous synthetic monitoring with percentile alarms detects regressions promptly. Daily review, monthly averages, and user reports all detect late.

246. A GenAI application must maintain quality while reducing its model cost. Which approach is appropriate?

Answer and explanation

Answer: C. Per-class routing validated by evaluation preserves quality at lower cost. The cheapest model degrades some classes, the most capable maximises cost, and random routing produces inconsistent quality.

247. A GenAI application's monitoring must distinguish a model problem from an application problem. Which approach is appropriate?

Answer and explanation

Answer: D. Per-stage instrumentation localises the fault. Total error rate, provider status, and request volume do not distinguish where the failure occurred.

248. A GenAI application's quality must be monitored continuously in production. Which approach is appropriate?

Answer and explanation

Answer: A. Sampling and scoring production interactions tracks quality as it actually is. Pre-deployment evaluation is a point in time, error rate covers failures, and complaints are partial and delayed.

249. A GenAI application's guardrail effectiveness must be monitored. Which approach is appropriate?

Answer and explanation

Answer: B. Tracking blocks and incorrect blocks makes both false positives and the overall rate visible. Counts of blocks or allowances alone show one direction, and latency is operational.

250. A GenAI application's token consumption must be monitored for anomalies. Which approach is appropriate?

Answer and explanation

Answer: A. Per-request token metrics with anomaly detection surface changes in consumption behaviour. Monthly bills are retrospective, fixed thresholds miss compositional change, and request counts do not track tokens.

251. A GenAI agent's behaviour must be monitored for unexpected tool usage. Which approach is appropriate?

Answer and explanation

Answer: D. Recording invocations with context and alarming on deviation detects unexpected behaviour. Totals lack context, weekly review is slow, and restricting to one tool removes capability.

252. A GenAI application's observability data must support investigating a specific user complaint. Which approach is appropriate?

Answer and explanation

Answer: C. Per-interaction records retrievable by a user-supplied identifier support investigation. Aggregates, short windows, and error-only retention all fail to cover the complaint.