Implementation and Integration

Domain 2: Implementation and Integration

75 practice questions for Domain 2 of the AWS Certified Generative AI Developer - Professional (AIP-C01) exam, which makes up 26% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 2: Implementation and Integration

26% of scored content · 75 practice questions

85. An agent must maintain conversation context across turns without the application resending the full transcript. Which solution meets these requirements?

Answer and explanation

Answer: D. Session state associates a conversation with an identifier so context persists without the caller resending everything. Resending the transcript grows cost and eventually exceeds the context window. Models retain nothing between requests. A static instruction cannot hold a conversation that has not happened yet.

86. An agent occasionally loops, calling the same tool repeatedly without converging on an answer. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: A, E. An iteration ceiling bounds cost and latency with defined behaviour, and a circuit breaker stops a repeatedly failing tool being called again within the same request. A longer timeout lets the loop continue. Higher temperature makes behaviour less predictable. More tools widen the space the agent can loop through.

87. A tool must be exposed to an agent with a standard interface so other agents can consume it without bespoke integration. Which solution meets these requirements?

Answer and explanation

Answer: C. A Model Context Protocol server exposes tools through a standard interface that any compatible client can discover and call, which is what removes bespoke integration. Implementing the tool inside each agent duplicates it. Documentation still leaves each team integrating separately. A bespoke API per consumer is the integration cost being avoided.

88. A workflow must pause for human review before an agent's proposed action is executed. Which solution meets these requirements?

Answer and explanation

Answer: A. A workflow task that waits for an approval callback blocks execution until a person responds, which is what a human gate requires. Acting first and notifying afterwards is not review. A prompt instruction cannot enforce a pause. Daily audit reviews actions already taken.

89. A complex request must be broken into reasoning steps with tool calls interleaved, and each step must be observable. Which solution meets these requirements?

Answer and explanation

Answer: B. An orchestrated workflow records each reasoning step and tool call as a discrete state with its input and output, which is what makes the loop observable. A single function hides the intermediate state. One response makes the steps unobservable and unbranchable. Logging only the final answer discards the path.

90. A GenAI feature has unpredictable, spiky traffic and a strict cost constraint. Which deployment approach is appropriate?

Answer and explanation

Answer: B. On-demand invocation bills per request, which suits unpredictable traffic under a cost constraint. Provisioning for the spike pays for reserved capacity that is mostly idle. A dedicated endpoint has the same problem. Provisioning for the average throttles during spikes while still paying continuously.

91. A workload mixes short classification requests with long generation requests, and cost must be reduced without harming quality. Which solution meets these requirements?

Answer and explanation

Answer: D. Model cascading matches capability to request type, and evaluating each class confirms the smaller model is adequate before the change ships. Routing everything to the smaller model degrades generation quality. Routing everything to the larger model is the cost being reduced. A uniform output cap truncates legitimate responses.

92. A GenAI capability must be consumed by many internal applications with consistent authentication, throttling, and logging. Which solution meets these requirements?

Answer and explanation

Answer: B. A central gateway applies the controls once regardless of the calling application, which is what consistency requires. Direct calls leave each application to implement the controls. A library can be bypassed by calling the API directly. Per-application copies of prompts and guardrails diverge immediately.

93. A GenAI component must be integrated into an existing order system without coupling the two release cycles. Which solution meets these requirements?

Answer and explanation

Answer: B. Event-driven integration lets each side deploy and scale independently and tolerates the other being unavailable. A synchronous call inside a transaction couples availability and latency. Embedding the component couples the release cycles directly. Polling the database couples the component to another system's schema.

94. A pipeline must deploy GenAI components with automated testing, security scanning, and the ability to roll back. Which solution meets these requirements?

Answer and explanation

Answer: C. Gating stages stop a regression or vulnerability progressing, and deploying by reference makes rollback immediate. Manual deployment after a meeting is slow and inconsistent. Ungated automatic deployment ships regressions. Deploying and monitoring means users encounter the problem.

95. A chat interface must display generated text as it is produced rather than after the full response completes. Which solution meets these requirements?

Answer and explanation

Answer: B. A streaming invocation returns tokens incrementally so the user sees output almost immediately. A larger token limit produces longer responses that take longer. Shorter responses reduce content rather than improving perceived responsiveness. Polling for a completed response delays display until generation finishes.

96. Model invocations intermittently fail with throttling during traffic spikes. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: C, E. Backoff with jitter spreads retries so clients do not converge on the same instant, and raising capacity absorbs the spike. A tight retry loop amplifies the overload. More output tokens increase the work per request. Disabling SDK retries removes the resilience that handles transient throttling.

97. Requests must be routed to different foundation models based on the content of each request. Which solution meets these requirements?

Answer and explanation

Answer: D. Content-based routing in an orchestration layer selects the model per request from configuration that can change without a code deployment. One model for everything is the situation being improved. Asking users to select a model exposes an implementation detail. Invoking every model multiplies cost.

98. An engineering team must reduce the time spent writing boilerplate integration code for a GenAI application. Which solution meets these requirements?

Answer and explanation

Answer: C. Amazon Q Developer generates and refactors code in the editor, and reviewing before committing keeps the engineer accountable for what ships. Copying and adapting propagates the original's defects. Writing from scratch is the effort being reduced. Waiting for a library blocks the work.

99. A document processing workflow must extract, classify, and summarize incoming files with minimal custom code. Which solution meets these requirements?

Answer and explanation

Answer: C. Bedrock Data Automation handles multimodal extraction and processing as a managed capability, which minimises the custom code. Chained Lambda functions are exactly the code being avoided. Asking a model to parse raw files directly is unreliable for scanned documents and tables. Manual processing does not scale.

100. An agent must decide between several specialised models depending on the subtask it is performing. Which approach is appropriate?

Answer and explanation

Answer: B. Routing based on measured capability per subtask matches the model to the work. The largest model for everything maximises cost. The cheapest degrades subtasks that need more. Leaving selection to the model's judgement introduces nondeterminism into a decision that should be governed.

101. A multi-agent system must prevent two agents from making conflicting changes to the same resource. Which approach is appropriate?

Answer and explanation

Answer: C. Conflicting concurrent changes require a serialization mechanism such as a single owner or a lock. A prompt instruction cannot coordinate independent processes. Running sequentially removes the parallelism the system exists for. A longer timeout permits retries without preventing the conflict.

102. An agent's tool must fail predictably when it receives a parameter it cannot process. Which approach is appropriate?

Answer and explanation

Answer: B. A structured error tells the agent what was wrong so it can correct the call, which is what makes recovery possible. Raw exception text is inconsistent and may leak internals. An empty result is indistinguishable from no data found. Retrying an invalid parameter fails identically.

103. A human review step must be inserted into an agentic workflow without blocking unrelated requests. Which approach is appropriate?

Answer and explanation

Answer: A. Suspending one execution while others proceed is what a callback pattern provides. Pausing everything blocks unrelated work. Polling in a loop consumes resources while waiting. Acting first and reversing later is not review.

104. An agent must combine outputs from several models into one answer. Which consideration applies?

Answer and explanation

Answer: D. Ensembling requires a defined rule for reconciling differing outputs, whether by voting, ranking, or precedence. First response measures speed. Concatenation presents contradictions to the user. Length is not a quality signal.

105. A generative workload must be deployed where GPU memory constrains how many concurrent requests each instance can serve. Which consideration applies?

Answer and explanation

Answer: A. Each in-flight request holds memory for its context in addition to the model weights, so both determine how many fit. vCPU count does not govern GPU memory. A longer timeout queues rather than adds capacity. Image size affects startup rather than concurrency.

106. A cascading deployment must route routine queries away from the largest model. Which step establishes that the cascade is safe?

Answer and explanation

Answer: D. Evaluating the smaller model on the specific class it will handle establishes that the cascade does not degrade those answers. A percentage split of all queries mixes the classes. Assuming routine queries are simple is the assumption being tested. Benchmarks describe other tasks.

107. A GenAI gateway must give the platform team visibility into which application made each model call. Which approach is appropriate?

Answer and explanation

Answer: B. Authenticating callers and recording identity per invocation attributes usage to an application. Hourly totals aggregate across callers. Self-reporting is inconsistent. The model identifier records what was called rather than by whom.

108. A GenAI capability must be integrated with an on-premises system that cannot be exposed to the internet. Which approach is appropriate?

Answer and explanation

Answer: B. Private connectivity keeps the integration off the internet while the model is invoked from within AWS. A public endpoint with IP restriction is still internet-exposed. A nightly copy serves batch use but not an integration. Running a foundation model on premises is a substantially different architecture than the requirement implies.

109. A CI/CD pipeline for a GenAI component must prevent a prompt or model change from degrading quality in production. Which stage is required?

Answer and explanation

Answer: A. A blocking evaluation against the production baseline prevents a regression reaching users. Reading a prompt does not reveal how the model responds to it. A canary limits exposure but still ships an unevaluated change. A documented rollback acts after the damage.

110. A client must handle a model response that is streamed incrementally and may end without completing. Which approach is appropriate?

Answer and explanation

Answer: D. A truncated stream is a failure, and distinguishing it from a complete response requires checking the completion signal. Assuming completeness presents truncated answers as whole. Buffering everything removes the benefit of streaming. Retrying on delay duplicates work that may still be in flight.

111. An application must observe latency across the retrieval, model invocation, and post-processing stages of each request. Which approach is appropriate?

Answer and explanation

Answer: A. Per-stage segments attribute latency so the slow stage is identified. A total duration shows that something is slow without saying what. The model's generation time covers one stage. Token counts describe volume.

112. An API fronting a foundation model must handle requests whose responses may exceed the client's expected size. Which approach is appropriate?

Answer and explanation

Answer: B. A token limit bounds the response and streaming delivers it incrementally, so neither side is surprised. A longer timeout accommodates size without bounding it. Silent truncation returns an incomplete answer presented as complete. Rejecting requests denies legitimate work.

113. A team must expose a GenAI capability to non-developers who need to compose their own workflows. Which approach is appropriate?

Answer and explanation

Answer: B. A visual builder lets non-developers compose steps without code, which is what the audience requires. API documentation, an SDK example, a command line tool, and direct API access all assume development capability.

114. An engineer must diagnose why a GenAI application returns inconsistent answers to the same question. Which source should be examined first?

Answer and explanation

Answer: D. Comparing the actual prompts sent establishes whether the inconsistency arises from differing input or from sampling, which determines the remedy. Infrastructure metrics and status pages describe availability. Error logs record failures rather than inconsistent successes.

115. An agent must persist its memory across sessions so a returning user is recognised. Which approach is appropriate?

Answer and explanation

Answer: D. Durable user-scoped storage persists across sessions. The context window resets. Process memory is lost on restart. Re-explaining is poor experience.

116. An agent's reasoning must be structured so each step is explicit and can be inspected. Which pattern applies?

Answer and explanation

Answer: B. The reason-act-observe loop makes each step explicit. A single prompt hides reasoning. Random tool selection is not reasoning. A fixed script is not agentic.

117. An agent must be prevented from taking an irreversible action without confirmation. Which safeguard is appropriate?

Answer and explanation

Answer: C. A confirmation gate on classified irreversible tools enforces the safeguard. An instruction is advisory. Read-only limits capability. Logging is after the fact.

118. An MCP server must be deployed for a tool with a long-running, stateful operation. Which hosting choice is appropriate?

Answer and explanation

Answer: B. ECS containers support long-running stateful operations. Lambda's duration and stateless model do not fit. S3 and CloudFront serve content rather than run tools.

119. An MCP server must be deployed for a lightweight stateless tool invoked occasionally. Which hosting choice is appropriate?

Answer and explanation

Answer: C. Lambda fits brief stateless occasional invocations at minimal cost. A continuous instance or service pays while idle. A stored procedure is not an MCP server.

120. A multi-agent system must coordinate a supervisor agent delegating to specialist agents. Which orchestration pattern applies?

Answer and explanation

Answer: D. Supervisor-specialist orchestration delegates and aggregates. Independent responses to everything are redundant. A single agent is not multi-agent. File-based communication is fragile.

121. A foundation model must be deployed for a workload where every invocation requires a GPU and traffic is continuous. Which deployment is appropriate?

Answer and explanation

Answer: A. Continuous GPU load suits a dedicated scaled endpoint. Lambda lacks GPU. Bedrock on-demand may suit but the question specifies a self-hosted model. Serverless inference does not support GPU workloads in the same way.

122. A large language model must be loaded onto an instance whose GPU memory is smaller than the model's full precision size. Which technique applies?

Answer and explanation

Answer: B. Quantization reduces the memory footprint. Loading half the model does not work. CPU memory does not substitute for GPU memory. Context window affects activation memory rather than weight size.

123. A GenAI feature must be embedded into a legacy application that exposes only a SOAP interface. Which integration is appropriate?

Answer and explanation

Answer: B. An adapter integrates without modifying the legacy system. Rewriting or replacing is disproportionate. Manual copying is not integration.

124. A GenAI capability must be consumed by teams in a subsidiary whose data cannot leave its jurisdiction. Which architecture is appropriate?

Answer and explanation

Answer: A. Residency requires processing within the jurisdiction. Sending data out, even encrypted or with deletion, violates the constraint.

125. A GenAI gateway must enforce a token budget per consuming application. Which implementation is appropriate?

Answer and explanation

Answer: D. Per-caller metering with enforcement applies the budget. Trust and combined budgets do not enforce per application. Monthly review is after the spend.

126. A batch of documents must be processed by a foundation model without holding an HTTP connection open for each. Which approach is appropriate?

Answer and explanation

Answer: D. Queue-based asynchronous processing decouples submission from completion. Synchronous loops and parallel open connections hold resources. Single-threading is slow.

127. An API client must handle a streaming response over a transport that does not support server push. Which approach is appropriate?

Answer and explanation

Answer: A. Chunked encoding or long polling delivers incrementally without server push. Waiting for the full response loses the streaming benefit. Model and length changes do not address transport.

128. A static routing configuration sends all requests to one model, and the team wants to route by request type without redeploying. Which approach is appropriate?

Answer and explanation

Answer: B. Configuration-driven routing changes without redeployment. Hardcoding requires redeployment. Separate applications multiply maintenance. User choice exposes implementation.

129. An API fronting a foundation model must handle a model timeout without leaving the client hanging. Which approach is appropriate?

Answer and explanation

Answer: C. A bounded timeout with a defined response keeps the client informed. No timeout hangs the client. An empty response is ambiguous. Indefinite retries hold resources.

130. A team must build a customer-facing UI for a GenAI feature quickly with authentication and hosting handled. Which approach is appropriate?

Answer and explanation

Answer: B. Amplify provides integrated authentication and hosting for rapid frontend development. From-scratch EC2 builds are slower. Model endpoints do not serve UIs. An API without a UI is not customer-facing.

131. A CRM must be enhanced with a GenAI feature that summarizes customer interactions. Which integration is appropriate?

Answer and explanation

Answer: A. Event-triggered generation written back through the API integrates seamlessly. Nightly export lags. Manual copying is poor workflow. Replacing the CRM is disproportionate.

132. A team must diagnose a GenAI-specific error pattern across many requests. Which approach is appropriate?

Answer and explanation

Answer: C. Aggregate log queries reveal frequency and triggers. Hand-reading does not scale. User reports are partial. Restarting does not diagnose.

133. An agent must be prevented from looping indefinitely on a task it cannot complete. Which approach is appropriate?

Answer and explanation

Answer: C. A step limit with a defined terminal outcome bounds the loop. Longer timeouts, unlimited execution, and restarts all permit or repeat the loop.

134. An agent's tool selection must be constrained to those appropriate for the current task. Which approach is appropriate?

Answer and explanation

Answer: D. Exposing only relevant tools removes inappropriate choices entirely. Instructions are advisory, post-hoc validation acts after the call, and ordering does not constrain.

135. An agent must hand off a task to a specialist agent and receive the result. Which approach is appropriate?

Answer and explanation

Answer: D. A structured request and response makes delegation reliable. Shared files are fragile, direct responses bypass the delegator, and concurrent work duplicates effort.

136. An agent's actions must be attributable to the user who initiated the task. Which approach is appropriate?

Answer and explanation

Answer: A. Propagating user identity attributes every action to its originator. Agent identity, no identity, and partial recording all break the audit trail.

137. An agent must be evaluated before being given permission to act autonomously. Which approach is appropriate?

Answer and explanation

Answer: C. Recording proposed actions without executing them evaluates judgement safely at scale. Full deployment risks real actions, definition review does not test behaviour, and a few examples are not representative.

138. A model endpoint must handle variable traffic without over-provisioning. Which approach is appropriate?

Answer and explanation

Answer: C. Auto scaling matches capacity to demand. Peak provisioning wastes capacity, minimum provisioning throttles, and on-demand invocation may not suit a self-hosted model.

139. A fine-tuned model must be deployed alongside the base model for comparison. Which approach is appropriate?

Answer and explanation

Answer: C. Traffic splitting with outcome recording compares them under identical conditions. Replacement removes the comparison, a separate application changes the population, and alternating confounds time with model.

140. A model's deployment must be rolled back quickly if quality degrades. Which approach is appropriate?

Answer and explanation

Answer: A. Retaining the previous version makes rollback a traffic shift. Redeploying and retraining both take time, and disabling the feature is an outage.

141. A GenAI capability must be exposed to several internal applications with consistent governance. Which approach is appropriate?

Answer and explanation

Answer: C. A shared service applies governance once for every consumer. Direct calls and shared credentials bypass governance, and per-application endpoints multiply both cost and policy surface.

142. A GenAI application must integrate with an enterprise identity provider. Which approach is appropriate?

Answer and explanation

Answer: D. Federated identity with derived authorization ties actions to real users. Local accounts duplicate the directory, shared accounts destroy attribution, and network location is not identity.

143. A GenAI feature must be added to an application without changing its existing release cycle. Which approach is appropriate?

Answer and explanation

Answer: D. A separate service releases independently. Embedding it couples the cycles, waiting delays delivery, and replacement is disproportionate.

144. Model invocations from several teams must be attributed for chargeback. Which approach is appropriate?

Answer and explanation

Answer: D. Per-caller token metering attributes cost accurately. Equal division ignores consumption, self-reporting is unreliable, and daily totals aggregate across teams.

145. A GenAI application must be deployed to several environments with environment-specific models. Which approach is appropriate?

Answer and explanation

Answer: C. Parameterization keeps one artifact across environments. Separate codebases and per-environment builds drift, and a single model may not suit every environment's cost or quality needs.

146. A GenAI service must enforce different rate limits for different consuming applications. Which approach is appropriate?

Answer and explanation

Answer: D. Per-caller quotas enforce differentiated limits. A single limit cannot differentiate, self-limiting is unenforced, and scaling absorbs cost rather than applying limits.

147. An API client must handle a model response that exceeds the client's processing time budget. Which approach is appropriate?

Answer and explanation

Answer: B. Incremental processing with a budget check bounds the client's exposure. Waiting and unlimited timeouts block, and retrying duplicates work already in flight.

148. A foundation model API's throttling responses must be handled correctly. Which approach is appropriate?

Answer and explanation

Answer: A. Backoff with jitter on throttling specifically is the correct handling. Immediate retries amplify, treating it as permanent loses recoverable requests, and increasing the rate worsens it.

149. An application must record the exact model version used for each response. Which approach is appropriate?

Answer and explanation

Answer: D. Capturing the version from the response records what actually served the request. Deployment configuration may differ, an unversioned name is ambiguous, and inference is unreliable.

150. An application must handle a foundation model API returning a content filter refusal. Which approach is appropriate?

Answer and explanation

Answer: B. Recognising a refusal and responding appropriately handles it correctly. Retrying unchanged produces the same refusal, treating it as an error misreports, and stripping the filter removes a safety control.

151. A GenAI feature must be integrated into an existing web application's user interface. Which approach is appropriate?

Answer and explanation

Answer: C. Streaming with clear AI attribution gives responsive feedback and appropriate transparency. Waiting, separate windows, and email all degrade the experience.

152. A GenAI application must let users correct an incorrect response. Which approach is appropriate?

Answer and explanation

Answer: C. Linked corrections feed evaluation and improvement. Resubmission does not capture the correction, unlinked logs cannot be diagnosed, and escalating everything does not scale.

153. A GenAI capability must be made available to business users through a workflow tool. Which approach is appropriate?

Answer and explanation

Answer: A. A defined workflow step lets business users compose without code. Documentation and direct access assume technical capability, and engineer-built workflows are a bottleneck.

154. An application must show users which sources a generated answer drew on. Which approach is appropriate?

Answer and explanation

Answer: C. Returning structured source references allows reliable display. Sources named within generated text may be fabricated, unlogged display is not possible, and a link to everything is not attribution.

155. A GenAI application's responses must be consistent across the web and mobile clients. Which approach is appropriate?

Answer and explanation

Answer: D. A shared backend produces identical behaviour for both clients. Per-client logic, different models, and separate caches all diverge.

156. An application must handle a user request that requires information the model cannot access. Which approach is appropriate?

Answer and explanation

Answer: A. Recognising the boundary and saying so is honest and actionable. Generating from training data risks fabrication, an unexplained error is unhelpful, and a larger model does not have the missing information.

157. A GenAI feature's rollout must be limited to a subset of users initially. Which approach is appropriate?

Answer and explanation

Answer: B. A per-user feature flag controls exposure precisely and allows quick reversal. Full deployment exposes everyone, a separate environment changes the population, and timing does not limit who is affected.

158. A GenAI application's errors must be distinguishable from its refusals in operational metrics. Which approach is appropriate?

Answer and explanation

Answer: B. Separate metrics let each be investigated appropriately, since a refusal is correct behaviour and an error is not. Conflating them obscures both, and totals show neither.

159. A GenAI application must degrade gracefully when its retrieval layer is unavailable. Which approach is appropriate?

Answer and explanation

Answer: C. Indicating the reduced grounding, or declining where grounding is essential, is honest degradation. Errors block all use, silent ungrounded generation misleads, and indefinite retries hold requests.