Deployment and Orchestration of ML and AI Workflows

Domain 3: Deployment and Orchestration of ML and AI Workflows

90 practice questions for Domain 3 of the AWS Certified Machine Learning Engineer - Associate (MLA-C02) exam, which makes up 24% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 3: Deployment and Orchestration of ML and AI Workflows

24% of scored content · 90 practice questions

154. An inference workload receives sporadic requests with long idle periods, and a cold start of a few seconds is acceptable. Which solution meets these requirements at the lowest cost?

Answer and explanation

Answer: C. Serverless inference scales to zero between requests and bills per invocation, which suits sporadic traffic tolerating a brief cold start. An always-on endpoint pays continuously through the idle periods. A multi-model endpoint still runs provisioned instances. Hourly batch transform cannot answer requests on demand.

155. A request payload of 500 MB must be processed by a model, with the result returned when ready rather than synchronously. Which solution meets these requirements?

Answer and explanation

Answer: B. Asynchronous inference accepts large payloads, queues requests, and notifies on completion, which matches both the payload size and the deferred response. Real-time and serverless endpoints impose much smaller payload limits and are built for synchronous responses. A multi-model endpoint addresses hosting many models rather than large payloads.

156. A model trained outside AWS in a supported framework format must be served through Amazon Bedrock alongside the company's other foundation models. Which solution meets these requirements?

Answer and explanation

Answer: C. Custom Model Import brings a supported externally trained model into Bedrock so it is invoked through the same API and governed by the same guardrails and logging. A SageMaker endpoint serves the model through a different API. Retraining discards the existing model. A container on ECS reintroduces infrastructure management and a separate invocation path.

157. Hundreds of per-customer models must be hosted cost-effectively, with only a small subset active at any time. Which solution meets these requirements?

Answer and explanation

Answer: C. A multi-model endpoint keeps artifacts in S3 and loads them into memory as requested, sharing one fleet across many models. An endpoint per model multiplies fixed cost by the number of customers. Batch transform cannot serve interactive predictions. Combining customers into one model changes the modelling approach rather than solving the hosting problem.

158. A foundation model must serve predictable high-volume traffic with reserved capacity and consistent latency. Which Amazon Bedrock deployment option meets these requirements?

Answer and explanation

Answer: D. Provisioned throughput reserves model units for a workload, giving predictable capacity and latency for sustained high volume. On-demand invocation is subject to shared quotas and throttling. Batch inference is asynchronous and unsuitable for interactive traffic. Cross-Region inference raises available capacity without reserving it.

159. An agentic application must invoke internal APIs as part of completing a task, and the credentials for those APIs must never reach the model. Which solution meets these requirements?

Answer and explanation

Answer: A. Tool execution happens in the application's own code path, so the credential stays in the backing function and the model receives only the result. A credential in the prompt is exposed to prompt extraction. Session attributes are part of the model's context. Storing credentials in an action group definition places them where the invocation path can surface them.

160. A RAG application must be configured so that retrieval returns passages relevant to a question that combines two topics. Which configuration change is most likely to help?

Answer and explanation

Answer: C. A broader candidate set raises the chance that passages covering both topics are retrieved, and reranking promotes the most relevant of them. Retrieving fewer passages narrows coverage further. Temperature affects sampling rather than retrieval. Output length does not change what context the model receives.

161. A new model version must receive live production traffic in parallel with the current version, with its responses recorded but not returned to users. Which solution meets these requirements?

Answer and explanation

Answer: D. Shadow testing replicates live requests to a shadow variant and logs its responses without serving them, so the new version is exercised on real traffic with no user impact. Weighted variants and a canary both return responses to real users. Replaying historical requests does not exercise the live serving path.

162. An endpoint update must roll out a new model with automatic reversion if the error rate rises. Which solution meets these requirements?

Answer and explanation

Answer: C. SageMaker supports blue/green endpoint updates with canary or linear traffic shifting and alarm-based automatic rollback. Delete and recreate causes an outage and cannot revert quickly. An in-place update exposes all traffic at once. A second endpoint with client changes pushes deployment concerns into every caller.

163. An inference workload must serve two models that use different framework versions behind one endpoint, each invokable directly. Which solution meets these requirements?

Answer and explanation

Answer: A. Multi-container endpoints host different containers behind one endpoint, allowing distinct frameworks and direct invocation of a chosen container. Multi-model endpoints share a single container image and cannot host two framework versions. An inference pipeline chains containers in sequence rather than serving them independently. Two endpoints do not share the endpoint the requirement specifies.

164. A model must be deployed where the inference code is written in a framework SageMaker AI does not provide a container for. Which solution meets these requirements?

Answer and explanation

Answer: A. SageMaker supports custom containers that implement its serving interface, which allows any framework to be hosted with the platform's scaling and monitoring. Rewriting the inference code is a large unnecessary effort. Deploying outside SageMaker forfeits its managed capabilities. Format conversion may not be possible and changes the model.

165. An asynchronous inference endpoint must notify a downstream system when each result is ready. Which solution meets these requirements?

Answer and explanation

Answer: C. Asynchronous inference supports SNS notification topics for successful and failed inferences, which is the intended completion signal. An invocation alarm reports volume rather than individual completion. Polling adds latency and cost where a notification exists. An endpoint creation event fires once, not per inference.

166. A foundation model application must fail over to a second Region if the primary model endpoint becomes unavailable. Which solution meets these requirements?

Answer and explanation

Answer: D. Cross-Region inference or explicit secondary-Region routing sends requests elsewhere when the primary is unavailable. A response cache cannot answer new questions. More retries against an unavailable Region fail more slowly. Additional throughput in the failed Region does not help.

167. A batch transform job must associate each output record with the input record that produced it. Which solution meets these requirements?

Answer and explanation

Answer: B. Batch transform can join the source record to the prediction in the output, which makes the association explicit. Matching files by name pairs files rather than records. Sorting relies on order being preserved through parallel processing. Matching by position is fragile for the same reason.

168. A model artifact must be deployed to production as the exact artifact that passed evaluation. Which solution meets these requirements?

Answer and explanation

Answer: C. Deploying the registered package means production runs the exact bytes that were evaluated, which retraining cannot guarantee because runs differ. Rebuilding a container may resolve dependencies differently. Rerunning the training script has the same problem.

169. An endpoint must serve a model whose inference requires a GPU, while a preprocessing step in the same request needs only CPU. Which solution meets these requirements?

Answer and explanation

Answer: C. Separating the steps lets each run on hardware matched to its need, so expensive GPU capacity is not spent on CPU preprocessing. Combining them on a GPU instance wastes GPU time. Running the model on CPU may not meet latency. A larger GPU instance compounds the waste.

170. A model endpoint must be deployed into a VPC so it can reach a private database during inference. Which configuration is required?

Answer and explanation

Answer: B. Attaching the endpoint to subnets in the VPC gives it private addressing that the database's security group can permit. Public accessibility exposes the database. A NAT gateway routes to the internet rather than to a VPC resource. Interface endpoints reach AWS services rather than a customer database.

171. An endpoint must be updated with a new model while its existing instances continue serving, and capacity must never fall below the current level. Which solution meets these requirements?

Answer and explanation

Answer: B. A blue/green update provisions the replacement fleet before shifting traffic, so capacity never falls. In-place replacement reduces capacity while instances are cycled. Reducing the count first deliberately lowers capacity. Delete and recreate causes a complete outage.

172. A serverless inference endpoint's cold starts are affecting a small number of latency-sensitive requests. Which solution meets these requirements?

Answer and explanation

Answer: B. Provisioned concurrency keeps serverless inference environments initialized, which removes cold starts for the provisioned portion. More memory shortens but does not eliminate initialization. Asynchronous inference changes the response model. Higher maximum concurrency allows more parallel requests without pre-initializing them.

173. A foundation model application must choose between on-demand invocation and provisioned throughput. Which factor should decide?

Answer and explanation

Answer: A. Provisioned throughput reserves capacity billed continuously, so it pays off only when utilization is high and predictable. Context window is a property of the model rather than the purchasing option. SDK availability and Regional presence do not bear on the capacity decision.

174. An inference endpoint must scale to zero when idle, but the application cannot tolerate a cold start. Which resolution is appropriate?

Answer and explanation

Answer: D. Scaling to zero means no environment is warm, so a first request after idle necessarily pays initialization; the requirements cannot both hold and one must give. More memory shortens rather than removes the cold start. Real-time endpoints do not scale to zero. Queueing hides the latency from the caller only if the response can be deferred.

175. An endpoint must serve a model that occasionally receives a request requiring far more computation than typical. Which approach is appropriate?

Answer and explanation

Answer: A. Handling the outlier on a separate asynchronous path keeps the synchronous fleet sized for typical load. Sizing for the heaviest request pays for that capacity on every instance continuously. A longer timeout holds connections and threads. Rejecting heavy requests denies legitimate work.

176. An inference container must be validated before it is registered for deployment. Which approach is appropriate?

Answer and explanation

Answer: D. Testing the serving contract locally with representative payloads validates that the container will serve correctly before any deployment. Deploying to production to find out is the risk being avoided. A successful build says nothing about runtime behaviour. Artifact size confirms presence rather than correctness.

177. A foundation model deployment must keep inference data within a specific country for regulatory reasons. Which approach is appropriate?

Answer and explanation

Answer: C. Data residency is controlled by where the request is served, so the Region must be within the jurisdiction and the service's handling confirmed. Encryption protects confidentiality without controlling location. A proxy changes the network path rather than where inference occurs. A contract documents an obligation without technically enforcing it.

178. An endpoint's traffic follows a predictable daily curve with a fivefold difference between peak and trough. Which solution meets these requirements?

Answer and explanation

Answer: D. Target tracking adjusts instance count to load, and scheduled scaling pre-warms capacity before a known peak so scaling lag does not affect users. A peak-sized fleet pays for idle capacity through the trough. An average-sized fleet fails at peak. Serverless suits sporadic rather than sustained high-volume traffic.

179. A team must determine the instance type and count for a new endpoint based on measured performance rather than estimation. Which solution meets these requirements?

Answer and explanation

Answer: B. Inference Recommender benchmarks the model across instance types and reports latency, throughput, and cost so the endpoint is sized on evidence. Reusing a previous type assumes the models have similar profiles. Starting large and scaling down wastes spend during the trial. Observing in production experiments on users.

180. An Amazon Bedrock knowledge base must be configured so queries return only documents belonging to the requesting business unit. Which solution meets these requirements?

Answer and explanation

Answer: C. Filtering at retrieval means restricted documents never enter the model's context, which enforces the boundary before generation. A system prompt is an instruction the model may not follow. Filtering the answer happens after restricted content has already influenced it. Encryption at rest does not stop an indexed chunk being retrieved.

181. A retrieval pipeline must keep a vector index current as source documents are revised throughout the day. Which solution meets these requirements?

Answer and explanation

Answer: B. Event-driven incremental ingestion updates only what changed, keeping the index current at a cost proportional to the change rather than corpus size. A nightly rebuild leaves the index stale through the day and grows expensive. Retrieving more stale passages does not make them current. A higher threshold cannot distinguish stale content from fresh.

182. An agent must retain conversational context across a multi-turn interaction without the application resending the entire transcript each time. Which solution meets these requirements?

Answer and explanation

Answer: A. Session state associates a conversation with an identifier so context persists across turns without the caller resending everything. Foundation models retain nothing between requests. Resending the full transcript is exactly what the requirement excludes and eventually exceeds the context window. A static instruction prompt cannot hold a conversation that has not happened yet.

183. A training workload requires several GPU instances that must communicate with very low latency during distributed training. Which solution meets these requirements?

Answer and explanation

Answer: A. A cluster placement group packs instances close together within one Availability Zone and Elastic Fabric Adapter bypasses the operating system network stack, which together give the lowest inter-node latency. Spreading across zones adds latency. A spread placement group deliberately separates instances. A NAT gateway carries outbound internet traffic.

184. A model must run on constrained hardware where GPU memory is insufficient for the full-precision weights. Which solution meets these requirements?

Answer and explanation

Answer: A. Lowering weight precision reduces the memory footprint substantially with a modest accuracy cost, which is the standard remedy when weights do not fit. A larger batch size increases rather than decreases memory demand. Epoch count affects training rather than inference memory. Provisioned concurrency addresses startup latency.

185. A preprocessing step, the model, and a postprocessing step must run in sequence within a single endpoint invocation. Which solution meets these requirements?

Answer and explanation

Answer: B. An inference pipeline chains containers inside one endpoint so the sequence executes in a single invocation without network hops between stages. Three endpoints add latency and client orchestration. A Lambda wrapper places the processing outside the endpoint. A multi-model endpoint serves independent models rather than a sequence.

186. A training job requires more GPU memory than a single instance provides, and the model architecture must not change. Which solution meets these requirements?

Answer and explanation

Answer: D. Model parallelism shards the model across devices, which lowers peak memory per device while preserving the architecture. Epoch count does not change per-step memory. The optimizer choice has a modest effect and cannot resolve a fundamental shortfall. Validation data does not occupy training memory.

187. An endpoint's GPU utilization is low while its invocation latency is high, and the model is small relative to the instance. Which solution meets these requirements?

Answer and explanation

Answer: A. Low GPU utilization with a small model indicates over-provisioned hardware, so right-sizing or sharing the GPU across models removes the waste. More or larger instances compound the over-provisioning. Provisioned concurrency is a serverless and Lambda concept rather than an endpoint remedy here.

188. An agentic workflow must scale its tool-calling backend independently of the model invocation layer. Which solution meets these requirements?

Answer and explanation

Answer: B. Separating the tool backend lets it scale to its own load profile, which differs from model invocation. Implementing tools in the invoking process couples their scaling. Model throughput governs inference rather than tool execution. Co-locating tools with the vector index couples two unrelated workloads.

189. A retrieval pipeline must re-embed and update only the documents that changed, rather than rebuilding the whole index. Which solution meets these requirements?

Answer and explanation

Answer: D. Event-driven incremental ingestion updates only what changed, keeping cost proportional to change rather than corpus size. A nightly rebuild is expensive and leaves the index stale through the day. Retrieving more stale vectors does not make them current. A higher threshold cannot distinguish stale content from fresh.

190. An agent must resume a multi-step task after a transient tool failure without repeating the steps that already succeeded. Which solution meets these requirements?

Answer and explanation

Answer: C. Persisted state records which steps completed so a resumed run continues rather than restarting. A higher iteration limit repeats work already done. Temperature changes sampling rather than recovery. Removing tools reduces capability without addressing recovery.

191. A vector index must serve a RAG application with predictable latency as the embedding count grows from one million to fifty million. Which consideration should drive the configuration?

Answer and explanation

Answer: B. Approximate nearest neighbour indexes trade recall against latency through their parameters, so both must be measured at the target scale. Retrieving more candidates increases work per query. Reducing dimensionality without checking recall degrades retrieval silently. Storage class does not apply to a live search index.

192. A knowledge base ingestion job must parse PDFs containing tables before the content is chunked and embedded. Which solution meets these requirements?

Answer and explanation

Answer: C. Table structure must be extracted before chunking or the relationships between cells are lost, and a document extraction service preserves that structure. Embedding a binary produces no usable semantic representation. Byte-offset chunking splits content arbitrarily. Treating a text PDF as an image discards the text layer that is already present.

193. A training cluster must be provisioned so that a transient capacity shortage in one instance type does not prevent the job starting. Which solution meets these requirements?

Answer and explanation

Answer: D. Allowing alternative compatible types means a shortage in one pool does not block the job. A longer run time does not create capacity. A quota increase raises the account limit rather than the available supply. Fewer instances lengthens the job and may still hit the shortage.

194. An inference workload must run close to a factory floor where network connectivity to an AWS Region is intermittent. Which solution meets these requirements?

Answer and explanation

Answer: D. Intermittent connectivity requires inference to run locally with later synchronization, which is what local AWS infrastructure provides. A Regional endpoint of any kind is unreachable during an outage. A longer timeout does not help when the network is down.

195. An agentic application's infrastructure must isolate each tenant's tool executions from one another. Which solution meets these requirements?

Answer and explanation

Answer: D. Distinct roles and execution contexts mean a defect or injection in one tenant's path cannot reach another's resources, which is isolation by permission rather than convention. A shared process filtering by tenant relies on code correctness. A prompt identifier is an instruction the model may not honour. Prefix separation without a permission boundary relies on the tool behaving.

196. A batch inference workload processes a large backlog overnight and must not compete with the real-time endpoint for capacity. Which solution meets these requirements?

Answer and explanation

Answer: D. Batch transform provisions its own compute for the duration of the job, so the real-time endpoint is unaffected. Parallel invocation of the endpoint is exactly the contention being avoided. Adding capacity to the shared endpoint still mixes the workloads. A lower scaling target changes responsiveness rather than separating them.

197. A knowledge base must support both keyword matching on product codes and semantic matching on descriptions. Which solution meets these requirements?

Answer and explanation

Answer: C. Hybrid search combines exact lexical matching, which finds rare product codes, with vector similarity, which handles paraphrased descriptions. Vector search alone can miss exact rare tokens regardless of candidate count. Lexical search alone cannot match paraphrase. A higher threshold discards results without adding lexical matching.

198. A GPU training fleet must scale with queue depth so instances are not idle between jobs. Which solution meets these requirements?

Answer and explanation

Answer: B. Queue depth is the leading indicator of work waiting, so scaling against it adds capacity before jobs queue and removes it when the queue drains. A peak-sized fleet idles most of the time. An average-sized fleet builds a backlog at peak. GPU utilization is a lagging signal that stays high while the queue grows.

199. An agentic workflow's tool invocations must be rate limited so a downstream system is not overwhelmed. Which solution meets these requirements?

Answer and explanation

Answer: A. Limiting at the invocation layer with backpressure bounds the load regardless of how many agent runs are in flight. A lower iteration limit reduces capability without bounding concurrent load. Scaling the downstream system moves the limit without controlling the caller. A longer timeout tolerates queuing rather than preventing overload.

200. A retrieval pipeline's ingestion job must handle documents that fail parsing without stopping the whole run. Which solution meets these requirements?

Answer and explanation

Answer: C. Quarantining failures with their error keeps the run progressing while preserving the failed documents for investigation. Stopping on the first failure blocks the entire corpus. Silent skipping loses the documents and the signal. Indefinite retries block on a document that will never parse.

201. A vector index and the application querying it are in different VPCs, and traffic must not traverse the internet. Which solution meets these requirements?

Answer and explanation

Answer: B. An interface endpoint or PrivateLink connection provides private connectivity between the VPCs without an internet path. Public accessibility with address restriction still traverses the internet. A NAT gateway routes outbound to the internet. Public subnets expose both components.

202. A team must decide how many vector index replicas to provision for a RAG application's read traffic. Which approach is appropriate?

Answer and explanation

Answer: B. Replica count should follow measured latency under realistic load at the percentile the application targets. Matching replicas to application instances assumes a relationship that does not hold. Maximum replicas wastes capacity. Waiting for user complaints means the shortfall is discovered in production.

203. An agent must be prevented from consuming unbounded cost when a task cannot be completed. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: C, D. An iteration ceiling and a per-session budget both bound the work a single request can consume, and a defined fallback gives the application predictable behaviour. A longer timeout permits more spend. Higher temperature makes behaviour less predictable. More tools widen the search space the agent can loop through.

204. A knowledge base's documents must be re-chunked because retrieval quality is poor at the current chunk size. Which approach is appropriate?

Answer and explanation

Answer: B. Chunk size trades precision against context and the optimum is corpus-dependent, so it must be measured against labelled questions. Halving or doubling on principle assumes a direction without evidence. A published configuration was tuned for a different corpus.

205. A vector index's recall has degraded after the approximate nearest neighbour parameters were tuned for speed. Which approach is appropriate?

Answer and explanation

Answer: A. Approximate index parameters trade recall for speed explicitly, so the remedy is to relax them and verify recall against labelled queries. Retrieving more candidates from an index that no longer surfaces the right ones does not recover them. Higher dimensionality changes the embedding rather than the index behaviour. More replicas add throughput rather than recall.

206. An agent's tool definitions must be discoverable and consistent across several applications in the organization. Which approach is appropriate?

Answer and explanation

Answer: B. Centrally published versioned definitions give one source of truth that applications reference, so a change propagates deliberately. Per-application definitions diverge immediately. A wiki records intent without enforcing the schema. Copying guarantees drift as each copy is modified.

207. An ingestion pipeline must process documents of widely varying size without some batches taking far longer than others. Which approach is appropriate?

Answer and explanation

Answer: C. Batching by content size equalises the work per batch when document sizes vary widely. Counting documents produces batches whose work varies with the size mix. One document per batch removes batching efficiency entirely. Sorting changes the order rather than the imbalance within each batch.

208. A GPU instance quota prevents a training job from launching in the preferred Region. Which approach is appropriate?

Answer and explanation

Answer: A. A quota is an account limit that can be raised on request, and an alternative Region provides a path while the request is processed. Permanently reducing the instance count changes the job to fit an administrative limit. CPU instances may make the job infeasible. Retrying does not raise a quota.

209. A retrieval pipeline must preserve the ability to reconstruct which document version produced a given answer. Which approach is appropriate?

Answer and explanation

Answer: B. Version identifiers carried on each chunk and returned with the context make the specific source version reconstructable for any answer. An index-level timestamp describes the batch rather than the document. Retaining a previous index does not identify which version answered a given query. Chunk counts carry no provenance.

210. An agentic workflow must be prevented from invoking a tool with parameters outside an expected range. Which approach is appropriate?

Answer and explanation

Answer: B. Validation in the tool implementation rejects invalid parameters regardless of what the model produces, which is enforcement rather than instruction. A description guides the model without binding it. Temperature affects sampling unpredictably. Logging records the problem after the tool has acted.

211. A trained model must be promoted to a production endpoint through a governed pipeline. (Place the steps in the correct order.)

Answer and explanation

Answer: B → C → A → D. Evaluation must precede registration so the reviewer has metrics to judge. Registration with a pending status is what creates the gate; registering as approved would remove it. Approval is the human decision the governance requires. Deployment references the approved package, which guarantees production runs the exact artifact that was evaluated rather than a rebuild. Reordering evaluation after registration would mean approving a model whose performance is unknown.

212. A model must be retrained and redeployed only after a human reviews its evaluation metrics. Which solution meets these requirements?

Answer and explanation

Answer: B. Registering with a pending status creates an auditable human decision point inside an automated pipeline. Deploying first and notifying afterwards removes the gate. Scheduled retraining with automatic deployment does the same. A manual notebook run is neither repeatable nor auditable.

213. Prompts used by a production application must be versioned, tested, and updated without redeploying the application. Which solution meets these requirements?

Answer and explanation

Answer: A. Prompt Management versions prompts centrally and lets an application reference a published version, so a prompt change is a governed publish rather than a deployment. Embedded prompts require redeployment. An object read at startup requires a restart and carries no version history. Environment variables have the same problem and diverge per environment.

214. A change to a production prompt must be validated against a set of expected behaviours before it is published. Which solution meets these requirements?

Answer and explanation

Answer: D. Executing the candidate prompt against a curated test set is the only option that measures behaviour before users see it. Publishing and watching feedback exposes users to the regression. Reading the text does not reveal how the model responds to it. A traffic split limits blast radius but still ships an unvalidated prompt.

215. A RAG knowledge base must be refreshed on a defined cycle as source systems publish updates, and the refresh must be coordinated with downstream evaluation. Which solution meets these requirements?

Answer and explanation

Answer: D. Coordinating ingestion, refresh, and evaluation in one pipeline means the index is never advanced without its quality being checked. Evaluating monthly leaves regressions undetected between runs. Full re-ingestion scales with corpus size rather than change. Per-application copies diverge immediately.

216. An agent definition must be promoted through development, test, and production with the ability to roll back to a previous behaviour. Which solution meets these requirements?

Answer and explanation

Answer: B. Versioning with an alias means a specific tested definition is promoted and rollback is repointing the alias. Console edits per environment drift immediately. Recreating at release time deploys something that was never tested as a unit. A shared agent gives no isolation between environments.

217. A fine-tuned foundation model must be deployed automatically when a new version passes evaluation, with the deployed version recorded. Which solution meets these requirements?

Answer and explanation

Answer: A. Registering and deploying by reference guarantees production runs the exact evaluated artifact and leaves an auditable link between the two. Deploying immediately skips evaluation. A scheduled overwrite deploys whatever exists regardless of quality. Fine-tuning in place leaves no version history to roll back to.

218. An ML workflow must run processing, training, evaluation, and conditional registration, and start when new data lands in Amazon S3. Which solution meets these requirements?

Answer and explanation

Answer: A. Pipelines expresses the steps including a condition for registration, and EventBridge triggers it on arrival. A cron job reintroduces server management and provides no step-level lineage. Embedding everything in one job loses the conditional structure and the audit trail. Manual execution is neither triggered nor repeatable.

219. A SageMaker AI pipeline must rerun only the steps whose inputs have changed since the previous execution. Which solution meets these requirements?

Answer and explanation

Answer: D. Step caching compares a step's inputs and configuration against previous executions and reuses the earlier result when they match, which avoids repeating expensive work. Collapsing to one step removes the granularity caching depends on. A schedule controls when the pipeline runs rather than what it repeats. Checking for outputs manually reimplements caching with more failure modes.

220. A pipeline must halt before deployment when a newly trained model scores worse than the model currently in production. Which solution meets these requirements?

Answer and explanation

Answer: A. A condition step reads the evaluation result and branches, which is how a quality gate is expressed inside a pipeline. Logging for later review does not stop the deployment. Deploying first exposes users to the regression. Early stopping governs the training run rather than the promotion decision.

221. Model training code must be versioned so a given model artifact can be traced to the exact commit that produced it. Which solution meets these requirements?

Answer and explanation

Answer: C. Recording the commit identifier and propagating it to the registered package creates a traceable link from artifact back to source. A timestamp identifies when rather than what. An author name identifies who. Storing a copy of the script proves what ran but not which reviewed revision it corresponds to.

222. A retraining pipeline must start when Model Monitor reports drift rather than on a fixed schedule. Which solution meets these requirements?

Answer and explanation

Answer: C. Monitoring violations are published as events, so an EventBridge rule ties retraining to evidence with no polling. A weekly run retrains regardless of need and may still be too slow. An hourly query reimplements the event path with added latency. More frequent monitoring detects sooner without triggering anything.

223. A prompt change must be validated against a curated set of expected behaviours before it reaches production. Which solution meets these requirements?

Answer and explanation

Answer: B. Executing the candidate against a curated test set is the only option that measures behaviour before users see it. Publishing and watching feedback exposes users to the regression. Reading the text does not reveal how the model responds. A traffic split limits blast radius but still ships an unvalidated prompt.

224. An agent's action group definitions must be promoted from test to production with the ability to revert to the previous behaviour. Which solution meets these requirements?

Answer and explanation

Answer: A. Versioning with an alias means a specific tested definition is promoted and reverting is repointing the alias. Direct edits in production drift from what was tested. Delete and recreate causes an outage and discards history. A shared agent gives no isolation between environments.

225. A fine-tuned model must be redeployed automatically whenever a new version clears evaluation, with the deployed version recorded. Which solution meets these requirements?

Answer and explanation

Answer: A. Registering and deploying by reference guarantees production runs the exact evaluated artifact and leaves an auditable link. Deploying immediately skips evaluation. A nightly overwrite deploys whatever exists regardless of quality. Fine-tuning in place leaves no version to roll back to.

226. A knowledge base must be refreshed when source documents are republished, and retrieval quality must be checked before the refreshed index serves traffic. Which solution meets these requirements?

Answer and explanation

Answer: D. Coordinating ingestion, refresh, and evaluation in one pipeline means the index never advances without its quality being checked. Monthly evaluation leaves regressions undetected between runs. Refreshing first and evaluating afterwards serves a degraded index in the meantime. Per-application copies diverge immediately.

227. An ML pipeline defined in code must be deployed identically to development and production accounts. Which solution meets these requirements?

Answer and explanation

Answer: D. Infrastructure as code with CI/CD gives reproducibility, review, and identical deployment across accounts. Console creation drifts between environments. Export and import is not a reliable reproduction mechanism. A shared pipeline removes the separation the environments exist to provide.

228. An automated test suite for a generative application must detect when a model update changes behaviour on previously correct cases. Which solution meets these requirements?

Answer and explanation

Answer: B. A regression set run on every change is what detects behaviour that previously worked and no longer does. Testing only the new capability misses regressions elsewhere. Provider benchmarks describe the base model rather than this application. Higher temperature adds variance and makes results harder to compare.

229. A pipeline step must wait for a human approval recorded outside AWS before continuing. Which solution meets these requirements?

Answer and explanation

Answer: D. A callback step pauses the execution and resumes when the external system returns the token, which is the supported integration pattern. A fixed wait resumes whether or not approval happened. Splitting into two pipelines loses the single execution record. Polling from a step consumes compute while it waits.

230. A model package must be prevented from deploying to production unless its container image has passed a vulnerability scan. Which solution meets these requirements?

Answer and explanation

Answer: C. Gating promotion on scan findings stops a vulnerable image reaching production. Scanning afterwards means it is already running. Push restrictions are access control rather than content verification. A weekly meeting arrives after the deployment.

231. A pipeline execution must be reproducible months later, including the container image it used. Which solution meets these requirements?

Answer and explanation

Answer: A. A digest identifies exact image contents and cannot be reassigned, which is what makes the execution reproducible. A mutable tag may point to different contents later. A recorded tag has the same problem. Rebuilding from a Dockerfile may resolve dependencies differently.

232. A model deployment pipeline must run in a production account while the pipeline itself is defined in a tooling account. Which solution meets these requirements?

Answer and explanation

Answer: A. A cross-account role assumed by the pipeline issues temporary credentials scoped to deployment, which is the standard auditable pattern. Long-term keys are a standing credential to manage. Administrator access grants far more than deployment requires. Duplicating the pipeline doubles maintenance.

233. A pipeline must fail fast when a data quality check does not pass, before expensive training begins. Which solution meets these requirements?

Answer and explanation

Answer: D. Ordering the check before training means a failure stops the run before the expensive step consumes resources. Running in parallel starts training regardless. Checking afterwards pays the training cost first. A manual pre-check relies on someone remembering.

234. An automated evaluation must compare a new model against the one currently serving production traffic. Which solution meets these requirements?

Answer and explanation

Answer: D. Comparing against the production model's recorded metrics on the same held-out set answers whether the candidate is better than what is serving. A fixed threshold ignores how the production model has evolved. The previous training run may not be what is deployed. Deploying first exposes users to a possible regression.

235. A retraining pipeline occasionally produces a model worse than its predecessor, and the failure is discovered only after deployment. Which change prevents this?

Answer and explanation

Answer: C. A blocking comparison prevents a worse model reaching production at all. More frequent retraining produces more opportunities for the same failure. Email notification does not stop the deployment. A canary limits exposure but still deploys a model already known to be comparable offline.

236. A pipeline must be triggered when a new labelled dataset version is registered rather than on a schedule. Which solution meets these requirements?

Answer and explanation

Answer: B. An event emitted on registration triggers the pipeline exactly when new data is available with no polling. Hourly checking adds latency and wasted executions. Manual starts depend on people. A cron job timed to an expected window runs whether or not registration happened.

237. A generative application's prompt, model version, and guardrail configuration must be promoted together as one unit. Which solution meets these requirements?

Answer and explanation

Answer: C. Behaviour depends on the combination, so promoting them as one versioned unit means what was tested is what ships. Independent promotion creates untested combinations. A fixed promotion order still produces intermediate untested states. Always-latest in every environment removes the distinction between environments.

238. A pipeline's training step must use a specific framework version, and an upgrade silently changed the result. Which practice prevents recurrence?

Answer and explanation

Answer: C. Pinning the version and recording it makes the training environment part of the reproducible configuration and prevents silent change. Always-latest is what allowed the silent change. Retraining on every upgrade reacts rather than controls. Manual comparison depends on noticing.

239. A model registry entry must carry enough information for a reviewer to approve or reject it without consulting the training team. Which content is appropriate?

Answer and explanation

Answer: C. A reviewer needs the evidence and the context: how it performed, what produced it, where it fails, and what it is for. Artifact location and job name identify the object without describing it. Instance type and duration describe cost. The submitter's name identifies who rather than what.

240. A CI/CD pipeline must be prevented from deploying a model whose bias metrics exceed an agreed limit. Which solution meets these requirements?

Answer and explanation

Answer: D. A blocking evaluation step prevents a model exceeding the limit reaching production at all. Computing afterwards means it is already serving. Recording metrics for reviewers depends on someone acting on them. A quarterly sweep leaves months of exposure.

241. An orchestration must run a long training step and a short data validation step, and the validation result must be available before training starts. Which ordering is appropriate?

Answer and explanation

Answer: A. The requirement states validation must complete before training, so the steps are sequential with training conditional on the result. Parallel execution starts training regardless. Validating afterwards records a problem already paid for. Sampling in parallel still starts training on unvalidated data.

242. A generative application's evaluation suite must run against a specific model version rather than whichever version is current. Which approach is appropriate?

Answer and explanation

Answer: A. An explicit version identifier makes the evaluation reproducible and attributable to a specific model. A family name evaluates a moving target. Running immediately after release does not pin the version for later reproduction. Inferring a version from a date is fragile and breaks when releases overlap.

243. Two teams deploy models into the same account and have begun overwriting each other's endpoint configurations. Which solution meets these requirements?

Answer and explanation

Answer: C. IAM conditions on resource names or tags prevent one team modifying another's resources, which enforces the separation rather than relying on coordination. A shared calendar depends on discipline. A separate Region separates the teams but for the wrong reason and complicates operations. More frequent deployment increases the conflict rate.