Operating, Monitoring, and Securing ML and AI Solutions

Domain 4: Operating, Monitoring, and Securing ML and AI Solutions

77 practice questions for Domain 4 of the AWS Certified Machine Learning Engineer - Associate (MLA-C02) exam, which makes up 24% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 4: Operating, Monitoring, and Securing ML and AI Solutions

24% of scored content · 77 practice questions

244. Several Amazon SageMaker capabilities must be matched to the purpose each one serves. (Match each capability to its purpose.)

  1. SageMaker Model Monitor
  2. SageMaker Clarify
  3. SageMaker Debugger
  4. SageMaker Model Cards
Answer and explanation

Answer: 1-D, 2-C, 3-B, 4-A. Model Monitor observes a deployed endpoint, while Debugger observes a training job, and the two are frequently confused because both watch a running process. Clarify serves both explainability and fairness and is the only one of these that computes per-prediction attributions. Model Cards produce documentation rather than measurement, which is why a governance requirement for documented intended use is not satisfied by Model Monitor, Clarify, or Debugger.

245. Case study. A bank has deployed a fraud detection model to a real-time SageMaker endpoint. Fraudulent transactions are roughly 0.4 percent of volume. Ground truth labels arrive from the investigations team about two weeks after each transaction. Over the past four months the proportion of transactions flagged has remained steady, but the investigations team reports that fewer flagged transactions are turning out to be fraudulent.

Which conclusion is best supported by the reported symptoms?

Answer and explanation

Answer: C. A steady flag rate with a falling true positive rate means the model still fires as often but is right less often, which is degraded precision caused by the underlying relationship changing. Under-provisioning affects latency rather than correctness. A model that never detected fraud would not have shown good precision for four months. A changed definition is possible but would be a labelling change the team would report, and the symptom described is a gradual decline rather than a step.

246. Case study. A bank has deployed a fraud detection model to a real-time SageMaker endpoint. Fraudulent transactions are roughly 0.4 percent of volume. Ground truth labels arrive from the investigations team about two weeks after each transaction. Over the past four months the proportion of transactions flagged has remained steady, but the investigations team reports that fewer flagged transactions are turning out to be fraudulent.

Which monitoring configuration would have detected this condition soonest?

Answer and explanation

Answer: C. Model quality monitoring is the only option that uses ground truth, which is what makes a precision decline visible. Data quality monitoring would catch a shift in the input distribution and is worth running, but a relationship change can occur while inputs look unchanged. Latency and invocation alarms report operational health and would have stayed green throughout.

247. Case study. A bank has deployed a fraud detection model to a real-time SageMaker endpoint. Fraudulent transactions are roughly 0.4 percent of volume. Ground truth labels arrive from the investigations team about two weeks after each transaction. Over the past four months the proportion of transactions flagged has remained steady, but the investigations team reports that fewer flagged transactions are turning out to be fraudulent.

Given the two-week delay before labels arrive, which retraining approach is most appropriate?

Answer and explanation

Answer: D. A rolling window incorporates the newest labelled examples so the model learns the changed relationship, and triggering on measured degradation ties retraining to evidence rather than the calendar. Retraining on the original dataset reproduces the stale model. Using the model's own predictions as labels reinforces its current errors. Waiting a year leaves a degrading model in production for far longer than the label delay requires.

248. An endpoint must be alerted when the statistical distribution of an input feature diverges from the training baseline. Which solution meets these requirements?

Answer and explanation

Answer: B. Model Monitor computes baseline statistics and constraints from training data and compares live traffic against them, alerting on drift. CPU alarms report infrastructure load. Config evaluates resource configuration. Debugger inspects training rather than production inputs.

249. A model's predictions must be compared against ground truth labels that arrive two weeks later. Which solution meets these requirements?

Answer and explanation

Answer: B. Model quality monitoring joins captured predictions with ground truth to compute accuracy metrics over time, which is the only option that uses labels. Data quality monitoring compares inputs without reference to outcomes. Feature attribution drift tracks which features drive predictions. Latency is an operational measure.

250. An agentic workflow intermittently returns incomplete responses, and the team must determine whether the cause is a tool failure or a truncated model stream. Which solution meets these requirements?

Answer and explanation

Answer: B. Per-step spans attribute an incomplete response to a specific tool failure or a truncated generation, which is the distinction required. A higher iteration limit allows more attempts without revealing the cause. Temperature changes sampling. A longer timeout tolerates slowness rather than identifying the failing step.

251. A team must monitor the output quality of a deployed foundation model application on a recurring basis. Which solution meets these requirements?

Answer and explanation

Answer: A. Bedrock evaluations score model output against a dataset and rubric, and running them on a schedule turns quality into a tracked metric. Invocation counts measure volume. Log volume detects errors rather than quality degradation. Reactive feedback review detects problems only after users report them.

252. An endpoint shows rising latency while invocation volume is unchanged and instance CPU is moderate. Which diagnostic step should be taken next?

Answer and explanation

Answer: D. SageMaker publishes ModelLatency and OverheadLatency separately, which distinguishes time spent in the model container from time in the serving infrastructure and directs the investigation. Training loss describes fit rather than runtime. Storage class affects artifact retrieval at startup rather than steady-state latency. Feature count is fixed and would not cause a rise.

253. Model Monitor must compare live endpoint traffic against a baseline, and no baseline currently exists. Which step must be performed first?

Answer and explanation

Answer: B. Model Monitor compares live data against a baseline computed from a reference dataset, so the baselining job must run before a schedule can evaluate anything. Data capture is also required but does not produce a baseline. A schedule does not infer its own baseline. An invocation alarm measures volume.

254. A monitoring schedule reports violations on a feature whose distribution legitimately changed after a product launch. Which action is appropriate?

Answer and explanation

Answer: C. A legitimate distribution change makes the old baseline wrong, so it should be recomputed against current data. Disabling monitoring removes detection entirely. Raising the threshold conceals future genuine drift in the same feature. Removing a useful feature degrades the model to silence an alert.

255. A team must detect when the features driving a model's predictions change, even though the input distributions remain stable. Which monitoring type applies?

Answer and explanation

Answer: C. Feature attribution drift tracks which features drive predictions and detects a shift the input distributions alone would not reveal. Data quality monitoring compares those distributions. Model quality requires ground truth. Latency is operational.

256. Captured inference data must be retained for monitoring but the endpoint handles far more traffic than monitoring requires. Which solution meets these requirements?

Answer and explanation

Answer: C. Data capture supports a sampling percentage, and monitoring needs a statistically sufficient sample rather than every request. Capturing everything and deleting later pays the storage and processing cost first. Application logs are not in the format Model Monitor consumes. Reducing traffic changes the workload to suit the tooling.

257. An agent's coordination between tools must be monitored so that repeated failures of one tool are detected before users are affected. Which solution meets these requirements?

Answer and explanation

Answer: B. Per-tool metrics localise a failing tool and allow an alarm before the aggregate error rate moves enough to notice. An overall error rate can stay within tolerance while one tool fails consistently. Weekly transcript review is slow. More retries mask the failure and increase latency.

258. An endpoint's monitoring must distinguish a genuine quality decline from normal weekly variation. Which approach is appropriate?

Answer and explanation

Answer: A. A band accounting for seasonality distinguishes expected weekly movement from a genuine decline. A threshold at the worst observed value fires almost never. Day-on-day comparison fires on every normal weekly cycle. Weekly manual review is slow and subjective.

259. A generative application must detect when its responses begin drifting from the tone and structure established at launch. Which approach is appropriate?

Answer and explanation

Answer: A. Scoring a recurring sample against the original rubric measures the quality that is claimed to be drifting. Response length is a weak proxy that can stay constant while tone changes. Invocation count and latency are operational. Monthly complaint review detects drift long after users experienced it.

260. A team must be alerted when an endpoint begins returning errors for a specific model variant while others are healthy. Which solution meets these requirements?

Answer and explanation

Answer: B. Per-variant metrics isolate a failing variant that an aggregate rate would dilute. An aggregate error rate can stay within tolerance while one variant fails entirely. Invocation count measures volume. Weekly configuration review does not detect runtime errors.

261. Captured inference data shows that a feature's values are increasingly arriving as nulls. Which conclusion follows?

Answer and explanation

Answer: D. Rising nulls in an input feature is an upstream data problem rather than a model problem, so the pipeline that produces the feature is where to look. Overfitting affects predictions rather than input completeness. Endpoints do not drop fields under load. Regenerating the baseline would normalise the fault rather than diagnose it.

262. An endpoint's monitoring reports must be retained for audit while the captured payloads themselves need not be. Which approach is appropriate?

Answer and explanation

Answer: C. Reports carry the audit evidence while payloads are the raw input, so separate retention periods match their different values. Retaining both for the full period pays to store data the audit does not require. Disabling capture prevents monitoring entirely. Regenerating reports requires the payloads to still exist.

263. A model's monitoring must cover a feature that is computed at inference time rather than supplied in the request. Which approach is appropriate?

Answer and explanation

Answer: B. A feature computed inside the endpoint is invisible to monitoring unless it is emitted, so emitting it makes it observable. Inferring its behaviour from raw fields assumes the computation is correct, which is what monitoring should verify. Recomputing offline reproduces the computation rather than observing what the model received. Excluding it leaves a blind spot in exactly the place where a computation bug would hide.

264. Inference costs for a production generative application are rising faster than request volume. Which cause should the team investigate first?

Answer and explanation

Answer: C. Generative inference is billed on tokens, so cost outpacing request volume means each request has grown more expensive, most often through prompt or context growth. User counts drive request volume rather than cost per request. Storage class and zone count are minor contributors.

265. An assistant answers many near-identical questions each day, and the team must reduce inference cost. Which solution meets these requirements?

Answer and explanation

Answer: D. A semantic cache eliminates inference entirely for repeated questions, which is the largest available saving when queries repeat. A higher token limit increases cost per request. Routing everything to the largest model maximises cost. Varying responses removes the repetition that makes caching effective.

266. A workload sends both simple lookups and complex reasoning requests to the same large foundation model. Which solution reduces cost without harming quality?

Answer and explanation

Answer: A. Model routing matches capability to task difficulty, and evaluation confirms the smaller model is adequate for the simple class before the change ships. Reducing context indiscriminately drops information complex requests need. Temperature affects sampling rather than cost. Truncating prompts removes context and degrades quality.

267. A long shared instruction block is resent on every request in a high-volume application. Which solution reduces both cost and latency?

Answer and explanation

Answer: B. Prompt caching reuses the processed representation of a static prefix across requests, cutting input token cost and time to first token. Compression does not reduce the tokens the model must process once decompressed. Output token limits do not affect input cost. Omitting the instruction on later requests removes it entirely, since the model is stateless.

268. A vector index backing a RAG application has grown large and storage costs are significant. Which solution reduces cost while preserving retrieval quality?

Answer and explanation

Answer: A. Storage scales with embedding dimension and index parameters, so tuning them against measured quality is the lever that reduces cost without blind loss. Deleting documents by age removes content that may still be needed. Archival storage cannot serve a live index. Retrieving fewer candidates reduces query cost rather than storage.

269. A batch scoring job runs nightly on Amazon Bedrock and can tolerate a delayed result. Which solution meets these requirements at the lowest cost?

Answer and explanation

Answer: B. Batch inference is designed for large asynchronous workloads and is priced below real-time invocation, which fits a nightly window with no latency constraint. Parallel real-time workers pay real-time pricing and risk throttling. Provisioned throughput reserves capacity continuously for work that runs once a day. Streaming changes delivery rather than pricing.

270. A training workload runs for several hours and tolerates interruption. Which solution reduces cost?

Answer and explanation

Answer: C. Managed spot training uses spare capacity at a substantial discount and resumes from checkpoints after interruption, which matches an interruption-tolerant job. A Savings Plan discounts committed usage rather than targeting interruption tolerance. Serverless inference is a serving option. A larger instance increases hourly cost.

271. An endpoint serving steady production traffic must be right-sized against observed utilization. Which solution meets these requirements?

Answer and explanation

Answer: B. Matching instance size to measured utilization while holding the latency target removes waste without risking performance. A low scaling target adds instances rather than reducing cost. Asynchronous inference changes the response model and is unsuitable for latency-sensitive traffic. A larger instance costs more while making utilization look better.

272. An embedding pipeline recomputes vectors for the entire corpus on every run, and its cost is growing with corpus size. Which solution meets these requirements?

Answer and explanation

Answer: A. Embedding cost scales with the number of documents processed, so processing only what changed makes cost proportional to change rather than corpus size. A smaller dimension reduces cost per document without stopping the full recomputation. Running less often leaves the index staler. A larger instance completes the same waste faster.

273. An agent makes several tool calls per request, and the team must identify which tool drives most of the cost. Which solution meets these requirements?

Answer and explanation

Answer: D. Attributing cost to a tool requires per-tool measurement, which trace attributes or metric dimensions provide. A service-level bill aggregates across tools. Latency and request counts describe volume and speed rather than cost distribution.

274. A vector store's cost is dominated by storage as the corpus grows, and retrieval quality must be preserved. Which solution should be evaluated first?

Answer and explanation

Answer: D. Storage scales directly with embedding dimension, so reducing it while measuring quality is the lever that cuts cost without blind loss. Deleting by age removes content that may still be needed. Fewer candidates reduces query cost rather than storage. Archival storage cannot serve a live index.

275. A generative feature's cost per successful business outcome must be tracked so optimizations do not degrade quality unnoticed. Which measure should be adopted?

Answer and explanation

Answer: B. Cost per successful outcome ties spend to value, so an optimization that saves money by failing more often is immediately visible. Total spend, invocation counts, and tokens per call are inputs that can all improve while the feature becomes less useful.

276. A team must reduce the inference cost of a workload where most requests are simple and a minority require complex reasoning. Which solution meets these requirements?

Answer and explanation

Answer: B. Routing by request difficulty matches capability to need, and validating the smaller model on the simple class prevents shipping a quality regression. Routing everything to the smallest model degrades the complex minority. A uniform output cap truncates legitimate answers. Caching complex responses helps only when they repeat.

277. A training workload's cost must be reduced, and the team can tolerate a longer wall-clock time. Which solution meets these requirements?

Answer and explanation

Answer: A. Spot capacity offers the largest discount for interruption-tolerant work, and tolerating a longer wall-clock time is exactly what makes it viable. A larger instance costs more per hour. A Savings Plan sized for peak commits to capacity used intermittently. A smaller validation split saves little and weakens evaluation.

278. An endpoint serving a steady baseline with occasional spikes must minimize cost without failing during the spikes. Which solution meets these requirements?

Answer and explanation

Answer: A. Baseline sizing with auto scaling pays for the steady load and adds capacity only when needed. Peak sizing pays for idle capacity continuously. Sizing to an average fails during spikes and wastes capacity otherwise. Asynchronous inference changes the response model and may not suit a latency-sensitive endpoint.

279. A RAG application's prompt has grown because the team increased the number of retrieved passages, and cost has risen accordingly. Which action should be taken first?

Answer and explanation

Answer: D. Passage count is the variable that grew, so measuring quality against it identifies how much context is actually needed. Output limits address a different part of the cost. A larger context window permits more tokens rather than fewer. Raising the threshold reduces passages without checking whether quality survives.

280. A team must compare the cost of self-hosting a model on an endpoint against invoking a managed foundation model. Which comparison is appropriate?

Answer and explanation

Answer: D. A provisioned endpoint is billed by time regardless of use while a managed model is billed by tokens, so the comparison only holds when both are expressed at expected utilization and volume. Comparing raw prices ignores utilization entirely. Benchmarks describe quality. Artifact storage is negligible in both cases.

281. Model artifacts and training datasets accumulate in Amazon S3 and are rarely accessed after a model is retired. Which solution meets these requirements?

Answer and explanation

Answer: A. Lifecycle transitions reduce storage cost while retaining the data needed to reproduce a retired model. Deletion makes the model unreproducible for audit. Instance storage is ephemeral. Compression helps marginally without addressing the storage class.

282. A development team leaves SageMaker AI notebook instances running overnight and at weekends. Which solution meets these requirements?

Answer and explanation

Answer: C. Automated stopping removes the cost of idle hours without depending on people. A request relies on everyone remembering every day. Smaller instances still pay for idle hours. Delete and recreate loses local state and installed packages.

283. A model endpoint's cost is acceptable but a second endpoint serving a rarely used model costs nearly as much. Which solution meets these requirements?

Answer and explanation

Answer: B. A multi-model endpoint loads the rarely used model on demand and shares one fleet, which removes its dedicated always-on cost. A smaller instance still runs continuously. Deleting and redeploying on request introduces minutes of latency. A scaling target cannot take a real-time endpoint below one instance.

284. An architecture review must determine whether a self-managed vector database or a managed service is cheaper for a growing corpus. Which comparison is appropriate?

Answer and explanation

Answer: A. The comparison must include operations and expertise, which is where self-managed options often cost more than the infrastructure line suggests. Comparing raw prices omits that. Maximum index size and latency are capability measures rather than cost.

285. Training costs are dominated by repeated hyperparameter tuning runs over the same dataset. Which solution meets these requirements?

Answer and explanation

Answer: A. Warm start seeds a tuning job with prior results, so fewer jobs are needed to reach a good configuration. More parallel jobs finish sooner at the same or higher total cost. Tuning one hyperparameter may miss the best configuration. More jobs increases cost.

286. A cost review must establish which of several ML workloads is growing fastest. Which approach is appropriate?

Answer and explanation

Answer: C. Tag-based attribution compared across periods shows which workload's spend is growing. Account totals aggregate every workload. Job counts ignore job size and duration. Instance types describe configuration rather than spend trajectory.

287. An inference workload's latency requirement can be met by two instance families at different prices. Which approach is appropriate?

Answer and explanation

Answer: B. Cost per inference at the required throughput is the comparison that matters, since a more expensive instance may serve enough more requests to cost less per inference. Hourly price ignores throughput. vCPU count and release recency do not determine cost efficiency for a specific model.

288. A team must decide whether to keep a rarely used model deployed or redeploy it on demand. Which comparison is appropriate?

Answer and explanation

Answer: B. The decision trades continuous idle cost against the delay of an on-demand deployment, so both must be quantified against what the use case tolerates. Artifact size and memory affect feasibility rather than the decision. Accuracy comparison against another model is unrelated. User count describes demand without establishing the tolerance for delay.

289. A training pipeline's cost has risen although the dataset and model are unchanged. Which cause should be investigated first?

Answer and explanation

Answer: B. Cost rising with unchanged inputs points to work being repeated that was previously skipped, which caching invalidation causes. Accuracy is a quality measure rather than a cost driver. Storage class affects a small fraction of pipeline cost. Fewer executions would reduce cost rather than raise it.

290. An agentic application's cost varies widely between requests, and the team must set a defensible budget. Which approach is appropriate?

Answer and explanation

Answer: C. A wide distribution makes the mean a poor planning figure, so budgeting against a percentile captures the tail that actually drives spend. The mean understates when the distribution is skewed. The cheapest request understates badly. The most expensive overstates and wastes budget.

291. A cost optimization must be validated as not having degraded model quality. Which approach is appropriate?

Answer and explanation

Answer: C. A cost change can degrade quality, so the evaluation suite must be run on both sides and compared alongside the saving. Waiting for complaints detects degradation after users experience it. Latency and successful responses confirm the system works rather than that it works as well.

292. Several teams share a foundation model and the organization must attribute inference cost to each. Which approach is appropriate?

Answer and explanation

Answer: C. Attribution requires usage to be identifiable at invocation, through tagging or separate profiles. Equal division ignores actual consumption. Headcount is unrelated to usage. Manual log review is slow and does not produce a billing attribution.

293. A model's inference cost must be reduced, and the team is considering quantization. Which consideration applies?

Answer and explanation

Answer: C. Quantization's accuracy cost varies by model and task, so it must be measured on the evaluation set before adoption. Assuming negligible impact is exactly the assumption that fails on some models. A parameter threshold is an arbitrary rule. Monitoring afterwards means degraded predictions have already been served.

294. An application must be prevented from returning content that violates the company's responsible AI policy and must not disclose personal data in its responses. Which solution meets these requirements?

Answer and explanation

Answer: B. A guardrail is evaluated by the service on both request and response and blocks or masks regardless of whether the model cooperates. A system prompt is advisory and is the first thing a prompt injection targets. Weekly sampling detects violations after users have seen them. A shorter output limit truncates content without filtering it.

295. An application running on a compute service that already assumes an IAM role must authenticate to Amazon Bedrock. Which credential approach should be used?

Answer and explanation

Answer: B. When a compute service already assumes a role, that role supplies automatically rotated temporary credentials with no secret to store. A Bedrock API key is useful where an IAM role is unavailable but introduces a credential to manage here. Long-term user keys in environment variables leak through logs. Root credentials must never be used programmatically.

296. A training job must be unable to make any outbound network call, including to package repositories. Which solution meets these requirements?

Answer and explanation

Answer: D. Network isolation prevents the training container making any outbound connection, so all dependencies must be present in the image. A NAT gateway provides the internet access being prevented. A security group denying egress is close but network isolation is the purpose-built control and also blocks the container from the metadata service. An S3 endpoint permits S3 traffic rather than blocking egress.

297. Model artifacts in Amazon S3 must be encrypted with a key the ML team controls, and every use of that key must be auditable. Which solution meets these requirements?

Answer and explanation

Answer: B. A customer managed key gives the team an author-controlled key policy and CloudTrail records of every use, and the SageMaker roles need decrypt permission in that policy. S3 managed keys expose no key policy or usage trail. A key in a script is neither protected nor rotatable. Block Public Access prevents public exposure without encrypting.

298. A CI/CD pipeline that builds ML container images must be checked for code and image vulnerabilities before deployment. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: C, E. Scanning the code and the resulting image, with gates on both, stops vulnerable artifacts progressing before any deployment. Manual review does not scale and misses dependency vulnerabilities. Runtime monitoring detects behaviour after the image is running. Push restrictions are access control rather than vulnerability detection.

299. A SageMaker AI endpoint must be reachable only from within the company's VPC, with no public path. Which solution meets these requirements?

Answer and explanation

Answer: B. An interface endpoint backed by PrivateLink places a network interface in the subnet so calls to the runtime never traverse the internet, and an endpoint policy narrows what may be invoked. Gateway endpoints exist only for Amazon S3 and DynamoDB. SageMaker endpoints are managed and are not deployed into customer subnets in that way. A NAT gateway routes traffic to the internet.

300. An execution role for a training job currently allows s3:* on all buckets. Which change applies least privilege?

Answer and explanation

Answer: C. Least privilege means naming the actions the job performs and the exact prefixes it reads and writes. A boundary permitting the same broad action narrows nothing. CloudTrail records activity without restricting it. A VPC endpoint changes the network path rather than the permissions.

301. A RAG application must prevent a retrieved document containing injected instructions from causing the model to take an unintended action. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: A, C. Indirect prompt injection is mitigated by separating untrusted retrieved content from trusted instructions and by limiting what any successful injection could actually do. A prompt instruction is itself part of the context the injection targets. Temperature adds randomness rather than a control. Removing the system prompt discards the instruction that shapes safe behaviour.

302. A model endpoint must be invocable only by a specific application role, and every invocation must be attributable. Which solution meets these requirements?

Answer and explanation

Answer: D. An IAM policy scoped to the endpoint ARN restricts who may invoke it, and CloudTrail records the calling identity for attribution. Network controls restrict reachability without identifying the caller. SageMaker endpoints do not use API keys for authorization. Data capture records payloads rather than enforcing access.

303. Training data containing personal information must remain encrypted with a key the ML team controls throughout training. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, D. Encrypting the data at rest with a customer managed key and encrypting the job's storage volumes and inter-container traffic with the same key protects the data throughout training. Network isolation is valuable but addresses egress rather than encryption. Changing the storage service does not by itself encrypt anything. A shorter run time does not protect the data.

304. A generative application must prevent users from extracting the personal data held in its retrieval corpus through crafted questions. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, D. Removing personal data before indexing means it cannot be retrieved at all, and a sensitive information guardrail catches anything that remains. A prompt instruction is advisory. Retrieving fewer passages reduces but does not eliminate exposure. Weekly review detects disclosures after users have seen them.

305. A CI/CD pipeline building ML container images must detect vulnerable dependencies before the image is promoted. Which solution meets these requirements?

Answer and explanation

Answer: D. Scanning in the pipeline with a gate stops a vulnerable image progressing. Scanning after deployment means it is already running. Promotion restrictions are access control rather than content verification. Weekly rebuilds refresh dependencies without detecting what is vulnerable.

306. An application invoking a foundation model from outside AWS must authenticate without a long-lived AWS access key. Which solution meets these requirements?

Answer and explanation

Answer: C. Federation exchanges an external identity or certificate for temporary credentials, removing the long-lived key. An IAM user with stored keys is exactly the credential being avoided. Sharing keys between workloads destroys attribution. Root credentials must never be used programmatically.

307. A guardrail must be applied consistently to every invocation of a foundation model, and developers must not be able to bypass it by calling the model directly. Which solution meets these requirements?

Answer and explanation

Answer: D. An IAM condition on the guardrail identifier is evaluated by the service, so an invocation without it is denied regardless of the calling code. A published standard relies on compliance. A shared library can be bypassed by calling the API directly. Weekly log review detects bypasses after they have returned output.

308. A SageMaker Studio environment must be prevented from exfiltrating data to an external Amazon S3 bucket. Which solution meets these requirements?

Answer and explanation

Answer: A. VPC-only mode with an endpoint policy conditioned on the organization restricts which buckets can be reached at all, which blocks writes to an external bucket. A NAT gateway provides the internet path being prevented. A read-only role on internal buckets says nothing about external ones. Data capture records activity rather than preventing it.

309. A model endpoint's inference payloads must be encrypted in transit between the client and the endpoint. Which statement is correct?

Answer and explanation

Answer: D. SageMaker endpoints are reached over TLS by default, and the practical risk is a client that disables certificate verification. No separate in-transit setting is required. VPC placement addresses network reachability. A KMS key protects data at rest.

310. An ML team must be prevented from launching training jobs outside the approved Regions. Which solution meets these requirements?

Answer and explanation

Answer: B. A service control policy caps permissions for every principal in the account and cannot be overridden locally. An IAM policy can be modified by an account administrator. A Config rule reports after the job has run. Documentation is not an enforced control.

311. A foundation model application must record every prompt and completion for a compliance review. Which solution meets these requirements?

Answer and explanation

Answer: A. Model invocation logging captures the full request and response payload with the model identifier, which is what a review of prompts and completions requires. CloudTrail management events record that an API was called without the payload. Request identifiers omit the content. Flow Logs capture network metadata.

312. A notebook environment must be prevented from installing packages from public repositories during a regulated project. Which solution meets these requirements?

Answer and explanation

Answer: C. Removing internet access and providing a curated internal repository makes public installation impossible while keeping the environment usable. An instruction relies on compliance. Weekly scanning detects after the fact. Restricting S3 writes addresses a different risk entirely.

313. An inference request may contain personal data that must not be written to the endpoint's captured data. Which solution meets these requirements?

Answer and explanation

Answer: C. Not capturing the data is the only approach that prevents it being written. Encryption protects the captured data from outside readers while it is still stored and readable by authorised principals. A lower sampling rate reduces volume rather than removing the data. Access restriction is control over stored data rather than prevention.

314. A team must confirm that a model endpoint's execution role cannot be used to reach resources beyond those the model requires. Which approach is appropriate?

Answer and explanation

Answer: A. Comparing granted permissions against those actually used, with Access Analyzer findings, establishes what can be removed. Sole attachment says nothing about scope. Managed policies are frequently broader than a specific workload needs. Session duration limits exposure time rather than scope.

315. A team must confirm that no model artifact in an account is stored unencrypted. Which solution meets these requirements?

Answer and explanation

Answer: A. A Config rule evaluates every bucket continuously and identifies those without default encryption for remediation. Quarterly manual review is slow and error-prone. Assuming existing buckets are compliant is the gap being checked. CloudTrail records writes rather than evaluating encryption configuration.

316. A fine-tuning job must use a dataset that may not leave a specific account. Which approach is appropriate?

Answer and explanation

Answer: D. Keeping the operation in the account holding the data means it is never copied out, and the service's handling confirms it is not retained elsewhere. Copying, sharing encrypted, and presigning all move the data out of the account in different ways.

317. An ML platform must prevent a notebook user from assuming a role with broader permissions than their own. Which solution meets these requirements?

Answer and explanation

Answer: D. A permissions boundary caps effective permissions and, applied to roles the notebook role can create or assume, closes the escalation path. A broad managed policy is the permission being escalated from. Instance type restrictions are unrelated. CloudTrail records the escalation rather than preventing it.

318. An application must verify that a foundation model response has not been altered between the service and the client. Which consideration applies?

Answer and explanation

Answer: C. TLS provides integrity as well as confidentiality for data in transit, so a correctly validating client already detects alteration. A separately published hash is not offered and would still travel over the same channel. Double encryption addresses confidentiality rather than integrity. Periodic log comparison detects alteration long after it occurred.

319. A guardrail's configuration must be auditable so a reviewer can establish which version was active for a given interaction. Which approach is appropriate?

Answer and explanation

Answer: C. Recording the guardrail version with each invocation lets a reviewer reconstruct exactly which policy applied. A modification date does not establish which version applied to a specific interaction. Updating in place destroys the history. A name in a configuration file describes intent rather than what was applied at run time.

320. An ML workload's security posture must be assessed against a recognised control framework with evidence collected automatically. Which solution meets these requirements?

Answer and explanation

Answer: C. Audit Manager maps controls to evidence sources and collects the evidence continuously against a named framework. Security Hub checks security configuration rather than mapping to an audit framework. Config evaluates configuration rules without assembling framework evidence. Inspector assesses software vulnerabilities.