ML Model and Foundation Model (FM) Development

Domain 2: ML Model and Foundation Model (FM) Development

78 practice questions for Domain 2 of the AWS Certified Machine Learning Engineer - Associate (MLA-C02) exam, which makes up 24% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 2: ML Model and Foundation Model (FM) Development

24% of scored content · 78 practice questions

76. A team must select a foundation model from Amazon Bedrock for a summarization workload with a strict latency budget. Which solution meets these requirements?

Answer and explanation

Answer: B. Model selection is an evidence question, and Bedrock evaluations compare candidates on a representative sample so quality can be weighed against latency and cost. Parameter count and recency are not evidence of fitness for one task. A large context window does not help a summarization workload bounded by latency and typically costs more per request.

77. A company must decide between prompt engineering, retrieval augmentation, and fine-tuning for a foundation model application. Which requirement points to fine-tuning?

Answer and explanation

Answer: D. Fine-tuning encodes durable behaviour, format, and tone into the weights, which is the gap prompting could not close. Weekly document changes and source citation both require supplying documents at query time, which is retrieval. Fine-tuning adds training cost and does not inherently reduce inference cost.

78. A RAG application must answer questions that require combining information from several documents rather than retrieving one passage. Which architecture pattern is appropriate?

Answer and explanation

Answer: B. Multi-document questions require several passages in context, and reranking raises the chance the right ones are among those supplied. A single passage cannot support synthesis across documents. Larger chunks dilute relevance and waste context. Fine-tuning on a corpus teaches style rather than making facts retrievable with citations.

79. A business problem could be solved with a purpose-built AI service, a fine-tuned foundation model, or a custom trained model. Which consideration should decide the approach?

Answer and explanation

Answer: C. Standard tasks are met most cheaply by a purpose-built service, while a genuine differentiator justifies customization on proprietary data. Service recency is not a fitness criterion. Existing experience is a real constraint but should not override whether the capability is commodity or differentiating. Artifact size is irrelevant.

80. An application must choose between a smaller foundation model and a larger one for a high-volume classification task. Which approach should decide the choice?

Answer and explanation

Answer: A. Routing a task to the smallest model that clears the quality bar is the standard optimization, and it requires measurement on labelled data rather than assumption. Larger models are not uniformly better on narrow tasks. Choosing the smaller model without measuring risks shipping inadequate quality. Context window is irrelevant to short classification inputs.

81. A business problem requires predicting a continuous value from structured tabular data with around 200,000 rows. Which modeling approach should be evaluated first?

Answer and explanation

Answer: D. Gradient boosted trees are the standard strong baseline for tabular regression and handle mixed feature types well. Fine-tuning a foundation model on tabular data is far more expensive with no expected advantage. A deep network usually underperforms trees on tabular data of this size. Clustering is unsupervised and does not predict a target.

82. A regulated lending decision requires that the reason for each individual prediction can be stated to the applicant. Which consideration should drive model selection?

Answer and explanation

Answer: C. An explainability requirement constrains model choice, either to an interpretable family or to one whose per-prediction attributions can be produced and defended. Discriminative performance, latency, and artifact size are all real considerations but none satisfies a requirement to explain an individual decision.

83. A team must choose between a purpose-built AI service and a custom trained model for extracting entities from documents. Which consideration should decide?

Answer and explanation

Answer: D. A managed service is the cheapest route for standard entity types, while domain-specific entities require custom training on labelled data, so the nature of the entities decides. Release recency is not a fitness criterion. Prior experience is a real constraint but should not override fitness. Artifact size is irrelevant.

84. An application must answer questions from a document corpus revised weekly, and answers must cite their source. Which approach should be selected?

Answer and explanation

Answer: D. Weekly revision and source citation both require documents supplied at query time, which is retrieval. Weekly fine-tuning or continued pre-training is expensive and leaves nothing to cite. A full corpus in every prompt exceeds the context window and inflates cost.

85. A RAG application must answer questions requiring several retrieval steps, where each step depends on what the previous step found. Which architecture pattern applies?

Answer and explanation

Answer: A. Questions whose second query depends on the first result require iterative retrieval, which single-pass retrieval cannot express regardless of candidate count. Fine-tuning memorizes rather than reasoning across steps. A longer context window does not help when the relevant documents are unknown until the first result is read.

86. An architect must weigh a managed foundation model against a self-hosted open-weights model for a sustained high-volume workload. Which comparison is appropriate?

Answer and explanation

Answer: C. The comparison must include hosting and operating cost against per-token cost at real volume, with quality measured identically for both. Parameter count does not determine fitness or cost. Regional availability matters for residency rather than economics. Licence length is meaningless, though licence terms themselves are relevant.

87. A latency-sensitive feature must respond within 300 milliseconds, and the highest quality model in evaluation averages 1.2 seconds. Which approach should be taken?

Answer and explanation

Answer: B. A hard latency budget constrains the candidate set, so the selection must be made among models that can meet it. Raising the budget changes the requirement rather than meeting it. Streaming improves perceived responsiveness but the complete answer still takes 1.2 seconds. Caching helps only for repeated requests.

88. A fine-tuning strategy must be chosen for a foundation model where the organization has abundant unlabelled domain text but few labelled examples. Which approach is appropriate?

Answer and explanation

Answer: B. Continued pre-training uses unlabelled text to build domain familiarity, and supervised fine-tuning then shapes behaviour with the scarce labels. Supervised fine-tuning alone wastes the abundant unlabelled resource. Reinforcement learning from human feedback requires preference labels. Discarding usable data before collecting more is unnecessary.

89. A forecasting problem requires predicting demand for thousands of items with related seasonal patterns. Which modeling approach is appropriate?

Answer and explanation

Answer: C. A global model learns patterns shared across related series and generalises to items with sparse history, which is why it is standard at this scale. Thousands of independent models are costly to train and maintain and cannot share structure. Classification discards the magnitude the forecast requires. Clustering groups items without predicting demand.

90. An anomaly detection problem has no labelled examples of the anomalies to be detected. Which modeling approach is appropriate?

Answer and explanation

Answer: D. Unsupervised anomaly detection learns the shape of normal data and scores deviation, which is what unlabelled anomalies require. A supervised classifier needs both classes. A regression model predicts values rather than flagging outliers, though residuals can be used as a signal. Waiting for labels leaves the problem unaddressed.

91. A team must decide whether to use a pre-trained vision model or train from scratch, given 4,000 labelled images. Which approach is appropriate?

Answer and explanation

Answer: D. Transfer learning from a pre-trained model is the standard approach at this data volume and reaches usable accuracy where training from scratch would not. Training from scratch on 4,000 images overfits badly. A purpose-built service may not recognise the required classes and must be evaluated first. Collecting far more data delays work that transfer learning makes feasible now.

92. An application must both classify support tickets and generate a suggested reply. Which approach is appropriate?

Answer and explanation

Answer: C. The two tasks have different characteristics, so the choice between one model and a specialised pair should be made by measurement rather than assumption. Adopting either without evaluation skips the comparison. A classifier cannot generate text. Cluster-level replies ignore the individual ticket.

93. A model must run on a device with intermittent connectivity and limited memory. Which consideration should drive model selection?

Answer and explanation

Answer: C. Device constraints bound the candidate set, so selection happens among models that fit, with accuracy validated against the requirement rather than maximised. The highest-accuracy model may not run at all. Context window is irrelevant to a constrained vision or tabular task. Recency is not a fitness criterion.

94. A RAG application must be compared against a fine-tuned model for the same task before one is adopted. Which comparison is appropriate?

Answer and explanation

Answer: D. The comparison must measure what the application needs: answer quality, how current the content is, and what each approach costs to run. Artifact and index sizes describe storage. Build times describe effort rather than outcome. Delivery speed is a constraint rather than a quality comparison.

95. A use case could be met by prompt engineering today or by fine-tuning after several weeks of data collection. Which approach is appropriate?

Answer and explanation

Answer: B. Shipping the cheaper approach first and measuring establishes whether the more expensive one is needed at all, which is the point of the escalation ladder. Delaying assumes a gap that has not been demonstrated. Two versions doubles maintenance and confuses users. Fine-tuning on insufficient data risks degrading the base model's behaviour.

96. An architecture must support swapping the underlying foundation model as better options become available. Which design choice supports this?

Answer and explanation

Answer: A. Keeping the surrounding assets model-independent means a replacement can be measured on the same evaluation set and adopted without rebuilding. Fine-tuning binds behaviour to one model. Model-specific syntax scattered through code makes replacement a rewrite. A long support window delays the problem rather than preparing for it.

97. A long-running training job on Spot capacity must survive interruption without restarting from the beginning. Which solution meets these requirements?

Answer and explanation

Answer: A. Checkpoints written to S3 let a resumed job continue from the last saved state after a Spot reclamation. A longer run time allows more attempts without preserving progress. A larger instance shortens exposure without making the job resumable. Automatic tuning runs multiple jobs and does not address interruption within one.

98. Hyperparameter tuning must stop unpromising training jobs early so budget is concentrated on better configurations. Which strategy meets these requirements?

Answer and explanation

Answer: C. Hyperband allocates a small budget to many configurations and progressively halts the weaker ones. Grid and random search run every configuration to completion. Bayesian optimization chooses configurations intelligently but does not itself terminate underperforming jobs early.

99. A foundation model must be adapted to produce output in a fixed internal report format, and the team has 3,000 curated examples. Which solution meets these requirements?

Answer and explanation

Answer: B. Several thousand curated pairs is the standard input to supervised fine-tuning, which makes the format intrinsic rather than resupplied per request. Continued pre-training injects broad knowledge at far greater cost and does not target a format. Two hundred in-prompt examples would exhaust the context window and inflate cost. Retrieval supplies facts rather than teaching an output format.

100. A RAG application returns relevant documents but the most useful passage is often ranked low among the candidates supplied to the model. Which solution meets these requirements?

Answer and explanation

Answer: C. A reranker scores each candidate against the query jointly, which reorders results far more accurately than vector similarity alone. Reducing candidates would discard the useful passage entirely. Embedding dimension is a property of the model rather than a tunable remedy. Temperature affects sampling, not how the model weights retrieved context.

101. Validation loss stops improving after epoch 12 while training loss continues to fall. Which action should be taken?

Answer and explanation

Answer: B. Divergence between training and validation loss is overfitting, and early stopping retains the checkpoint from before the divergence. Training longer deepens the overfit. More capacity increases it further. Removing the validation split eliminates the only signal that would reveal the problem.

102. A model must be adapted with a task-specific prompt rather than a training run, and the behaviour must be consistent across requests. Which solution meets these requirements?

Answer and explanation

Answer: B. A versioned template fixes the instruction structure while allowing the variable parts to change, which is what makes behaviour consistent and governable. Caller-composed prompts guarantee divergence. Higher temperature increases variance. Raw input with no instruction leaves the behaviour entirely to the model.

103. A fine-tuning run on a small dataset produces a model that performs worse on general tasks than the base model. Which cause explains this?

Answer and explanation

Answer: D. Narrow fine-tuning can overwrite general capability, which shows as degraded performance outside the fine-tuning domain. Underfitting would leave the model close to the base rather than worse on general tasks. Leakage inflates evaluation scores rather than degrading general ability. Class imbalance affects the fine-tuned task rather than pre-trained capability.

104. A team must adapt a foundation model with limited compute and cannot afford to update all of its weights. Which approach is appropriate?

Answer and explanation

Answer: C. Parameter-efficient methods train a small adapter while the base weights remain frozen, which is what makes adaptation feasible on limited compute. Full fine-tuning on less data still updates every weight. Continued pre-training is more expensive still. More in-prompt examples consume context on every request and do not adapt the model.

105. A fine-tuning dataset is small, and the team must estimate whether more labelled examples would improve the result. Which approach is appropriate?

Answer and explanation

Answer: A. A learning curve across data fractions shows whether quality is still rising with data volume, which is the question being asked. A single run gives one point with no trend. Epoch tuning addresses training duration rather than data sufficiency. Adding synthetic data changes the dataset before the question is answered.

106. A distributed training job scales poorly, and profiling shows GPUs idle while waiting for input batches. Which solution meets these requirements?

Answer and explanation

Answer: B. Idle GPUs waiting on input is an input-bound job, so the data pipeline is the constraint. More GPUs multiply the idling. A higher learning rate changes convergence rather than throughput. A smaller model completes steps faster and starves sooner.

107. A model must be customized so that it answers in the company's house style without any training run. Which approach is appropriate?

Answer and explanation

Answer: D. Prompt engineering with instructions and examples changes behaviour with no training run, which is the constraint stated. Fine-tuning and continued pre-training both require training. Higher temperature increases variability rather than enforcing a style.

108. An embedding model used for retrieval performs poorly on the company's domain terminology. Which approach is most likely to improve retrieval?

Answer and explanation

Answer: C. Poor retrieval on domain terminology is an embedding problem, so adapting or replacing the embedding model and measuring against labelled pairs addresses the cause. Retrieving more passages returns more of the same poorly ranked candidates. The generation model cannot repair retrieval. Larger chunks dilute relevance.

109. Two candidate fine-tuned models must be compared fairly before one is promoted. Which approach is appropriate?

Answer and explanation

Answer: B. A fair comparison requires the same held-out data, metrics, and inference settings for both candidates. Different holdouts are not comparable. Training data measures memorization. Training loss is not comparable across runs with different data or hyperparameters.

110. A training run must be reproducible so the same data and configuration produce the same model. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: B, D. Seeds control the stochastic elements and recording data version, image digest, and hyperparameters pins everything else that determines the result. More epochs do not remove randomness. Instance size does not guarantee identical numerics and may change them. Checkpointing does not affect the final model when training completes.

111. A hyperparameter tuning job must explore a learning rate spanning several orders of magnitude. Which configuration is appropriate?

Answer and explanation

Answer: D. Logarithmic scaling samples evenly across orders of magnitude, which is how learning rates should be explored. Linear scaling concentrates nearly all samples at the top of the range. Three candidates is a coarse search. Fixing the learning rate abandons the most influential hyperparameter.

112. A training job must be stopped automatically when a rule detects that gradients have vanished. Which solution meets these requirements?

Answer and explanation

Answer: D. Debugger evaluates rules against captured tensors during training and can stop the job when one fires. Profiler reports resource utilization rather than model-level training problems. A run time limit ends the job regardless of cause. A lower learning rate is a possible remedy rather than a detection mechanism.

113. A fine-tuning run must be compared against the base model on the same evaluation set to confirm it improved. Which practice is appropriate?

Answer and explanation

Answer: A. A valid comparison requires identical data and inference settings for both models. Published benchmarks measure different tasks under different conditions. Training data measures memorization. Training loss is not comparable between a base model and a fine-tuned one.

114. A training job's per-epoch time increases steadily as training progresses, while the dataset size is constant. Which cause should be investigated first?

Answer and explanation

Answer: B. Time per epoch rising with constant data points to a resource or accumulation problem such as growing memory use or an unbounded structure. A learning rate schedule changes convergence rather than epoch duration. Validation split size is constant across epochs. Convergence affects quality rather than time per epoch.

115. A team must adapt a model's behaviour using preference data indicating which of two responses a human preferred. Which technique applies?

Answer and explanation

Answer: C. Preference data expresses a relative judgement, and preference-based alignment uses both the preferred and rejected responses to learn the distinction. Discarding the rejected response loses the contrast that carries the signal. Continued pre-training treats the text as a corpus. Prompt examples do not use the ranking.

116. A distributed training job produces different results on each run despite identical data and hyperparameters. Which cause is most likely?

Answer and explanation

Answer: B. Variation across runs with identical inputs points to unfixed seeds and non-deterministic kernels, particularly in distributed settings. Epoch count, class balance, and learning rate all affect the outcome but would affect every run identically.

117. A retrieval component's embedding model must be upgraded, and existing vectors were produced by the previous model. Which action is required?

Answer and explanation

Answer: A. Embeddings from different models occupy different spaces and cannot be compared, so the corpus must be re-embedded. Mixing vectors from two models produces meaningless similarity scores. No scaling factor reconciles different learned spaces. A threshold change cannot fix incomparable vectors.

118. A fine-tuning job must not expose the training data to the model provider. Which consideration applies when selecting the approach?

Answer and explanation

Answer: C. The service's documented data handling determines the answer, and managed customization keeps customer data within the account and out of base model training. Assuming the worst without checking leads to unnecessary architecture. Data must be decrypted to train on it. A smaller dataset reduces volume rather than resolving the question.

119. A training run must resume from where it stopped after a deliberate interruption, preserving optimizer state. Which solution meets these requirements?

Answer and explanation

Answer: B. Resuming correctly requires the optimizer state as well as the weights, since momentum and adaptive learning rates are part of the training trajectory. Weights alone resume with a reset optimizer, which changes the trajectory. Restarting from the beginning discards progress. Fewer epochs changes the training rather than making it resumable.

120. A fine-tuned model must be evaluated for whether it retained the base model's general capabilities. Which approach is appropriate?

Answer and explanation

Answer: B. Detecting capability loss requires measuring capabilities outside the fine-tuning task and comparing against the base model. Task-specific evaluation cannot reveal degradation elsewhere. Training loss is not comparable between the two. Self-assessment is not evidence.

121. A team must decide the batch size for a training job on a fixed GPU configuration. Which consideration applies?

Answer and explanation

Answer: A. Larger batches use the hardware efficiently, and the learning rate must be adjusted with them because the gradient noise scale changes. Very small batches underuse the GPU and lengthen training. A batch of one is the extreme case of that. A published value was tuned for different hardware and a different model.

122. A retrieval component must be improved, and the team can either change the embedding model or add reranking. Which approach establishes which to do?

Answer and explanation

Answer: D. Measuring each change independently establishes which contributes and whether both are needed. Applying both together makes the contributions indistinguishable. Choosing by which seems foundational or cheaper is a guess about where the problem lies.

123. A customization must make a model consistently refuse a category of request that it currently sometimes answers. Which approach is most reliable?

Answer and explanation

Answer: B. A requirement for consistent refusal needs a deterministic control, and a guardrail evaluates every request and response at the service layer. Fine-tuning raises the refusal rate without guaranteeing it. A stronger instruction remains probabilistic and is what an injection targets. Lower temperature reduces variation without removing the failure mode.

124. Several evaluation metrics must be matched to the situation in which each is the right choice. (Match each metric to its situation.)

  1. Recall
  2. Precision
  3. Root mean squared error
  4. Area under the ROC curve
Answer and explanation

Answer: 1-B, 2-C, 3-D, 4-A. Recall measures how many true positives were found and is what a screening problem optimises when a missed case is the worst outcome. Precision penalises false alarms and is the metric to raise when investigation is costly. Root mean squared error squares the error, so large deviations dominate, which distinguishes it from mean absolute error. Area under the ROC curve integrates across all thresholds, which is why it is used for comparison rather than for describing one operating point.

125. The quality of generated summaries must be measured automatically against reference summaries during development. Which solution meets these requirements?

Answer and explanation

Answer: D. ROUGE measures overlap between generated and reference summaries and is the standard automatic metric for summarization. Accuracy requires discrete class labels that free text does not provide. Comparing lengths measures verbosity rather than quality. Latency is an operational measure.

126. A translation model's output must be compared against professional reference translations. Which metric is conventionally used?

Answer and explanation

Answer: C. BLEU measures n-gram precision against reference translations and is the conventional machine translation metric. F1 and area under the ROC curve are classification metrics. Comparing sentence lengths measures neither adequacy nor fluency.

127. Two candidate responses convey the same meaning in different words, and the evaluation must recognise them as equivalent. Which metric is most appropriate?

Answer and explanation

Answer: B. BERTScore compares contextual embeddings, so paraphrases score highly where surface-overlap metrics do not. BLEU penalises different wording that carries the same meaning. Exact match rejects any paraphrase. Edit distance measures character changes rather than meaning.

128. Generated responses must be assessed at scale on helpfulness and faithfulness, where no reference answer exists. Which solution meets these requirements?

Answer and explanation

Answer: A. An LLM judge scores open-ended qualities at scale, and calibrating its rubric against human ratings is what makes the scores trustworthy. BLEU compares against reference text, and a prompt is not a reference answer. Response length and refusal counts measure neither helpfulness nor faithfulness.

129. A subjective quality judgement must be incorporated into an evaluation pipeline that otherwise runs automatically. Which solution meets these requirements?

Answer and explanation

Answer: D. Sampling into a human review step captures the judgement automated metrics miss while keeping the pipeline scalable, and recording the ratings allows the automated metrics to be calibrated. Reviewing every response does not scale. Abandoning subjective measurement leaves a known gap. Self-rating by the same model is not an independent judgement.

130. A RAG system returns fluent answers that are sometimes unsupported by the retrieved passages. Which evaluation should be applied?

Answer and explanation

Answer: C. Separating retrieval accuracy from answer faithfulness identifies whether the right passages were found and whether the answer is grounded in them, which are distinct failures requiring different fixes. Fluency is already adequate by the description. Latency and passage counts say nothing about groundedness.

131. A classification model must be compared across all decision thresholds rather than at one operating point. Which metric summarises this?

Answer and explanation

Answer: A. Area under the ROC curve integrates performance across all thresholds, which is what threshold-independent comparison requires. Accuracy and F1 at a fixed threshold describe one operating point. Training log loss measures fit to training data rather than discriminative ability.

132. A confusion matrix shows many false negatives, few false positives, and high true negatives. Which metric captures the weakness?

Answer and explanation

Answer: D. Many false negatives means genuine positives are being missed, which is exactly what low recall measures. Low precision would show as many false positives. High true negatives indicates specificity is good. Accuracy can remain high on an imbalanced dataset despite poor recall.

133. A fraud model's business cost of a missed case is roughly twenty times the cost of a false alarm. Which approach to threshold selection is appropriate?

Answer and explanation

Answer: A. When error costs are asymmetric the operating point should minimize expected cost, which will favour recall at this ratio. Maximizing accuracy on an imbalanced problem tends to favour the majority class. A default threshold encodes no business information. Equalizing precision and recall is arbitrary unless the costs happen to be equal.

134. A regression model's errors must be reported in the target's original units, and a small number of extreme outliers should not dominate the measure. Which metric is appropriate?

Answer and explanation

Answer: C. Mean absolute error is in the target's units and weights every error equally, so outliers do not dominate. Root mean squared error squares the error, which is the opposite of what is required. R-squared is a proportion rather than an error magnitude. Mean absolute percentage error becomes unstable when actual values approach zero.

135. Two models must be compared on an imbalanced dataset where the positive class is rare and is the class of interest. Which comparison is most informative?

Answer and explanation

Answer: D. The precision-recall curve focuses on the positive class and is the more informative summary when positives are rare. Accuracy is dominated by the majority class. Area under the ROC curve can look strong even when performance on rare positives is poor, because the large negative class inflates it. Training time is not a quality measure.

136. An evaluation must determine whether a RAG system's retrieval stage is finding the passages that contain the answer. Which metric applies?

Answer and explanation

Answer: C. Context recall measures whether the relevant passages were retrieved, which isolates the retrieval stage. Fluency assesses the generated text. Latency and token count are operational measures.

137. An LLM-as-a-judge evaluation reports consistently high scores while users report poor answers. Which action should be taken?

Answer and explanation

Answer: A. A judge measures whatever its rubric encodes, so divergence from user perception means the rubric needs calibrating against human judgement. A larger judge applies the same flawed rubric. A smaller evaluation set reduces coverage. Consistent application of a wrong rubric produces consistently wrong scores.

138. A summarization model must be evaluated on whether its summaries contain claims absent from the source document. Which evaluation approach applies?

Answer and explanation

Answer: A. Unsupported claims are a faithfulness failure, which is assessed by checking whether each claim is entailed by the source. ROUGE measures overlap with a reference and can score a hallucinated summary well if the wording matches. Compression ratio and reading level describe form rather than truth.

139. A team must decide whether a model's performance difference between two customer segments is meaningful or the result of sample size. Which approach is appropriate?

Answer and explanation

Answer: C. Confidence intervals show whether an observed difference is distinguishable from sampling noise, which is the question asked. Comparing point estimates ignores uncertainty. Adjusting a threshold to equalize metrics conceals the difference. Combining segments removes the comparison entirely.

140. A model's evaluation must reflect how it will perform on data from a future period rather than a random sample. Which approach is appropriate?

Answer and explanation

Answer: B. A temporally held-out period mirrors deployment, where the model predicts on data that did not exist at training time. A random sample mixes future and past. Cross-validation on training data measures fit rather than forward performance. Stratified sampling balances classes without respecting time.

141. A generative application must be evaluated for whether its outputs differ systematically in quality across demographic groups. Which approach is appropriate?

Answer and explanation

Answer: C. Detecting a systematic difference requires measuring each group separately and comparing, which a stratified evaluation set provides. An aggregate mean hides differences between groups. Removing references prevents the measurement rather than performing it. Self-assessment by the model is not evidence.

142. A monitoring dashboard must show whether RAG answer quality is degrading over time in production. Which combination of signals is most informative? (Select TWO.)

Answer and explanation

Answer: B, C. A recurring question set reveals retrieval regressions and sampled faithfulness reveals generation regressions, which are the two stages that determine answer quality. Token count, latency, and corpus size are operational measures that can stay flat while quality falls.

143. A model's validation metric is materially better than its test metric on a held-out set drawn from the same period. Which cause is most likely?

Answer and explanation

Answer: B. Repeated use of a validation set during tuning fits the model to that set, which is why a fresh test set scores lower. A different distribution is excluded by the question's stipulation that both come from the same period. Underfitting would show as poor performance on both. Metric choice would affect both sets equally.

144. An evaluation must establish whether a new model version is genuinely better than the current one on live traffic. Which approach is appropriate?

Answer and explanation

Answer: A. A controlled split isolates the version as the variable and a defined sample size makes the comparison statistically meaningful. Comparing against the previous week confounds the change with everything else that differed. Offline scores predict but do not establish live performance. A reviewer panel is useful qualitative input rather than a controlled comparison.

145. An evaluation of a RAG application must separate failures caused by the retriever from those caused by the generator. Which approach is appropriate?

Answer and explanation

Answer: B. Supplying the correct passages removes retrieval as a variable, so the gap between that result and end-to-end performance is attributable to the retriever. Proportional attribution of end-to-end failures is guesswork. Removing context entirely tests a different task. Retrieval latency is operational rather than a quality measure.

146. A team must decide how many examples the evaluation set needs to detect a two percentage point difference between models. Which approach is appropriate?

Answer and explanation

Answer: A. Detecting a specified effect size reliably requires a sample sized by a power calculation. A convenience sample may be far too small to distinguish two points from noise. Training set size is unrelated to evaluation power. An arbitrary round number encodes no statistical reasoning.

147. A confusion matrix for a three-class problem must be summarised into one metric that treats each class equally regardless of its frequency. Which metric applies?

Answer and explanation

Answer: B. Macro averaging computes the metric per class and averages them equally, so a rare class counts as much as a common one. Micro averaging and accuracy both weight by frequency and are dominated by the largest class. A ROC curve computed for one class describes that class alone and says nothing about the classes it excludes.

148. A human evaluation must produce consistent ratings across several reviewers assessing generated text. Which practice is appropriate?

Answer and explanation

Answer: C. A rubric with anchored examples aligns reviewers, and agreement measurement reveals whether it has worked. Unanchored numeric judgement produces reviewer-specific scales. A single reviewer removes disagreement by removing the check on it. Averaging without measuring agreement conceals whether the ratings mean the same thing.

149. An evaluation report must communicate a model's suitability to a non-technical approver. Which content is most appropriate?

Answer and explanation

Answer: D. An approver needs to know what the model will do in the business, under what conditions that was established, and where it fails. Architecture, hyperparameters, loss curves, and raw metric tables are engineering detail that does not answer the approval question.

150. A model that scores well offline performs noticeably worse in production on the same metric. Which cause should be investigated first?

Answer and explanation

Answer: D. A quality gap between offline and online on the same metric most often traces to training and serving skew in how features are computed. Instance type affects latency rather than prediction quality. A corrupted artifact would fail rather than degrade gracefully. Traffic volume affects throughput rather than per-prediction quality.

151. An evaluation must establish whether a model's improvement on a benchmark reflects genuine capability or memorization of the benchmark. Which approach is appropriate?

Answer and explanation

Answer: C. Contamination is established by checking overlap and confirmed by evaluating on data constructed after the model was trained. Repeating the same benchmark reproduces the same contamination. Adding examples from the same source may add more contaminated data. Comparing against other models says nothing about whether all of them are contaminated.

152. A team must report a model's performance to stakeholders who will use it to set an operational threshold. Which presentation is most useful?

Answer and explanation

Answer: C. Choosing an operating point requires seeing the trade-off across thresholds, which is what a precision and recall curve shows. A single value at a default threshold presents one arbitrary point. Area under the curve summarises across thresholds without showing the trade-off at each. Training accuracy measures fit.

153. A RAG evaluation must measure whether the system declines to answer when the corpus does not contain the answer. Which approach is appropriate?

Answer and explanation

Answer: B. Abstention can only be measured by asking questions the corpus cannot answer and checking the system declines. An answerable-only set never tests the behaviour. Retrieval always returns something regardless of relevance. Reported confidence is often poorly calibrated and does not establish abstention.