Data Preparation for ML and AI

Domain 1: Data Preparation for ML and AI

75 practice questions for Domain 1 of the AWS Certified Machine Learning Engineer - Associate (MLA-C02) exam, which makes up 28% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 1: Data Preparation for ML and AI

28% of scored content · 75 practice questions

1. Training data held in an Amazon RDS database must be made available to a SageMaker AI training job without the job querying the database directly. Which solution meets these requirements?

Answer and explanation

Answer: C. SageMaker training jobs read most efficiently from S3, and exporting to a columnar format gives parallel reads without holding a database connection for the run. Mounting a database directory is not how a relational engine exposes data. Querying during training couples the job to database availability and throughput. A read replica still requires the job to query a database rather than read training data.

2. Several thousand small image files must be read repeatedly by distributed training jobs with the lowest possible per-file latency. Which solution meets these requirements?

Answer and explanation

Answer: C. FSx for Lustre links to an S3 bucket and presents the data as a high-performance file system, which removes the per-object request overhead that dominates when many small files are read repeatedly. Reading each object individually each epoch incurs that overhead every time. EBS volumes must be populated per instance and are not shared. DynamoDB is not designed for bulk binary training data.

3. A Retrieval Augmented Generation application must store several million embeddings with metadata filtering and sub-100ms retrieval. Which solution meets these requirements?

Answer and explanation

Answer: B. OpenSearch provides approximate nearest neighbour indexing at this scale together with metadata filtering in the same query. DynamoDB has no vector similarity index. Athena over S3 performs no vector search and has seconds-scale latency. A binary column has no similarity operator, although Amazon RDS for PostgreSQL with the pgvector extension would be a valid alternative.

4. A team must choose a vector store for a RAG application with modest embedding volume, where the operational data already resides in a relational database. Which solution meets these requirements with the LEAST additional infrastructure?

Answer and explanation

Answer: D. pgvector adds vector similarity search to an existing PostgreSQL instance, so no new service is introduced and embeddings join relational data in one query. A dedicated OpenSearch domain is appropriate at large scale but is additional infrastructure. MemoryDB offers vector search at higher cost for data that must fit in memory. Scanning objects in S3 performs no indexed similarity search.

5. Text, image, and audio assets must be ingested and stored for a multimodal AI application, with the raw assets retained for reprocessing. Which solution meets these requirements?

Answer and explanation

Answer: A. S3 holds large binary assets of any modality durably and cheaply, and a catalog makes them discoverable for reprocessing. DynamoDB items are capped well below typical media sizes. EFS is a shared file system priced for active workloads rather than a durable asset store. Discarding raw assets makes reprocessing with a different model impossible.

6. A streaming feature pipeline must ingest events continuously and compute windowed aggregations before they are written for training. Which solution meets these requirements?

Answer and explanation

Answer: B. Flink provides stateful windowed aggregation over a stream, which is what a windowed feature computation requires. Firehose delivers records without computing aggregations across a window. A per-message Lambda has no windowing state. A scheduled scan is batch processing and loses the streaming semantics.

7. An ingestion job reading from Amazon DynamoDB is throttled during peak periods, and the table's capacity is already at its configured ceiling. Which cause should the team investigate first?

Answer and explanation

Answer: B. Throttling while table capacity is not exhausted indicates a hot partition, because DynamoDB applies limits per partition as well as per table. A missing permission produces an access error rather than throttling. Point-in-time recovery does not consume read capacity. Eventually consistent reads consume less capacity, so using them would reduce rather than cause throttling.

8. A training dataset must be reproducible so an audit can establish exactly which records produced a given model. Which solution meets these requirements?

Answer and explanation

Answer: B. Versioning the data and recording which versions were used links a model artifact to the exact records that produced it. Overwriting in place destroys the association. A record count cannot identify which records were used. Keeping only the latest copy makes past models unreproducible.

9. An ingestion pipeline must move 20 TB of training images from an on-premises NFS share into Amazon S3 while preserving file metadata. Which solution meets these requirements?

Answer and explanation

Answer: B. DataSync transfers between NFS and S3 with metadata preservation and built-in integrity verification. The CLI copies objects without preserving NFS metadata. Mounting a bucket does not provide true file semantics or metadata fidelity. Transfer Acceleration speeds uploads without preserving metadata.

10. Training data arrives as a continuous stream and must be buffered into Amazon S3 in Parquet partitioned by ingestion date, with no servers to manage. Which solution meets these requirements?

Answer and explanation

Answer: C. Firehose buffers streaming records, converts them to Parquet using a catalog schema, applies dynamic partitioning, and writes to S3 with no infrastructure to operate. A Kinesis stream is transport and cannot itself write formatted output. Glue streaming and EMR both introduce job or cluster management.

11. A dataset must be shared with a partner account for model training without copying it. Which solution meets these requirements?

Answer and explanation

Answer: C. Cross-account read requires both object permission and decrypt permission on the key, and grants access in place with no duplication. Copying creates a second copy that drifts. Public access exposes the data to everyone. Presigned URLs per object does not scale to a training dataset.

12. An ingestion process must detect when a source system stops delivering data, rather than silently training on a stale dataset. Which solution meets these requirements?

Answer and explanation

Answer: A. Detecting absence requires a metric recording success and an alarm treating missing data as breaching. An error rate alarm stays silent when the job never runs. Weekly review is slow. Retries handle transient failure without detecting a source that has stopped.

13. An audio dataset for a speech model must be stored so that both the raw files and their derived transcripts remain associated. Which solution meets these requirements?

Answer and explanation

Answer: C. A manifest or catalog holding the object key alongside the transcript makes the association explicit and queryable. Matching counts across buckets is not an association. Audio metadata fields are limited and lost on format conversion. Regenerating transcripts each time is expensive and may not reproduce the original.

14. A storage decision must be made for a training dataset read once per epoch by a distributed job, where cost matters more than per-file latency. Which solution meets these requirements?

Answer and explanation

Answer: C. S3 is the cheapest durable option and suits sequential reads when a streaming input mode is used, which fits a cost-sensitive job. FSx for Lustre gives the lowest latency at higher cost, which the requirement deprioritises. EBS requires populating a volume per instance. EFS with provisioned throughput costs more than S3 for this pattern.

15. A training job reading many small objects from Amazon S3 spends most of its time waiting on requests rather than computing. Which solution meets these requirements?

Answer and explanation

Answer: A. Per-object request overhead dominates when files are small, and consolidating into larger shards removes most of it. More instances multiply the request count without reducing per-object overhead. A nearer Region reduces latency slightly but the overhead remains. Versioning adds stored objects and no throughput.

16. An ML team must query a catalogued training dataset in Amazon S3 with SQL during exploration, without provisioning infrastructure. Which solution meets these requirements?

Answer and explanation

Answer: B. Athena runs SQL against catalogued S3 data serverlessly, which suits exploration with no infrastructure. Redshift, RDS, and EMR all require provisioning and a load step before any query runs.

17. A dataset must be retained for seven years for audit but is read only when a model is re-examined. Which solution meets these requirements at the lowest cost?

Answer and explanation

Answer: D. Archival storage classes cost a fraction of Standard and suit data read once or twice over years. Deletion makes the model unreproducible for audit. Standard pays frequent-access rates for cold data. EFS is priced for active file workloads rather than long-term archive.

18. An ingestion pipeline must record which source system produced each training record so lineage can be traced. Which solution meets these requirements?

Answer and explanation

Answer: D. Attaching source metadata at ingestion and preserving it in the catalog makes lineage queryable for any record. A log entry records that the job ran rather than tagging individual records. Separate buckets carry origin implicitly but break when a record is copied. Inferring from schema fails when sources share a schema.

19. Streaming events must be retained long enough that a consumer failure of up to three days does not lose data. Which solution meets these requirements?

Answer and explanation

Answer: A. Records expire at the retention period, so extending it to cover the outage prevents loss, and a lag alarm warns before that point. Shard count governs throughput rather than retention duration. Smaller events do not extend the time window. A second consumer helps availability but both could fail together.

20. An ML platform must expose one dataset to several teams with different permissions, without copying it. Which solution meets these requirements?

Answer and explanation

Answer: B. Lake Formation grants permissions at table and column level over one physical copy, enforced across query engines. Per-team copies duplicate storage and drift. A single bucket policy operates on objects and cannot exclude a column. Presigned URLs do not scale to a dataset accessed continuously.

21. A multimodal dataset must be stored so training jobs can select only the modality they need without reading the others. Which solution meets these requirements?

Answer and explanation

Answer: B. Partitioning by modality lets a job read only the relevant prefixes, which reduces both time and cost. Interleaving forces every job to read everything. One object per record has the same problem at record granularity. Separate accounts add access complexity without changing what must be read.

22. A feature pipeline writes to the online store for serving and the offline store for training, and the two must not diverge. Which solution meets these requirements?

Answer and explanation

Answer: D. A single ingestion populating both stores is what keeps them consistent by construction. Separate writes from two paths diverge. A nightly copy leaves the online store stale within the day. Current online values cannot reconstruct the history training requires.

23. A team must decide where to store labelled images used by several training jobs across two accounts. Which consideration should drive the decision?

Answer and explanation

Answer: C. A storage decision balances how the data is read, who must reach it, and what it costs, and optimizing one dimension alone typically degrades another. Lowest price often means highest retrieval latency or cost. Lowest latency is usually the most expensive. Familiarity is a constraint rather than a criterion.

24. An ingestion job repeatedly fails when the source system returns a burst of records larger than the job's configured memory. Which solution meets these requirements?

Answer and explanation

Answer: C. Bounded batching with backpressure makes memory use independent of burst size, which is the durable fix. Sizing to the largest observed burst fails on the next larger one. The source system is usually outside the team's control. Retrying repeats the same failure.

25. A training dataset stored in Amazon S3 must be readable by a training job in a second account without the object owner losing control. Which solution meets these requirements?

Answer and explanation

Answer: C. Granting cross-account read in place keeps ownership and lifecycle with the producer while the consumer reads directly, and KMS-encrypted objects also require decrypt permission. Copying creates a second copy that drifts and doubles storage. Public access exposes the data to everyone. Transferring ownership gives away the control the requirement preserves.

26. A pipeline must detect when an upstream system delivers a file whose schema has changed unexpectedly. Which solution meets these requirements?

Answer and explanation

Answer: D. Validating against a registered schema at ingestion detects the change before downstream work begins and preserves the file for investigation. Inspecting afterwards means the data is already loaded. A crawler silently adopts the new schema rather than flagging it. Letting training fail wastes the compute and reports the problem late.

27. An ML workload must read the same dataset from two Regions with low latency in each. Which solution meets these requirements?

Answer and explanation

Answer: C. Cross-Region replication gives each Region a local copy, which is what removes the cross-Region read latency and transfer cost. Reading across the boundary incurs both on every read. Concentrating in one Region leaves the other slow. EFS file systems are Regional and cannot be mounted across Regions.

28. Several models must consume the same engineered features, and training data must reflect the feature values as they were at the time of each historical event. Which solution meets these requirements?

Answer and explanation

Answer: C. Feature Store's offline store supports point-in-time correct retrieval, which prevents training on values that did not exist when the event occurred. Joining latest values leaks future information into training. A nightly refresh has the same leakage problem with no historical record. Recomputing in each job duplicates logic and risks divergence between training and serving.

29. Text and image assets must be converted into numerical representations so they can be compared for semantic similarity across modalities. Which solution meets these requirements?

Answer and explanation

Answer: C. A multimodal embedding model places semantically related text and images near each other in one vector space, which is what cross-modal similarity requires. One-hot encoding and resizing produce representations that are not comparable across modalities. TF-IDF and histograms occupy unrelated feature spaces. Tokenization and base64 encoding produce no semantic representation.

30. Documents must be prepared for a RAG application so that a fact spanning a section boundary can still be retrieved intact. Which solution meets these requirements?

Answer and explanation

Answer: B. Overlap ensures a fact split by a boundary is wholly contained in at least one chunk, and retained metadata lets the answer cite its section. Fixed chunks with no overlap create exactly the boundary problem described. Whole documents exceed the context window and retrieve poorly. Discarding short chunks loses content without addressing boundaries.

31. Customer names and account numbers must be removed from support transcripts before the transcripts are used to fine-tune a model. Which solution meets these requirements?

Answer and explanation

Answer: A. Detection followed by redaction or tokenization removes the sensitive values before training, so the model cannot memorize and later emit them. Encryption protects storage while the plaintext still reaches training. A lower learning rate reduces but does not eliminate memorization. Shuffling changes order without removing anything.

32. A dataset of prompt and completion pairs must be prepared for fine-tuning a foundation model. Which solution meets these requirements?

Answer and explanation

Answer: A. Supervised fine-tuning requires records in the service's expected prompt and completion structure, validated so malformed examples do not train, with an evaluation holdout to measure the result. A concatenated text file is the input to continued pre-training, which is a different technique. Raw logs contain noise and no clear supervision signal. Selecting only long examples biases the dataset.

33. A categorical feature contains 40,000 distinct values and must be represented numerically for a model. Which solution meets these requirements?

Answer and explanation

Answer: D. High-cardinality categoricals are handled with embeddings or hashing so dimensionality stays manageable and related categories can share structure. One-hot encoding produces 40,000 sparse columns. Integer label encoding implies a false ordering between unrelated categories. Removing the feature discards signal that may be predictive.

34. Text must be prepared for a transformer model that expects token identifiers rather than raw strings. Which solution meets these requirements?

Answer and explanation

Answer: A. A transformer expects the sub-word vocabulary it was trained with, so the matching tokenizer must be used or the identifiers are meaningless to the model. Whitespace splitting produces a different vocabulary. TF-IDF produces a fixed-length numeric vector suited to classical models. Character code points are not the model's token space.

35. Numeric features spanning very different ranges are degrading a distance-based algorithm's performance. Which solution meets these requirements?

Answer and explanation

Answer: D. Distance-based algorithms are dominated by large-range features, and fitting the scaler on the training split alone prevents test statistics leaking into training. Fitting on the full dataset before splitting is the classic leakage error. One-hot encoding applies to categorical data. Removing features discards information.

36. A feature must be computed identically in the training pipeline and the online serving path. Which solution meets these requirements?

Answer and explanation

Answer: C. Sharing one definition or one served value is what prevents training and serving skew, where the two paths drift apart. Separate implementations diverge even when documented. Approximating at serving time introduces the skew deliberately. Inferring training values retrospectively cannot reproduce what the model actually learned from.

37. Text records vary widely in length, and the model requires fixed-length input. Which solution meets these requirements?

Answer and explanation

Answer: B. Truncation and padding to a length informed by the distribution is the standard approach and keeps every record usable. Discarding long records biases the dataset. Concatenating unrelated records creates examples the model must not learn from. Splitting into sentences destroys the record-level label relationship.

38. A pipeline must apply a transformation learned from the training data to new records at inference time. Which solution meets these requirements?

Answer and explanation

Answer: A. Persisting the fitted transformer guarantees inference applies exactly what training applied. Refitting on inference data changes the transformation and introduces skew. Hardcoded parameters drift from the artifact and are not updated when the model is retrained. Omitting the transformation feeds the model inputs unlike its training distribution.

39. An image dataset must be augmented to improve robustness without collecting new photographs. Which approach is appropriate?

Answer and explanation

Answer: A. Augmentation creates plausible variants that teach invariance, provided each transformation preserves the label. Duplicates add no new signal. A flip applied where orientation carries meaning corrupts the label. Lower resolution discards information rather than augmenting.

40. A pipeline must convert free-text product descriptions into features for a classical classifier without a neural embedding model. Which solution meets these requirements?

Answer and explanation

Answer: C. TF-IDF produces a fixed-length numeric representation weighting informative terms, which is the standard classical text feature. One-hot encoding whole strings produces a near-unique column per record. Character counts discard the content. Alphabetical label encoding imposes a meaningless ordering.

41. A dataset must be prepared for model distillation, where a smaller model learns from a larger one. Which solution meets these requirements?

Answer and explanation

Answer: D. Distillation trains a student on the teacher's outputs, so generating those outputs over a representative input set is the data preparation step. Training on original labels is ordinary training rather than distillation. Reducing precision is quantization, a different technique. A random data subset does not transfer the teacher's behaviour.

42. A pipeline must anonymize a dataset so that individuals cannot be re-identified, while records remain usable for aggregate analysis. Which approach is appropriate?

Answer and explanation

Answer: C. Re-identification usually proceeds through combinations of quasi-identifiers such as postcode, date of birth, and sex, so those must be generalized or suppressed alongside the direct identifiers. Removing only the name leaves the combination intact. Per-record random values break aggregation by individual. Encryption with an accessible key is reversible.

43. A chunking strategy must be selected for a corpus of legal contracts with deeply nested clause structure. Which approach is most appropriate?

Answer and explanation

Answer: B. Structured documents retrieve best when chunks respect their own boundaries and carry the hierarchy as metadata, so a retrieved clause can be located in context. Fixed-token chunking splits clauses arbitrarily. Single sentences lose the surrounding clause. A whole contract exceeds the context window and retrieves imprecisely.

44. Numeric features contain values recorded in different units across source systems. Which step must occur before scaling or model training?

Answer and explanation

Answer: A. Values in different units are not comparable, so they must be converted to one unit before any statistical transformation. Standardization applied to mixed units scales an incoherent distribution. Removing one source discards data. Adding the unit as a feature leaves the model to learn a conversion it should not have to.

45. A categorical feature's values in production include categories that appeared after the encoder was fitted. Which solution meets these requirements?

Answer and explanation

Answer: C. A reserved unknown bucket lets inference proceed predictably when a new category appears, which is inevitable. Refitting on production data creates training and serving skew. Rejecting requests makes the service brittle. Mapping to the most frequent category silently misrepresents the input.

46. A date field must be turned into features a tree model can use to capture weekly and seasonal patterns. Which approach is appropriate?

Answer and explanation

Answer: A. Decomposing a date into calendar components exposes the cyclical structure a model can split on. A single epoch offset encodes ordering but hides weekly and seasonal cycles. One-hot encoding every date produces a column per day with no generalization. Trees can use derived temporal features effectively.

47. A transformation must be applied to 8 TB of tabular data before training, using a distributed engine and no cluster management. Which solution meets these requirements?

Answer and explanation

Answer: C. Glue provides managed distributed Spark with no cluster to operate, which suits this volume. A single Processing instance is not distributed. Lambda has a fifteen-minute limit and limited memory per invocation. Athena can transform with SQL but is limited for complex row-level logic at this scale.

48. An imbalanced training set must be rebalanced without discarding the majority class information. Which approach is appropriate?

Answer and explanation

Answer: D. Oversampling the minority or weighting the loss addresses the imbalance while retaining every majority record. Undersampling discards majority information, which the requirement excludes. Unsupervised training abandons the labels. Duplicating everything preserves the same ratio.

49. A text corpus contains near-duplicate documents that would over-weight certain content during fine-tuning. Which solution meets these requirements?

Answer and explanation

Answer: C. Near-duplicates differ slightly, so similarity-based clustering is required to find them, and keeping one representative removes the over-weighting. Byte-identical matching misses them entirely. Fewer epochs reduces exposure to all data equally. Shuffling changes order without changing frequency.

50. A feature engineering step must be reusable across several models maintained by different teams. Which solution meets these requirements?

Answer and explanation

Answer: C. A shared feature store distributes one computed definition, which eliminates divergence between teams. Documentation and copied code both drift as teams modify their versions. Independent derivation guarantees inconsistency.

51. Images of varying dimensions must be prepared for a model expecting a fixed input size, and aspect ratio must be preserved. Which approach is appropriate?

Answer and explanation

Answer: A. Resizing with padding preserves the aspect ratio while producing fixed dimensions. Stretching distorts the content. Centre cropping discards content at the edges, which may carry the subject. Discarding images biases the dataset toward one shape.

52. A tokenizer must be selected for a fine-tuning dataset containing substantial non-English text. Which consideration applies?

Answer and explanation

Answer: C. A tokenizer with poor coverage of a language splits its words into many sub-word pieces, inflating token counts, cost, and effective context use. Vocabulary size alone does not indicate coverage of a specific language. A whitespace tokenizer does not match the model's expected token space. The tokenizer must match the model rather than being chosen independently.

53. A feature pipeline must be updated, and the change must not silently alter the meaning of an existing feature consumed by deployed models. Which approach is appropriate?

Answer and explanation

Answer: B. Versioning the feature means deployed models continue to receive the semantics they were trained on until they are migrated deliberately. Redefining in place changes the input distribution of every consumer silently. Notifying after deployment is too late. Holding the change indefinitely blocks the improvement for everyone.

54. An embedding step must process a large corpus within a cost budget, and the corpus contains many duplicate passages. Which approach is appropriate?

Answer and explanation

Answer: D. Embedding cost scales with the number of passages processed, so removing duplicates first avoids paying for the same computation repeatedly. Deduplicating afterwards pays for every duplicate first. A smaller dimension reduces unit cost without avoiding redundant work. Smaller batches spread the same total cost.

55. A tabular dataset must be prepared for supervised training without leaking information between splits. (Place the steps in the correct order.)

Answer and explanation

Answer: B → C → A → D. Splitting comes first because any statistic computed before the split is contaminated by data the model must not see. Fitting transformations on the training split alone is what prevents that leakage. The same fitted transformation is then applied to the other splits so they are represented consistently without influencing the fit. Training follows. Fitting a scaler on the whole dataset before splitting is the most common form of this error and inflates validation scores.

56. Before fine-tuning on support transcripts, the team must confirm that one customer segment is not over-represented. Which solution meets these requirements?

Answer and explanation

Answer: B. Clarify computes pre-training bias metrics such as class imbalance and difference in proportions across a chosen facet, which is a representation check on the dataset. Debugger inspects tensors during training rather than assessing a dataset. Comparing outputs after launch measures the consequence rather than the cause. Adding data without changing its composition does not correct the proportion.

57. A multimodal dataset combining numeric, text, and image assets must be assessed for distributional bias across all three modalities. Which solution meets these requirements?

Answer and explanation

Answer: D. Bias must be assessed against a consistent facet across every modality, because a dataset can be balanced in its numeric features while its images or text are not. Measuring only numeric features leaves the other modalities unexamined. Equal record counts per modality is not what balance across a facet means. Measuring only a combined vector obscures which modality carries the imbalance.

58. Prompt and response pairs collected for fine-tuning must be validated so unsafe or malformed examples are not learned. Which solution meets these requirements?

Answer and explanation

Answer: D. Screening before training prevents the model learning from unsafe or malformed examples at all. Evaluating afterwards means the behaviour is already in the weights. A lower learning rate reduces but does not eliminate the influence. Diluting a harmful example does not remove what it teaches.

59. A dataset contains duplicate records, missing values in several columns, and extreme values in one numeric feature. Which sequence of cleaning steps is appropriate?

Answer and explanation

Answer: D. Deduplicating first prevents duplicates distorting the statistics used for imputation, and extreme values must be judged as errors or genuine observations before being altered. Dropping every record with any missing value discards large amounts of usable data. Removing everything beyond two standard deviations deletes legitimate observations by rule. Replacing extremes with the mean fabricates values before their nature is established.

60. A model trained on a dataset that included a field recorded only after the outcome was known performs almost perfectly in validation and fails in production. Which cause explains this?

Answer and explanation

Answer: A. Near-perfect validation collapsing in production is the signature of leakage, where a feature encodes the outcome and cannot exist at inference. Underfitting produces poor validation performance as well. Class imbalance affects which class is predicted rather than inflating validation scores. Concept drift develops over time rather than appearing immediately at launch.

61. A time series dataset must be split for training and validation without leaking future information. Which solution meets these requirements?

Answer and explanation

Answer: A. Time series must be split by time so the model never sees future observations during training, which mirrors how it will be used. A random split places future records in training. Stratifying by target ignores temporal ordering. Random-fold cross-validation likewise mixes future and past.

62. A dataset's label distribution differs substantially between the training set and recent production data. Which conclusion follows?

Answer and explanation

Answer: B. A shifted label distribution means the training data is unrepresentative, and retraining on current data is the remedy. Overfitting is a different failure visible as a train and validation gap. Changing the metric conceals the shift. Resampling production to match training distorts what the model actually sees.

63. A team must decide whether to impute or drop records with missing values in a feature that is missing for 40 percent of rows. Which consideration is most important?

Answer and explanation

Answer: D. Missingness that correlates with the target is itself a signal, and imputing it away destroys information while dropping the rows biases the sample. Data type affects the imputation method rather than the decision. Affordability matters only after the mechanism is understood. Schema position is irrelevant.

64. A bias assessment must be repeated on every retraining run rather than performed once at project start. Which solution meets these requirements?

Answer and explanation

Answer: D. Embedding the analysis in the pipeline means every version is assessed and the metrics are recorded alongside it. A quarterly review leaves versions unassessed between cycles. Reusing earlier metrics assumes the data has not changed. Aggregate accuracy can hold steady while group-level disparity grows.

65. A deduplication step must identify records describing the same entity where the text differs slightly. Which approach is appropriate?

Answer and explanation

Answer: D. Near-duplicates require similarity matching with a threshold validated against known pairs, since exact matching misses them. Exact matching by definition cannot catch slight differences. Identifier-based removal assumes an identifier exists and is consistent. Leaving duplicates over-weights those entities during training.

66. An outlier in a sensor feature is traced to a recording fault rather than a genuine reading. Which action is appropriate?

Answer and explanation

Answer: B. Values known to be faulty are errors rather than observations, so removing or correcting them with a documented rule is correct. Retaining known errors trains the model on noise. Capping without investigation treats errors and genuine extremes identically. Mean replacement fabricates plausible values for records known to be wrong.

67. A content safety screen must be applied to a fine-tuning dataset assembled from public web text. Which approach is appropriate?

Answer and explanation

Answer: B. Screening every record before training prevents unsafe content entering the weights, and recording counts documents what was removed. A fifty-record sample cannot characterise a large corpus. Screening outputs afterwards means the behaviour is already learned. Built-in safety can be eroded by fine-tuning on unsafe data.

68. A model's error rate is acceptable overall but materially worse for one customer segment. Which action should be taken first?

Answer and explanation

Answer: C. Disparate performance usually traces to representation or distribution in the training data, so that is the first thing to examine. More capacity may fit the majority better still. Adjusting a threshold per segment changes the operating point without addressing the cause and raises fairness questions of its own. Excluding the segment conceals the problem.

69. A validation rule must confirm that every record in a training dataset conforms to the expected schema before training starts. Which solution meets these requirements?

Answer and explanation

Answer: B. An automated ruleset evaluated before training catches violations deterministically and stops the run. Manual sampling is inconsistent. Relying on the training job to fail wastes the compute already consumed and may not catch every violation. Reviewing logs afterwards is too late.

70. Ground truth labels produced by several annotators disagree on a proportion of records. Which approach is appropriate?

Answer and explanation

Answer: B. Agreement measurement identifies where the labelling task itself is ambiguous, and consolidation plus clearer guidance addresses the cause. Taking the first label is arbitrary. Removing disputed records discards the hardest and often most informative cases. Averaging label encodings produces a value no annotator chose.

71. A bias metric must be selected to assess whether a model's positive predictions are distributed proportionally across groups. Which consideration applies?

Answer and explanation

Answer: D. Fairness metrics encode different definitions that cannot all be satisfied simultaneously, so the choice must follow an agreed definition. Choosing the most favourable metric is motivated reasoning. Averaging incompatible definitions produces a meaningless number. Computation speed is irrelevant to correctness.

72. A validation step must confirm that a fine-tuning dataset's prompts and completions are correctly paired. Which approach is appropriate?

Answer and explanation

Answer: C. Structural validation catches missing fields and a stratified sample catches mispairing that structure cannot detect. A record count confirms volume rather than content. Valid JSON confirms format. A length check catches one specific defect.

73. An image dataset must be checked for label noise before it is used for training. Which approach is appropriate?

Answer and explanation

Answer: A. High-confidence misclassifications are where the model disagrees strongly with the label, which is the most efficient place to find label errors. A random sample finds errors only in proportion to their rarity. More epochs makes the model fit the noise. File size is unrelated to label correctness.

74. A dataset's quality checks pass but the resulting model behaves oddly on a subset of inputs the checks did not cover. Which improvement is appropriate?

Answer and explanation

Answer: A. Quality rules should grow from observed failures, which is how the check set becomes a record of what has gone wrong before. Tightening unrelated checks rejects valid data without catching this defect. Diluting the subset leaves the behaviour unaddressed. Removing it from evaluation hides the failure rather than fixing it.

75. A bias assessment must consider a facet that is not recorded in the dataset. Which approach is appropriate?

Answer and explanation

Answer: B. Assessing a facet requires the data, and where it cannot be collected the limitation should be documented rather than worked around. Inferring a protected characteristic from proxies creates a new privacy problem and an unreliable label. Skipping silently leaves a known gap undisclosed. Substituting a different facet and calling it equivalent misrepresents the assessment.