88 practice questions for Domain 1 of the AWS Certified Data Engineer - Associate (DEA-C01) exam, which makes up 34% of its scored content. Your answers count towards one score and one timer for the whole exam.
Domain 1: Data Ingestion and Transformation
1. A pipeline must ingest clickstream events at roughly 50,000 records per second, preserve ordering within each user session, and allow several independent consumers to read the same records. Which solution meets these requirements?
Answer and explanation
Answer: B. Kinesis sustains this ingestion rate, guarantees ordering within a shard so a session partition key keeps each session ordered, and lets multiple consumers read the same stream independently. A standard SQS queue provides no ordering and delivers each message to one consumer. SNS pushes notifications without retention or replay. Listing an S3 bucket is polling rather than streaming ingestion.
2. A Kinesis Data Streams consumer is falling behind and the iterator age is rising while each shard is at its read throughput limit. Which solution meets these requirements?
Answer and explanation
Answer: B. A rising iterator age with shards at their throughput ceiling means the stream lacks parallelism, so adding shards raises capacity and allows more concurrent consumers. Longer retention buys time before data expires without helping the consumer keep up. Smaller records raise the record count for the same volume. Polling more often cannot exceed the per-shard read limit.
3. Several consumers read the same Kinesis Data Stream and compete for the shared per-shard read throughput. Which solution meets these requirements?
Answer and explanation
Answer: A. Enhanced fan-out provisions dedicated throughput per consumer per shard and pushes records over HTTP/2, so consumers do not compete. More shards raise total capacity but the sharing behaviour is unchanged. Firehose is a delivery service rather than a consumer model for custom applications. Polling less often reduces throughput further.
4. An on-premises Oracle database must be replicated continuously into Amazon S3 with minimal downtime at cutover. Which solution meets these requirements?
Answer and explanation
Answer: A. DMS performs an initial full load and then applies ongoing change data capture, which is the standard pattern for continuous replication with a short cutover. A full-load-only task leaves a growing delta and implies a long outage. Snowball moves data physically and supplies no continuous change feed. AppFlow integrates SaaS applications rather than on-premises relational databases.
5. Data must be ingested from Salesforce into Amazon S3 on a schedule without writing custom API integration code. Which solution meets these requirements?
Answer and explanation
Answer: B. AppFlow provides managed connectors to SaaS applications including Salesforce and transfers data on a schedule with no custom code. A Glue job calling the API is exactly the custom integration being avoided. DMS migrates relational and NoSQL databases rather than SaaS applications. A polling Lambda function is also custom code.
6. An Amazon Data Firehose delivery stream must write records to Amazon S3, and the team must control how often a batch is delivered. Which combination of settings controls this? (Select TWO.)
Answer and explanation
Answer: B, E. Firehose flushes when either the configured buffer size or the buffer interval is reached, whichever comes first, and these are the only two delivery triggers. Source shard count affects ingestion throughput rather than delivery batching. Storage class and catalog version have no bearing on when a batch is written.
7. Records occasionally arrive twice in a streaming pipeline because the producer retries on timeout. Which solution meets these requirements?
Answer and explanation
Answer: A. Streaming systems provide at-least-once delivery, so consumers must tolerate duplicates by keying on a deterministic identifier. Retention governs how long records persist rather than whether they duplicate. Disabling producer retries loses records when a request genuinely fails. Adding DISTINCT to every query pushes cost onto readers and leaves the stored data duplicated.
8. An AWS DataSync task copying an on-premises NFS share to Amazon S3 must verify that every transferred file matches the source. Which solution meets these requirements?
Answer and explanation
Answer: C. DataSync provides a built-in verification option that compares source and destination data integrity as part of the task. Comparing object counts detects missing files but not corrupted ones. Versioning retains copies without verifying them. A second run confirms the file list matches without validating contents.
9. An ingestion job reading from an Amazon DynamoDB table is receiving throttling errors during peak periods. Which combination of steps meets these requirements? (Select TWO.)
Answer and explanation
Answer: A, C. Backoff with jitter spreads retries so clients do not converge on the same instant, and raising or removing the capacity ceiling absorbs the peak. A tight retry loop amplifies the throttling. Larger items consume more read capacity units. Disabling SDK retries removes the resilience that handles transient throttling.
10. An existing Apache Kafka workload must move to AWS with minimal change to its producers and consumers. Which solution meets these requirements?
Answer and explanation
Answer: D. MSK runs Apache Kafka itself, so existing clients work with configuration changes rather than rewrites. Kinesis has a different API and would require rewriting clients. SQS is a queue with different delivery semantics. EventBridge routes events and is not a Kafka-compatible broker.
11. A streaming pipeline must invoke an AWS Lambda function for each batch of records read from an Amazon Kinesis Data Stream. Which solution meets these requirements?
Answer and explanation
Answer: D. An event source mapping polls the stream on the function's behalf and invokes it with configurable batches. Scheduled self-polling reimplements the mapping with more failure modes. Routing through SNS adds a component and loses stream ordering. Writing to S3 first converts a streaming pipeline into a batch one.
12. An AWS Glue job must process only the files added to an Amazon S3 prefix since its previous run. Which solution meets these requirements?
Answer and explanation
Answer: D. Job bookmarks persist state about what a job has already processed so subsequent runs pick up only new data. A crawler updates the catalog with new partitions but does not stop the job reprocessing old files. More workers make a full reprocess faster rather than avoiding it. Data Quality evaluates records against rules and performs no incremental tracking.
13. A Spark job joins a 2 TB fact table with a 40 MB dimension table, and runtime is dominated by shuffling data between executors. Which solution meets these requirements?
Answer and explanation
Answer: D. When one side of a join is small enough to fit in memory, broadcasting it lets each executor join locally and removes the shuffle. A read timeout affects error handling rather than shuffle volume. CSV output is larger and slower to write than Parquet. Fewer executors reduce parallelism and lengthen the job.
14. An ETL script written in pandas processes a 3 GB dataset and does not require distributed processing. Which solution meets these requirements at the lowest cost?
Answer and explanation
Answer: D. Python shell jobs run a single-process script on a fraction of a DPU, which suits modest datasets. A ten-worker Spark job provisions far more capacity than 3 GB requires. An EMR cluster carries cluster overhead for a task a single process handles. A streaming job applies to continuous data rather than a bounded dataset.
15. Source records arrive as deeply nested JSON, and analysts query only four top-level fields with filters on event date. Which combination of steps meets these requirements? (Select TWO.)
Answer and explanation
Answer: A, B. A columnar format lets the query engine read only the four needed columns, and partitioning by the filtered column prunes whole partitions before scanning. Gzipped JSON remains row-oriented and is not splittable in a way that avoids reading unused fields. A single unpartitioned CSV forces a full scan. Workgroups manage query isolation and cost controls without reducing bytes scanned.
16. A data lake table receives frequent row-level updates and deletes, and analysts need consistent snapshot reads while writes are in progress. Which solution meets these requirements?
Answer and explanation
Answer: B. Iceberg maintains metadata defining a consistent snapshot, so readers see a stable view while writers commit, and it supports row-level updates and deletes. Plain Parquet with a crawler has no transactional guarantee. CSV partitions offer neither ACID semantics nor efficient updates. A hand-maintained manifest reimplements a fraction of a table format without atomicity.
17. A pipeline must merge incoming change records into an existing dataset in Amazon S3, updating rows that exist and inserting those that do not. Which solution meets these requirements?
Answer and explanation
Answer: A. Open table formats implement row-level upserts through a MERGE operation, which is what change data capture into a lake requires. Full overwrite is expensive and loses history. Appending everything pushes resolution cost onto every reader. A crawler updates metadata and performs no merge.
18. An AWS Glue job fails intermittently with an out-of-memory error on one executor while the others finish quickly. Which cause should the engineer investigate first?
Answer and explanation
Answer: B. One executor failing while others idle is the signature of skew, where a single key concentrates the data. More workers do not help when work is concentrated on one partition. A permission problem produces access errors rather than an out-of-memory failure. A missing partition produces a schema or data resolution failure.
19. Parquet files in a data lake must be compressed for frequent Amazon Athena queries where query speed matters more than the smallest file size. Which solution meets these requirements?
Answer and explanation
Answer: B. Snappy decompresses quickly and works well with Parquet's columnar structure, which is why it is the common default for query workloads. Maximum-level gzip trades substantial CPU for a modest size gain. bzip2 is slower still. Uncompressed files increase the bytes Athena scans and therefore cost.
20. A staging layer receives records appended row by row, and the source system adds fields every few weeks. Which file format is most appropriate?
Answer and explanation
Answer: C. Avro is row-oriented, which suits append-heavy writes, and it carries a schema supporting evolution rules. Parquet and ORC are columnar and optimised for analytical reads rather than row-by-row writes. Fixed-width text has no schema metadata and no evolution support.
21. An AWS Glue job must read from an on-premises PostgreSQL database over a private connection. Which combination of steps meets these requirements? (Select TWO.)
Answer and explanation
Answer: D, E. A Glue connection defines the JDBC target and network placement so the job runs elastic network interfaces in the specified subnet, and Secrets Manager supplies the credentials at run time. Exposing the database publicly contradicts the private requirement. An S3 endpoint does not reach a JDBC source. Embedded credentials place a secret in code.
22. A transformation must mask email addresses and normalize inconsistent date formats, and the analysts performing the work prefer not to write code. Which solution meets these requirements?
Answer and explanation
Answer: B. DataBrew offers a visual, no-code interface with built-in transformations for masking and date normalisation, and its recipes can be reused in jobs. A Glue Spark script and an EMR notebook both require code. An Athena view expresses SQL rather than a visual preparation workflow.
23. An ETL job fails partway through and leaves partial output in the target Amazon S3 prefix, which downstream consumers then read. Which solution meets these requirements?
Answer and explanation
Answer: D. Writing to a staging location and switching to it only on success means consumers never observe partial output. A longer timeout addresses one failure cause without making output atomic. Retrying compounds partial writes. Smaller files do not change the visibility of an incomplete run.
24. An Amazon EMR cluster must process a nightly batch that tolerates interruption, at the lowest cost that keeps the job reliable. Which solution meets these requirements?
Answer and explanation
Answer: B. Core nodes hold HDFS data and should be stable, while task nodes provide compute only and can be reclaimed without losing data, which captures most of the saving safely. Putting the primary node on Spot risks losing the whole cluster. All On-Demand forgoes the saving. A single instance abandons distributed processing.
25. A workflow must run an AWS Glue job, wait for it to succeed, then run two independent jobs in parallel, and alert on any failure. Which solution meets these requirements?
Answer and explanation
Answer: A. Step Functions expresses sequencing, parallel branches, retries, and catch handlers as an explicit state machine with visible execution history. EventBridge rules trigger targets on events but express no parallel branch or unified failure handling. A Glue job invoking others hides orchestration in application code. An SQS queue orders messages without expressing dependencies.
26. A team must orchestrate data pipelines with existing Apache Airflow DAGs without operating the scheduler, workers, or web server. Which solution meets these requirements?
Answer and explanation
Answer: A. MWAA runs the Airflow scheduler, workers, and web server as a managed service while keeping DAG compatibility. Installing Airflow on EC2 leaves patching, scaling, and availability with the team. Glue workflows are not Airflow and cannot run existing DAGs. Flink on EMR is a different processing framework.
27. An orchestration must fan out processing across 500 partitions in parallel and continue when a small number of partitions fail. Which solution meets these requirements?
Answer and explanation
Answer: A. A distributed Map state processes a large collection in parallel with configurable concurrency and failure tolerance. A Parallel state requires each branch to be defined and does not scale to 500. A Choice state directs flow rather than iterating. A Wait state introduces delay.
28. A pipeline must start automatically when new objects land in an Amazon S3 prefix, and must not poll for them. Which solution meets these requirements?
Answer and explanation
Answer: C. An EventBridge rule on the object created event starts the pipeline as objects arrive with no polling. A listing function polls by definition. An Airflow sensor also polls, occupying a worker slot while it waits. An hourly schedule adds latency and runs whether or not data arrived.
29. A pipeline step must notify a downstream team when a stage completes, and the notification must reach both an email list and a queue for automated processing. Which solution meets these requirements?
Answer and explanation
Answer: B. SNS delivers each published message to every subscriber, which supports email and a queue from one publish. A queue delivers each message to a single consumer and email recipients cannot poll it. Sending separately from the pipeline couples it to each destination. Writing to S3 requires both destinations to poll.
30. An AWS Lambda function in a data pipeline must be prevented from consuming the account's entire concurrency pool during a spike. Which solution meets these requirements?
Answer and explanation
Answer: B. Reserved concurrency both guarantees and caps a function's concurrency so it cannot exhaust the account pool. Provisioned concurrency pre-initializes environments without capping scale. Memory and timeout affect execution characteristics rather than concurrent execution count.
31. A serverless data pipeline of Lambda functions, a Step Functions state machine, and a DynamoDB table must be packaged for repeatable deployment. Which solution meets these requirements?
Answer and explanation
Answer: D. SAM provides shorthand syntax for serverless resources and deploys them as one stack, which makes deployment repeatable. Console creation is not repeatable. A CLI script recreates infrastructure-as-code without state tracking or rollback. Generating a template from manual resources may not reproduce the environment reliably.
32. An Amazon Redshift query that aggregates a large fact table runs slowly, and the execution plan shows a full scan of every partition. Which solution meets these requirements?
Answer and explanation
Answer: B. A sort key lets Redshift use zone maps to skip blocks whose value range falls outside the filter, which removes most of the scan. More nodes scan the same data faster at higher cost. SELECT * increases the data read. A longer timeout permits a slow query to finish rather than making it fast.
33. Transformation logic that runs inside Amazon Redshift must be reusable, parameterized, and callable from several pipelines. Which solution meets these requirements?
Answer and explanation
Answer: A. A stored procedure encapsulates parameterized logic inside the database where it can be versioned and called consistently. Copying SQL into each pipeline guarantees drift. An Athena view queries S3 data rather than executing logic inside Redshift. A Lambda wrapper adds a component for logic the database can hold.
34. A data engineering team must manage pipeline code in source control with feature branches merged after review. Which set of Git actions supports this workflow?
Answer and explanation
Answer: D. Branching, committing to the branch, and merging after review is the standard workflow that supports review before code reaches the main line. Committing directly to main removes the review point. Copying directories abandons version history. A repository per change makes history unusable.
35. An AWS Lambda function in a pipeline must write temporary files larger than its ephemeral storage allows. Which solution meets these requirements?
Answer and explanation
Answer: C. Mounting EFS gives a Lambda function shared file system storage well beyond its ephemeral limit. Memory allocation does not increase ephemeral disk. The /tmp directory is the ephemeral storage that is already insufficient. A longer timeout does not create more space.
36. A Kinesis Data Streams producer is receiving throughput exceeded errors on one shard while other shards are underused. Which cause should be investigated first?
Answer and explanation
Answer: A. One hot shard with others idle is the signature of a skewed partition key. Retention governs how long records persist. Consumer speed affects iterator age rather than producer throughput. On-demand mode scales shards but still applies per-shard limits to a hot key.
37. An AWS DMS task performing ongoing replication is falling behind the source database. Which combination of steps meets these requirements? (Select TWO.)
Answer and explanation
Answer: C, D. Lag usually traces to an undersized replication instance or a source purging change logs before they are read. A full load only task abandons ongoing replication. Reducing target capacity slows the task further. Restarting repeats the full load and loses the change position.
38. An Amazon Data Firehose delivery stream must transform records before writing them to Amazon S3. Which solution meets these requirements?
Answer and explanation
Answer: A. Firehose invokes a transformation function on each buffered batch before delivery, which transforms in flight with no separate job. A post-delivery Glue job adds latency and a second pass. Producer-side transformation pushes the logic into every producer. Dynamic partitioning organises output paths rather than transforming records.
39. An ingestion pipeline must detect duplicate records arriving from a source that retries on timeout. Which approach is appropriate?
Answer and explanation
Answer: C. A deterministic identifier lets the pipeline recognise a resent record cheaply. Comparing against all existing records scales poorly. Ignoring duplicates corrupts aggregates. A shorter timeout increases retries.
40. A batch ingestion job reads from an API that limits requests per second. Which approach is appropriate?
Answer and explanation
Answer: A. Matching the request rate to the limit with backoff on the occasional throttle respects the API and completes reliably. Parallel copies multiply requests against the same limit. A single large call may exceed the API's payload or time limits. Immediate retries amplify the throttling.
41. An Amazon MSK cluster's consumers must resume from where they stopped after a restart. Which mechanism provides this?
Answer and explanation
Answer: D. Committed offsets are what a restarted consumer resumes from. Retention keeps records available but does not record position. Acknowledgement settings govern producer durability. Replication factor governs fault tolerance.
42. A Glue job's output has many partitions containing a handful of tiny files each. Which cause should be investigated first?
Answer and explanation
Answer: B. Many writers into many partitions produce many small files, which is a write pattern problem addressed by coalescing or repartitioning before the write. Worker size affects speed. Source compression affects reading. Bookmarks affect which input is processed.
43. A transformation must produce a slowly changing dimension with history preserved. Which write pattern is appropriate?
Answer and explanation
Answer: C. Closing the existing row and inserting a new one preserves history with unambiguous validity periods. Overwriting destroys history. Appending without closing leaves two rows both appearing current. Delete and insert loses the previous value.
44. A Spark job must join a 5 TB table with a 3 TB table on a key that is evenly distributed. Which approach is appropriate?
Answer and explanation
Answer: A. Two large tables with an even key are joined by shuffling both on the key, which is the correct pattern when neither fits in executor memory. Broadcasting 3 TB exceeds memory. Collecting to the driver is worse. A single partition removes all parallelism.
45. A Glue job must apply a transformation that uses a Python library not included in the Glue runtime. Which approach is appropriate?
Answer and explanation
Answer: B. Glue supports additional Python modules supplied as job parameters, which install at start. Manual installation on workers is not a supported path. Rewriting discards a working library. Switching platforms is disproportionate.
46. Records in a data lake table must be updated in place without rewriting whole partitions. Which table format provides this?
Answer and explanation
Answer: C. Open table formats implement row-level updates through metadata and merge-on-read or copy-on-write semantics. Plain Parquet, CSV, and JSON all require rewriting the affected files or partitions.
47. A transformation must handle a source column whose type changed from integer to string mid-way through the historical data. Which approach is appropriate?
Answer and explanation
Answer: A. Casting to a common type unifies the schema across both eras. Dropping rows discards history. Two columns push the problem onto every consumer. Waiting for the source may never resolve.
48. A Step Functions workflow must retry a Glue job on transient failure but stop after three attempts. Which configuration is appropriate?
Answer and explanation
Answer: A. Retry configures which errors are retried, how often, and how many times. Catch handles an error after retries are exhausted. A Wait state delays without retrying. Parallel copies triple the work.
49. An Airflow DAG on Amazon MWAA must not start a new run while the previous run is still executing. Which configuration is appropriate?
Answer and explanation
Answer: B. Limiting active runs to one prevents a new run beginning while one executes. A longer interval reduces the chance without preventing overlap. Fewer workers slows all DAGs. A sensor adds a wait without expressing the constraint.
50. A pipeline must run daily but skip execution when the input for that day has not arrived. Which approach is appropriate?
Answer and explanation
Answer: D. An explicit check with a clean skip records that the run was intentionally not performed. Failing on an expected condition generates noise. Processing incomplete data produces wrong output. Waiting indefinitely holds the workflow open.
51. A workflow step depends on the output of two independent upstream steps. Which Step Functions structure is appropriate?
Answer and explanation
Answer: B. A Parallel state runs the independent steps concurrently and completes when both finish, so the dependent step follows naturally. Sequential states serialise work that could run concurrently. Separate machines lose the coordination. A Map state iterates over a collection rather than running distinct steps.
52. A Lambda function in a pipeline must process a batch from a queue and must not lose records if it fails partway. Which approach is appropriate?
Answer and explanation
Answer: C. Partial batch failure reporting returns only the failed records for retry while successful ones are acknowledged. Whole-batch handling retries everything on any failure. One record per invocation loses batching efficiency. A longer timeout does not address failure mid-batch.
53. A SQL transformation must compute a running total ordered by date within each account. Which construct is appropriate?
Answer and explanation
Answer: B. A window function computes a running aggregate per row within an ordered partition. GROUP BY collapses rows to one per account. A self join is possible but expensive and hard to read. A per-account subquery returns the grand total rather than a running one.
54. A pipeline's infrastructure must be deployed identically to development and production accounts. Which approach is appropriate?
Answer and explanation
Answer: B. One template with parameters gives reproducible deployment. Console creation drifts. Export and import is not a reliable reproduction mechanism. Shared resources remove the separation.
55. A data engineer must query an Amazon Redshift cluster from a Lambda function without managing database connections. Which approach is appropriate?
Answer and explanation
Answer: C. The Data API removes connection management and suits serverless callers. Per-invocation JDBC connections are slow and can exhaust the cluster. A pool inside a function does not persist reliably. Athena queries S3 rather than Redshift.
56. A pipeline's code must be reviewed before it reaches the main branch. Which Git workflow supports this?
Answer and explanation
Answer: A. Feature branches with pull requests insert review before code reaches main. Direct commits remove the gate. Copying abandons version control. Tags record without gating.
57. A Kinesis Data Streams application must read the same records as another consumer without sharing read throughput. Which solution meets these requirements?
Answer and explanation
Answer: B. Enhanced fan-out provisions dedicated throughput per consumer per shard. More shards raise total capacity while the sharing behaviour remains. Retention governs persistence. Copying the stream duplicates ingestion cost.
58. An AWS DMS migration must validate that the target data matches the source after the load. Which configuration is appropriate?
Answer and explanation
Answer: B. DMS data validation compares rows and reports mismatches. Row counts detect missing rows but not corrupted values. Logging records task activity. Running twice doubles the work without validating.
59. An ingestion pipeline must read from an on-premises Kafka cluster during a phased migration. Which approach is appropriate?
Answer and explanation
Answer: C. Topic replication bridges the clusters during migration without changing producers. Direct internet connection exposes the cluster. Daily file export loses the streaming semantics. Dual writes require changing every producer.
60. A Firehose delivery stream must write to Amazon S3 in Parquet using a schema from the Data Catalog. Which configuration is required?
Answer and explanation
Answer: B. Firehose record format conversion uses a catalog schema to write Parquet in flight. A scheduled conversion job adds a second pass. Producers send records rather than columnar files. Compression reduces size without changing format.
61. A streaming ingestion pipeline must not lose records if the destination is briefly unavailable. Which configuration is appropriate?
Answer and explanation
Answer: A. A backup destination captures records that fail delivery for later reprocessing. A longer buffer delays delivery without handling failure. A slower producer and larger destination reduce the chance without preventing loss.
62. A Glue job must process a dataset that grows each day without reprocessing prior days. Which configuration is appropriate?
Answer and explanation
Answer: C. Job bookmarks persist processing state between runs. A date filter works but requires the partition scheme to align exactly and is maintained manually. Deleting source files destroys the raw layer. Fewer workers does not limit which data is read.
63. A Spark transformation produces one output file per partition, and downstream readers require files of a consistent size. Which approach is appropriate?
Answer and explanation
Answer: C. Repartitioning controls the number and size of output files. More executors produce more small files. Uncompressed output changes size without controlling the count. Changing source partitioning affects reading rather than writing.
64. A pipeline must convert a nested JSON structure into a flat table for analytical queries. Which transformation is appropriate?
Answer and explanation
Answer: C. Flattening or relationalizing produces queryable columns. A single string column requires parsing on every query. Compression reduces size without changing structure. More memory does not make nested JSON columnar.
65. A transformation must deduplicate records while keeping the most recent version of each key. Which approach is appropriate?
Answer and explanation
Answer: C. Ranking by timestamp within the key identifies the most recent version. DISTINCT removes identical rows rather than selecting a version. Keeping the first encountered depends on arbitrary ordering. Taking column maxima produces a record that never existed.
66. A Glue job's transformation logic must be unit tested before deployment. Which approach is appropriate?
Answer and explanation
Answer: D. Testable functions run against sample data give fast feedback. Deploying to inspect output is slow. Running against production data risks live systems. Code review does not execute the logic.
67. A pipeline must process several partitions concurrently with a bounded number in flight. Which Step Functions construct is appropriate?
Answer and explanation
Answer: D. A Map state iterates a collection with a configurable concurrency limit. A Parallel state requires each branch to be defined. A Choice state directs flow. A fixed sequence does not scale to a variable partition count.
68. A workflow must wait for a Glue job to complete before starting the next step. Which integration pattern is appropriate?
Answer and explanation
Answer: D. The run-a-job pattern blocks the state until the job completes. Proceeding immediately ignores the dependency. A fixed wait may be too short or too long. A polling loop reimplements the pattern with more cost.
69. An orchestration must record which input files each pipeline run processed. Which approach is appropriate?
Answer and explanation
Answer: C. Recording the manifest with the execution links the run to specific inputs. A timestamp and a file count do not identify which files. A later bucket listing reflects current contents rather than what the run read.
70. A Lambda function processing a Kinesis stream must not block the shard when one record repeatedly fails. Which configuration is appropriate?
Answer and explanation
Answer: B. A retry limit with an on-failure destination moves the failing record aside so the shard progresses. Timeout and memory affect execution rather than the retry loop. A batch size of one isolates the record but still retries it indefinitely without a limit.
71. A SQL query must return the difference between each row's value and the previous row's value ordered by date. Which construct is appropriate?
Answer and explanation
Answer: D. LAG returns a prior row's value within an ordered window. GROUP BY aggregates. A self join can achieve it but is more expensive and harder to read. An overall average is a different calculation.
72. A data pipeline's Python code must be shared across several Glue jobs without duplication. Which approach is appropriate?
Answer and explanation
Answer: C. A shared library referenced by each job avoids duplication and drift. Copying guarantees divergence. A database is not a code distribution mechanism. Rewriting in SQL duplicates the logic in another form.
73. A pipeline must ingest data from a SaaS application on a schedule without writing API integration code. Which solution meets these requirements?
Answer and explanation
Answer: D. AppFlow provides managed SaaS connectors on a schedule with no code. A custom function or Glue job is the integration code being avoided, and manual export does not scale.
74. An ingestion job must read a compressed file format that cannot be split across workers. Which consequence applies?
Answer and explanation
Answer: D. Non-splittable compression forces single-worker processing, which is why splittable formats are preferred at scale. The file is readable, does not require JSON conversion, and is not faster.
75. A streaming source produces records faster than the configured shard count can accept. Which symptom appears first?
Answer and explanation
Answer: B. Producers are throttled first when ingestion exceeds shard capacity. Rising iterator age indicates a slow consumer, records are not silently discarded, and retention is a configured setting.
76. An ingestion pipeline must preserve the original raw files alongside the processed output. Which approach is appropriate?
Answer and explanation
Answer: D. A separate landing zone preserves raw data for reprocessing with independent lifecycle management. Overwriting and deleting lose the source, and a shared prefix makes the layers indistinguishable.
77. A DMS task must migrate a source table that has no primary key. Which consideration applies?
Answer and explanation
Answer: B. Change data capture needs a unique row identifier to apply updates and deletes. The table can be full-loaded, full-load-only may not meet the requirement, and engines need not match.
78. An ingestion pipeline must handle a source that occasionally sends records out of order. Which approach is appropriate?
Answer and explanation
Answer: A. Ordering by an event timestamp handles late and out-of-order arrival. Rejecting loses data, buffering indefinitely is impractical, and retention governs persistence.
79. A pipeline must ingest data from a database without affecting the source system's performance. Which approach is appropriate?
Answer and explanation
Answer: A. A read replica or log-based capture avoids load on the primary. Querying the primary with limits or longer timeouts still adds load, and a larger instance absorbs rather than avoids it.
80. An ingestion pipeline must confirm that a delivered file is complete before processing begins. Which approach is appropriate?
Answer and explanation
Answer: C. A marker file written after the data signals completeness. Creation events fire before a multipart write completes, fixed waits guess, and expected sizes are often unknown in advance.
81. A Lambda function processing pipeline events must handle a downstream service being temporarily unavailable. Which approach is appropriate?
Answer and explanation
Answer: B. Backoff with a dead-letter destination handles transient and persistent failure separately. Immediate retries amplify load, discarding loses data, and a longer timeout holds resources.
82. A SQL transformation must assign a sequential number to rows within each group ordered by date. Which construct is appropriate?
Answer and explanation
Answer: D. ROW_NUMBER assigns a sequence within an ordered partition. COUNT aggregates, a self join is expensive, and a total is a different value.
83. A pipeline's Python code must handle a schema field that is sometimes absent. Which approach is appropriate?
Answer and explanation
Answer: D. Defaulting handles optional fields predictably. Assuming presence raises errors, rejecting loses valid records, and string conversion does not address absence.
84. A data pipeline's SQL must be version controlled and reviewed like application code. Which approach is appropriate?
Answer and explanation
Answer: A. Repository-stored SQL applied through a pipeline is versioned and reviewed. Console edits, documents, and local files are none of those.
85. A transformation must convert a string column containing dates in several formats into a date type. Which approach is appropriate?
Answer and explanation
Answer: D. Trying known formats with quarantine for the remainder converts what can be converted and preserves the rest for investigation. A single format drops valid data, string storage pushes parsing to readers, and defaulting fabricates values.
86. A pipeline's transformation logic must be shared between a batch job and a streaming job. Which approach is appropriate?
Answer and explanation
Answer: A. A shared library keeps the logic identical. Separate implementations diverge, chaining the jobs changes the architecture, and converting the streaming job abandons its purpose.
87. A Redshift stored procedure must handle an error without aborting the entire transaction. Which construct is appropriate?
Answer and explanation
Answer: A. Exception handling catches errors and defines the response. Frequent commits change transaction boundaries, an initial rollback is meaningless, and a timeout governs duration.
88. A transformation must produce the same result when it is re-run over the same input. Which property is required?
Answer and explanation
Answer: B. Determinism means the same input produces the same output regardless of when it runs. Destination idempotency governs repeated writes, and partitioning and compression affect layout and size.