Data Operations and Support

Domain 3: Data Operations and Support

72 practice questions for Domain 3 of the AWS Certified Data Engineer - Associate (DEA-C01) exam, which makes up 22% of its scored content. Your answers count towards one score and one timer for the whole exam.

Domain 3: Data Operations and Support

22% of scored content · 72 practice questions

146. An Amazon MWAA environment shows a DAG stuck in the queued state with tasks that never start. Which cause should the engineer investigate first?

Answer and explanation

Answer: A. Tasks queued but never starting is characteristic of exhausted worker capacity or a concurrency limit preventing scheduling. A syntax error prevents the DAG appearing at all rather than queueing. Bucket versioning affects DAG file history. The Airflow version does not by itself block queued tasks.

147. A data preparation step defined visually by analysts must run automatically as part of a nightly pipeline. Which solution meets these requirements?

Answer and explanation

Answer: C. A DataBrew recipe is built visually and runs as a scheduled job, so the analysts' work executes directly. QuickSight prepares data for visualisation rather than producing a pipeline output. Reimplementation introduces translation error and delay. An Athena view expresses SQL rather than a visual recipe.

148. An application must invoke AWS Glue and Amazon Athena operations programmatically from Python inside an existing service. Which solution meets these requirements?

Answer and explanation

Answer: D. The AWS SDK exposes the service APIs natively with typed responses, retries, and credential handling. Shelling out to the CLI adds a process boundary and requires parsing text. An API Gateway wrapper adds a component for calls the SDK already makes. A scheduled job removes the programmatic control described.

149. A pipeline must run an AWS Glue job at 02:00 each day and also whenever a manifest file arrives in Amazon S3. Which solution meets these requirements?

Answer and explanation

Answer: B. EventBridge supports both time-based schedules and event patterns, so one job can be started by either trigger without polling. Polling the bucket adds latency and cost. Duplicating the job doubles maintenance. An Airflow sensor polls and occupies a worker slot while waiting.

150. An AWS Lambda function must transform each object as it lands in Amazon S3, and objects sometimes exceed the function's ephemeral storage. Which solution meets these requirements?

Answer and explanation

Answer: A. Streaming avoids materialising the whole object, and mounting EFS provides working storage well beyond the ephemeral limit. Memory allocation does not increase ephemeral disk beyond its configured size. A longer timeout does not create space. Batch size governs how many events arrive rather than the size of one object.

151. A recurring Amazon Athena query must be run each morning and its results written to a known Amazon S3 location for downstream consumers. Which solution meets these requirements?

Answer and explanation

Answer: D. A scheduled function that starts the query execution automates the run and Athena writes results to the workgroup's output location. Manual execution depends on a person. A view moves the query to consumers rather than producing a scheduled output. A crawler updates metadata and runs no query.

152. Business analysts must build interactive dashboards over data queried from Amazon Athena, with per-user access control. Which solution meets these requirements?

Answer and explanation

Answer: A. QuickSight connects natively to Athena and applies its own user and group permissions to datasets and dashboards. Spreadsheet exports produce uncontrolled copies. OpenSearch duplicates the data into a search engine. CloudWatch dashboards display operational metrics rather than business data.

153. Analysts repeatedly write the same complex three-table join and must be given a consistent way to query it in Amazon Athena. Which solution meets these requirements?

Answer and explanation

Answer: B. A view encapsulates the join logic under one name so every analyst queries it consistently and the definition is maintained in one place. A CTAS table becomes stale between runs and duplicates storage. Personal query history is not shared. Workgroups isolate execution and cost rather than sharing logic.

154. A data engineer must explore a large dataset interactively with Apache Spark without provisioning a cluster. Which solution meets these requirements?

Answer and explanation

Answer: C. Athena notebooks run Apache Spark serverlessly, which suits interactive exploration with no cluster to provision. An EMR cluster is exactly the provisioning being avoided. A Glue Python shell job runs a single process without Spark and is not interactive. Lambda has a fifteen-minute limit and is not a Spark environment.

155. A dataset contains duplicate rows and inconsistent capitalization in a category column, and analysts must clean it without writing code. Which solution meets these requirements?

Answer and explanation

Answer: A. DataBrew provides a visual recipe with built-in cleansing transformations and no code. A Glue Spark script and a Redshift stored procedure both require code. An Athena view masks the problem at query time without cleaning the stored data.

156. A report must present a seven-day rolling average of daily order volume from data held in Amazon Redshift. Which solution meets these requirements?

Answer and explanation

Answer: A. A window function with a moving frame computes a rolling average directly in the query. Returning raw daily totals pushes the calculation into the reporting tool inconsistently. Seven separate queries is the same calculation done by hand. Exporting to compute an average introduces a pipeline for something SQL does natively.

157. Analysts must be able to compare current month revenue against the same month last year, grouped by product line. Which solution meets these requirements?

Answer and explanation

Answer: B. Aligning two periods in one aggregate query gives a single comparable result set that any reporting tool can consume. Manual comparison of two outputs is error-prone and not repeatable. Spreadsheet comparison moves the data out of the platform. A table per month fragments the dataset and complicates every query.

158. A team must be alerted when a nightly pipeline has not completed by 06:00. Which solution meets these requirements?

Answer and explanation

Answer: C. Detecting the absence of a success within a window requires a metric that records completion and an alarm that treats missing data as breaching. A dashboard depends on someone looking. Notifying on start says nothing about completion. A longer timeout delays failure rather than alerting on it.

159. An Amazon Athena query that previously ran in seconds now takes minutes after the table accumulated millions of small files. Which solution meets these requirements?

Answer and explanation

Answer: B. Many small files force a large number of expensive file-open operations, and compacting them restores throughput. A longer timeout allows the slow query to finish without making it faster. Additional columns do not affect file-count overhead. SELECT * scans more data and would be slower still.

160. A team must be notified when an AWS Glue job fails or exceeds its expected duration, without polling the console. Which solution meets these requirements?

Answer and explanation

Answer: B. Glue emits job state change events to EventBridge, and a rule matching failed or timed-out states can publish to SNS immediately with no polling. A dashboard requires a person to look. An S3 notification cannot signal a job that produced no output. A polling function adds cost and latency where an event path exists.

161. An engineer must diagnose a failed Apache Spark job on an Amazon EMR cluster that has already been terminated. Which combination of steps meets these requirements? (Select TWO.)

Answer and explanation

Answer: A, E. Archived logs survive termination and the history server exposes stage and task detail including skew and failures. Flow Logs capture network metadata rather than application behaviour. Config records configuration state. Billing detail reports cost.

162. A Kinesis consumer must be monitored to confirm it is keeping pace with the stream. Which metric should the team alarm on?

Answer and explanation

Answer: D. Iterator age measures how far behind the newest record a consumer is reading, so it directly indicates whether the consumer keeps pace. Shard count is a capacity setting. Incoming bytes describe producer volume. CPU utilization may be high or low for reasons unrelated to lag.

163. Pipeline costs have risen sharply and the team must attribute the increase to specific jobs. Which solution meets these requirements?

Answer and explanation

Answer: A. Tagging resources per job or team and grouping spend by tag attributes cost to a specific job, which is what identifying the cause requires. Halving worker counts is a blind change that may break jobs. Archiving active data breaks queries. Deleting log groups trims a small cost and loses diagnostic history.

164. An Amazon Redshift cluster runs both short interactive queries and long ETL jobs, and the interactive queries are being starved. Which solution meets these requirements?

Answer and explanation

Answer: A. Workload management assigns queries to queues with their own concurrency and memory, so interactive queries are not blocked behind long ETL work. A larger cluster adds capacity without isolating the workloads. A materialized view helps specific queries without resolving contention. A uniform shorter timeout would kill the ETL jobs.

165. Logs from AWS Glue, Amazon EMR, and AWS Lambda must be searchable together during incident investigation. Which solution meets these requirements?

Answer and explanation

Answer: C. CloudWatch Logs ingests from all three compute types and Logs Insights queries across log groups in one place. Downloading logs is slow and not repeatable. Unindexed buckets provide no query layer. Reviewing consoles separately prevents correlating events across services.

166. An auditor must be able to establish which principal started a particular AWS Glue job run. Which solution meets these requirements?

Answer and explanation

Answer: A. CloudTrail records each API call with the calling identity, timestamp, and parameters. Application logs record what the job did rather than who started it. The console shows the run without reliably attributing the caller. Config records configuration changes rather than invocations.

167. A batch must be rejected when more than 2 percent of rows have a null in a required column, and the pipeline must stop. Which solution meets these requirements?

Answer and explanation

Answer: C. A ruleset expresses the completeness threshold declaratively and returns a result the orchestration can branch on to halt the run. A dashboard reports without stopping anything. Logging records the problem without acting. A weekly review is far too slow to prevent bad data propagating.

168. A pipeline must confirm that a numeric column's values fall within an expected range before the data is published. Which solution meets these requirements?

Answer and explanation

Answer: C. A range rule evaluated before publishing prevents out-of-range data reaching consumers. A dashboard reports after publication. Sampling ten rows is neither systematic nor sufficient. Relying on a load failure detects the problem late and may fail the whole load rather than reporting the specific violation.

169. An engineer must understand the distribution, cardinality, and null rates of an unfamiliar dataset before designing a pipeline for it. Which solution meets these requirements?

Answer and explanation

Answer: B. A DataBrew profile job computes distribution, cardinality, null rates, and other column statistics, which is exactly data profiling. A table definition describes types without describing values. A crawler infers schema rather than profiling content. Reading a thousand rows is a sample that misses distribution across the whole dataset.

170. A Glue job must be started whenever a Glue crawler finishes successfully. Which approach is appropriate?

Answer and explanation

Answer: A. A Glue workflow expresses the dependency natively with an event trigger. A fixed schedule assumes the crawler duration. A Lambda intermediary adds a component for a native capability. Continuous polling wastes resources.

171. A Glue DataBrew recipe must be applied to a new dataset each day automatically. Which approach is appropriate?

Answer and explanation

Answer: A. A scheduled or event-triggered DataBrew job runs the recipe without manual action. Manual runs depend on a person. Converting to Spark abandons the visual recipe. Manual application is the automation being replaced.

172. An Amazon Redshift stored procedure must run nightly without an external scheduler. Which approach is appropriate?

Answer and explanation

Answer: B. Redshift supports scheduled queries that run SQL including procedure calls with no external component. Lambda and Step Functions are valid but are external schedulers the requirement excludes. Manual runs depend on a person.

173. An Athena query over a large table is scanning far more data than the filtered result requires. Which cause should be investigated first?

Answer and explanation

Answer: C. Unpartitioned or row-oriented data forces full scans regardless of the filter. A workgroup limit caps scanning rather than reducing it. SELECT * increases columns read but partitioning is the larger lever. Results bucket location affects output rather than scanning.

174. A QuickSight dashboard must refresh its data automatically from a Redshift source. Which approach is appropriate?

Answer and explanation

Answer: D. A scheduled SPICE refresh or direct query keeps the dashboard current automatically. CSV export and manual refresh depend on a person. Rebuilding is disproportionate.

175. A data engineer must run an ad hoc SQL query across data in both Amazon S3 and Amazon Redshift. Which approach is appropriate?

Answer and explanation

Answer: C. Spectrum lets one Redshift query join local tables with external S3 data. Loading first duplicates storage. Exporting Redshift data is the reverse and duplicates too. Spreadsheet combination is manual.

176. A pipeline's Glue jobs are completing successfully but taking progressively longer each week. Which cause should be investigated first?

Answer and explanation

Answer: D. Gradual slowdown with successful completion points to growing input or accumulating small files. A version change would be a step rather than a trend. Role changes cause failures rather than slowness. Output storage class does not affect processing time.

177. A data pipeline's cost must be attributed to the business team that owns each dataset. Which approach is appropriate?

Answer and explanation

Answer: B. Cost allocation tags attribute spend to the owner precisely. Estimation and equal shares are inaccurate. Attributing everything to the platform team hides the real consumers.

178. A Redshift query is slow, and the execution plan shows a nested loop join. Which cause should be investigated first?

Answer and explanation

Answer: B. A nested loop join results from a missing or non-equality join condition. More nodes make a bad plan faster at higher cost. VACUUM and WLM affect other aspects of performance.

179. A streaming pipeline's Lambda consumer is being invoked but the iterator age keeps rising. Which cause should be investigated first?

Answer and explanation

Answer: C. Rising iterator age with active invocation means the consumer processes slower than records arrive. Retention affects expiry rather than lag. Higher memory typically speeds processing. Additional consumers do not help one consumer's lag.

180. A pipeline must alert when a Glue job runs longer than its expected duration. Which approach is appropriate?

Answer and explanation

Answer: C. A timeout with an event route alerts when the duration is exceeded. Morning review is retrospective. More workers change duration without alerting. CPU utilization does not indicate duration.

181. A data quality check must run on a table after each load and record its results over time. Which approach is appropriate?

Answer and explanation

Answer: D. A ruleset evaluated after each load with published metrics gives a consistent record over time. Manual inspection is inconsistent. Waiting for consumers means problems are found late. A one-time profile does not track change.

182. A pipeline must quarantine records that fail a quality rule rather than failing the whole batch. Which approach is appropriate?

Answer and explanation

Answer: A. Quarantining preserves the failing records for investigation while good records proceed. Failing the batch blocks good data. Silent dropping loses the records and the signal. Automatic correction fabricates values.

183. A quality rule must confirm that a column's values fall within an expected range, and the range shifts seasonally. Which approach is appropriate?

Answer and explanation

Answer: B. A learned baseline adapts to seasonal variation. A wide fixed range misses genuine anomalies. Disabling seasonally removes protection. An annual average is wrong for most of the year.

184. A pipeline must run when a specific object key pattern appears in a bucket, ignoring other objects. Which configuration is appropriate?

Answer and explanation

Answer: D. An event pattern filters before invocation, which avoids unnecessary executions. Filtering inside the pipeline pays for every execution. A separate bucket requires the producer to change. Polling adds latency and cost.

185. A daily pipeline must not run when the previous day's run has not completed. Which approach is appropriate?

Answer and explanation

Answer: D. An explicit in-progress check prevents overlap deterministically. Relying on typical duration fails when a run is slow. More concurrency permits the overlap. A shorter runtime reduces the chance without preventing it.

186. An Athena query must write its results to a specific location with encryption applied. Which configuration is appropriate?

Answer and explanation

Answer: B. Workgroup settings apply the location and encryption to every query consistently. Per-query client configuration drifts. Copying and manual encryption are post-hoc and unreliable.

187. A recurring analytical query must run faster without changing the underlying data. Which approach is appropriate?

Answer and explanation

Answer: D. Materializing the result turns a repeated computation into a lookup. A longer timeout permits slowness. More frequent execution increases cost. Fewer displayed columns does not reduce the work if the query still computes them.

188. A data engineer must join data held in Amazon Redshift with data in an Amazon RDS database without moving either. Which capability is appropriate?

Answer and explanation

Answer: B. Federated query reads the external database in place within a Redshift query. UNLOAD and export both move data. A Glue job also moves data and produces a third copy.

189. A Glue job fails with an out-of-memory error on the driver rather than an executor. Which cause should be investigated first?

Answer and explanation

Answer: B. Driver out-of-memory typically indicates collecting data to the driver. Executor count affects distributed processing. Compression affects reading. Bookmarks affect which data is processed.

190. An Athena query fails with an error about too many open partitions. Which cause should be investigated first?

Answer and explanation

Answer: A. Excessive partition counts usually follow from partitioning on a high-cardinality column. Bucket capacity, scan limits, and column counts produce different errors.

191. A Redshift cluster's queries have slowed, and the tables have received large volumes of updates and deletes. Which maintenance operation is appropriate?

Answer and explanation

Answer: B. VACUUM reclaims space and restores sort order after heavy modification, and ANALYZE refreshes the statistics the planner uses. More nodes scan the degraded tables faster. Rebuilding is disproportionate. Fewer concurrent queries masks the problem.

192. A pipeline's cost must be tracked per pipeline rather than per service. Which approach is appropriate?

Answer and explanation

Answer: A. Tag-based grouping attributes spend per pipeline. Service grouping aggregates across pipelines. An account per pipeline is heavy for attribution alone. Runtime-based estimation ignores resource size and data volume.

193. A data quality rule must confirm that a foreign key value exists in the referenced dataset. Which rule type is appropriate?

Answer and explanation

Answer: C. Referential integrity compares values against a reference set. Completeness, uniqueness, and range rules each check a different property.

194. A data quality score must be tracked over time so gradual degradation is visible. Which approach is appropriate?

Answer and explanation

Answer: C. Publishing a metric makes the trend visible and alarmable. Reviewing output and logging record point results without a trend. Alerting only on total failure misses gradual decline.

195. A pipeline must be triggered by the completion of another pipeline in a different account. Which approach is appropriate?

Answer and explanation

Answer: C. Cross-account event delivery triggers the dependent pipeline on actual completion. Polling adds latency, shared credentials are a standing secret, and offset schedules guess at duration.

196. A pipeline must run only after all files for a batch have arrived, and the file count varies. Which approach is appropriate?

Answer and explanation

Answer: C. A manifest declares the batch explicitly regardless of file count. Counting in the pipeline requires knowing the expected total, fixed waits guess, and hourly processing splits batches arbitrarily.

197. An automation must start a Glue job only when the previous run has finished. Which approach is appropriate?

Answer and explanation

Answer: C. Checking run state prevents overlap deterministically. Schedule intervals fail when a run is slow, more workers reduce but do not prevent overlap, and higher concurrency permits it.

198. A scheduled pipeline must skip execution on non-business days. Which approach is appropriate?

Answer and explanation

Answer: C. A calendar check within the workflow handles holidays and weekends in one place. Per-day schedules are unmanageable, running regardless wastes resources, and manual disabling depends on a person.

199. A pipeline must notify a downstream team only when its output differs materially from the previous run. Which approach is appropriate?

Answer and explanation

Answer: D. Threshold-based comparison notifies on material change. Notifying every run produces noise, failure-only misses successful but anomalous runs, and a weekly schedule is decoupled from the runs.

200. An orchestration must run a step only when a previous step produced output. Which approach is appropriate?

Answer and explanation

Answer: A. Branching on a returned result expresses the condition in the workflow with visible history. Exiting inside the step hides the decision, manual checks depend on a person, and delays guess.

201. A pipeline must reprocess a specific historical date range on demand without changing its scheduled behaviour. Which approach is appropriate?

Answer and explanation

Answer: B. Parameterization supports both scheduled and ad hoc runs from one definition. A separate pipeline duplicates logic, changing the schedule disrupts normal operation, and editing code for a run is error-prone.

202. An automated pipeline must stop and alert rather than proceed when its input volume is far below normal. Which approach is appropriate?

Answer and explanation

Answer: D. A volume check against a baseline catches a truncated or missing delivery before bad data propagates. Processing regardless produces wrong output, waiting indefinitely blocks, and schedule changes do not detect the condition.

203. A pipeline's steps must run in different accounts with the orchestration defined in one place. Which approach is appropriate?

Answer and explanation

Answer: D. A central workflow with per-step cross-account roles keeps orchestration in one place with scoped access. Per-account workflows fragment the definition, replication moves data unnecessarily, and administrator access over-grants.

204. An automation must apply the same transformation to data landing in several buckets. Which approach is appropriate?

Answer and explanation

Answer: B. One parameterized workflow handles every source. Per-bucket workflows duplicate logic, consolidation adds a copy step, and a schedule loses the event-driven trigger.

205. A pipeline must record whether each run processed new data or found nothing to do. Which approach is appropriate?

Answer and explanation

Answer: C. A recorded outcome distinguishes a genuine no-op from a failure and makes the pattern visible. Failing on an expected condition generates noise, silent success hides it, and logs alone are not alarmable.

206. An automation must prevent two pipeline runs from writing to the same output partition simultaneously. Which approach is appropriate?

Answer and explanation

Answer: D. A lock serialises writes to the same partition. Resolving afterwards leaves a window of inconsistency, schedule changes reduce but do not prevent overlap, and new partitions per run changes the output layout.

207. A pipeline must be redeployed to a new environment with its schedules disabled until it is validated. Which approach is appropriate?

Answer and explanation

Answer: C. A parameterized enabled state deploys safely and is turned on deliberately. Deploying enabled risks an immediate run, manual addition drifts from the definition, and validating elsewhere does not validate this environment.

208. An orchestration must limit how many pipeline runs execute concurrently across an account. Which approach is appropriate?

Answer and explanation

Answer: B. A concurrency limit or semaphore bounds parallel execution while allowing queuing. Schedule changes reduce the chance without bounding it, higher quotas permit more concurrency, and full serialization removes useful parallelism.

209. A pipeline must pass a large intermediate result between steps without exceeding the workflow's payload limit. Which approach is appropriate?

Answer and explanation

Answer: B. Passing a location keeps payloads small regardless of data size. Compression delays the limit, splitting complicates the workflow, and the payload limit is a service constraint.

210. An automation must retry a failed pipeline run from the failed step rather than from the beginning. Which approach is appropriate?

Answer and explanation

Answer: D. Discrete checkpointed states support resumption. Rerunning from the start repeats completed work, retry counts govern attempts rather than resumption, and fewer larger steps make resumption coarser.

211. A pipeline's cost has risen after a source system began sending significantly more data. Which response is appropriate?

Answer and explanation

Answer: D. Confirming need and right-sizing addresses the change directly. Reducing frequency delays processing, sampling discards data, and storage class affects a small part of the cost.

212. A Glue job's logs must be retained longer than the default for compliance. Which approach is appropriate?

Answer and explanation

Answer: A. Log group retention or S3 delivery with lifecycle governs how long logs persist. Timeout, bookmarks, and worker count are unrelated to retention.

213. A pipeline's performance must be compared before and after a change. Which approach is appropriate?

Answer and explanation

Answer: B. Comparing distributions across representative runs distinguishes the change from normal variation. Two runs may differ for other reasons, a diff shows intent, and monthly cost aggregates many factors.

214. An Athena workgroup must prevent a single query from scanning more data than expected. Which configuration is appropriate?

Answer and explanation

Answer: D. A per-query limit cancels an individual runaway query. A monthly limit governs aggregate spend, concurrency governs parallel queries, and a longer timeout permits more scanning.

215. A Redshift cluster's disk usage is growing although the data volume is stable. Which cause should be investigated first?

Answer and explanation

Answer: C. Unreclaimed space from updates and deletes grows disk usage with stable data. More nodes add capacity without addressing it, compression reduces size generally, and backups are stored separately.

216. A pipeline must alert when its output row count deviates materially from the historical pattern. Which approach is appropriate?

Answer and explanation

Answer: A. An anomaly band adapts to the historical pattern including seasonality. A zero check catches only total failure, weekly review is slow, and a fixed threshold set once goes stale.

217. A data quality rule must be applied to a column only when another column has a particular value. Which approach is appropriate?

Answer and explanation

Answer: C. A conditional rule evaluates only the applicable rows. Applying unconditionally produces false failures, splitting the dataset changes the pipeline, and removing the rule loses the check.