66 practice questions for Domain 4 of the AWS Certified Data Engineer - Associate (DEA-C01) exam, which makes up 18% of its scored content. Your answers count towards one score and one timer for the whole exam.
Domain 4: Data Security and Governance
218. An AWS Glue job must connect to an Amazon RDS database using credentials that rotate automatically without changing the job. Which solution meets these requirements?
Answer and explanation
Answer: A. Secrets Manager rotates the credential on a schedule and the Glue connection resolves it at run time, so the job is unchanged. A String parameter is neither encrypted nor rotated. Job parameters are visible to anyone who can describe the job. Embedded credentials place a secret in source control.
219. An Amazon Redshift cluster in a private subnet must be queried from an AWS Lambda function without internet traversal. Which solution meets these requirements?
Answer and explanation
Answer: C. A function attached to the VPC reaches the cluster over private addressing, and the cluster's security group must permit the function's security group. Public accessibility exposes the cluster. A NAT gateway routes outbound internet traffic and is unnecessary within a VPC. An internet gateway route creates public exposure.
220. Consumers in other AWS accounts must reach a data service in the producer VPC with no routed network path between the VPCs. Which solution meets these requirements?
Answer and explanation
Answer: A. PrivateLink exposes one service through interface endpoints so consumers reach that service and nothing else, with no routed path between the networks. Peering and transit gateway create general routed connectivity. A public endpoint with allow-lists is internet-exposed and does not authenticate callers.
221. An Amazon S3 bucket shared by many teams has a bucket policy that has grown too large to maintain. Which solution meets these requirements?
Answer and explanation
Answer: C. Access Points decompose a monolithic bucket policy into a named endpoint and policy per team, which is the documented answer to a policy reaching its size limit. A bucket per team multiplies management. Consolidating statements postpones the limit. IAM policies help but leave the resource policy problem when cross-account access must be controlled at the bucket.
222. A Lambda function in a data pipeline must call Amazon S3 and AWS Glue without long-term credentials on any host. Which solution meets these requirements?
Answer and explanation
Answer: A. An execution role supplies temporary, automatically rotated credentials the SDK retrieves with no secret stored anywhere. Environment variables holding keys leak through logs and console output. An encrypted object still resolves to long-term keys in the function. Parameter Store protects the value at rest but the credential remains long-term.
223. Analysts must query a data lake table but must not be able to see the salary column. Which solution meets these requirements?
Answer and explanation
Answer: A. Lake Formation grants permissions at database, table, and column level and enforces them across Athena, Redshift Spectrum, and EMR from one place. A filtered copy duplicates storage and drifts. A bucket policy operates on whole objects and cannot exclude a column. Instruction is not an enforced control.
224. Analysts in each region must see only the rows for their own region in a shared table. Which solution meets these requirements?
Answer and explanation
Answer: B. Lake Formation data filters apply row-level restrictions at query time per principal, enforced consistently across engines. Per-region tables duplicate storage and complicate cross-region analysis. A single view with a fixed WHERE clause cannot differentiate between principals. Prefix-based policies work only if the data is physically partitioned that way and still cannot express per-principal rules within a prefix.
225. A managed IAM policy grants more permissions than a data engineering role requires. Which solution meets these requirements?
Answer and explanation
Answer: C. A customer managed policy naming the required actions and resources is least privilege and is the intended approach when a managed policy does not fit. An inline deny on top of a broad grant is fragile and hard to reason about. CloudTrail detects misuse after it occurs. A boundary permitting the same actions narrows nothing.
226. Database users in Amazon Redshift must be granted access by team rather than individually. Which solution meets these requirements?
Answer and explanation
Answer: A. Groups and roles let privileges be granted once and inherited by members, which is how access scales with team size. Per-user grants multiply as people join and leave. A cluster per team fragments the data. Shared credentials destroy attribution.
227. Data written by an AWS Glue job in one account must be readable by an Amazon Athena query in another account, and the objects are encrypted with a customer managed AWS KMS key. Which combination of steps meets these requirements? (Select TWO.)
Answer and explanation
Answer: A, B. Cross-account reads of KMS-encrypted objects require both object permission through the bucket policy and decrypt permission in the key policy. S3 managed keys cannot be shared cross-account in this way. Copying duplicates storage and creates drift. Public access exposes the data to everyone.
228. An Amazon Redshift cluster must refuse any client connection that is not encrypted in transit. Which solution meets these requirements?
Answer and explanation
Answer: A. The require_SSL parameter forces clients to connect over SSL and refuses those that do not. Encryption at rest protects stored data rather than the connection. Subnet placement and security groups control reachability rather than encryption. Audit logging records connections after the fact.
229. National identity numbers must be rendered unusable to analysts while records remain joinable across tables. Which solution meets these requirements?
Answer and explanation
Answer: B. Deterministic tokenization produces the same surrogate for the same input, so joins still work while the original value is not exposed. Truncation leaves partial identifying information and breaks uniqueness. Giving analysts the key defeats the protection. Removing the field removes the join capability the requirement preserves.
230. A compliance requirement states that the encryption key protecting analytics data must be controllable and its use auditable by the company. Which solution meets these requirements?
Answer and explanation
Answer: D. A customer managed key exposes a key policy the company authors, supports controllable rotation, and records every cryptographic operation in CloudTrail. An AWS managed key exposes no key policy. S3 managed keys provide no key usage trail. An unencrypted bucket fails the encryption requirement outright.
231. An auditor must query two years of AWS API activity across an organization using SQL, without building an ingestion pipeline. Which solution meets these requirements?
Answer and explanation
Answer: B. CloudTrail Lake stores events in a managed event data store with configurable retention and supports SQL queries with no pipeline to build. An S3 trail with Athena requires building and maintaining partitioning and table definitions. CloudWatch Logs retention at this scale is costly and Logs Insights is not SQL. OpenSearch requires provisioning and an ingestion path.
232. Application logs from an Amazon EMR cluster must survive cluster termination for later analysis. Which solution meets these requirements?
Answer and explanation
Answer: A. An EMR log URI archives application and step logs to S3 continuously, so they remain available after the cluster is gone. Node storage disappears with the cluster. Manual download depends on remembering and fails on unexpected termination. Detailed monitoring reports metrics rather than application logs.
233. Very large volumes of pipeline logs must be analysed to produce an audit report covering a full quarter. Which solution meets these requirements?
Answer and explanation
Answer: B. S3 with Athena or EMR handles very large log volumes economically and supports the aggregate queries an audit report requires. Reading a quarter of logs in the console is impractical. Local processing does not scale to this volume. SNS delivers notifications rather than supporting analysis.
234. Application logs written by an AWS Glue job must be centralized so they can be searched alongside other pipeline logs. Which solution meets these requirements?
Answer and explanation
Answer: A. CloudWatch Logs centralizes log data from Glue alongside other services, and Logs Insights queries across log groups. Local storage disappears with the job. Mixing logs into the output dataset corrupts the data product. Reading the run detail page does not support searching across runs or services.
235. Data in a specific Amazon S3 bucket must not be replicated or backed up into Regions outside the company's jurisdiction. Which solution meets these requirements?
Answer and explanation
Answer: C. A service control policy with a region condition blocks the API calls that would create resources or replicas elsewhere and cannot be overridden inside the account. A bucket policy on the source does not prevent a backup service writing a copy elsewhere. A Config rule reports after the configuration exists. Disabling replication today does not prevent it being re-enabled.
236. Amazon Redshift data must be shared with an analytics team in another AWS account without copying it. Which solution meets these requirements?
Answer and explanation
Answer: C. Redshift data sharing exposes live data from a producer cluster to a consumer cluster across accounts with no data movement. Unload and load creates a stale second copy. A cross-account login gives access to the same cluster rather than sharing data. A restored snapshot is a point-in-time duplicate.
237. Personally identifiable information must be located across a data lake before analysts are granted broader access. Which solution meets these requirements?
Answer and explanation
Answer: B. Macie applies managed and custom data identifiers to discover and classify sensitive data such as PII across S3. A crawler infers schema and column names without classifying values. Sampling rows is neither systematic nor complete. CloudTrail records who accessed objects rather than what they contain.
238. A governance team must be able to review every configuration change made to the data lake's resources over the past year. Which solution meets these requirements?
Answer and explanation
Answer: D. Config records the configuration state of resources over time and presents a timeline of changes, which is what reviewing configuration history requires. CloudTrail records the API calls that caused changes rather than the resulting configuration state. Access logging records object requests. Job bookmarks track processing position.
239. A Glue job must connect to an Amazon RDS database in a private subnet. Which network configuration is required?
Answer and explanation
Answer: C. A Glue connection places the job's network interfaces in the VPC with private addressing. A NAT gateway routes outbound to the internet. A public address exposes the database. An internet gateway route is not a private path.
240. A pipeline's Lambda functions must authenticate to Amazon Redshift without a stored password. Which approach is appropriate?
Answer and explanation
Answer: B. IAM authentication issues temporary credentials from the role with no stored password. Environment variables, configuration files, and code all store a long-term secret.
241. An Amazon MSK cluster must authenticate producers without managing certificates. Which approach is appropriate?
Answer and explanation
Answer: A. IAM access control authenticates with the producer's IAM identity and no certificate management. TLS certificates require issuance and rotation. Plaintext with security groups authenticates a network location. A shared credential destroys attribution.
242. Analysts must query a Lake Formation table but see only rows for their own department. Which approach is appropriate?
Answer and explanation
Answer: C. Data filters apply row-level restrictions per principal across query engines. Per-department tables duplicate storage. Saved queries can be edited. S3 prefixes work only if the data is physically separated that way.
243. A Glue job's role must read from one bucket and write to another, and nothing else. Which policy is appropriate?
Answer and explanation
Answer: C. Naming the specific actions and prefixes is least privilege. Wildcard actions and managed full access grant far more. GetObject on all buckets is both too broad and missing the write.
244. A Redshift user must query a view but not the underlying tables. Which approach is appropriate?
Answer and explanation
Answer: A. Granting on the view alone works because the view executes with its owner's permissions on the underlying tables. Granting on the tables exposes them. Superuser is far too broad. Copying data creates a stale duplicate.
245. A Glue job's intermediate data on worker disks must be encrypted. Which configuration is required?
Answer and explanation
Answer: B. A Glue security configuration covers encryption of data on worker disks and in bookmarks. Output bucket encryption covers results. Catalog encryption covers metadata. A source key covers the source.
246. A column containing card numbers must be masked in query results for most users while a compliance role sees the full value. Which approach is appropriate?
Answer and explanation
Answer: B. Dynamic data masking returns masked values by default and full values to permitted roles from one table. A separate table duplicates the data. Hashing at load prevents the compliance role seeing it. Removing the column loses it.
247. Data in transit between a Glue job and an Amazon Redshift target must be encrypted. Which configuration is required?
Answer and explanation
Answer: B. Requiring SSL on the connection encrypts the transfer. At-rest encryption protects stored data. Subnet placement controls reachability. S3 encryption covers a different destination.
248. An Athena query's results must be encrypted with a key the security team controls. Which configuration is appropriate?
Answer and explanation
Answer: B. Workgroup result encryption with a customer managed key protects the results with the team's key. Source encryption protects inputs. There is no query editor encryption setting. IAM restriction is access control rather than encryption.
249. An audit must establish which user ran a specific Athena query and what it returned. Which sources provide this?
Answer and explanation
Answer: B. CloudTrail records the identity that started the query and the workgroup history holds the query and results. Athena does not write application logs to CloudWatch. Flow logs record network metadata. The catalog does not record query execution.
250. Redshift must log every user query for a compliance review. Which configuration is required?
Answer and explanation
Answer: D. Redshift audit logging with the user activity log captures each query. CloudTrail records management API calls rather than SQL. Enhanced VPC routing governs network path. Performance insights reports metrics.
251. Access to objects in a data lake bucket must be logged with reliable attribution of the calling principal. Which approach is appropriate?
Answer and explanation
Answer: A. CloudTrail data events record object-level access with the calling identity reliably. Server access logging is best-effort with less reliable attribution. Flow logs capture network metadata. Storage Lens reports usage metrics.
252. A data lake must prevent any table from being shared with an account outside the organization. Which approach is appropriate?
Answer and explanation
Answer: C. A policy-level restriction prevents external grants regardless of who attempts them. An instruction relies on compliance. Quarterly review finds sharing after it occurred. Encryption without key sharing blocks access but does not prevent the grant.
253. Personal data in a data lake must be located and classified before analysts receive broader access. Which approach is appropriate?
Answer and explanation
Answer: A. Macie classifies content with managed identifiers. Column names are unreliable. Sampling is incomplete. Owner recollection is unverified.
254. A dataset must be shared with a partner organization for analysis without the partner receiving a copy. Which approach is appropriate?
Answer and explanation
Answer: B. Clean Rooms and datashares let a partner analyse without a copy leaving control. Export, bucket read access, and public publication all give the partner a copy.
255. A data governance requirement states that every dataset must have a documented owner and retention period. Which approach is appropriate?
Answer and explanation
Answer: C. Required catalog metadata with enforcement keeps the record where the data is and prevents omissions. A spreadsheet, prefix naming, and wiki all drift and are unenforced.
256. An Athena query must run with permissions belonging to the querying user rather than a shared role. Which approach is appropriate?
Answer and explanation
Answer: D. Per-user roles with Lake Formation grants attribute and scope access individually. A shared role and a service account destroy attribution. Direct bucket access bypasses the catalog's permissions.
257. A Redshift cluster must authenticate users from a corporate identity provider without local database passwords. Which approach is appropriate?
Answer and explanation
Answer: B. Federation issues temporary credentials from the corporate identity. Per-employee passwords duplicate the directory. A shared user destroys attribution. A spreadsheet is not credential management.
258. Lake Formation permissions must apply to a table's data regardless of which query engine reads it. Which mechanism provides this?
Answer and explanation
Answer: B. Lake Formation permissions are enforced by the integrated engines consistently. A bucket policy operates on objects and cannot express column or row restrictions. Per-analyst IAM policies must be maintained individually. Per-engine copies drift.
259. A data engineer must grant a team access to a table while hiding two sensitive columns. Which approach is appropriate?
Answer and explanation
Answer: B. Column-level permissions exclude the columns at the catalog level for every engine. A view is workable but must be maintained and can be bypassed if the base table is accessible. A copy duplicates storage and drifts. An instruction is not enforcement.
260. An EMR cluster must encrypt data written to its local disks. Which configuration is required?
Answer and explanation
Answer: B. An EMR security configuration with local disk encryption covers data written to node storage. Output bucket encryption covers results. A key on a role grants permission rather than enabling encryption. In-transit encryption covers network traffic.
261. A pipeline must tokenize an identifier so the same input always produces the same token across runs. Which property is required?
Answer and explanation
Answer: D. Deterministic tokenization preserves joinability. Random tokens and per-record salted hashes break joins. Truncation leaves partial identifying information and may collide.
262. An audit must reconstruct which Glue job run produced a specific output file. Which approach is appropriate?
Answer and explanation
Answer: A. Embedding the run identifier makes each file traceable. A start time and a modification timestamp are ambiguous when runs overlap or files are copied. A file count does not identify individual files.
263. Redshift audit logs must be retained for several years at low cost. Which approach is appropriate?
Answer and explanation
Answer: B. S3 delivery with lifecycle transitions retains logs cheaply for years. System tables retain a limited window. Cluster storage is expensive for archival. Spreadsheet export does not scale.
264. A dataset shared with a partner must exclude rows belonging to other customers. Which approach is appropriate?
Answer and explanation
Answer: C. A row-level filter enforces the restriction at query time without a copy. Asking the partner to filter relies on them. A filtered copy drifts and duplicates storage. Bucket access exposes everything.
265. A governance requirement states that every dataset must record the legal basis for processing personal data. Which approach is appropriate?
Answer and explanation
Answer: C. Required catalog metadata with enforcement keeps the record discoverable and complete. A separate document, file names, and code comments all drift and are not enforced.
266. A pipeline must authenticate to a third-party API using a credential that rotates. Which approach is appropriate?
Answer and explanation
Answer: A. Secrets Manager retrieval at run time picks up rotation without redeployment. Job parameters are visible to anyone who can describe the job, and code and environment variables both embed the value.
267. An EMR cluster must authenticate users against the organization's directory. Which approach is appropriate?
Answer and explanation
Answer: A. Kerberos integration authenticates users against the directory. Local users do not scale, shared credentials destroy attribution, and disabling authentication removes the control.
268. A pipeline running outside AWS must obtain temporary credentials without a stored access key. Which approach is appropriate?
Answer and explanation
Answer: D. Roles Anywhere exchanges a certificate for temporary credentials with no stored key. Stored keys are the long-term credential being avoided, sharing destroys attribution, and root credentials must never be used.
269. An Athena query must run under the identity of the user who submitted it rather than a shared role. Which approach is appropriate?
Answer and explanation
Answer: A. Per-user roles with Lake Formation grants attribute and scope access individually. A shared role and a common broad role both destroy attribution, and direct bucket access bypasses the catalog.
270. A pipeline component must access a resource in a VPC while running outside it. Which approach is appropriate?
Answer and explanation
Answer: C. VPC attachment or a private endpoint provides private connectivity. Public accessibility and public addressing expose the resource, and a NAT gateway routes outbound to the internet.
271. A Lake Formation permission must apply to a table's future partitions as well as its current ones. Which consideration applies?
Answer and explanation
Answer: D. Table-level grants cover partitions as they are added. Per-partition granting is unnecessary, grants are not frozen at grant time, and partitions do not have a separate model.
272. A Redshift user must be prevented from querying a table outside business hours. Which consideration applies?
Answer and explanation
Answer: C. Database grants have no time dimension, so the restriction belongs in the access path. Workload management governs resources, and row-level security filters rows rather than times.
273. An S3 bucket policy must permit a Glue job while denying every other principal. Which approach is appropriate?
Answer and explanation
Answer: B. An explicit allow with a scoped deny achieves the restriction, and the lockout risk must be considered. Denying all blocks the job too, permitting the account is too broad, and removing the policy relies solely on IAM.
274. A pipeline must write data that can be decrypted only by a specific downstream account. Which approach is appropriate?
Answer and explanation
Answer: C. A key policy granting decrypt only to the intended account enforces the restriction cryptographically. Service keys cannot be shared this way, a key stored with the data defeats the purpose, and bucket policies alone are not encryption.
275. A column containing free text must be scanned for sensitive values before the dataset is shared. Which approach is appropriate?
Answer and explanation
Answer: C. Automated discovery classifies content systematically. Manual sampling is incomplete, column names are unreliable, and hashing destroys the column's utility without confirming what it contained.
276. An audit must establish which principal modified a Glue job's script. Which source is appropriate?
Answer and explanation
Answer: A. CloudTrail records the modification with its caller. Current content shows the result, run history shows executions, and logs show runtime output.
277. Audit logs for a data platform must be queryable without moving them out of Amazon S3. Which approach is appropriate?
Answer and explanation
Answer: A. Athena queries the logs in place. Loading into Redshift or a database moves them, and local searching does not scale.
278. An audit must cover a period longer than CloudTrail's default event history retains. Which approach is required?
Answer and explanation
Answer: B. Retention beyond the default requires a trail or event data store configured in advance. Event history covers a limited window, support cannot supply past events, and configuration state does not reconstruct API activity.
279. An audit must show which datasets a specific user queried over the past month. Which source is appropriate?
Answer and explanation
Answer: A. Query history records what was queried and by whom. Access counts aggregate, IAM policy shows what is permitted, and catalog metadata describes the datasets.
280. Audit evidence must be protected from modification by the team that produces it. Which approach is appropriate?
Answer and explanation
Answer: A. A separate account with Object Lock places the evidence beyond the producing team's reach. Restricted permissions in their own account can be changed, their key does not prevent deletion, and retention alone does not protect.
281. An audit must confirm that a pipeline's output was not modified after it was produced. Which approach is appropriate?
Answer and explanation
Answer: D. A recorded checksum detects modification. Size and timestamp can be preserved through a change, and a successful run says nothing about later modification.
282. A dataset must be shared with a partner in a way that can be revoked immediately. Which approach is appropriate?
Answer and explanation
Answer: D. Owner-controlled grants are revoked in one action. A copy cannot be recalled, long-lived presigned URLs remain valid until expiry, and indefinite bucket access is revocable but grants more than necessary.
283. A data platform must record the purpose for which each dataset may be used. Which approach is appropriate?
Answer and explanation
Answer: C. Catalog metadata keeps the permitted purpose discoverable where access decisions are made. Agreements, code comments, and file names are all detached from the grant process.