Securing Spark Pipelines

Securing BookNest's Spark Pipelines

BookNest's jobs read customer emails and order history, which makes them personal data. The baseline for its production Spark 129 :

Security baseline for BookNest's Spark jobs
Area BookNest's rule
Cluster access Kubernetes 5,150 or a managed platform; no ports reachable from outside the VPC
Internal traffic spark.authenticate, network and local-disk encryption on (automatic secrets on Kubernetes)
Data access Per-job service accounts or roles with read-only access to inputs, write access to one output prefix
Tables Governed catalog with column masks on email and row filters for regional analysts
Secrets Workload identity first; otherwise a secret manager; never in code, conf files or arguments
UIs and logs History Server behind single sign-on; redaction regex extended; event logs in a restricted bucket
Dependencies Pinned Spark image and Python packages, scanned in CI (Package Managers and DevOps)

Most of the list is configuration, not code, which is why the chapter's jobs keep credentials and cluster settings out of the Python files: the same orders_daily.py runs on a laptop, on Kubernetes and on a managed platform, and each environment supplies its own secure configuration.