Versus Platform Engineers

Data Engineers versus Platform Engineers

A platform engineer builds the shared infrastructure other teams deploy onto: Kubernetes 5,150 clusters, CI/CD pipelines, networking, identity and access management, logging and monitoring. Package Managers and DevOps covers that ground with Git 1,932 , Docker 514 , Kubernetes and continuous integration. In larger organizations a data platform team specializes it for data: it runs the Kafka 129 clusters, the Spark-on-Kubernetes setup, the orchestration service, the object store and the catalog, and offers them as self-service tools.

The difference is the customer. A platform engineer serves other engineers and measures success by uptime, deployment speed and cost per workload. A data engineer serves the business through data and measures success by whether the right tables are correct and fresh. When an Airflow 129 scheduler crashes, the platform team fixes the service; when a task in it produces duplicate orders, the data engineer fixes the pipeline.

You will cross this line often in this book, because a single workstation has no platform team. You will run Kafka, Airflow and Trino 403,499 in Docker yourself (Apache Kafka and Managed Cloud Kafka-Lakehouses, Data Quality and Governance) and see why companies eventually hand that work to a dedicated team or a managed cloud service.