Most of the data stack's core engines are open source: Spark 129 , Kafka 129 , Airflow 129 , Iceberg 129 , Trino 403,499 and DuckDB 61,228 . Most of the money is made by companies selling them as managed services, or selling proprietary features around an open core. "Open source" is therefore a spectrum, and the license is the first thing to check.
| License family | Examples (verified on GitHub 29 , October 2026) | What it allows |
|---|---|---|
| Permissive | Apache 2.0: Kafka, Airflow, Trino; MIT: DuckDB | Use, modify, sell, host as a service |
| Copyleft | AGPL-3.0: MinIO 30,943 (repository now archived) | Use freely; share changes, even over a network |
| Source-available | Business Source License 1.1: Redpanda 64,862 ; Elastic License 2.0: Airbyte 128,715 | Use and read; no competing hosted service |
| Proprietary | Snowflake, BigQuery 1 , Confluent Cloud 17,245 | Use under a paid contract only |
Source-available licenses are not open source by the Open Source Initiative's definition, though the code is on GitHub. Redpanda's license, for example, forbids offering it as a streaming service to third parties and converts each release to Apache 2.0 four years after it ships. For a company running the software for itself, these restrictions rarely bite; for one that sells a platform, they decide everything. Licenses also change, so recheck them at each major upgrade. Check what you already have installed, straight from the packages' metadata:
# Which licenses did you just install? Read them from each package's own metadata.
from importlib.metadata import metadata
for pkg in ("duckdb", "pyarrow", "pandas", "pyspark", "dbt-core", "lxml", "protobuf"):
try:
m = metadata(pkg)
except Exception:
print(f"{pkg:20} (not installed)"); continue
lic = m.get("License-Expression") or m.get("License") or ""
if not lic or len(lic) > 40:
lic = "; ".join(c.split(" :: ")[-1] for c in m.get_all("Classifier") or []
if c.startswith("License ::")) or "(see LICENSE file)"
print(f"{pkg:20} {m['Version']:10} {lic}")duckdb 1.5.5 MIT License pyarrow 25.0.1 Apache-2.0 pandas 2.3.3 BSD License pyspark 4.2.0 Apache-2.0 dbt-core 1.12.5 Apache-2.0 lxml 6.1.3 BSD-3-Clause protobuf 7.36.2 3-Clause BSD License
Every package in the book's environment is permissively licensed. The metadata is only a declaration, though: for anything you ship or host, read the repository's LICENSE file, which is what binds you.
Beyond the license, weigh the trade-offs that matter in production:
Open source gives you control, no license fee and the freedom to run anywhere. You supply the operations: upgrades, security patches, scaling and the 3 a.m. pager.
Commercial products give you support, service-level agreements, security certifications and features that open-source editions often lack. You pay, often by usage, and accept the vendor's roadmap and price changes.
The usual compromise is a managed service for an open-source engine: Amazon MSK 24 or Aiven for Kafka, Google's Managed Service for Apache Airflow, Databricks 2,717 for Spark. You keep the open APIs and data formats, so you can leave, and someone else runs the servers.