Open Source vs Commercial

Open Source versus Commercial Trade-offs

Most of the data stack's core engines are open source: Spark 129 , Kafka 129 , Airflow 129 , Iceberg 129 , Trino 403,499 and DuckDB 61,228 . Most of the money is made by companies selling them as managed services, or selling proprietary features around an open core. "Open source" is therefore a spectrum, and the license is the first thing to check.

License families you meet in the data stack
License family Examples (verified on GitHub 29 , October 2026) What it allows
Permissive Apache 2.0: Kafka, Airflow, Trino; MIT: DuckDB Use, modify, sell, host as a service
Copyleft AGPL-3.0: MinIO 30,943 (repository now archived) Use freely; share changes, even over a network
Source-available Business Source License 1.1: Redpanda 64,862 ; Elastic License 2.0: Airbyte 128,715 Use and read; no competing hosted service
Proprietary Snowflake, BigQuery 1 , Confluent Cloud 17,245 Use under a paid contract only

Source-available licenses are not open source by the Open Source Initiative's definition, though the code is on GitHub. Redpanda's license, for example, forbids offering it as a streaming service to third parties and converts each release to Apache 2.0 four years after it ships. For a company running the software for itself, these restrictions rarely bite; for one that sells a platform, they decide everything. Licenses also change, so recheck them at each major upgrade. Check what you already have installed, straight from the packages' metadata:

Reading the declared licenses of installed Python packages
# Which licenses did you just install? Read them from each package's own metadata.
from importlib.metadata import metadata
for pkg in ("duckdb", "pyarrow", "pandas", "pyspark", "dbt-core", "lxml", "protobuf"):
    try:
        m = metadata(pkg)
    except Exception:
        print(f"{pkg:20} (not installed)"); continue
    lic = m.get("License-Expression") or m.get("License") or ""
    if not lic or len(lic) > 40:
        lic = "; ".join(c.split(" :: ")[-1] for c in m.get_all("Classifier") or []
                        if c.startswith("License ::")) or "(see LICENSE file)"
    print(f"{pkg:20} {m['Version']:10} {lic}")
Output
duckdb               1.5.5      MIT License
pyarrow              25.0.1     Apache-2.0
pandas               2.3.3      BSD License
pyspark              4.2.0      Apache-2.0
dbt-core             1.12.5     Apache-2.0
lxml                 6.1.3      BSD-3-Clause
protobuf             7.36.2     3-Clause BSD License

Every package in the book's environment is permissively licensed. The metadata is only a declaration, though: for anything you ship or host, read the repository's LICENSE file, which is what binds you.

Beyond the license, weigh the trade-offs that matter in production:

The usual compromise is a managed service for an open-source engine: Amazon MSK 24 or Aiven for Kafka, Google's Managed Service for Apache Airflow, Databricks 2,717 for Spark. You keep the open APIs and data formats, so you can leave, and someone else runs the servers.