This chapter's open stack and the managed platforms share their foundations, open table formats and the Iceberg 129 REST protocol, and differ in who does the work:
| Layer | Open source, as run here | Managed equivalent | What you trade |
|---|---|---|---|
| Object storage | MinIO 30,943 (source build, AGPL) | S3, Cloud Storage, ADLS | Operations for a per-GB fee |
| Table format | Iceberg, Delta, Hudi 129 | The same formats | Nothing: files stay portable |
| Catalog | REST fixture, Nessie, Lakekeeper | Unity, Glue, Open Catalog | Control for managed security |
| Maintenance | Spark 129 procedures you schedule | Automatic, metered | Effort for a line item |
| Engines | Trino 403,499 , Spark, DuckDB 61,228 | Photon, Snowflake, BigQuery 1 | Tuning for speed and support |
| Governance | DataHub 559,249 , Trino rules, Presidio | Built-in catalogs and policies | Integration work for one console |
The open stack wins on cost at small scale and on control at any scale: everything in this chapter ran on one 4-CPU host, cost nothing in licences, and can move to any cloud. It loses on people: Self-Hosted Cost Model found the administrator's time to be most of a self-hosted bill, and Governance and Security and Lakehouse Costs showed how many separate pieces (classification, access rules, encryption, erasure) someone must wire together and keep wired. The managed platforms sell exactly that integration, at the price of a bill that grows with use and features that do not travel.
For BookNest the sensible path is staged. Start open, as this chapter did, while data fits on a few servers and the team learns what a table format really does. Move storage first when operations hurt, keeping Iceberg tables and a REST catalog so engines remain interchangeable. Adopt a managed platform for the parts where it clearly earns its fee, usually maintenance and governance, and keep the exit open by refusing anything that stores tables in a format only one engine can read.