Platform Consolidation

From Best-of-Breed Tools to Platform Consolidation

The original promise of the modern data stack was best of breed: pick the strongest product in every layer and connect them through the warehouse. By the early 2020s the downside was visible. A mid-sized team might pay a dozen vendors, each with its own pricing model, login, permissions and failure modes. Integration work moved from writing pipelines to wiring products together, and a schema change in one tool could break three others. Cost was hard to predict, because several layers billed by usage, and usage grew with every new dashboard.

The market has since swung toward consolidation. Large platforms now cover most layers themselves: Databricks 2,717 and Snowflake have added ingestion, transformation, cataloging and machine learning around their engines; Microsoft Fabric 6 bundles ingestion, Spark 129 , a warehouse, real-time analytics and Power BI on one storage layer; and each cloud provider offers the integrated set in the table above. Independent tools have responded by merging, partnering or narrowing their focus.

Consolidation trades flexibility for simplicity. One vendor means one bill and fewer integration gaps, but also deeper lock-in and less leverage on price. The counterweight is open formats: when data sits in Parquet 129 files managed by an open table format such as Apache Iceberg 129 , several engines can read the same tables, and switching a layer no longer means migrating the data (Lakehouses, Data Quality and Governance). Choosing Tools and Cost turns this into practical advice on choosing tools and estimating their total cost.