The Iceberg vs Delta vs Hudi War Is Over: Why That Matters More Than the Lakehouse Hype Did
For three years the data infrastructure trade press carried headlines that read like a religious schism. Apache Iceberg, Delta Lake, and Apache Hudi were presented as competing answers to the same question: which open table format would the industry standardize on. Vendors picked sides. Conference keynotes argued for one over the other. Procurement decks included a row labelled "table format," as if this were the architecturally consequential decision.
That war is over. It ended in 2025, quietly, with a sequence of vendor commitments that left no doubt about the outcome. Snowflake committed unambiguously to Iceberg as a first-class native format. Databricks open-sourced Unity Catalog and converged Delta Lake's metadata layer toward Iceberg interoperability through Delta Universal Format. AWS, Google, Cloudera, Dremio, and Starburst all standardized on Iceberg as their interop layer. Apache Hudi remains technically capable and has a real user base, but it is no longer in the running as the format the industry will converge on. The de facto answer is Iceberg, and the practical answer is "all three formats can be read by an Iceberg-aware engine, so the choice no longer matters the way it used to."
This is the part the trade press got right. The part it has mostly missed is that this outcome makes the table-format choice the wrong layer to argue about. Lock-in did not disappear. It moved up the stack. The fight that matters in 2026 is happening one layer above, at the catalog, and one layer above that, at the compute engine and the governance plane. Enterprises that treat "we picked Iceberg" as the architecture decision are walking into the same trap the previous decade of data warehouse buyers walked into, with a different vocabulary.

What Actually Changed in 2025
The substantive shifts are worth naming because they explain why the table-format layer commoditized.
Snowflake's Iceberg commitment moved from an experimental feature to a mainstream product surface. Customers can now create Iceberg tables in Snowflake-managed storage and treat Iceberg as the default open format for data shared between Snowflake and other engines. This was the strongest signal the war was settled, because Snowflake had the most to lose from giving up a proprietary format.
Databricks open-sourced Unity Catalog and shipped Delta Universal Format, which writes Delta tables in a way that lets Iceberg readers consume them without translation. The strategic message: Delta is no longer being defended as a moat. It is being repositioned as a high-quality implementation that interoperates with Iceberg readers, with the real differentiation moving to Unity Catalog as the governance layer.
AWS extended Iceberg support across S3 Tables, Glue, Athena, EMR, and Redshift. Google completed integration across BigQuery, BigLake, and Dataproc. Microsoft Fabric standardized on Delta with full Iceberg interop. Cloudera, Dremio, and Starburst were Iceberg-first from the start. By the end of 2025 no major data platform lacked a credible Iceberg story.
The technical convergence is real. All three formats now solve substantially the same problems with substantially similar designs: ACID transactions over object storage, schema evolution, time travel, hidden partitioning, and metadata-driven query planning. The functional surface, for the workloads most enterprises run, is the same.

Where Lock-In Actually Lives Now
Lock-in is a property of the layer where switching costs concentrate. Five years ago that layer was the table format itself: reading Delta required Databricks tooling, reading proprietary Snowflake formats required Snowflake compute. With Iceberg as the interop layer, the table format is no longer the lock-in point. Three other layers are.
The first is the catalog. Iceberg tables are not portable on their own. They are portable when paired with a catalog that the next engine can read. When a team writes Iceberg tables into a vendor-managed catalog and three years of partition history, schema evolution, table renames, and access policies accumulate inside it, switching engines means migrating the catalog. That migration is harder than migrating the data, because the catalog encodes operational history that the data files alone do not represent. Catalog choice is now the highest-leverage architectural decision in the lakehouse stack.
The second is the governance and access-control plane. Unity Catalog, Polaris, AWS Lake Formation, Gravitino, and Nessie all implement governance differently. Permissions modelled in one system do not translate cleanly to another. Row-level security, column masking, and lineage capture are vendor-specific in their semantics even when the storage format is open. Switching governance planes means rebuilding the compliance regime, not just pointing a new engine at the same files.
The third is the compute engine and its tooling. Engines differ in performance, query optimizer maturity, materialized view support, and time-travel performance. They differ even more in surrounding tooling: notebooks, ML platform integration, BI connectors, and orchestration. Most of the productivity that platform teams attribute to "Snowflake" or "Databricks" comes from this layer rather than the table format. Replacing the engine means replacing all of it.
The argument is not that these layers are necessarily bad to commit to. It is that the commitment is real, the switching cost is high, and the open-format story does not soften it. An organization that says "we are open because we use Iceberg" while running on a single vendor's catalog, governance, and engine has the same lock-in profile it had in 2018, with a different vocabulary attached.

This is a Premium Article
Sign up for a Premium membership to read this article and get full access to strategic intelligence on technology and business.
Already a member? Sign in