Herding Cat(alogue)s: The Apache Iceberg™ Catalog Landscape
Session Abstract
Apache Iceberg™ relies on catalogs, but with many options available, choosing the right one is challenging. This session explores the Iceberg catalog landscape, covering core functions, the impact of REST catalogs, and comparisons of Hive, JDBC, Snowflake, and Apache Polaris™ to help you select the best fit for your needs.
Session Description
Apache Iceberg™ is nothing without its catalog layer; without them, we can’t discover and interact with Iceberg tables or ensure consistency across multiple writers. But with so many Iceberg catalog implementations out there, how can you possibly choose?
In this session, we’ll herd the cat(alogue)s for you and break down the Iceberg catalog landscape. We’ll begin with the basics: why catalogs are essential to Iceberg and the general requirements of a catalog. From there, we’ll examine the introduction of the REST catalog and its impact on Iceberg’s ecosystem. Finally, we’ll survey a number of the most widely used Iceberg catalogs—including HIVE, JDBC, Snowflake, Apache Polaris (incubating) and more––analyzing their metadata storage approaches, table discovery mechanisms, and integration with both open-source and commercial platforms. Along the way, we’ll highlight the strengths, trade-offs, and key considerations for each option.
By the end of this session, you’ll have a clear understanding of how Iceberg catalogs function, what differentiates them, and how to select the best one for your needs—whether you’re optimizing for performance, compatibility, or scalability.