The problem solved by an open catalog
Without a shared metadata model, every engine creates its own tables, permissions and maintenance procedures. Data is copied, users lose track of the authoritative version, and changing compute platforms becomes a storage migration.
Iceberg represents table state through snapshots and metadata, supports schema and partition evolution, and enables atomic changes. REST Catalog lets clients use a consistent API without embedding one vendor's catalog implementation.
Separate four layers
Object storage holds data and metadata files. The catalog knows namespaces, tables and current metadata pointers. Compute engines plan and execute work. The control plane supplies identity, policy, observability and lifecycle management.
This separation enables independent engine choice, but it does not remove version compatibility and write-path testing.
Design access before onboarding users
Decide who may list namespaces, read metadata, retrieve file access and commit new table state. Prefer short-lived credentials scoped to a request or table. The catalog should not return broad, permanent object-store keys.
Row and column policy, masking and audit often sit above the table format. Select the component that will enforce them consistently for all engines.
Test multi-engine behaviour
The matrix should cover read, append, overwrite, schema evolution, partition evolution, concurrent writes, snapshot rollback, credential expiry and catalog outage. Test data types, name casing, time zones and delete semantics separately.
Do not call the platform open merely because two products can display one table. They must interpret changes consistently and avoid corrupting state under concurrency.
Govern maintenance and cost
Define small-file compaction, snapshot expiry, orphan-file detection, log retention and quotas. Every maintenance operation needs a window, owner and protection against deleting data still referenced by another workflow.
Track files per table, manifest size, planning time, commit conflicts, snapshot age, unreferenced volume and query cost.
Migrate by data product
Begin with one dataset that has a clear owner and consumer group. Compare query results, performance, cost, permissions and recovery. Move further domains in waves, preserving a read period on the old estate and a clear cutover point. An open format reduces coupling; data contracts and operations make it governable.

