You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
For external catalogs (Hive, JDBC, Kafka, ...), Gravitino writes its own entity ID into the source object as a marker (gravitino.identifier in properties; for JDBC, in the comment). On load, the marker is used to recover the Gravitino ID. Without a marker, the loader falls back to the entity stored at the same path. See StringIdentifier, schema import, and table import.
Gravitino's tags, policies, ownership, and grants hang off that ID. So which ID an external object resolves to decides which governance metadata it gets. Today that decision relies on a marker that anyone can copy, strip, or edit.
Problems
No.
Scenario
What goes wrong today
1
Two catalogs over the same source, or a copied table carrying the marker
Import does put(overwrite=true) with the marker ID, which is an upsert on the PK (ON CONFLICT (table_id) DO UPDATE SET catalog_id, schema_id, table_name, deleted_at=0). Loading c2.db.tmovesc1.db.t's live row, with its tags, owner, and grants, into c2; loading from c1 moves it back. No error is raised.
2
dropCatalog(force) then recreate the same name
Surviving source objects still carry the old IDs. The new catalog reuses them, and the upsert can even revive a soft-deleted row.
3
Out-of-band drop + create at the same path
The same-path fallback attaches the old registration to a different object. Grants made for the old dataset apply to the new one.
4
Out-of-band rename or move
The store keeps the old path, and the object at the new path conflicts or gets re-imported. Name-based authz plugins (e.g. Ranger) keep policies on the old path.
5
Kafka
The Gravitino ID is derived from the topic UUID (convertToGravitinoId), so two catalogs over one cluster collide by construction (see problem 1).
6
Columns
Columns are matched by name only. An out-of-band column rename drops column tags and policies; a re-added column with the same name inherits them.
7
Stale registrations
A missing external object is never detected (dropTable returning false keeps the entity), so tag and policy listings return dangling objects. list* does not import either, so the store only knows objects that have been loaded.
8
Writing into user-owned metadata
Needs write privilege on the source, is visible to every engine, can be truncated by comment-length limits, and gets copied by clone and backup tools.
The marker only exists for one use: recovering a registration after the external DDL succeeded but the store write did not.
Proposal: C. The Gravitino ID is allocated and stored only in Gravitino. It is never read from or derived from the source. Fingerprints are used only to detect replacement, not to follow renames. Renames keep the ID only when Gravitino performs them, when a trusted event proves them, or when an admin adopts them. When a replacement is detected, grants and ownership are not inherited; descriptive metadata stays on the tombstone and can be restored.
Rollout
Phase 0: bug fixes, no storage change (can start now)
Import never upserts by a foreign ID. Reuse a marker ID only if it belongs to a live entity in the same catalog whose old path is gone (a verified rename). Otherwise allocate a new ID. Use insert-or-fail and never reset deleted_at. This fixes problems 1 and 2.
Kafka: allocate IDs with idGenerator and keep the topic UUID as a fingerprint (problem 5).
Add lifecycle tests: catalog recreate, two catalogs over one source, copied marker, same-path replacement, out-of-band column rename.
Phase 1: fingerprint + stop writing markers
Store the provider fingerprint, filled lazily on load. On a mismatch at the same path, tombstone the old entity and create a new one (problems 3 and 6).
Add a config switch to stop writing new markers. Existing markers are left in place and only read as a Phase 0 hint.
Phase 2: reconciliation
A periodic per-catalog inventory marks missing objects MISSING, then DELETED, with a fail-safe threshold. It also imports objects that were never loaded (problem 7).
Gravitino-initiated DDL writes a PENDING store record before the external DDL (outbox). This replaces the marker's recovery role.
Admin APIs: list unreconciled objects, adopt (old → new), restore.
Phase 3 (optional): HMS notification events and similar feeds, so that out-of-band renames can keep their IDs.
Questions
Should two catalogs over one source have separate IDs and governance? (Proposed: yes.)
On a detected replacement, should grants and ownership be dropped while tags stay restorable? (Proposed: yes.)
Is Phase 0 acceptable as an immediate fix before we agree on the storage format?
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Background
For external catalogs (Hive, JDBC, Kafka, ...), Gravitino writes its own entity ID into the source object as a marker (
gravitino.identifierin properties; for JDBC, in the comment). On load, the marker is used to recover the Gravitino ID. Without a marker, the loader falls back to the entity stored at the same path. See StringIdentifier, schema import, and table import.Gravitino's tags, policies, ownership, and grants hang off that ID. So which ID an external object resolves to decides which governance metadata it gets. Today that decision relies on a marker that anyone can copy, strip, or edit.
Problems
put(overwrite=true)with the marker ID, which is an upsert on the PK (ON CONFLICT (table_id) DO UPDATE SET catalog_id, schema_id, table_name, deleted_at=0). Loadingc2.db.tmovesc1.db.t's live row, with its tags, owner, and grants, intoc2; loading fromc1moves it back. No error is raised.dropCatalog(force)then recreate the same nameconvertToGravitinoId), so two catalogs over one cluster collide by construction (see problem 1).dropTablereturningfalsekeeps the entity), so tag and policy listings return dangling objects.list*does not import either, so the store only knows objects that have been loaded.The marker only exists for one use: recovering a registration after the external DDL succeeded but the store write did not.
How other platforms do it
(platform, platform_instance, name)fail_safe_threshold(75%)service.db.schema.tablemarkDeletedTablessoft-delete; restorabledb.table@clusterWhat they have in common:
Options
(catalog_id, type, path)table-uuid, Kafka topic ID, PGoid, HiveTable.id/createTime)Proposal: C. The Gravitino ID is allocated and stored only in Gravitino. It is never read from or derived from the source. Fingerprints are used only to detect replacement, not to follow renames. Renames keep the ID only when Gravitino performs them, when a trusted event proves them, or when an admin adopts them. When a replacement is detected, grants and ownership are not inherited; descriptive metadata stays on the tombstone and can be restored.
Rollout
Phase 0: bug fixes, no storage change (can start now)
deleted_at. This fixes problems 1 and 2.idGeneratorand keep the topic UUID as a fingerprint (problem 5).Phase 1: fingerprint + stop writing markers
Phase 2: reconciliation
adopt(old → new),restore.Phase 3 (optional): HMS notification events and similar feeds, so that out-of-band renames can keep their IDs.
Questions
All reactions