[DISCUSS] Trusted discovery with automatic classification #13575
laserninja
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
We are developing a solution for automatic PII classification and would like to explore a broader capability in Gravitino: trusted discovery with automatic classification across catalogs.
The goal is to help users find appropriate data, understand its sensitivity and reliability, and obtain the access they need. PII classification is a concrete starting use case. Gravitino's existing tags, ownership, catalog integrations, and governance APIs provide foundations for this workflow.
Proposed capabilities
1. Permission-aware discovery across catalogs
Provide a unified search experience over asset names, column names, descriptions, owners, and effective classification tags. Support structured filters for catalog, asset type, business domain, sensitivity, and available trust signals. Apply discovery permissions consistently to results, counts, and suggestions. Start with tables and columns, with a model that can extend to filesets, models, and other assets.
2. Pluggable automatic classification
Support classification algorithms and integrations through a common interface. Depending on the data and use case, a classifier could use:
Classifiers can inspect metadata and, where authorized, bounded samples of data. Define scan scope, scheduling, incremental rescans, and resource limits independently of the detection algorithm. Classification workers should publish findings through a shared contract; raw PII samples should stay out of catalog metadata.
3. Classification lifecycle and governance
Record the target asset/column, classification, detector/version, scan time, confidence where meaningful, and review state. Support configurable thresholds for automatic acceptance or human review, preserve manual decisions, and handle conflicting findings, rescans, and schema changes explicitly.
Map accepted classifications to existing tags so discovery and policy consumers can use them. Keep provenance and freshness visible. Classification identifies sensitive data; access-control integrations enforce the resulting policies.
4. Trust signals and access workflows
Show ownership, documentation, classification provenance, freshness, and certification where available in search results and asset pages. Let users route access requests to an owner or an existing approval system. Expose the same discovery and classification capabilities through APIs and MCP.
Initial delivery
Deliver one complete table/column workflow:
Build the shared discovery API in coordination with existing search work. Additional detectors, richer trust signals, access-request integrations, and other asset types can follow as separate increments.
Questions for the community
Related work
Issue #13576 tracks the initial result-ingestion and review slice of this broader proposal.
All reactions