DataHub v1.7.0
Requirements
- CLI / Python SDK: 1.7.0
- Helm Chart: 1.1.0
Full upgrade guidance, including every breaking change and migration step: Updating DataHub — v1.7.0.
Upgrade path / ZDU: You must upgrade to v1.6.0 before upgrading to v1.7.0 — do not skip 1.6.0. Deploy v1.6.0 with Helm chart 1.0.3, let system-update complete, then upgrade to v1.7.0 with Helm chart 1.1.0. Enable Elasticsearch/OpenSearch ZDU (global.datahub.systemUpdate.zdu) with the 1.1.0 chart on a subsequent OpenSearch/Elasticsearch version bump — not during the v1.6.0 install.
Feature highlights
UI and experience
- Metrics and Semantic Models — first-class
metricandsemanticModelentities with dedicated pages, a metrics home/sidebar experience, autocomplete, modular summary tabs, and lineage wiring (#18134, #18350–#18407, #18442–#18459, #18462, #18482, #18701, and related). - Logical Models UI — create, link, edit, and delete logical models from the UI, in addition to API/SDK paths (#18498).
- Data Product lineage — data products participate directly in the lineage graph (#18463).
- Multi-language (i18n) default on —
I18N_ENABLEDdefaults to on in OSS; UI follows the browser locale. New Beta locales include French, Italian, Norwegian Bokmål, Swedish, Hungarian, and Finnish (#18285, #18282, #18221, #18222, #18263, #18265, #18339, #18520). - Lineage graph — always the latest lineage experience;
LINEAGE_GRAPH_V2/LINEAGE_GRAPH_V3flags removed.
Ingestion — new sources
- Cube — semantic-layer connector (#17964)
- AWS Kinesis — Kinesis Data Streams and Amazon Data Firehose (#17592)
- MicroStrategy — BI connector (#18158)
- Open Data Contract Standard (ODCS) — contracts from S3, GCS, HTTP, and Git (#17331, #18474, #18477)
- ThoughtSpot (#17400)
- SAP Datasphere (#17802)
- DocumentDB platform — opt-in
platform: documentdbon the MongoDB source for AWS DocumentDB (#17443)
Ingestion — major connector improvements
- Hex — major in-place upgrade: table- and column-level lineage from Hex APIs, Project → Component links, run history, optional AI context documents; Components ingested as Chart entities (see breaking changes) (#17376).
- Snowflake — Semantic Views can emit first-class
semanticModel/metric/ logical-dataset entities (semantic_views.emit_semantic_model_entities; OSS default off / auto-resolve) (#18395, #18509). - Databricks Unity Catalog — Lakehouse Federation (foreign catalogs); usage/ops/queries from
system.query.historyvia the shared SQL parsing aggregator; ML model ingestion controls fixed (#18213, #17971, #18220, and related). - Matillion — foldered container hierarchy, environment-scoped lineage, run history at pipeline and component levels, corrected console links (#17927).
- Glue — Lake Formation resource-link schema resolution (default on), column-level LF tags, cross-account platform instances (#17963, #17812).
- Redshift — multi-line SQL no longer dropped from lineage/usage; per-query popularity stats;
table_patternapplied to SQL-parsing path (#18542, #18001, #18065). - BigQuery — table stats from
INFORMATION_SCHEMA.PARTITIONS(see breaking changes); usage window fields consolidated to top-level (#18367, #18133). - Power BI / Mode — column-level lineage preserves original upstream column casing (#18181).
- Spark — Apache Spark 4.x support (Scala 2.13 agent); OpenLineage 1.50 with full shading for EMR/DataZone coexistence (#14911).
- S3 / ABS — profile data-lake files without PySpark; optional
emit_folders_onlyfor object-store folder cataloging (#18347, #18599, #18437). - Airbyte — Public API stream namespace recovery (#18727).
- Kafka — profiling support (#14367).
- Great Expectations — GX Core 1.x action path (#18706).
Search, auth, and metadata
- View authorization overhaul — entity types restricted by default when VBAC is on;
VIEW_UNRESTRICTED_*overlays; documents view-restricted by default; schemaField can inherit VIEW from parent dataset (#18612, #18664, and related). - Structured properties — ES field-name collision rejection; keyword max-length validation; type-mismatch reindex detection.
- Configurable search entity-type defaults —
SEARCH_*_ENTITY_TYPESenv overlays for GraphQL search/autocomplete/browse defaults. - File upload / object storage — config path moved to
datahub.objectStorage(see breaking changes); documentation file attach/download guide.
Operations and platform
- Secrets caller guard —
SECRET_SERVICE_CALLER_GUARD_MODEdefaults to ENFORCE; human PATs can no longer decrypt UI secrets via GraphQL. - Primary storage read pool — optional Ebean/Cassandra read pool for entity-aspect reads (
EBEAN_READ_POOL_*/CASSANDRA_READ_POOL_*). - Ebean transaction conflicts — stable retryable 503 /
DATABASE_TRANSACTION_CONFLICTinstead of opaque 500s on deadlock exhaustion. - Docker tags — floating
:headremoved; coordinated:quickstartand immutable:sha-*tags. - Optional Loki log shipping —
LOG_AGGREGATOR_ENDPOINTfor core services and the frontend. - jose4j shipped for Kafka SASL/OAUTHBEARER JWT validation.
- Airflow plugin — Airflow 2.x dropped; Airflow 3.0+ required. Prefect plugin requires Prefect 3.x.
- Orchestration plugins — Airflow / Dagster / Prefect / GX default emit mode is ASYNC.
- Built-in column classifier removed —
DataHubClassifier/acryl-datahub-classifyno longer shipped.
Breaking changes
Review the full Breaking Changes section in Updating DataHub before upgrading. Summary of items that may require action:
| Area | What changed | Who is affected |
|---|---|---|
| Must install v1.6.0 first | Do not skip 1.6.0; ZDU enablement uses Helm 1.1.0 after 1.6.0 + system-update | All upgraders from pre-1.6.0 |
| Secrets ENFORCE | Human PATs/browser can no longer decrypt UI secrets; use datahub-actions / system client or AUDIT temporarily |
Anyone using user PATs for getSecretValues |
| Airflow 2 dropped | Plugin requires Airflow 3.0+ | Airflow 2.x deployments — pin plugin <= 1.6.0 or upgrade Airflow |
| Prefect 3 required | datahub-prefect requires Prefect 3.x |
Prefect 2.x users |
| Classifier removed | Built-in DataHubClassifier gone; recipes with classification.enabled: true fail fast |
Classification-enabled recipes |
| Plugin emit ASYNC | Airflow/Dagster/Prefect/GX default emit is async | Operators needing sync/raise-on-reject — set SYNC_PRIMARY |
| Hex Components → Chart | Component entity type/URNs change; tags/policies may need reapply | Hex workspaces with Components |
| Spark OL 1.50 trimmers | Partition dirs stripped from FS/object-store dataset names by default | Spark lineage without path_spec_list / file_partition_regexp |
| Power BI / Mode CLL casing | Upstream column paths keep source casing | Re-ingest; remove lowercase workarounds on SQL sources |
| BigQuery table stats | Stats from PARTITIONS; empty/external/views/snapshots lose some timestamps |
Set use_legacy_table_stats: true to restore |
| Object storage YAML | datahub.s3 → datahub.objectStorage |
Custom YAML overrides (env vars largely unchanged) |
| View authorization | Restricted-by-default + VIEW_UNRESTRICTED_*; documents restricted |
Deployments with VIEW_AUTHORIZATION_ENABLED=true |
| Lineage graph flags | LINEAGE_GRAPH_V2 / V3 removed |
Anyone still setting those env vars |
Docker :head |
Use quickstart / immutable sha-* / release tags |
Compose and production pin practices |
| Workunit processors | Helper functions → processor classes; several renames | Custom ingestion code calling old helpers |
| Relationship edge uniqueness | One aspect per (source, dest, relationship) signature (#18845) |
Custom / plugin entity registries |
| Logical parent auth | Edit Entity required on child and parent when linking | Logical-model operators |
Potential downtime: Structured-property Elasticsearch type-mismatch reindex when both system-update flags are on — see Updating DataHub — v1.7.0.
Deprecations: Hex lineage time/page-size recipe fields; BigQuery usage.* window / formatting fields migrated to top-level — see the v1.7.0 Deprecations section in Updating DataHub.
Contributors
Thank you to everyone who contributed to v1.7.0. For the complete changelog, compare v1.6.0...v1.7.0.