4.1.3
Release Date: July 14, 2026
Behavior Changes
CTASnow preserves explicitly declaredVARCHAR(N)column lengths instead of widening them toVARCHAR(MAX). Existing tables are unaffected; new tables created withCTASwill enforce the declared length on subsequent writes. #73498- Querying
sys.fe_memory_usageorsys.fe_lockswithout theOPERATE ON SYSTEMprivilege now returns a clear access-denied error instead of a misleading node-lookup failure. #73567 FILES()and broker/stream load no longer apply a session timezone shift toINT64Parquet timestamps written withisAdjustedToUTC=false; those timestamps are now treated as wall-clock values and loaded as-is. Data loaded from such files before v4.1.3 may differ from data loaded after upgrading; reload if consistency is required. #73674- Successfully committed multi-table transaction stream load jobs are now correctly shown as
VISIBLEininformation_schema.loadsandSHOW STREAM LOADinstead of remaining stuck atPREPARING. #74386 - Connector incremental scan range scheduling now consistently reuses the deployed fragment's driver layout, preventing scan ranges from being incorrectly assigned to non-existent drivers. #74674
LIKEconstant-folding now matches MySQL 8 backslash-escape semantics, correcting cases where patterns such as'a\\\\b'previously returned opposite results. #74814- Routine Load now supports the
property.kafka_partition_discoveryproperty, which allows partition auto-discovery to continue even whenkafka_partitionsandkafka_offsetsare specified to seed exact starting offsets. Whenproperty.kafka_default_offsetsis not set, the default starting offset for partitions discovered after the job already has consuming progress changes fromOFFSET_ENDtoOFFSET_BEGINNING— and this applies to all auto-discovery jobs, not only those using the new property. #74729 - Non-group-by aggregates are now pushed down through
UNION ALLbranches before merging, reducing network transfer and memory usage for queries that aggregate over a union. #73930 - IVM maintenance queries are now re-derived from the current view definition at each refresh instead of using the frozen query text stored at
CREATEtime; existing MVs automatically benefit from rewriter bug fixes without needing to be recreated. #74881 - Sample-based tablet pre-split now spreads pre-split shards across all compute nodes (
SPREADplacement) instead of packing them onto the source tablet's worker (PACKplacement), improving load parallelism. #75514 ALTER TABLE ... MODIFY COLUMNthat changes only the column comment now takes the lightweight metadata-only path instead of spawning a full schema-change job, and this now works on Primary Key columns as well. #75325FLOORandCEILare now treated as non-reserved keywords and can be used as column names without quoting. #75241SHOW FUNCTIONSoutput now always includes theisolationproperty (sharedorisolated) in the Properties column for UDFs and UDAFs. #75255- The default value of
lake_vacuum_min_batch_delete_sizeis raised from 100 to 200, improving vacuum throughput on object storage by batching more stale-file deletions perDeleteObjectsrequest. #74304 - Iceberg REST catalog tables with vended credentials are now cached and their credentials are refreshed in the background, eliminating per-
getTable()GetDataAccesscalls that caused AWS Lake Formation rate limiting. #75431 - IVM
bitmap_union,hll_union, andpercentile_unionaggregate states are now stored once in the materialized view instead of twice (visible column + hidden__AGG_STATE_column), halving storage for those sketch types. #75760 - Incremental materialized views now support
bitmap_agg,hll_union,percentile_union, andbitmap_unionaggregate functions, enabling exact distinct-count and sketch-based aggregations to be maintained incrementally. #75587 #75610 - Sample-based tablet pre-split tablet count is now rounded up to the nearest multiple of the active compute-node count for even distribution, and bounded below by a minimum tablet size to avoid excessive fragmentation on small loads. #75360 #75584
Improvements
- The
ngram_searchfunction now accepts a non-constant needle argument. #74675 - Added an HTTP authentication framework controlled by the
enable_http_authFE configuration, gating authentication and RBAC enforcement on all external HTTP endpoints. #73822 - Added refresh and placement observability columns (
refresh_warehouse,refresh_resource_group,refresh_mode,refresh_type,last_refresh_details) toinformation_schema.materialized_views. #74342 - Added opt-in lazy refresh of the external statistics cache on journal replay, controlled by a new FE config, to prevent a slow or stuck external metastore from stalling FE journal replay or startup. #74371
VARCHARlength increase is now allowed on range-distribution (shared-data) sort-key columns via fast schema evolution without a data rewrite. #74698- Added a stack-trace dump when a shared-data transaction log write exceeds a configurable threshold, making slow
put_txn_log/put_combined_txn_logcalls easier to diagnose. #74704 - Tablet pre-split meta-tier footer readers now support
DATE,DATETIME,DECIMAL,VARCHAR, and ORCTIMESTAMPsort keys, reducing the number of loads that must fall back to data-tier sampling. #74710 #74739 #74792 #74902 #74955 #75186 #75209 #75427 #75697 - Sample-based tablet pre-split now applies to
INSERT INTO ... SELECT ... FROM <OLAP table>loads, and also to column-listINSERTstatements that include all sort-key columns. #74828 #75345 - Added Adler-32 checksum protection for shared-data tablet metadata and transaction log files, enabling silent corruption to be detected on read. #74924
- Added the
txn_max_committed_pending_publish_msFE metric per database, reporting the age of the oldest committed-but-not-yet-published transaction to help detect stalled version publishing. #75025 - Tablet split/merge is now triggered in real time from the publish-version response, reducing the lag between a load completing and automatic split/merge being initiated. #75010
- Optimized condition-update compare phase for lake primary-key tables by routing no-SST condition-merge tasks to the
pk_index_executionthread pool. #74572 - Scoped lake schema-change and rollup job locks to the table level instead of the whole database, reducing lock contention on concurrent operations for other tables in the same database. #75087
- Narrowed several database-level write locks to table-scoped intensive write locks in shared-nothing mode, reducing lock contention during BE report callbacks and cooldown operations. #74521 #74523
- Avro Routine Load now supports native
MAPandSTRUCTtarget columns. #74901 - Range-colocate tablet stability gating now waits for StarOS placement convergence before marking a group stable, ensuring colocate joins achieve host-local execution. #75290 #75656 #75883
- Improved CBO statistics for external tables: the optimizer now estimates row counts from Iceberg manifests without full file enumeration, corrects Hive/Hudi row-count underestimation for Parquet/ORC compression, adds async row-count statistics for JDBC connectors, and provides NDV estimation fallback for Iceberg and external connectors when Puffin stats are unavailable. #75280 #75082 #75083 #75092 #75097 #75382 #75474
- Iceberg manifest column statistics are now cached selectively for clustered columns only, reducing FE heap consumption for wide tables with many data files. #75395
- External table statistics collection now supports persistent predicate-column tracking across FE restarts and HA failovers, enabling auto-ANALYZE to target the correct columns. #75653
- Added structured
[ExternalStats]log lines covering the full lifecycle of external-table statistics collection from scheduling through execution. #75335 #75529 SHOW ANALYZE STATUSnow includes partition, column, and snapshot metadata in the Properties column for external table statistics jobs. #75630- The statistics source (
TABLE_METADATA,ANALYZE, orNONE) for each external table is now exposed in the query runtime profile. #75253 - Added support for partition filter requirement and partition count limit for Iceberg and Delta Lake external tables (previously only available for Hive, Hudi, and Paimon). #75790
- Supports sub-1% sampling ratios in
TABLE SAMPLEand histogramANALYZE, fixing failures on large tables where the computed ratio would truncate to zero. #74551 - Added the
jemalloc_confBE configuration item, making jemalloc runtime options visible viainformation_schema.be_configs. #75344 - Added the
compaction_chunk_reset_memory_tracker_threshold_percentBE configuration to reduce memory usage during Primary Key compaction in shared-nothing mode by releasing retained chunk capacity. #75091 - Upgraded staros to v4.1.1, including persistent
datacache.enableacross restarts, per-worker-group shard warmup timeout override, and improved S3 retry jitter. #75204 - Optimized SQL credential redaction on the audit hot path by skipping the regex scan when no credential marker is present in the SQL string. #74812
- Expression-driven on-demand lazy column loading for the Parquet scanner reduces unnecessary I/O on multi-branch
ORqueries. #74886 ds_hll_count_distinct/DataSketchesHllnow produces stable cardinality estimates by using the composite estimator instead of the order-dependent HIP estimator. #75053
Security
- [CVE-2026-45416] [CVE-2026-44249] [CVE-2026-45673] Upgraded Netty to 4.1.135.Final to fix SNI handler heap exhaustion (DoS), IPv6 subnet filter bypass, and DNS cache poisoning. #74668
- [CVE-2026-54512] [CVE-2026-54513] Upgraded
jackson-databindto 2.21.4 to fix two deserialization vulnerabilities. #75373 - [GHSA-2r2c-cx56-8933] [GHSA-47qp-hqvx-6r3f] Excluded
org.jline:jline-remote-telnetfrom Hadoop transitive dependencies to remediate unauthenticated Telnet-server DoS vulnerabilities. #75066 - [CVE-2026-39822] Updated pprof prebuilt to fix a vulnerability in the pprof binary. #76248 #74669
- Fixed SQL injection in
information_schema.task_runswhere a single quote in a predicate value could escape the literal boundary. #75520 tencent.cos.access_key,tencent.cos.secret_key, andiceberg.catalog.jdbc.passwordare now masked inSHOW CREATE CATALOGoutput. #74696- Fixed an out-of-bounds read in
url_decodewhen the input ends with a truncated percent-escape sequence. #75139 - Fixed
HyperLogLog::deserializeaccepting out-of-rangeSPARSEregister indices, which could corrupt heap memory and crash the BE on malformed input. #75521 - Fixed
bar()rejecting negative width values, which previously allowed unbounded string growth and BE memory exhaustion. #75143
Bug Fixes
The following issues have been fixed:
add_filespopulated Iceberg file bounds with Parquet physical encoding bytes instead of logical typed values, causing incorrect file-level min/max pruning (for example, onDECIMALcolumns). #69207ApplyTuningGuideRulethrewUnsupportedOperationExceptionwhen traversing plan nodes whose input lists were built as immutableList.of(...). #70785INSERT OVERWRITEtwo-phase re-plan could produce stale lambda-argument column reference IDs from the first planning session, causingexpr_type does not match slot_typeerrors. #73273- Partial updates on tables with GIN (inverted) indexes caused queries to hang indefinitely or fail when the GIN-indexed column was omitted from the update. #73773
- Lake PCU (partial column update) crashed or silently corrupted data when a schema drift occurred between the rowset schema and the tablet schema. #74005
- A combined
ALTER TABLEon an external Iceberg table with multiple schema clauses incorrectly re-executed all previously queued actions on every clause dispatch. #74036 PartitionedSpillerWritercrashed withSIGSEGVwhen thenum_rowssnapshot exceeded the actual chunk row count during partition-flush and resource-group-cancellation interleaving. #74081- BE process could exit unexpectedly during startup (typically right after deployment) because SIGPIPE was not ignored in BE signal initialization. #74424
- Parquet temporary dict-code columns leaked to upper layers when a struct VARCHAR subfield fill was skipped due to row-range filtering, causing type mismatches. #74452
SELECT ... INTO OUTFILErecordedReturnRows=0in the audit log instead of the actual exported row count. #74467TabletChecker.doCheck()threwIllegalMonitorStateExceptioninblockingAddTabletCtxToSchedulerdue to a lock type mismatch, causing entire checker rounds to abort silently. #74596information_schema.COLUMNSalways returnedNULLforDATETIME_PRECISION, breaking MySQL-protocol clients that derive column size from that field. #74623- MV refresh failed with
Duplicate keywhen the query joined two tables with the same unqualified name across different databases or catalogs. #74730 - Spillable hash join probe crashed under certain conditions. #74978 #75140
- Iceberg
truncateandbuckettransform functions crashed the BE withSIGFPEwhen the width or bucket count argument was zero. #74998 mod()andpmod()crashed the BE withSIGFPEwhen the dividend wasTYPE_MINand the divisor was-1. #74980histogram()crashed the BE withSIGFPEwhenbucket_numwas zero or negative. #75041encode_fingerprint_sha256crashed withSIGSEGVwhen all input rows wereNULL. #75042LIKEpatterns containing the single-char wildcard_returned incorrect results when evaluated via a GIN inverted index. #75551- AND-only
MATCHqueries against a GIN inverted index returned a spurious error when the target segment was empty. #75161 - CLucene
match_allqueries returned incorrect results; resolved by upgrading the CLucene dependency. #75180 - Vector index rewrite registered a synthetic distance column directly on the shared table schema, causing
Multiple entries with same keyerrors for unrelated concurrent queries on the same table. #74785 - Join reorder pruning could prune columns still referenced by scan predicates, causing statistics estimation to throw
missing statistic of col. #74791 avg(DISTINCT x)was incorrectly rewritten via a sum/count materialized view, silently dropping theDISTINCTand returning wrong results when duplicates existed. #75071ALTER TABLE ... MODIFY COLUMN ... AFTER <nonexistent_col>threw an internalNullPointerExceptioninstead of a clean semantic error. #75073SHOW CREATE ROUTINE LOADemitted a spurious leading comma before the first load-description clause for jobs without aCOLUMNS TERMINATED BYclause. #75522SECURITY INVOKERviews with CTEs could fail privilege checks with an NPE when the CTE name was mistaken for a real table reference. #74813ReduceCastRuleaborted query planning with aSemanticExceptionwhen a date/datetime boundary literal shifting would overflow the representable range (for example,<= '9999-12-31'). #75036SplitJoinORToUnionRuleemitted duplicate rows when a join condition used null-safe-equal (<=>) disjuncts. #75038Tracersshared across parallel metadata-preparation threads for multi-table external queries causedIllegalStateExceptionunderenable_profile=true. #74746- Partition consumer errors in
ChunksPartitionerwere silently discarded, allowing a partitioned TopN to return partial or wrong results without surfacing an error. #74693 - BE vacuum tasks continued running as zombies after the FE caller's timeout elapsed, exhausting the
RELEASE_SNAPSHOTthread pool and collapsing vacuum throughput. #74694 - An autovacuum race could momentarily compute a
minActiveTxnIdone greater than an in-flight transaction, causing the BE to delete still-needed combined transaction logs and permanently wedging publish. #74906 - Queries were incorrectly marked as canceled after finishing successfully due to a race between FE EOS-cancel and BE stage-2 deploy. #75009
- BE crashed with
SIGSEGVinAggTopNRuntimeFilterUpdaterImplwhen the aggregate TopN runtime filter build key was aConstColumn. #74809 #74941 array_map/transformsilently droppedNULLrows and returned wrong row counts when all non-null arrays were empty. #75141LARGEINT/DECIMAL128literals above 2^64 were silently truncated to 64 bits in JIT-compiled expressions. #75137- UTF-8 string functions (
split,split_part,str_to_map) read beyond the end of the string when the last character had a truncated or invalid multi-byte lead byte with an empty delimiter. #75068 parse_json()silently returnedNULLfor malformed JSON even inALLOW_THROW_EXCEPTIONSQL mode instead of failing the query. #74976- Strict-mode numeric narrowing casts incorrectly raised overflow errors on
NULLrows whose slot data was undefined. #74903 - Pre-1970 Parquet
INT64timestamps with a non-zero sub-second part were decoded to garbage values due to negative truncating-division remainder. #75207 - Pre-1970 ORC
TIMESTAMPvalues had their sub-second component dropped on load. #75432 - ORC stripe min/max timestamp statistics were decoded incorrectly for pre-1970 and sub-second bounds, causing data files to be incorrectly pruned. #75543
- Nested
INT96Parquet timestamps insideARRAY,MAP, orSTRUCTcolumns lost one session-timezone offset on load. #74868 - Parquet
UINT_32values were sign-extended instead of zero-extended when loaded into aBIGINTcolumn, silently storing negative values for high-bit unsigned integers. #75002 HiveDataSourcedestructor caused a heap-use-after-free by destroying_pool(and itsExprnodes) before_scanner_ctx(which holds predicates referencing those nodes). #74818- Reading gzip-compressed JSON Hive external tables with OpenX SerDe failed with
UTF8_ERRORwhen a multi-byte UTF-8 character straddled an 8 MB decompression buffer boundary. #74827 - ADLS2
ListPathscrashed withSIGSEGVon non-HNS accounts due to missing JSON fields that the client accessed unconditionally. #75166 unnestcrashed or returned wrong results when multipleUNNESToperators shared the same input array column and consumed different subfields. #75012 #75445 #76002query_mem_limitwas not enforced duringunnestexecution, allowingunnestover large arrays to OOM-kill the BE instead of failing the query. #75179TopNwithRANKboundary dropped one row when the rank limit fell exactly on a chunk boundary. #75045- Column pruning after
PushDownDistinctAggregateRulecould generate an empty analytic (window) operator, causing planning or execution errors. #74810 EliminateSortColumnWithEqualityPredicateRuleset the row limit only on the scan operator without setting a global limit, causingCOUNT(*)over a limited sub-query to return more rows than expected under concurrency. #74983- Lake primary-key persistent index rebuild used wrong segment iterator positions in segment-range mode, causing incorrect key-range filter application. #74887 #75206
DROP PERSISTENT INDEXmodifiedrebuildPindexVersionwithout a table lock;RestoreJobpost-restore mutated MV base-table info under only a DB READ lock;FinalizeCreateTableActionpassed a DB-level lock across iterator creation. #74968dumpImagecould strand the global meta lock indefinitely if acquiring a per-database lock threw mid-loop. #75488- Multi-statement stream load leaked one
TxnStateCallbackFactoryentry per transaction, growing without bound and eventually exhausting FE heap. #75188 information_schema.task_runsrow count for histogram statistics could overflowprimary_key_limit_size(128 bytes) when catalog, database, table, or partition names were long. #75735- BE JVM metrics emitted invalid Prometheus
# TYPElines (label sets inside the metric name), causing Prometheus to abort the entire scrape. #75240 SHOW PARTITIONSandinformation_schema.partitions_metareported all physical partitions' bucket counts as the table-level default instead of each partition's actual bucket count on shared-data tables. #75734SHOW PROC '.../index_schema/<id>'returned the base table schema for all rollup indexes on shared-data (CLOUD_NATIVE) tables. #76069ALTER TABLE ... MODIFY COLUMNno-op clauses were incorrectly routed to the lightweight comment path, causingMODIFY COLUMN COMMENT can not be combined with other alter operationserrors in batchALTER TABLEstatements. #75736isCommentOnlyModificationcould misidentify key/aggregation columns as comment-only changes due to incorrectisKey/aggregationTypenormalization. #75545ALTER VIEWcould commit a cyclic view definition that caused subsequentSELECTto throwStackOverflowError. #75033OrderedPartitionExchangercaused a heap-use-after-free when a downstream consumer mutated the previous chunk whileaccept()still held a pointer to it. #75279- NLJoin crashed when the build-side's slot descriptor was non-nullable but the runtime state was nullable. #75343 #75788
CAST(json/variant AS struct)crashed the BE at fragment prepare when a struct field name was not a parseable JSON path. #75355- Dict-decode for nested dictionary expressions could produce incompatible dictionary translations between producer and consumer fragments, causing
Dict Decode failederrors at runtime. #75246 - Schema change crashed with a null-pointer dereference when
get_rowset_by_versionreturnednullptrand thegtidcomparison was placed before the null check. #74855 - Shared-data cluster snapshots became unrestorable after a tablet split/merge because the snapshot manager did not consider reshard jobs when deciding whether to reap parent tablet metadata. #75638
- File-bundling vacuum incorrectly flagged zero-row bundled segments of sibling tablets as non-shared, causing their bundle file to be deleted while other tablets still referenced it. #75689
- Shared-data table compaction publish dropped rollup/synchronous-MV indexes that became visible after the compaction transaction began, leaving the bundle file without those indexes. #76105
- Shared-data persistent-index compaction incorrectly deleted pass-through-reused SSTable files when a compaction was dropped during tablet reshard. #75726
NOT NULL-to-nullable flat-JSON column schema evolution caused aCHECKcrash in the compaction read path. #75680count_combineover a nullable column crashed the BE withSIGSEGVin the streaming pre-aggregation pass-through path. #75298- Java UDFs failed to load on JDK 21+ due to a reflective
DirectByteBufferconstructor lookup that was removed in JDK 21. #75666 CTASinto a Unified catalog (Hive metastore) always failed because the parser did not support theENGINEclause inCREATE TABLE AS SELECT. #75771JoinTuningGuidefeedback-driven join rebuild lostpredicateCommonOperators, causingInputDependenciesCheckervalidation failures on plans with common sub-expression reuse. #75773- Query cache normalization crashed with
Preconditions.checkStateon tables with sub-partitions where empty sub-partitions were pruned before the version list was built. #75789 replayFromJsonsilently skipped session variables stored under a legacy alias, causing query dump replay to fall back to default values. #75813- Iceberg
_row_idvirtual column returned incorrect values for data files with more than one Parquet row group due to double-counting of the row-group start offset. #75758 - Iceberg DELETE/UPDATE planner could not locate the target scan node because it matched by synthetic table ID instead of physical table identity, losing the base snapshot ID and conflict-detection filter. #76013
FragmentContext::set_final_statuscrashed withSIGSEGVwhen acancel_plan_fragmentRPC arrived beforePipelineExecutorSet::start()was called. #75030QueryContextcould be reclaimed whileFragmentExecutorwas still tearing down a fragment, causing heap-use-after-free inResGuard::reset(). #74978StringSearch::_patternwas uninitialized, allowing a default-constructedsearch()to dereference an uninitialized pointer. #75614DATETIMEmicroseconds were rendered using the JVM default locale's digit set, causing non-ASCII digits on locales such as Arabic or Persian that broke boundary value parsing in tablet pre-split. #75001- Partition row counts could be written as zero into
_statistics_.column_statisticsafterINSERT OVERWRITEbefore tablet stats were refreshed, causing the optimizer to collapse partition cardinality estimates. #74801 enable_statistic_collect_on_first_loadtable-level override could not enable first-load statistics collection when the global config disabled it. #74794PushDownNonGroupedAggregateBelowUnionproduced nullable outputs with non-nullable declared types when a union branch had no input rows, causing BECHECKfailures. #76101