Support for DType::DecFloat #10039
moshap-firebolt
started this conversation in
Feature Requests
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Vortex supports
Decimal- which is fixed precision numbers - same(precision, scale)for entire columnand it supports binary floating numbers.
We propose to add support for
Decimal Floating(akaDecFloat) - decimals whose scale varies per value.Such type exists in several data systems, but no columnar format atm can store them losslessly:
Decimal128DECFLOAT(16|34)DECFLOATNUMBERParquet's
DECIMALfixes scale per column; Arrow'sDecimal32/64/128/256likewise; ORC and Icebergthe same. So these columns get rounded into a fixed
(p,s), silently degraded toDOUBLE, orstringified.
How the ecosystem copes today: MongoDB's own Spark connector infers the type per sampled value and
falls back to
DataTypes.DoubleTypewhen the merge would exceed precision 38 — silently turning amonetary type into binary floating point. Trino's MongoDB connector samples to infer
(p,s).ClickHouse's MongoDB engine has no mapping at all.
Relevant prior art for the type itself:
DECFLOAT, defined as IEEE 754-2008decimal64/decimal128(COMPARE_DECFLOAT,TOTALORDER,NORMALIZE_DECFLOATfunctions)DECFLOAT,Firebird
DECFLOATLogical DType
DType::DecFloat { max_digits, nullability }, wheremax_digitsin {16, 34} corresponding toIEEE
decimal64anddecimal128. Both are fixed-width, so arrays stay fixed-width.Canonical Array
A
DecFloatArrayholding a common exponent plus a coefficient array plusPatchesforvalues whose exponent differs from the common one:
This is ALP's structure applied to decimals - the encoder signature Vortex already has:
ALP exists because most real
f64values are decimal-shaped, so it factors out a common decimalexponent and stores integers. For a decimal-float type that isn't a heuristic — it's the type's
literal semantics.
Two ways the decimal case is simpler than ALP for floats:
field of the value, so "does this match the common exponent?" is a comparison, not a trial
encoding. Compression is one pass, no search.
exist only to absorb exponent diversity, never precision loss.
DecimalBytePartsapplies to the coefficients unchanged: a signed i64 MSP plus u64 lower parts,compressed independently. A page whose coefficients all fit 64 bits stores a single i64 part.
Scalar
Scalar::DecFloatcarrying(coefficient: i128, exponent: i16), with the IEEE special values(
±Infinity,NaN,-0) represented in the coefficient/exponent encoding rather than as separatevariants, matching how the IEEE interchange format does it.
Comparison needs both IEEE orderings, which ANSI SQL already establish as the expected API
surface: numeric comparison (cohort-insensitive, so
2.00 == 2.0) for==/<, and a separatetotal-order predicate for sorting/
TOTALORDERsemantics.Why decomposed rather than 16 opaque bytes
FixedSizeBinary(16)holding raw IEEE bytes would be simpler, but gives up three things:puts exponent bits ahead of significand bits, so byte-wise comparison disagrees with numeric order
across differing exponents. Byte-wise min/max on an opaque column is therefore silently wrong.
nothing, while an opaque blob is incompressible.
that unlocks coefficient-only arithmetic.
Measurements
We benchmarked decimal128 arithmetic across available implementations (i9-13900K, GCC 11.4,
-O3 -march=native, one translation unit so every contender sees bit-identical values and the sameloop shape; 2M values × 5 reps). Slowdown vs native
double, money-shaped data:decQuad(DPD)Against the same operations on raw coefficients with the exponent hoisted to a batch scalar — what
the decomposed layout enables:
int64coefficient4.6–21× faster than the best scalar library, within 1.3–1.7× of hardware
double.int64coefficient compare is faster than
doublecompare.From the generated assembly, the coefficient kernels vectorize to packed SIMD when three conditions
hold together: 64-bit coefficients, no in-loop overflow check, and a non-widening multiply. The
layout above is what makes all three knowable from page metadata rather than per value.
The scalar libraries can't reach this because their fast paths are necessarily per-element — the
API is scalar, so the uniformity check is re-paid on every call and can never be hoisted out of a
loop. Storing the decomposition turns it into a page-header read.
Open questions / caveats
Non-normalization is required, and it complicates patches. IEEE decimal128 is deliberately
non-normalized:
2.00(200E-2) and2.0(20E-1) are distinct representations that compareequal, and the BSON spec requires that round-tripping "MUST not change its value or representation in
any way". So a conforming implementation cannot canonicalize on write — meaning a column of
2.00and
2.0carries two exponents despite being numerically uniform. The patch mechanism should staycheap rather than assume near-zero patch rates. (Snowflake's
DECFLOATnormalizes and consequentlycannot round-trip BSON; worth being deliberate about which side to land on.)
Clamping. The quantum exponent caps below the value range (6111 vs 6144 for decimal128), so
values near the top of the range are stored with the coefficient padded and the exponent reduced.
Lossless, but the encoder has to do it rather than reject.
ExtDTypeis a viable interim and would work today with a struct storage type, the wayuuidand
datetimeare handled — but it leaves the encoding on the table, since the compressor can't knowthe struct's two fields are an exponent and a coefficient.
BID vs DPD. BSON mandates BID, and Intel DFP (which MongoDB vendors as
src/third_party/IntelRDFPMathLib20U1) is BID-native. A BID coefficient is a plain two's-complementinteger, so it feeds
DecimalBytePartswith no transcoding. DPD would require conversion per value.All reactions