Skip to content

Releases: basic-automation/weftdb

WeftDB 0.1.0 — pre-beta baseline release

Choose a tag to compare

@github-actions github-actions released this 09 Oct 00:56
b9a8ee4

First public release, and the baseline that later releases' upgrade, rollback and
compatibility checks run against. WeftDB is pre-beta: the API will change, and the
version number says so. This release ships as pre-built binaries on the GitHub release
only; the library crates are not published to crates.io (build them from the v0.1.0
tag).

The Changed, Fixed and Security entries are for anyone who built WeftDB from source
before this release. A bare "0.1" in them means splimes 0.1.

Known limitations

  • Pre-beta. The HTTP API, the Rust API, the configuration and the on-disk format can
    change in any 0.x release. Nothing is covered by a stability promise before 1.0.
  • No authentication, authorization or TLS. weft-server serves an unauthenticated,
    plain-HTTP API; see
    SECURITY.md.
    Do not expose it to an untrusted network: keep the default loopback bind, or put it
    behind something that terminates TLS and authenticates callers.
  • Crash consistency is incomplete. A 2xx from a storage ingest survives a crash of
    the weft-server process, but not necessarily a power loss or an operating-system
    crash: the index commit is fsynced, the segment frame it points at is not.
    Maintenance (reconcile, split, squash, compaction) still rewrites committed frames in
    place, so a crash during a rewrite can tear one. And two seals into one aspect at the
    same moment can be given the same segment id, so that one replaces the other.
    docs/design/crash-consistency.md
    lists the 43 windows it verified. This release closes one of them, the deleted MVCC
    log (see Fixed). The durability work after it closes or mitigates the rest, though a
    legacy (rows-mode) database keeps some of them by design.
  • The next release changes the store layout. It moves a store to layout v2, in
    place, and the upgrade is one-way: once a newer release has written to a store, do not
    run v0.1.0 on it again. Before upgrading, stop weft-server and copy the whole store
    directory; POST /api/v1/storage/backup copies only the control plane, not the
    segment frames.
  • The Linux binaries need glibc 2.34 or newer, and weft-tui needs 2.39 or newer.
    They are built on Ubuntu 24.04. Ubuntu 22.04, Debian 12 and RHEL 9 run weft-server
    and weft-bench but not weft-tui; on an older glibc, build from source.
  • The macOS and Windows binaries are not signed. Gatekeeper and SmartScreen warn
    before running them. Check each archive against SHA256SUMS instead.
  • The bit-sliced codec is off in the release binaries. None of them is built with
    bitsliced-codec, so they never write the bit-sliced value codec
    (WEFT_SEGMENT_TRANSPOSED_MAX_OVERHEAD is ignored, with a warning) and cannot read a
    segment written with it; for a store that has such segments, build weft-server from
    source with --features bitsliced-codec. experimental-codecs is compiled into
    weft-bench only, for its advisory size estimates; nothing writes those codecs to
    disk.

Added

  • Interpolation as a first-class query. Ask for any resolution and get back a
    continuous series, reconstructed with the spline method you choose (linear,
    quadratic, cubic, polynomial), computed on SIMD/parallel CPU or GPU (wgpu).
  • Provenance labelling. Every returned point is marked raw, interpolated or
    extrapolated, so a synthetic value is never silently mistaken for an observed one.
  • Declared precision. Values are logically BigDecimal. Each aspect declares a
    physical encoding (F64, F32, ScaledI64, ScaledI128, Decimal128,
    BigDecimalText) and an error bound; a value the encoding cannot represent within
    that bound is rejected rather than quietly rounded.
  • Typed columnar segment store (.weftseg) on the measurement hot path, with
    bit-packing and delta-of-delta timestamp coding (an opt-in bit-sliced value codec sits
    behind the bitsliced-codec feature; see Changed), over a Turso (libSQL) control
    plane for catalog, metadata and the segment index.
  • Downsampling with mergeable partial reductions (.weftpart sidecars), including
    time-weighted averages, over an epoch-aligned bucket grid.
  • Apache Arrow interchange — sealed segments to RecordBatch and back, plus Arrow
    IPC and Parquet bytes.
  • HTTP API (weft-server) covering health/readiness, ingest, query, interpolation,
    storage management, backup/restore and a restore drill, with OpenTelemetry tracing.
  • Terminal UI (weft-tui) for exploring and administering an instance.
  • Weft-Bench (weft-bench), a reproducible, correctness-gated benchmark harness.
  • Analytics pipeline (weft-orchestration) chaining batching, pattern extraction,
    event detection, correlation and signal generation.
  • Pre-built weft-server, weft-tui and weft-bench archives for Linux (x86_64 and
    aarch64), macOS (x86_64 and arm64) and Windows (x86_64), with a SHA256SUMS file,
    attached to the GitHub release.
  • Interpolation grids requested over HTTP are capped at 10,000,000 points per
    request by default (weft_server::MAX_INTERPOLATE_OUTPUT_POINTS), on every
    /api/v1/interpolate* endpoint; set WEFT_MAX_INTERPOLATE_POINTS to change it. The
    variable is read once at startup, and a value that is not a positive integer stops
    weft-server from starting. A larger grid is a 400 that names its size and the
    limit, refused before anything is allocated; before, a fine resolution over a long
    range tried to allocate all of it.
  • weft-server calibrates the interpolation backends at startup. Once the listener
    is bound, it runs splimes::calibrate() once in the background on a blocking thread:
    that starts the GPU if there is one, times the single-thread, rayon and GPU backends,
    and sets where Backend::Auto switches between them. It takes several seconds and
    never delays serving or fails startup; until it finishes, requests interpolate on the
    CPU with splimes' default thresholds. The adapter (or why there is none) and the
    thresholds are logged. A CPU/software adapter (llvmpipe, lavapipe, WARP) is not
    calibrated unless WEFT_GPU_CALIBRATE=force; WEFT_GPU_CALIBRATE=0 skips calibration.
    Without it, splimes 1.0 never uses the GPU.
  • Weft-Bench calibrates like weft-server. An interpolation run (line protocol or
    --synthetic) calls splimes::calibrate() once before anything is timed, skipping a
    CPU/software adapter as the server does (it has no counterpart to
    WEFT_GPU_CALIBRATE=force), so the numbers come from the engine the server runs on
    that machine; --no-gpu-calibrate turns it off. The calibration, the GPU and
    Backend::Auto's thresholds are printed and recorded in the report's
    metadata.engine (bench schema v16) and its HTML view.
  • Third-party attribution. A root NOTICE, and a THIRD-PARTY-NOTICES file shipped in
    the weft-physical-type and weft-reduce packages, credit the two Apache-2.0 projects
    whose code is adapted here: the Chimp/Chimp128 codecs (from the authors' reference
    implementation) and the DDSketch quantile sketch (from Datadog's sketches-java), with the
    upstream NOTICE text and the license. The Chimp128 docs no longer credit DuckDB as the
    source.
  • Contribution terms. Every pull request takes one of two routes, the contributor's
    choice: a DCO sign-off (git commit -s) on every commit, or the WeftDB Individual
    Contributor License Agreement
    (CLA.md, version 1,
    adapted from the Apache Software Foundation's ICLA with Justin Icenhour as the
    recipient), signed once by a pull request comment. The contribution-terms check
    passes a pull request when either holds. Contributions are licensed
    MIT OR Apache-2.0. See
    CONTRIBUTING.md.
  • Inputs::register_dictionary_if_absent registers a dictionary as
    set_dictionary_metadata does unless a complete registration of that name exists,
    checking inside its own write transaction, and returns whether it registered it. A
    registration that cannot be read counts as existing and is left alone, and a write
    that loses an MVCC conflict is tried once more after a short random delay. The check
    sees registrations committed before its transaction began: one committed while it
    runs, like a second writer registering the same new dictionary at the same moment, is
    not seen, both are written, and reads pick the newest.

Changed

  • Crates renamed for publication. The core crate is now weftdb (was database,
    a name already taken on crates.io) and the pipeline crate is weft-orchestration
    (was database_orchestration). Import paths change accordingly:
    use database::… becomes use weftdb::…, and use database_orchestration::…
    becomes use weft_orchestration::….
  • The workspace builds on stable Rust (MSRV 1.95). The
    #![feature(stmt_expr_attributes)] gate is gone, so nightly is no longer required to
    build, test or depend on any crate. Only cargo fmt still uses nightly, because
    rustfmt.toml sets nightly-only options.
  • splimes moved to its own repository (basic-automation/splimes)
    and is now a crates.io dependency. It is released on its own schedule, and a stable
    WeftDB waits on a stable splimes.
  • splimes 0.1 → 1.0. WeftDB depends on splimes = "1", resolved to 1.0.0 from
    crates.io. splimes' migration guide
    lists the results that change; through WeftDB:
    • large inputs keep their method (0.1 swapped cubic for quadratic from 2,500 input
      points, and cubic and quadratic for linear abov...
Read more