Repository navigation
v0.1.0
First public release.
Added
- Derived columns declared in SQL:
bee.add_columnsets a prompt, a model, the source columns and an output type (enum,text,boolean,integer,numeric,jsonb,vector). Inserts and updates of the sources enqueue a job in the same transaction; the value is written back by a worker. - Backends:
llm(structured output with self-reported confidence),decision(typed questions to TypeSafe Jev with class probabilities; thedecisioncolumns of a row share one call, up to 16 questions),embedding(batched vectors, needs pgvector) andcustom(jobs claimed by your own worker). - Versions and lineage:
bee.update_columncreates a new version and recomputes only the rows computed with an older one. Every value is kept inbee.resultwith its model, prompt version, confidence, tokens, cost and latency; a value computed while the row or the version changed is recorded under the version it was computed with and the row is queued again. - Human overrides: a value written by hand is never overwritten (
override_policypin, oruntil_source_change);bee.unpinor setting the cell to NULL gives it back to the model. Writing the value a cell already holds is not an override. - Review queue: values below
confidence_thresholdare listed inbee.needs_review;low_confidence_policyholdkeeps them out of the column. - Budgets per column (
budget_usdper day, month or in total), enforced when jobs are claimed, withbee.budgets,bee.spentandbee.cost_by_column. Token prices per column (input_usd_per_mtok,output_usd_per_mtok) for providers that report no cost. - Retries with exponential backoff, dead jobs in
bee.dead_jobsandbee.retry_dead, abandoned jobs reclaimed after a timeout. - Backfill in chunks: declaring a column on a table of a million rows takes a fraction of a second and does not block writers.
bee.disableandbee.enablefor bulk loads. - Lineage retention per column (
lineage_retention_days), pruned in batches bybee.prune. - Least-privilege worker: the
bee_workerrole can only call the contract functions and read the views; it has no access to user tables. - Worker contract:
bee.claim_jobs,bee.complete_job,bee.fail_job,bee.reclaim_stale,bee.prune, and thebee_jobsNOTIFY channel. Any process that speaks it is a valid worker. - Two ways to install, generated from the same SQL files:
pgbee installon any PostgreSQL 15 to 18, managed services included, orCREATE EXTENSION pgbeefrompgbee extension-fileson self-hosted servers, withALTER EXTENSION pgbee UPDATEandpg_dumpsupport. - Reference worker
pgbee(Python 3.13):pgbee runwith one lane per backend so fast models never wait for slow ones,--backendsto split backends across machines,--drainand--max-secondsfor cron and serverless schedulers,pgbee status,pgbee install,pgbee extension-files. - Model providers: OpenRouter (all backends, reported cost) or any OpenAI-compatible endpoint such as OpenAI, Azure OpenAI, Ollama or vLLM (
llmandembedding, standard parameters only), chosen withOPENROUTER_API_KEY,OPENAI_API_KEY,OPENAI_BASE_URLorPGBEE_PROVIDER. - Docker images:
ghcr.io/essedev/pgbee-workerandghcr.io/essedev/pgbee-postgres(PostgreSQL 17 and 18 with the extension available), and adocker composequickstart. - Benchmarks anyone can rerun without cost in
bench/: a failure test against two app-side designs, scale on a million rows, lanes against a single loop.
Upgrading
- Nothing to upgrade from. This release ships extension version 0.14 (SQL files 0001 to 0014). Install with
pgbee install, or withCREATE EXTENSION pgbeeafter copying the files frompgbee extension-files(also attached to the GitHub release). - On managed services (tested on Neon) connect the worker directly, not through a transaction pooler: the pooler drops the notifications that wake the worker, which then only polls every
PGBEE_POLL_INTERVAL_SECONDS.