Grounded v0.2.0
Image
ghcr.io/ncecere/grounded:v0.2.0
ghcr.io/ncecere/grounded@sha256:2f896d385feb201d03c3d2a38b6bb1d77ac22f54abfb82093c54894a9a7d3bfd
Deploy by digest. Verify the signature (keyless, GitHub Actions OIDC):
cosign verify ghcr.io/ncecere/grounded@sha256:2f896d385feb201d03c3d2a38b6bb1d77ac22f54abfb82093c54894a9a7d3bfd \
--certificate-identity-regexp '^https://github.com/ncecere/grounded/\.github/workflows/.+@refs/tags/v0.2.0$' \
--certificate-oidc-issuer https://token.actions.githubusercontent.comThe OCR sidecar (optional, components/ocr-tesseract), signed the same way:
ghcr.io/ncecere/grounded-ocr:v0.2.0
ghcr.io/ncecere/grounded-ocr@sha256:1f17eb7d7b979460b91b6673278585230dcc7b33965505d413a2c6aa38cd9099
Release assets
grounded-v0.2.0-<os>-<arch>.sbom.spdx.json: the image's SBOM (SPDX JSON) for each platform, as attached to the image.grounded-v0.2.0.digest.txt: the image reference by digest.checksums.txt: SHA-256 of the files above (sha256sum -c checksums.txt).
The SBOM and build provenance are also attached to the image as attestations, which the signature covers:
docker buildx imagetools inspect ghcr.io/ncecere/grounded@sha256:2f896d385feb201d03c3d2a38b6bb1d77ac22f54abfb82093c54894a9a7d3bfd --format '{{ json .SBOM }}'.
Grounded v0.2.0
v0.2.0 is about measuring answer quality and running many teams. Teams can now test their knowledge bases and agents with evaluation sets. Platform admins can map identity-provider groups to teams, price model use and enforce monthly budgets, and read scanned documents with OCR. Every new feature is optional, and all but evaluations are off until an admin turns them on. The full list of changes is in the changelog, and the plan with the owner's decisions is in v0.2.0.md.
v0.2.0-rc.1 ran on the reference install before this release; v0.2.0 is the same code, apart from the release job and documentation. These notes cover both.
Highlights
- Evaluations (A2). A knowledge base and an agent have an Evaluations tab for team editors and above. A set is a list of test questions, each with the documents a good result finds (picked documents, URL prefixes or filenames) and, for answers, phrases it must mention. Questions are typed, imported from CSV or ragbench's URL-judged JSONL, or added from your own conversations and the agent's Try it panel. A retrieval check (recall@k and MRR, nearly free) or a full-answer check (cites an expected document, mentions the phrases, doesn't refuse wrongly; SystemOne support rate when on) shows what failed and what came back instead, with a score over time and a comparison of two runs. Sets can run automatically after a publish, a profile migration switch, or nightly when documents changed, and a drop notifies the team's editors. Nothing from other people's conversations is copied (ADR-0010 is unchanged). Platform admins can turn evaluations off under Admin → Limits → Evaluations. Guide:
operations/evaluations.md. - Costs and budgets (E2). Off by default. Admins enter dated prices per model and unit, and choose Track only (spend per team, agent and model, with a daily chart and CSV) or Enforce (monthly team budgets in the platform's time zone, a warning at 80%, and at 100% the team's chats, searches and ingestion pause until an admin raises the budget, grants an extension, or the month ends). Teams can override the platform's mode, so one team can pilot Enforce. Team owners and admins see their spend; members only see the banner. Runbook:
operations/costs.md. - OCR for scanned documents (B4). Off by default. Admin → Parsing turns it on with one backend: the new Tesseract sidecar image
ghcr.io/ncecere/grounded-ocr(Kustomize componentocr-tesseract), Apache Tika's-fullimage, or a vision model from the catalog. Only pages without a text layer are read, and PNG, JPEG and single-page TIFF uploads become one-page documents. OCR is bounded per document, per worker and per team per day (ocr_pages_per_day, where documents wait rather than fail), can be turned off per source, and documents skipped as scanned before can be retried together. A vision model's OCR is priced per token; Tesseract and Tika pages are counted, not priced. Runbook:operations/ocr.md. - SSO group mapping (E1). Admin → Group mapping maps an identity-provider group to a team role. At each sign-in, memberships the mapping created are added, raised, lowered or removed. Memberships made by hand, by invite or by Assign owner are never touched, and a rule never removes a team's last owner. Every rule shows a dry run before it's saved. Runbook:
operations/sso-groups.md. - Search in the command palette (E15). ⌘K finds agents, knowledge bases, sources and your conversations across all your teams (no longer the first ten), and, for platform staff, teams, users, models, connections, embedding profiles and shared sources, through the new
GET /v1/search. - Many small fixes. A deleted agent's conversations open read-only; a raised crawl limit wakes waiting crawls at once; revoked API keys open from audit links; one date rule across the app; an agent Settings tab; knowledge-base top-k inherited by agents until overridden; clearer feedback buttons; clickable table rows with one button each; toasts announced as status messages. See the changelog.
- Release and CI. Each release now attaches SPDX SBOMs, digest files and
checksums.txt. The authorization matrix runs as its own CI job, and a tag build fails unless the image reports exactly its tag: v0.1.0's binary reportedv0.1, and v0.2.0 reportsgrounded v0.2.0 (<commit>).
Requirements
Unchanged from v0.1.0: Kubernetes 1.30+, PostgreSQL 17 with pgvector 0.8+ (and the vector, citext, btree_gin and pg_trgm extensions), Valkey or Redis 7+, S3-compatible storage, an OpenAI-compatible gateway and an OIDC provider. New optional pieces:
| Optional | Notes |
|---|---|
| Tesseract OCR | The ocr-tesseract component runs ghcr.io/ncecere/grounded-ocr (Tesseract 5, common languages built in; derive an image for more). Or use Tika's -full image, or a vision model. |
| A groups claim | For SSO group mapping, your OIDC provider must send the user's groups in a claim (OIDC_GROUPS_CLAIM, default groups; Authentik sends it with the profile scope). |
Upgrading from v0.1.0
v0.2.0 is a rolling upgrade from v0.1.0 with no downtime. Take and validate a backup first, as for any upgrade (operations/upgrades.md).
- Bump the base and the image together. Change every
?ref=v0.1.0in your overlay to?ref=v0.2.0, and pin the new digest ofghcr.io/ncecere/grounded. If you still vendor the base, re-vendor atv0.2.0, or switch to the remote base now. The image is public, so theprivate-registrycomponent and its pull secret can go. - Migrations.
00030_search_indexesto00034_evaluationsonly add tables, columns, indexes and allowed values, so v0.1.0 pods keep working while the new ones start. They run in each new pod's init container, as before. - New settings, all optional:
OIDC_GROUPS_CLAIM(SSO group mapping),OCR_TESSERACT_URL,OCR_TIMEOUT,OCR_MAX_PAGES_PER_DOCUMENTandOCR_CONCURRENCY(OCR),EVALUATION_CONCURRENCY(evaluation runs), andRETENTION_EVALUATION_RUNS_DAYS..env.exampledocuments each. - OCR, if you want it. Add the
ocr-tesseractcomponent and pinghcr.io/ncecere/grounded-ocrby digest in your overlay, like the main image. It passes therestrictedPod Security Standard. Then turn OCR on in Admin → Parsing and press Test. - What changes for users straight away. Only evaluations: team editors see an Evaluations tab. Costs, OCR and group mapping stay off until an admin sets them up. Evaluation runs are deleted after 180 days by default (retention kind
evaluation_runs), unlike other kinds, which keep everything until configured.
Downgrading isn't supported. To go back, restore a backup taken before the upgrade.
Known limitations
v0.2.0 is pre-1.0; the v0.1.0 limitations about performance, capacity and availability still apply. New in this release:
- Budgets are checked with a short cache, so a team can overshoot its budget by about 30 seconds of use. Documents already being ingested when a budget runs out finish; only pending ones wait. Re-embedding during a profile migration isn't checked against budgets.
- Costs before v0.2.0. Usage that retention had already purged before the upgrade is only kept per UTC day, so budget months and report days before the upgrade are exact only to the UTC day.
- OCR reads only the first page of a multi-page TIFF (with a warning), and the per-document page cap is an environment variable, not an admin setting. Tesseract's accuracy depends on the scan's resolution; a vision model reads forms and tables better.
- Evaluations check an agent's retrieval and answers, not its tools beyond knowledge-base search. Questions come only from editors, imports and their own conversations: people can't yet share a failed question with the team.
- Not in v0.2.0 (roadmap candidates, not commitments): cross-encoder reranking (A1b, parked until a rerank model is available), citation marks per claim (A13), SCIM (E4), per-agent budgets, per-page prices for Tesseract and Tika, and multi-page TIFF.
Verifying the images
Images are published for linux/amd64 and linux/arm64 to ghcr.io/ncecere/grounded and, new in v0.2.0, ghcr.io/ncecere/grounded-ocr. The tags are v0.2.0, v0.2 and latest-release; release candidates get only their own tag, such as v0.2.0-rc.1. There is no latest tag. Tags can move, so deploy by digest.
Each image is built only in CI (.github/workflows/image.yml), scanned with Trivy, and signed with cosign using keyless signing tied to the workflow's GitHub identity. Look up the digest for a tag, then verify it (the same for grounded-ocr):
docker buildx imagetools inspect ghcr.io/ncecere/grounded:v0.2.0 # prints the index digest
cosign verify ghcr.io/ncecere/grounded@sha256:<digest> \
--certificate-identity-regexp '^https://github.com/ncecere/grounded/\.github/workflows/' \
--certificate-oidc-issuer https://token.actions.githubusercontent.comFor a stricter check, use --certificate-identity https://github.com/ncecere/grounded/.github/workflows/image.yml@refs/tags/v0.2.0 in place of the regular expression. docker run --rm ghcr.io/ncecere/grounded@sha256:<digest> version prints grounded v0.2.0 (<commit>).
Checksums, SBOM and provenance
-
The images are the release artifacts:
ghcr.io/ncecere/groundedand the optional OCR sidecarghcr.io/ncecere/grounded-ocr, built, scanned, signed and tagged the same way. No binary archives are published. Each image's digest is its checksum; the release lists both, ingrounded-v0.2.0.digest.txtandgrounded-ocr-v0.2.0.digest.txt. GitHub provides the tag's source archives. -
Release assets. New in v0.2.0: the GitHub release has the image's SPDX SBOM for each platform (
grounded-v0.2.0-linux-amd64.sbom.spdx.json,grounded-v0.2.0-linux-arm64.sbom.spdx.json), the digest files and achecksums.txtwith the SHA-256 of each. Check the downloads withsha256sum -c checksums.txt. -
SBOM and provenance attestations. The files are copies of the SPDX SBOM and SLSA provenance (
mode=max) attached to the image's index on ghcr.io as BuildKit attestations. The cosign signature covers the index, so the attestations are the signed source. To read them from the image:IMAGE=ghcr.io/ncecere/grounded@sha256:<digest> docker buildx imagetools inspect "$IMAGE" --format '{{ json (index .SBOM "linux/amd64").SPDX }}' > sbom.spdx.json docker buildx imagetools inspect "$IMAGE" --format '{{ json (index .Provenance "linux/amd64").SLSA }}' > provenance.json
-
Dependencies and scans.
security/dependencies.mdlists the direct Go and npm dependencies with their licences, and the latestgovulncheck,npm auditand Trivy results.
Reporting problems
Report bugs and requests as GitHub issues. Report vulnerabilities privately, as SECURITY.md describes.