Skip to content

Releases: Jonathanlight/context_os

v4.2.0

Choose a tag to compare

@github-actions github-actions released this 29 May 01:13
8b70f7e

v4.2.0 — First-time-user UX fixes

Three frictions reported on the v4.1 scaffolding flow are fixed.

✨ New / changed

Change Why
ctx init is now recursive (4 levels deep, --depth N or --no-recursive to opt out) v4.1 only scanned the root, so monorepos with frontend/ + api/ saw only the root manifest. Vendored dirs are skipped (node_modules, .venv, vendor, dist, build, target, ...). Each rationale line is prefixed with the relative path.
ctx eval-init <name> [--target rag|anthropic_skill] First-time users hit "Invalid value for SUITE_FILE" because the suite is a hand-written artefact. ctx eval-init scaffolds one with a realistic case so ctx eval ... --dry-run works immediately.
Friendlier ctx eval error Missing suite file → actionable hint pointing at ctx eval-init, not Typer's generic "File does not exist".
README rewrite (EN + FR) New Input / Output columns on the CLI table + 11 numbered worked examples covering create → init → compile → lint → audit → eval-init → eval → upgrade. Every command's --help now includes an Examples: block.

🧪 Quality

  • 1178 passed, 1 skipped (Anthropic key absent)
  • mypy --strict clean across 178 files
  • ruff check + ruff format clean
  • 11 new tests: recursive walk (depth cap, skip list, symlinks), ctx eval-init for both targets, friendly eval error

📦 Install / upgrade

pip install --upgrade context-os-ctx==4.2.0
# or, if already installed:
ctx upgrade

🔗 PRs

  • #82 — feat(ux): recursive ctx init + ctx eval-init + README rewrite
  • #83 — release: v4.2.0

Full changelog: CHANGELOG.md § 4.2.0

v4.1.0

Choose a tag to compare

@github-actions github-actions released this 29 May 00:16
b5962a5

v4.1.0 — Scaffolding release

Turns ContextOS from "a linter for existing .ctx files" into "the way you start a new project's .ctx".

✨ New commands

Command Purpose
ctx create <project> --lang python,fastapi,react,... Scaffold a starter .ctx from a project name + comma-separated language list. Baseline rules (TDD-001, SEC-001, DOC-001) + per-language rules + matching [stack]. Flags: --domain, --role, --title, --output, --force, --list-languages.
ctx init [path] [--dry-run] Walk an existing repo, read pyproject.toml / package.json / composer.json / go.mod / Cargo.toml / pom.xml / build.gradle / pubspec.yaml / mix.exs / Gemfile / *.csproj, and feed the detected stack to the same builder ctx create uses. --dry-run prints without writing.
ctx upgrade [--check] [--pre] Query the PyPI JSON API for the newest context-os-ctx and pip install --upgrade via sys.executable (so the upgrade hits the same interpreter ctx runs under).

📚 Language catalog (waves 1-3, 90+ entries)

  • Backend — Python (+ FastAPI, Django, Flask, Litestar, Starlette), PHP (+ Symfony, Laravel, Slim, Hyperf, Doctrine), TypeScript/Node (+ NestJS, Express, Fastify, Hono, Elysia, AdonisJS, Bun), Go (+ Gin, Echo, Fiber, Chi), Rust (+ Axum, Actix), Ruby (+ Rails, Sinatra), Elixir/Phoenix, .NET/ASP.NET Core, Java/Kotlin (+ Spring Boot, Ktor, Quarkus, Micronaut), Dart, Scala, Swift, Clojure, Haskell
  • Frontend — React, Next.js, Remix, Vue, Nuxt, Svelte, SvelteKit, Angular, SolidJS, Qwik, Astro, Preact, Lit
  • HTML-first — HTMX, Hotwire (Turbo+Stimulus), Livewire, Alpine.js
  • CSS/UI — Tailwind 4, shadcn/ui
  • Mobile — Flutter, React Native, Kotlin Multiplatform, Ionic, SwiftUI, Jetpack Compose
  • Desktop — Tauri 2, Electron, Compose Multiplatform
  • ML / Data — PyTorch, TensorFlow, JAX, scikit-learn, LangChain, LlamaIndex
  • DB / ORM — PostgreSQL, SQLAlchemy, Prisma, Drizzle, Doctrine
  • Systems — C, C++, Zig
  • Infra / DevOps — Docker, Kubernetes, Terraform, GitHub Actions
  • API — GraphQL, gRPC, OpenAPI 3.1

Each entry carries an opinionated 1-3 rule starter set, a [stack].required block, and a category + wave tag.

🪄 Slug aliases

Type Next.js, c#, ts, spring boot, nextjs, reactnative, tailwindcss, nest … they all normalise to canonical registry keys.

🧪 Quality

  • 1167 passed, 1 skipped in pytest (38 new tests for templates / builder / detector / CLI)
  • mypy --strict clean across 177 files
  • ruff check + ruff format clean

📦 Install / upgrade

pip install --upgrade context-os-ctx==4.1.0
# or, if a previous version is already installed:
ctx upgrade

🔗 PRs

  • #80 — feat(scaffold): ctx create + ctx init + ctx upgrade
  • #81 — release: v4.1.0

Full changelog: CHANGELOG.md § 4.1.0

v4.0.2

Choose a tag to compare

@github-actions github-actions released this 28 May 22:46
31b1393

Changelog

All notable changes to ContextOS are documented in this file.

The format is based on Keep a Changelog
and this project adheres to Semantic Versioning.

[Unreleased]

Nothing yet.

[4.0.2] — 2026-05-29

README-only re-publish to refresh the PyPI project page.

Fixed

  • README logo now resolves on PyPI. v4.0.1 was uploaded with the
    logo referenced as logo.png (relative path); PyPI's Camo proxy
    cached that as a dead URL because PyPI doesn't serve repo assets.
    v4.0.1's PyPI README was frozen at upload time so the absolute
    raw.githubusercontent.com URL that landed on develop after the
    publish couldn't reach PyPI without a fresh version bump. This
    release pushes the corrected README so the logo renders inline on
    the project page.
  • No code changes vs v4.0.1.

[4.0.1] — 2026-05-28

Hotfix release. The shipped feature set is identical to v4.0.0 — this
bump exists so the freshly-renamed distribution can land on PyPI
under context-os-ctx without colliding with the v4.0.0 git tag
(which was cut before the rename and would carry the wrong name:
in pyproject.toml).

Changed

  • PyPI distribution name renamed from context-os to
    context-os-ctx. The context-os name on PyPI was already
    reserved by an unrelated project; context-os-ctx is free and is
    now the distribution name shipped from this repository.
    • Import name is unchanged (import contextos).
    • CLI binary is unchanged (ctx).
    • Install commands across the README, getting-started guide,
      editor / eval docs, the lint-action default, and the local
      publish script all reference the new name.

Fixed

  • release.yml (PyPI publish): now uses a Detect publish mode
    step that prints which credential path will be attempted.
  • vscode-publish.yml: dropped environment: vscode (which
    required manual setup on the repo) and replaced
    if: secrets.VSCE_PAT != '' with the env-indirection pattern
    via steps.detect.outputs.should_publish, which is the canonical
    way to use secrets in step-level conditionals. Added
    workflow_dispatch for manual runs; non-v* ref names skip
    the version-sync step cleanly.
  • docs.yml: the GitHub Pages deploy steps now continue-on-error
    when Pages isn't enabled on the repo. The build still verifies
    on every push; the deploy is best-effort until the operator
    enables Pages in Settings → Pages → Source = GitHub Actions.
  • Repo-wide ruff format pass.

[4.0.0] — 2026-05-28

📦 Adoption & visualization. ContextOS ships to PyPI and the
VSCode Marketplace via configurable release workflows, surfaces
audit and eval results as self-contained HTML pages, and includes
ctx fix for auto-applying the four structured code-actions across
a repository.

Major version bump because the conceptual surface expands again —
from "lint + evaluate" to "lint + evaluate + auto-fix + ship." No
backwards-incompatible API changes; every v3.x consumer keeps
working unchanged.

Added

Phase 8 — Adoption & visualization (PRs #68–#73)

  • PyPI publish unblock (PR #68) — release.yml documents the
    two supported paths (API token vs trusted publishing), prints a
    workflow notice naming which path runs, falls through from one
    to the other when only one is configured. New
    scripts/publish-to-pypi.sh manual escape hatch.
  • VSCode Marketplace workflow (PR #69) — vscode-publish.yml
    triggered on the same v* tags; bumps package.json to match
    the tag, builds, runs vsce publish when VSCE_PAT secret is
    present, degrades to build-only otherwise.
  • HTML audit report (PR #70) — ctx audit --html renders a
    self-contained page with severity filter buttons, per-file
    accordion sections (sorted by path), cross-artifact + skipped
    blocks, summary footer. Hand-rolled with html.escape at every
    interpolation — no jinja2 dep — and inline CSS+JS so the output
    is one drop-in file. --json and --html are mutually
    exclusive.
  • HTML eval report (PR #71) — ctx eval --html renders cases
    as a filterable table with PASS / FAIL chips, pass-rate progress
    bar, token total badge, expected/actual columns, inline error
    notes for provider exceptions.
  • ctx fix command + structured fixes F001 / X001 / S005
    (PR #72) — new fix module with compute_fix(text, diag) → TextEdit | None dispatcher routing by diag.code. Four
    fixes ship today: X003 (strip trailing ?), F001 (sentence-
    case ALL CAPS title), X001 (strip TODO/FIXME markers at title
    start), S005 (prepend # <title> to SKILL.md body lacking H1).
    Dry-run by default; --apply writes the new content. Walks
    directories via the audit scanner so the set of fixed files
    matches what ctx audit would lint.
  • Docs (this PR) — docs/dashboard.md covering both HTML
    reports + the ctx fix workflow with safety properties and a
    pre-commit hook example. docs/release.md covers the PyPI +
    VSCode Marketplace setup walkthroughs.

New CLI subcommands

Command What it does
ctx fix <target> Auto-apply structured fixes; dry-run by default, --apply writes

Extended CLI

Command New flag Purpose
ctx audit --html, --output Self-contained HTML report; --json and --html mutually exclusive
ctx eval --html Self-contained HTML report; same mutual exclusion

New workflows

  • .github/workflows/vscode-publish.yml — Marketplace publication
    on v* tag push.

Documented limitations

  • The LSP code_actions.py module still has its own X003
    implementation. A subsequent PR will unify both paths on
    contextos.fix.structured. Documented in the fix package
    docstring.
  • PyPI publishing still requires a manual one-time setup step
    (either create the PYPI_API_TOKEN secret OR configure a
    Trusted Publisher on pypi.org). Documented in docs/release.md.
  • VSCode Marketplace publishing requires a Marketplace publisher
    account + a Personal Access Token in the VSCE_PAT secret.
    Documented in docs/release.md.
  • No multi-edit fixes today — each diagnostic gets at most one
    TextEdit. Multi-step refactors (e.g. moving a URL out of a
    title into links) need additional plumbing.

[3.0.0] — 2026-05-28

🎯 Functional evaluation. ContextOS no longer only validates
structure (does the skill have trigger phrasing, does the RAG
config have a freshness policy) — it now validates behavior. Run
your skills against real Anthropic models and your RAG corpus
against real OpenAI embeddings, score the results, gate CI on
regressions.

Major version bump because the conceptual surface widens: the same
toolchain you use to lint a CLAUDE.md now scores whether your
skills actually fire. No backwards-incompatible API changes — every
v2.x consumer keeps working.

Added

Phase 7B — Live evaluation (PRs #62–#67)

  • .eval.toml format + AST (PR #62) — EvalSuite with target
    literal (anthropic_skill | rag), SkillCase (prompt +
    expected_skill + tags), RagCase (query + expected_sources +
    top_k bounded 1–100). Cross-target model validator rejects
    mismatched case lists. Parser reuses ContextOSParseError so
    eval-suite mistakes carry file:line:column + suggestion.
  • Anthropic Skills evaluator (PR #63) — SkillRoutingProvider
    Protocol + MockSkillProvider (deterministic for tests) +
    AnthropicSkillProvider (lazy SDK import, Haiku 4.5 default
    model). Skills routing emulated via the Messages API tool-use
    feature: each SkillDocument → tool definition with description
    = trigger signal; tool-use block's name = picked skill slug.
    SkillEvalRunner captures per-case errors so a flaky provider
    doesn't waste the whole run.
  • RAG retrieval evaluator (PR #64) — RagRetrievalProvider
    Protocol + MockRagProvider + EmbeddingRagProvider with
    eager-stacked, row-normalized cosine matrix (each retrieve()
    is one matmul). User supplies the embedding callable
    (EmbedQueryFn = Callable[[str], list[float]]) and pre-indexed
    Chunk list. ContextOS does not ship an embedding service
    or an indexer. Pass criterion = OR semantics on
    expected_sources.
  • ctx eval CLI (PR #65) — dispatches on suite.target,
    --dry-run uses Mock providers (every case passes by
    construction), --json / --output for CI consumption,
    --skills-dir walks SKILL.md recursively (sorted-path order for
    deterministic tool ordering), --rag-chunks loads pre-indexed
    chunks from JSON, --rag-embed-model selects the OpenAI
    embedding model. All eval-side imports lazy so [eval] extras
    don't slow down ctx --version.
  • ctx eval-diff for regression detection (PR #66) — compare
    two eval result JSON files; classify case transitions into five
    buckets (regression / improvement / new_failure / new_pass /
    removed). Default CI gate: exit 1 on regression only;
    --fail-on-new-failure flag opt-in for stricter gating. Sticky
    comment shape compatible with PR review workflows.
  • Docs page (this PR) — docs/eval.md covering architecture,
    .eval.toml grammar, dry-run quick start, live Skills + RAG
    setup, BYO embedding service via Python API, CI workflow with
    eval-diff, troubleshooting table.

New CLI subcommands

Command What it does
ctx eval <suite.eval.toml> Run an eval suite; --dry-run for mock-driven smoke
ctx eval-diff <baseline> <current> Compare two eval JSON outputs; exit 1 on regression

New optional-dependencies group

  • [eval] = ["anthropic>=0.40", "numpy>=1.26", "openai>=1.0"].

Documented limitations

  • Skills routing is emulated via the Messages API tool-use
    feature, not the actual Skills product (which lives in Claude app
    / Claude Code, not the public API). Tool selection is the closest
    approximation; the gap is documented in
    `anthropic_prov...
Read more

v4.0.1

Choose a tag to compare

@github-actions github-actions released this 28 May 21:32
0bfbcbd

Changelog

All notable changes to ContextOS are documented in this file.

The format is based on Keep a Changelog
and this project adheres to Semantic Versioning.

[Unreleased]

Nothing yet.

[4.0.1] — 2026-05-28

Hotfix release. The shipped feature set is identical to v4.0.0 — this
bump exists so the freshly-renamed distribution can land on PyPI
under context-os-ctx without colliding with the v4.0.0 git tag
(which was cut before the rename and would carry the wrong name:
in pyproject.toml).

Changed

  • PyPI distribution name renamed from context-os to
    context-os-ctx. The context-os name on PyPI was already
    reserved by an unrelated project; context-os-ctx is free and is
    now the distribution name shipped from this repository.
    • Import name is unchanged (import contextos).
    • CLI binary is unchanged (ctx).
    • Install commands across the README, getting-started guide,
      editor / eval docs, the lint-action default, and the local
      publish script all reference the new name.

Fixed

  • release.yml (PyPI publish): now uses a Detect publish mode
    step that prints which credential path will be attempted.
  • vscode-publish.yml: dropped environment: vscode (which
    required manual setup on the repo) and replaced
    if: secrets.VSCE_PAT != '' with the env-indirection pattern
    via steps.detect.outputs.should_publish, which is the canonical
    way to use secrets in step-level conditionals. Added
    workflow_dispatch for manual runs; non-v* ref names skip
    the version-sync step cleanly.
  • docs.yml: the GitHub Pages deploy steps now continue-on-error
    when Pages isn't enabled on the repo. The build still verifies
    on every push; the deploy is best-effort until the operator
    enables Pages in Settings → Pages → Source = GitHub Actions.
  • Repo-wide ruff format pass.

[4.0.0] — 2026-05-28

📦 Adoption & visualization. ContextOS ships to PyPI and the
VSCode Marketplace via configurable release workflows, surfaces
audit and eval results as self-contained HTML pages, and includes
ctx fix for auto-applying the four structured code-actions across
a repository.

Major version bump because the conceptual surface expands again —
from "lint + evaluate" to "lint + evaluate + auto-fix + ship." No
backwards-incompatible API changes; every v3.x consumer keeps
working unchanged.

Added

Phase 8 — Adoption & visualization (PRs #68–#73)

  • PyPI publish unblock (PR #68) — release.yml documents the
    two supported paths (API token vs trusted publishing), prints a
    workflow notice naming which path runs, falls through from one
    to the other when only one is configured. New
    scripts/publish-to-pypi.sh manual escape hatch.
  • VSCode Marketplace workflow (PR #69) — vscode-publish.yml
    triggered on the same v* tags; bumps package.json to match
    the tag, builds, runs vsce publish when VSCE_PAT secret is
    present, degrades to build-only otherwise.
  • HTML audit report (PR #70) — ctx audit --html renders a
    self-contained page with severity filter buttons, per-file
    accordion sections (sorted by path), cross-artifact + skipped
    blocks, summary footer. Hand-rolled with html.escape at every
    interpolation — no jinja2 dep — and inline CSS+JS so the output
    is one drop-in file. --json and --html are mutually
    exclusive.
  • HTML eval report (PR #71) — ctx eval --html renders cases
    as a filterable table with PASS / FAIL chips, pass-rate progress
    bar, token total badge, expected/actual columns, inline error
    notes for provider exceptions.
  • ctx fix command + structured fixes F001 / X001 / S005
    (PR #72) — new fix module with compute_fix(text, diag) → TextEdit | None dispatcher routing by diag.code. Four
    fixes ship today: X003 (strip trailing ?), F001 (sentence-
    case ALL CAPS title), X001 (strip TODO/FIXME markers at title
    start), S005 (prepend # <title> to SKILL.md body lacking H1).
    Dry-run by default; --apply writes the new content. Walks
    directories via the audit scanner so the set of fixed files
    matches what ctx audit would lint.
  • Docs (this PR) — docs/dashboard.md covering both HTML
    reports + the ctx fix workflow with safety properties and a
    pre-commit hook example. docs/release.md covers the PyPI +
    VSCode Marketplace setup walkthroughs.

New CLI subcommands

Command What it does
ctx fix <target> Auto-apply structured fixes; dry-run by default, --apply writes

Extended CLI

Command New flag Purpose
ctx audit --html, --output Self-contained HTML report; --json and --html mutually exclusive
ctx eval --html Self-contained HTML report; same mutual exclusion

New workflows

  • .github/workflows/vscode-publish.yml — Marketplace publication
    on v* tag push.

Documented limitations

  • The LSP code_actions.py module still has its own X003
    implementation. A subsequent PR will unify both paths on
    contextos.fix.structured. Documented in the fix package
    docstring.
  • PyPI publishing still requires a manual one-time setup step
    (either create the PYPI_API_TOKEN secret OR configure a
    Trusted Publisher on pypi.org). Documented in docs/release.md.
  • VSCode Marketplace publishing requires a Marketplace publisher
    account + a Personal Access Token in the VSCE_PAT secret.
    Documented in docs/release.md.
  • No multi-edit fixes today — each diagnostic gets at most one
    TextEdit. Multi-step refactors (e.g. moving a URL out of a
    title into links) need additional plumbing.

[3.0.0] — 2026-05-28

🎯 Functional evaluation. ContextOS no longer only validates
structure (does the skill have trigger phrasing, does the RAG
config have a freshness policy) — it now validates behavior. Run
your skills against real Anthropic models and your RAG corpus
against real OpenAI embeddings, score the results, gate CI on
regressions.

Major version bump because the conceptual surface widens: the same
toolchain you use to lint a CLAUDE.md now scores whether your
skills actually fire. No backwards-incompatible API changes — every
v2.x consumer keeps working.

Added

Phase 7B — Live evaluation (PRs #62–#67)

  • .eval.toml format + AST (PR #62) — EvalSuite with target
    literal (anthropic_skill | rag), SkillCase (prompt +
    expected_skill + tags), RagCase (query + expected_sources +
    top_k bounded 1–100). Cross-target model validator rejects
    mismatched case lists. Parser reuses ContextOSParseError so
    eval-suite mistakes carry file:line:column + suggestion.
  • Anthropic Skills evaluator (PR #63) — SkillRoutingProvider
    Protocol + MockSkillProvider (deterministic for tests) +
    AnthropicSkillProvider (lazy SDK import, Haiku 4.5 default
    model). Skills routing emulated via the Messages API tool-use
    feature: each SkillDocument → tool definition with description
    = trigger signal; tool-use block's name = picked skill slug.
    SkillEvalRunner captures per-case errors so a flaky provider
    doesn't waste the whole run.
  • RAG retrieval evaluator (PR #64) — RagRetrievalProvider
    Protocol + MockRagProvider + EmbeddingRagProvider with
    eager-stacked, row-normalized cosine matrix (each retrieve()
    is one matmul). User supplies the embedding callable
    (EmbedQueryFn = Callable[[str], list[float]]) and pre-indexed
    Chunk list. ContextOS does not ship an embedding service
    or an indexer. Pass criterion = OR semantics on
    expected_sources.
  • ctx eval CLI (PR #65) — dispatches on suite.target,
    --dry-run uses Mock providers (every case passes by
    construction), --json / --output for CI consumption,
    --skills-dir walks SKILL.md recursively (sorted-path order for
    deterministic tool ordering), --rag-chunks loads pre-indexed
    chunks from JSON, --rag-embed-model selects the OpenAI
    embedding model. All eval-side imports lazy so [eval] extras
    don't slow down ctx --version.
  • ctx eval-diff for regression detection (PR #66) — compare
    two eval result JSON files; classify case transitions into five
    buckets (regression / improvement / new_failure / new_pass /
    removed). Default CI gate: exit 1 on regression only;
    --fail-on-new-failure flag opt-in for stricter gating. Sticky
    comment shape compatible with PR review workflows.
  • Docs page (this PR) — docs/eval.md covering architecture,
    .eval.toml grammar, dry-run quick start, live Skills + RAG
    setup, BYO embedding service via Python API, CI workflow with
    eval-diff, troubleshooting table.

New CLI subcommands

Command What it does
ctx eval <suite.eval.toml> Run an eval suite; --dry-run for mock-driven smoke
ctx eval-diff <baseline> <current> Compare two eval JSON outputs; exit 1 on regression

New optional-dependencies group

  • [eval] = ["anthropic>=0.40", "numpy>=1.26", "openai>=1.0"].

Documented limitations

  • Skills routing is emulated via the Messages API tool-use
    feature, not the actual Skills product (which lives in Claude app
    / Claude Code, not the public API). Tool selection is the closest
    approximation; the gap is documented in
    anthropic_provider.py:9.
  • OpenAI is the only built-in embedding provider in the CLI.
    Voyage / Cohere / local-model users drop down to the Python API
    (one EmbeddingRagProvider(chunks, embed_query) call). Documented
    in docs/eval.md.
  • No indexer ships with ContextOS — users feed pre-indexed
    chunks via --rag-chunks chunks.json. The chunks file is
    whatever the user's pipeline produces; we validate the shape and
    cosine over what's there.
  • No HTML eval report; the JSON renderer is the bridge today.

[2.1.0] — 2026-05-28

🛠️ Editor integration. ContextOS now ships as a language server
(`...

Read more