Skip to content

v3.27.0 - Current Copilot models, reliable plan gates and GA stack presets

Choose a tag to compare

@srnichols srnichols released this 01 Oct 19:39
· 177 commits to master since this release

v3.27.0 - Current Copilot models, reliable plan gates and GA stack presets

Security

  • Refresh dependencies with advisories published since 3.26.7. The workspace lock had sharp 0.35.3, affected by the libheif advisory GHSA-rgj7-g3m4-5g8c (high). The standalone MCP lock that installs into projects already pinned 0.35.4. Both locks had fast-uri 3.1.7 and ip-address 10.7.0, which reach runtime through the MCP SDK (moderate), plus brace-expansion. The development lock also had Vitest 4.1.10. All are updated within their declared ranges, and both audits now report zero vulnerabilities.

Changed

  • Model defaults follow GitHub Copilot's 2026-09 lineup. Copilot retired Claude Sonnet 4.6 on 2026-09-01. It retires Claude Opus 4.7, Gemini 3.5/3.6 Flash and Kimi K2.7 Code on 2026-10-02, and GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini, Gemini 3.7 Flash and Grok 4.5 on 2026-10-19. Plan Forge defaulted to several of them. New defaults:

    • --quorum=power: claude-opus-5.5 + gpt-6-astra + grok-4.7, reviewer claude-opus-5.5.
    • --quorum=speed: claude-sonnet-5.5 + gpt-6-luna + gemini-3.8-flash, reviewer claude-sonnet-5.5. All three members are Copilot-served.
    • Default quorum: claude-opus-5.5 + gpt-6-sol + grok-4.7.
    • modelRouting.default when unset, and the value setup writes: claude-opus-5.5.
    • Default escalation chain: auto → claude-opus-5.5 → gpt-6-astra.
    • Watcher analyze model: claude-opus-5.5.
    • --with-grok / quorum.includeGrok add-in: grok-4.7.
    • gh-copilot worker default and the cost-estimate base model: claude-sonnet-5.5.

    Grok members still call the xAI API and need XAI_API_KEY. The quorum defaults, schema defaults and tool descriptions now come from orchestrator/constants.mjs, and a contract test fails if a default is unpriced, scheduled for retirement, or restated differently in setup or pforge doctor. MODEL_PRICING adds Claude Opus 5.5 / Sonnet 5.5 / Fable 5.1, GPT-6 Astra / 6.1 Sol / 6 Sol / 6 Luna, Gemini 3.8 / 3.7 Flash, Grok 4.7 / 4.6 and MAI-Code-1.1-Flash. Claude Sonnet 5 is now $2/$10; its launch price became permanent.

  • Stack presets target the latest GA releases.

    • Runtimes and images: .NET 10, Node 24 LTS / TypeScript 7, Python 3.14, Java 25 LTS / Spring Boot 4.1 / JUnit 6, Go 1.27, Swift 6.4.
    • Services: PostgreSQL 18, Redis 8, Dapr 1.18.
    • Tooling: GitHub Actions majors (checkout v7), Terraform 1.16 / AzureRM 5.7 / AzAPI 2.13, PowerShell 7.6, Pester 6.
    • Bicep: the newest stable API versions that Bicep 0.47's bundled types can validate. Newer versions raise BCP081, which fails deployments that set failOnStdErr.

    Code samples were updated for the API changes in those releases:

    • Express 5 async handlers, Prisma 7 config and adapters, and HTTPX ASGITransport.
    • Spring Boot 4 @MockitoBean, @AutoConfigureTestRestTemplate and the tools jarmode, plus Testcontainers 2.
    • xUnit v3 ValueTask lifecycles.
    • JDK 25 structured concurrency.
    • PostgreSQL 18's data directory.
    • Bicep samples that did not compile.
    • azure/setup-azd@latest, which never existed; it is now @v2.

    Fluent UI Blazor stays on 4.x until its samples are checked against v5, and Vapor stays on 4.x because Vapor 5 is still beta. The Swift iOS CI sample uses GitHub's xcode-27 runner image, which is still in preview, because no GA macOS image ships Xcode 27. The PHP and Rust presets, and eight Swift files, still contain Go content; the rewrite is tracked in #292.

Fixed

  • The testbed tools no longer default to one contributor's folder. On Windows, when testbed.path was unset, resolveTestbedPath returned E:\GitHub\plan-forge-testbed, a path from the maintainer's machine, instead of the documented error. The testbed tools now use, in order:

    1. a testbedPath argument;
    2. testbed.path from .forge.json, with relative paths resolved against the project root rather than the server's working directory;
    3. a plan-forge-testbed clone next to the project.

    If none of these exists, they return ERR_TESTBED_PATH_REQUIRED on every platform.

  • Shipped examples no longer name a specific project.

    • The forge_watch and forge_watch_live examples used E:/GitHub/Rummag; they now use /path/to/my-app.
    • Crucible's linked-bugs prompt and the forge_crucible_submit bugId example used another project's tracker prefix (RMG-0035); they now use BUG-0035.
    • The manual, EVENTS.md samples and the GitHub-status screenshot now use neutral paths.

    A guard test now fails if Plan Forge's own source contains a path into a contributor's checkout.

  • Forge-Master no longer fails every turn now that GitHub Models has been retired. GitHub shut down GitHub Models (models.github.ai) on 2026-07-30. The host now answers with a plain-text 200 OK, and the default githubCopilot provider turned that into a JSON SyntaxError on every turn. Forge-Master no longer auto-selects that provider. It tries ANTHROPIC_API_KEY, then OPENAI_API_KEY, then XAI_API_KEY, and defaults to claude-sonnet-5.5, gpt-6-sol or grok-4.7 respectively. An explicit reasoningProvider: "githubCopilot" now fails immediately with a message that names those keys.

  • Forge-Master's Anthropic provider reaches the API. It posted to /v1/v1/messages and sent Copilot-style dotted IDs such as claude-sonnet-4.6, which the Anthropic API rejects. It now posts to /v1/messages and sends the hyphenated ID (claude-sonnet-5-5). OpenAI GPT-6 requests that carry tools now set reasoning_effort: "none", which Chat Completions requires for function calling on GPT-6.

  • Cost estimates price Gemini, Kimi and MAI quorum legs as Copilot requests. These families have no direct-API route in Plan Forge and always run through gh-copilot. The estimator still treated them as an unknown provider and priced them per token.

  • The dashboard keeps a configured model that is missing from its dropdown. The Settings model list loaded a .forge.json value it didn't list (a retired or custom model) as no selection, and saving could overwrite it. The value is now added to the list as "(current)".

  • Secret scanning and staged-diff classification handle release-sized diffs (#288, #289, #290, #291). forge_secret_scan, its REST endpoint (POST /api/secret-scan/run), forge_diff_classify and the LiveGuard secret check read git diff through Node's default 1 MiB buffer. Large diffs failed with spawnSync git ENOBUFS, and forge_diff_classify treated that failure as an empty, clean diff (severity: "none", totalAdded: 0). All four now use a shared reader with a 64 MiB limit. A larger diff returns an error saying the changes were not scanned, and forge_diff_classify also reports any other git failure instead of a clean result.

  • pforge diff / forge_diff and pforge analyze read the scope contract correctly (#283, #286). In pforge.ps1, the Forbidden Actions and In Scope sections ran to the end of the plan, so every later slice scope was reported FORBIDDEN. -like also read a backticked [parallel-safe] tag as a character class that matched nearly every file. In pforge.sh, both sections were always empty, so nothing was ever forbidden. Both shells now stop each section at the next heading. They consider only single-word hints, so prose such as git push --force is ignored. Hints match literally with * as the only wildcard, and a bare word such as true matches only a whole path segment. Bash compiles each hint once, and analyze no longer aborts on plans without SHOULD criteria or on a clean working tree.

  • pforge analyze completes under Windows PowerShell 5.1 with bracketed paths (#284). Test files under route directories such as app/[id]/ made Get-Content -Raw fail with "A parameter cannot be found that matches parameter name 'Raw'". Files are now read with -LiteralPath. The gate lint inside analyze also runs on Windows now; its bare E:/… module import had been failing silently.

  • Scope lists survive a formatter's blank line (#282). A blank line between **Scope** (files in scope): and its bullets, which Prettier adds, left the slice with an empty scope. The list now parses, and the lock hash is the same with or without that line.

  • The lock hash covers every scope and gate the parser accepts (#285). computeLockHash now hashes exactly the lines that parseSlices turns into slice scopes and gates, instead of rescanning for the canonical labels only. Changing a gate or scope under **Validation Gate:**, **Scope (files in scope):**, **Files**: or a similar label now invalidates the hash.

  • Gate lint checks the slices that will actually run (#281). A slice that declares a Validation Gate but parses no runnable command, such as a gate inside a ```text fence, is now a lint error, and run-plan refuses to launch it. [manual] gates remain exempt. A slice with no numbered tasks produces a warning, because workers receive only numbered items.

  • Hook launchers no longer depend on file associations or open visible windows (#287). By default, each hook runs bash …sh. The windows key (VS Code) and powershell key (Copilot CLI/SDK) both run a hidden, non-interactive powershell -NoProfile … -File …ps1. Five PowerShell hooks stored their payload in the automatic $input variable, and the next git call emptied it. They now use $hookInput, which restores forbidden-path and pre-deploy denials, unfinished-code warnings, and Stop-hook re-entry protection. Neither forbidden-path hook had ever denied an edit, so both now follow the pforge diff rules and match repo-relative paths. They enforce the plan named in .forge/active-plan; otherwise the only plan In Progress; otherwise the only plan HARDENED or Ready for execution. When several plans qualify, they enforce nothing, so a stale plan's rules no longer block unrelated edits. Deny messages are valid JSON even when a hint contains a backslash.

  • A gate marker no longer arms the next slice's fences. After an inline gate (**Validation Gate**: \npm test`) or a prose-only gate, the parser carried its gate state across the next slice heading. The next slice's first shell fence, often an illustrative example inside its tasks, then ran as part of that slice's gate. Gate capture now stops at every slice and plan-level heading. The 96 plans in docs/plans` parse and hash unchanged.

  • The diff-classify PreCommit chain entry classifies the whole staged diff. Both wrappers passed the diff to Node through an environment variable. In PowerShell, the diff's line breaks were lost, so the hook allowed every commit on Windows. A diff over the platform's variable-size limit never reached Node, and only the first 3,000 lines were checked. A shared check-diff-classify.mjs now reads the staged diff itself through the bounded reader and classifies every line. If git cannot produce the diff, the commit is blocked.

  • pforge update from Bash now delivers hook scripts. pforge.sh listed only top-level files in templates/.github/hooks, so the scripts under hooks/scripts/, including every #287 fix, never reached Bash-updated projects. It now recurses, like pforge.ps1.

  • Bash pforge update no longer fails after replacing itself. pforge update and pforge self-update overwrite pforge.sh while it is running. Bash reads a script as it goes, so after a successful update it resumed reading the new file at a stale offset and stopped with syntax error near unexpected token ';;' and exit code 2. Anything that checked the exit code treated the update as failed. The command dispatch is now one block that bash reads completely before running anything. The fix applies from the next update on, because each update runs the updater the project already has.

  • Bash pforge analyze runs the gate lint. pforge.sh analyze reports and scores the gate lint the same way pforge.ps1 does.

Upgrade Notes

  • Node requirements (>=20.19.0) and the Copilot SDK routing default (routing.copilotSdk: "off") are unchanged from 3.26.7.
  • The lock hash now covers exactly the lines the parser turns into scopes and gates, so a hardened plan gets a new hash if any of those lines were previously unprotected. This includes non-canonical labels such as **Validation Gate:** or **Files**:, a blank line after the scope label, bold-prefixed bullets inside a scope list, and a gate fence that does not directly follow its marker. In this repository, 4 of 12 hardened plans changed hash, 2 of them using only canonical labels. If run-plan reports LOCK_HASH_MISMATCH for a plan that has not been edited, review its scope and gate lines, then re-stamp lockHash with Step 2. A plan keeps its hash if it uses the canonical labels, has a plain path list directly under the scope label, and has its gate fence directly under the gate marker.
  • run-plan now rejects a slice whose declared Validation Gate has no runnable command. Move gate commands into a shell-tagged or untagged fence, or mark checks that must be manual with [manual].
  • Bash users updating from 3.26.x: the 3.26.x updater prints syntax error near unexpected token ';;' and exits with code 2 after Update complete, even though the update finished. It also skips the hook scripts under .github/hooks/scripts/. Run pforge self-update --force once more. The 3.27.0 updater then delivers those scripts and exits cleanly. A plain pforge self-update reports Already current and does nothing. PowerShell updates complete in one run.
  • Git Bash projects that keep their own root VERSION file (#297): on Windows, pforge self-update can mistake that file for Plan Forge's version and report Already current or DOWNGRADE. Run pforge update --from-github --force instead (twice when coming from 3.26.x), or update from PowerShell with pwsh -NoProfile -File .\pforge.ps1 self-update --yes --force.
  • pforge update installs the new hook launchers and scripts. Projects that disabled hooks to work around #287 can re-enable them after updating. Confirm in the editor that hook scripts no longer open as tabs. The forbidden-path hook now actually denies edits, so if several hardened plans are queued, write .forge/active-plan to name the plan it should enforce.
  • Existing .forge.json files are not rewritten. If modelRouting, quorum.models, quorum.reviewerModel or escalationChain name a retired Copilot model (claude-opus-4.7, claude-sonnet-4.6, gpt-5.5, gpt-5.4, gpt-5.4-mini, gpt-5-mini, Gemini 3.5–3.7 Flash, kimi-k2.7-code), replace it. pforge doctor lists the configured quorum models.
  • The testbed tools no longer assume E:\GitHub\plan-forge-testbed on Windows. Set testbed.path in .forge.json, or keep the testbed cloned next to the project as ../plan-forge-testbed.
  • Forge-Master setups that relied on GITHUB_TOKEN or gh auth login alone need ANTHROPIC_API_KEY, OPENAI_API_KEY or XAI_API_KEY in the environment.
  • pforge update only adds preset files a project does not have yet. It never overwrites existing instruction, agent or prompt files, because they may be customized. To adopt the refreshed preset versions, merge them from presets/<stack>/.github/, or delete a file you have not customized and re-run pforge update.