Skip to content

Releases: raandree/CopilotAtelier

v6.1.0-preview0003

v6.1.0-preview0003 Pre-release
Pre-release

Choose a tag to compare

@raandree raandree released this 05 Oct 13:58
9b6a341

[v6.1.0-preview0003]

Added

  • Add contributor calibration: every agent pitches answers and questions at your familiarity with each knowledge area (new, familiar, or expert; familiar until you say otherwise), and in new and familiar areas illustrates an abstract finding with a concrete example; every calculated or reconstructed result names its sources and method. Every technical question offers a recommended answer and a not sure, you pick option; a delegated answer becomes an assumption flagged for expert review, counts only when you write it yourself, and never authorizes an irreversible, destructive, or security-relevant action; a question that authorizes such an action comes without the option. A level changes wording and depth, never warnings, tests, or reviews, and is never written to a repository. See How Much Explanation You Want.
  • Add docs/SECURITY-REVIEW.md, the independent security review of contributor calibration with every finding, its severity, and its resolution.
  • Add the /simpler and /deeper Prompts, which re-explain the last answer one familiarity level simpler or deeper and keep that level for its knowledge area for the rest of the session.

Changed

  • Let grill-me, software-architect, and gilb-requirements-engineering accept not sure, you pick: the recommended answer, or for a numeric target a level derived from Past, Record, or a verified benchmark, is recorded as an assumption flagged for expert review instead of blocking the interview.
  • Document in agent-evals that the Copilot backend's content filter blocks some harmless eval prompts, why an embedded chat transcript makes it worse, and how to keep a comparison fair: situation as system context, FinishReason logged per call, retries, and arms compared only on cases complete in both. See harness prerequisites.
  • Document in agent-evals that compared arms must run at the same time, because backend behavior drifts within hours, and that an LLM judge's disagreement can expose a mislabelled reply, which justifies a label correction only for an objective error. See micro-tests.

Fixed

  • Deliver the SessionStart context to Copilot SDK (agent host) and Copilot CLI sessions. Add-SessionContext wrote it only under hookSpecificOutput, which those hosts ignore; it now also writes the top-level additionalContext they read.
  • Keep the session clock when a Copilot SDK chat resumes. The runtime reruns the SessionStart hook with source: resume when it reloads a chat, and Add-SessionContext restarted the clock there, so the closing elapsed line and the turn count started over and the injected start time moved; it now keeps a readable clock of the same session.
  • Keep the long-running-job-monitor heartbeat state readable while it is rewritten. Start-JobHeartbeat.ps1 wrote the state file in place, so -Stop, -TouchStatus, or a wake read at the same moment could find it empty or cut off and fail; it now renames a complete file over it, also on Windows PowerShell 5.1.
  • Stop the wording of a commit message from choosing the release version. GitVersion.yml raised it on words such as "major", "breaking", or "add" anywhere in a message, which is how review text made 6.0.0 a major release; now only the Conventional Commit type in the subject line (feat:, fix:, perf:, !), a line that starts with BREAKING CHANGE:, or a literal +semver: override raises it above the branch's default increment.
  • Send an authorized push to the user's own terminal. The PreToolUse block message, AGENTS.md, and the hooks README told the agent to set COPILOT_ATELIER_ALLOW_REMOTE=1 for the command, but each host starts the hook with its own environment, so a variable set in an agent terminal never reaches the guard. They now also warn never to persist the variable, for example with setx: every host started afterwards would run without the guard.
  • Show the PreToolUse guard's reason to the model in Copilot SDK chats on Windows. That host ran the cross-platform command launcher inside an outer PowerShell that reported the block as 1, so the model saw only "hook errored". A new powershell launcher, which the SDK host prefers on Windows, passes exit 2 on, and the guard prints the reason as one JSON object on standard output, which that host merges into the deny on exit 2. The object now carries only the top-level permissionDecision and permissionDecisionReason: the SDK runtime in VS Code 1.140.0 drops a PreToolUse object that also carries hookSpecificOutput, so the model read only "hook exited with code 2". VS Code still reads the reason from standard error.

Security

  • Block a push in VS Code Local chats on Windows. VS Code runs a hook's windows launcher as the -Command text of an outer Windows PowerShell, which reported the guard's exit code 2 as 1, and VS Code treats every exit other than 2 as a warning, so the PreToolUse guard only warned there. The launcher now ends with a statement that passes the inner exit code on. Found by the post-release review of the v6.0.0 hook launcher fix.
  • Look for the hook scripts under USERPROFILE before HOME. Since v6.0.0 the launchers tried HOME first, so a HOME that a tool such as Git for Windows pointed at a writable tree could run a planted script instead of the deployed guard, and a HOME on an unreachable network share delayed the guard past its 20-second timeout, which the Copilot SDK host treats as allow.
  • Scan the raw payload text whenever the PreToolUse guard cannot walk the payload field by field. A payload that was not valid JSON has been allowed since v6.0.0, and the walk stopped four levels deep, so a push nested more than four levels, or more than the 100 levels Windows PowerShell parses (1,024 in PowerShell 7), went through. The walk now reaches 64 levels, and whatever it cannot reach is scanned as raw text, with JSON escapes decoded and without JSON punctuation, so an argument array such as ["git","push"] is caught as well, before the call is allowed. The raw text scan also joins the command-bearing fields the way the walk does, so a command split across fields, such as {"command":"git","args":["push"]}, is caught there too.
  • Block a push whose arguments come before the command, such as {"args":["push"],"command":"git"}, the order a serializer that sorts its keys writes. Since v6.0.0 the PreToolUse guard joined the command-bearing fields only in the order they appear, which read push git; it now also joins each object's executable fields before its argument fields, and for a payload it cannot walk, the whole payload's fields in that order and in reverse.
  • Block a tool call the PreToolUse guard cannot inspect within five seconds. Some of its patterns slow down quadratically on one long line of git words: 64 KB of them took 21.5 seconds, past the 20-second hook timeout, which the Copilot SDK host treats as allow. Parsing and walking a payload cannot be interrupted either, and 300,000 small objects kept the guard busy past the timeout too. The whole decision now has a time limit: a payload is parsed only up to 1 MB and walked only up to 20,000 fields and nested objects, the rest is scanned as raw text, a payload over 4 MB is blocked unscanned, and a payload not inspected within the limit is blocked. Ordinary commands still take well under a second.

v6.1.0-preview0002

v6.1.0-preview0002 Pre-release
Pre-release

Choose a tag to compare

@raandree raandree released this 04 Oct 20:45
25d8233

[v6.1.0-preview0002]

Added

  • Add contributor calibration: every agent pitches answers and questions at your familiarity with each knowledge area (new, familiar, or expert; familiar until you say otherwise), and in new and familiar areas illustrates an abstract finding with a concrete example; every calculated or reconstructed result names its sources and method. Every technical question offers a recommended answer and a not sure, you pick option; a delegated answer becomes an assumption flagged for expert review, counts only when you write it yourself, and never authorizes an irreversible, destructive, or security-relevant action; a question that authorizes such an action comes without the option. A level changes wording and depth, never warnings, tests, or reviews, and is never written to a repository. See How Much Explanation You Want.
  • Add docs/SECURITY-REVIEW.md, the independent security review of contributor calibration with every finding, its severity, and its resolution.
  • Add the /simpler and /deeper Prompts, which re-explain the last answer one familiarity level simpler or deeper and keep that level for its knowledge area for the rest of the session.

Changed

  • Let grill-me, software-architect, and gilb-requirements-engineering accept not sure, you pick: the recommended answer, or for a numeric target a level derived from Past, Record, or a verified benchmark, is recorded as an assumption flagged for expert review instead of blocking the interview.
  • Document in agent-evals that the Copilot backend's content filter blocks some harmless eval prompts, why an embedded chat transcript makes it worse, and how to keep a comparison fair: situation as system context, FinishReason logged per call, retries, and arms compared only on cases complete in both. See harness prerequisites.
  • Document in agent-evals that compared arms must run at the same time, because backend behavior drifts within hours, and that an LLM judge's disagreement can expose a mislabelled reply, which justifies a label correction only for an objective error. See micro-tests.

Fixed

  • Deliver the SessionStart context to Copilot SDK (agent host) and Copilot CLI sessions. Add-SessionContext wrote it only under hookSpecificOutput, which those hosts ignore; it now also writes the top-level additionalContext they read.
  • Keep the session clock when a Copilot SDK chat resumes. The runtime reruns the SessionStart hook with source: resume when it reloads a chat, and Add-SessionContext restarted the clock there, so the closing elapsed line and the turn count started over and the injected start time moved; it now keeps a readable clock of the same session.
  • Keep the long-running-job-monitor heartbeat state readable while it is rewritten. Start-JobHeartbeat.ps1 wrote the state file in place, so -Stop, -TouchStatus, or a wake read at the same moment could find it empty or cut off and fail; it now renames a complete file over it, also on Windows PowerShell 5.1.
  • Stop the wording of a commit message from choosing the release version. GitVersion.yml raised it on words such as "major", "breaking", or "add" anywhere in a message, which is how review text made 6.0.0 a major release; now only the Conventional Commit type in the subject line (feat:, fix:, perf:, !), a line that starts with BREAKING CHANGE:, or a literal +semver: override raises it above the branch's default increment.
  • Send an authorized push to the user's own terminal. The PreToolUse block message, AGENTS.md, and the hooks README told the agent to set COPILOT_ATELIER_ALLOW_REMOTE=1 for the command, but each host starts the hook with its own environment, so a variable set in an agent terminal never reaches the guard. They now also warn never to persist the variable, for example with setx: every host started afterwards would run without the guard.
  • Pass the PreToolUse guard's exit code 2 on in Copilot SDK chats on Windows. That host ran the cross-platform command launcher inside an outer PowerShell that reported the block as 1, so the model saw only "hook errored". A new powershell launcher, which the SDK host prefers on Windows, passes exit 2 on, and the guard also prints the reason as one JSON object on standard output, which the GitHub hooks reference says that host reads on exit 2. The SDK runtime in VS Code 1.140.0 does not: the call is denied, but the model reads only "hook exited with code 2".

Security

  • Block a push in VS Code Local chats on Windows. VS Code runs a hook's windows launcher as the -Command text of an outer Windows PowerShell, which reported the guard's exit code 2 as 1, and VS Code treats every exit other than 2 as a warning, so the PreToolUse guard only warned there. The launcher now ends with a statement that passes the inner exit code on. Found by the post-release review of the v6.0.0 hook launcher fix.
  • Look for the hook scripts under USERPROFILE before HOME. Since v6.0.0 the launchers tried HOME first, so a HOME that a tool such as Git for Windows pointed at a writable tree could run a planted script instead of the deployed guard, and a HOME on an unreachable network share delayed the guard past its 20-second timeout, which the Copilot SDK host treats as allow.
  • Scan the raw payload text whenever the PreToolUse guard cannot walk the payload field by field. A payload that was not valid JSON has been allowed since v6.0.0, and the walk stopped four levels deep, so a push nested more than four levels, or more than the 100 levels Windows PowerShell parses (1,024 in PowerShell 7), went through. The walk now reaches 64 levels, and whatever it cannot reach is scanned as raw text, with JSON escapes decoded and without JSON punctuation, so an argument array such as ["git","push"] is caught as well, before the call is allowed. The raw text scan also joins the command-bearing fields the way the walk does, so a command split across fields, such as {"command":"git","args":["push"]}, is caught there too.
  • Block a push whose arguments come before the command, such as {"args":["push"],"command":"git"}, the order a serializer that sorts its keys writes. Since v6.0.0 the PreToolUse guard joined the command-bearing fields only in the order they appear, which read push git; it now also joins each object's executable fields before its argument fields, and for a payload it cannot walk, the whole payload's fields in that order and in reverse.
  • Block a tool call the PreToolUse guard cannot inspect within five seconds. Some of its patterns slow down quadratically on one long line of git words: 64 KB of them took 21.5 seconds, past the 20-second hook timeout, which the Copilot SDK host treats as allow. Parsing and walking a payload cannot be interrupted either, and 300,000 small objects kept the guard busy past the timeout too. The whole decision now has a time limit: a payload is parsed only up to 1 MB and walked only up to 20,000 fields and nested objects, the rest is scanned as raw text, a payload over 4 MB is blocked unscanned, and a payload not inspected within the limit is blocked. Ordinary commands still take well under a second.

v6.1.0-preview0001

v6.1.0-preview0001 Pre-release
Pre-release

Choose a tag to compare

@raandree raandree released this 02 Oct 20:25
7a48abe

[v6.1.0-preview0001]

Added

  • Add contributor calibration: every agent pitches answers and questions at your familiarity with each knowledge area (new, familiar, or expert; familiar until you say otherwise), and in new and familiar areas illustrates an abstract finding with a concrete example; every calculated or reconstructed result names its sources and method. Every technical question offers a recommended answer and a not sure, you pick option; a delegated answer becomes an assumption flagged for expert review, counts only when you write it yourself, and never authorizes an irreversible, destructive, or security-relevant action; a question that authorizes such an action comes without the option. A level changes wording and depth, never warnings, tests, or reviews, and is never written to a repository. See How Much Explanation You Want.
  • Add docs/SECURITY-REVIEW.md, the independent security review of contributor calibration with every finding, its severity, and its resolution.
  • Add the /simpler and /deeper Prompts, which re-explain the last answer one familiarity level simpler or deeper and keep that level for its knowledge area for the rest of the session.

Changed

  • Let grill-me, software-architect, and gilb-requirements-engineering accept not sure, you pick: the recommended answer, or for a numeric target a level derived from Past, Record, or a verified benchmark, is recorded as an assumption flagged for expert review instead of blocking the interview.
  • Document in agent-evals that the Copilot backend's content filter blocks some harmless eval prompts, why an embedded chat transcript makes it worse, and how to keep a comparison fair: situation as system context, FinishReason logged per call, retries, and arms compared only on cases complete in both. See harness prerequisites.
  • Document in agent-evals that compared arms must run at the same time, because backend behavior drifts within hours, and that an LLM judge's disagreement can expose a mislabelled reply, which justifies a label correction only for an objective error. See micro-tests.

Fixed

  • Deliver the SessionStart context to Copilot SDK (agent host) and Copilot CLI sessions. Add-SessionContext wrote it only under hookSpecificOutput, which those hosts ignore; it now also writes the top-level additionalContext they read.
  • Keep the long-running-job-monitor heartbeat state readable while it is rewritten. Start-JobHeartbeat.ps1 wrote the state file in place, so -Stop, -TouchStatus, or a wake read at the same moment could find it empty or cut off and fail; it now renames a complete file over it, also on Windows PowerShell 5.1.
  • Stop the wording of a commit message from choosing the release version. GitVersion.yml raised it on words such as "major", "breaking", or "add" anywhere in a message, which is how review text made 6.0.0 a major release; now only the Conventional Commit type in the subject line (feat:, fix:, perf:, !), a line that starts with BREAKING CHANGE:, or a literal +semver: override raises it above the branch's default increment.
  • Send an authorized push to the user's own terminal. The PreToolUse block message, AGENTS.md, and the hooks README told the agent to set COPILOT_ATELIER_ALLOW_REMOTE=1 for the command, but each host starts the hook with its own environment, so a variable set in an agent terminal never reaches the guard. They now also warn never to persist the variable, for example with setx: every host started afterwards would run without the guard.
  • Show the reason when the PreToolUse guard blocks a call in a Copilot SDK chat on Windows. That host ran the cross-platform command launcher inside an outer PowerShell that reported the block as 1, so the model saw only "hook errored". A new powershell launcher, which the SDK host prefers on Windows, passes exit 2 on, and the guard also prints the reason as one JSON object on standard output, which that host reads on exit 2.

Security

  • Block a push in VS Code Local chats on Windows. VS Code runs a hook's windows launcher as the -Command text of an outer Windows PowerShell, which reported the guard's exit code 2 as 1, and VS Code treats every exit other than 2 as a warning, so the PreToolUse guard only warned there. The launcher now ends with a statement that passes the inner exit code on. Found by the post-release review of the v6.0.0 hook launcher fix.
  • Look for the hook scripts under USERPROFILE before HOME. Since v6.0.0 the launchers tried HOME first, so a HOME that a tool such as Git for Windows pointed at a writable tree could run a planted script instead of the deployed guard, and a HOME on an unreachable network share delayed the guard past its 20-second timeout, which the Copilot SDK host treats as allow.
  • Scan the raw payload text whenever the PreToolUse guard cannot walk the payload field by field. A payload that was not valid JSON has been allowed since v6.0.0, and the walk stopped four levels deep, so a push nested more than four levels, or more than the 100 levels Windows PowerShell parses (1,024 in PowerShell 7), went through. The walk now reaches 64 levels, and whatever it cannot reach is scanned as raw text, with JSON escapes decoded and without JSON punctuation, so an argument array such as ["git","push"] is caught as well, before the call is allowed. The raw text scan also joins the command-bearing fields the way the walk does, so a command split across fields, such as {"command":"git","args":["push"]}, is caught there too.
  • Block a push whose arguments come before the command, such as {"args":["push"],"command":"git"}, the order a serializer that sorts its keys writes. Since v6.0.0 the PreToolUse guard joined the command-bearing fields only in the order they appear, which read push git; it now also joins each object's executable fields before its argument fields, and for a payload it cannot walk, the whole payload's fields in that order and in reverse.
  • Block a tool call the PreToolUse guard cannot inspect within five seconds. Some of its patterns slow down quadratically on one long line of git words: 64 KB of them took 21.5 seconds, past the 20-second hook timeout, which the Copilot SDK host treats as allow. Parsing and walking a payload cannot be interrupted either, and 300,000 small objects kept the guard busy past the timeout too. The whole decision now has a time limit: a payload is parsed only up to 1 MB and walked only up to 20,000 fields and nested objects, the rest is scanned as raw text, a payload over 4 MB is blocked unscanned, and a payload not inspected within the limit is blocked. Ordinary commands still take well under a second.

v6.0.1-preview0001

v6.0.1-preview0001 Pre-release
Pre-release

Choose a tag to compare

@raandree raandree released this 02 Oct 18:10

[v6.0.1-preview0001]

Fixed

  • Deliver the SessionStart context to Copilot SDK (agent host) and Copilot CLI sessions. Add-SessionContext wrote it only under hookSpecificOutput, which those hosts ignore; it now also writes the top-level additionalContext they read.
  • Keep the long-running-job-monitor heartbeat state readable while it is rewritten. Start-JobHeartbeat.ps1 wrote the state file in place, so -Stop, -TouchStatus, or a wake read at the same moment could find it empty or cut off and fail; it now renames a complete file over it, also on Windows PowerShell 5.1.
  • Stop the wording of a commit message from choosing the release version. GitVersion.yml raised it on words such as "major", "breaking", or "add" anywhere in a message, which is how review text made 6.0.0 a major release; now only the Conventional Commit type in the subject line (feat:, fix:, perf:, !), a line that starts with BREAKING CHANGE:, or a literal +semver: override raises it above the branch's default increment.
  • Send an authorized push to the user's own terminal. The PreToolUse block message, AGENTS.md, and the hooks README told the agent to set COPILOT_ATELIER_ALLOW_REMOTE=1 for the command, but each host starts the hook with its own environment, so a variable set in an agent terminal never reaches the guard. They now also warn never to persist the variable, for example with setx: every host started afterwards would run without the guard.
  • Show the reason when the PreToolUse guard blocks a call in a Copilot SDK chat on Windows. That host ran the cross-platform command launcher inside an outer PowerShell that reported the block as 1, so the model saw only "hook errored". A new powershell launcher, which the SDK host prefers on Windows, passes exit 2 on, and the guard also prints the reason as one JSON object on standard output, which that host reads on exit 2.

Security

  • Block a push in VS Code Local chats on Windows. VS Code runs a hook's windows launcher as the -Command text of an outer Windows PowerShell, which reported the guard's exit code 2 as 1, and VS Code treats every exit other than 2 as a warning, so the PreToolUse guard only warned there. The launcher now ends with a statement that passes the inner exit code on. Found by the post-release review of the v6.0.0 hook launcher fix.
  • Look for the hook scripts under USERPROFILE before HOME. Since v6.0.0 the launchers tried HOME first, so a HOME that a tool such as Git for Windows pointed at a writable tree could run a planted script instead of the deployed guard, and a HOME on an unreachable network share delayed the guard past its 20-second timeout, which the Copilot SDK host treats as allow.
  • Scan the raw payload text whenever the PreToolUse guard cannot walk the payload field by field. A payload that was not valid JSON has been allowed since v6.0.0, and the walk stopped four levels deep, so a push nested more than four levels, or more than the 100 levels Windows PowerShell parses (1,024 in PowerShell 7), went through. The walk now reaches 64 levels, and whatever it cannot reach is scanned as raw text, with JSON escapes decoded and without JSON punctuation, so an argument array such as ["git","push"] is caught as well, before the call is allowed. The raw text scan also joins the command-bearing fields the way the walk does, so a command split across fields, such as {"command":"git","args":["push"]}, is caught there too.
  • Block a push whose arguments come before the command, such as {"args":["push"],"command":"git"}, the order a serializer that sorts its keys writes. Since v6.0.0 the PreToolUse guard joined the command-bearing fields only in the order they appear, which read push git; it now also joins each object's executable fields before its argument fields, and for a payload it cannot walk, the whole payload's fields in that order and in reverse.
  • Block a tool call the PreToolUse guard cannot inspect within five seconds. Some of its patterns slow down quadratically on one long line of git words: 64 KB of them took 21.5 seconds, past the 20-second hook timeout, which the Copilot SDK host treats as allow. Parsing and walking a payload cannot be interrupted either, and 300,000 small objects kept the guard busy past the timeout too. The whole decision now has a time limit: a payload is parsed only up to 1 MB and walked only up to 20,000 fields and nested objects, the rest is scanned as raw text, a payload over 4 MB is blocked unscanned, and a payload not inspected within the limit is blocked. Ordinary commands still take well under a second.

v6.0.0

Choose a tag to compare

@raandree raandree released this 30 Sep 18:29

[v6.0.0]

Fixed

  • Custom agents lost web fetch, search, questions, the browser, and the session tools in VS Code agent-host (Copilot SDK) sessions, because the runtime drops every VS Code tool name it cannot resolve (github/copilot-cli#4594). Every agent now declares the runtime name next to each VS Code name, and the agents that are not contained get a common web, search, question, and session-tool baseline. Contained agents gain only grep, glob, and ask_user for tools they already had.
  • The Copilot CLI variant of software-engineer mapped web and search to the web and search aliases, which enable no tool. It now emits web_fetch, grep, and glob.
  • Block-RemoteMutation now allows the tool call with a warning when the hook payload is not valid JSON, as its message always said. It exited 1, which the Copilot SDK host treats as a denial, so a payload schema change would have blocked every tool call.
  • Hook launchers failed on Windows when HOME was unset, which blocked every tool call in Copilot SDK sessions and dropped the SessionStart context.
  • Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
  • Bound evaluation definitions, samples, and replies to 1 MiB per file by default (-MaxInputBytes can opt into larger inputs), and provide stable timeout diagnostics. See grading contracts.
  • Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match contains literally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts.
  • Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
  • Set the plugin manifest to the released version in the automated changelog pull request, so a release rollover no longer fails its own manifest version check (CI run 36549550887). Pushes to main that change only CHANGELOG.md or plugin.json are validated without republishing the Customization module.

v6.0.0-preview0004

v6.0.0-preview0004 Pre-release
Pre-release

Choose a tag to compare

@raandree raandree released this 29 Sep 18:32

[v6.0.0-preview0004]

Fixed

  • Custom agents lost web fetch, search, questions, the browser, and the session tools in VS Code agent-host (Copilot SDK) sessions, because the runtime drops every VS Code tool name it cannot resolve (github/copilot-cli#4594). Every agent now declares the runtime name next to each VS Code name, and the agents that are not contained get a common web, search, question, and session-tool baseline. Contained agents gain only grep, glob, and ask_user for tools they already had.
  • The Copilot CLI variant of software-engineer mapped web and search to the web and search aliases, which enable no tool. It now emits web_fetch, grep, and glob.
  • Block-RemoteMutation now allows the tool call with a warning when the hook payload is not valid JSON, as its message always said. It exited 1, which the Copilot SDK host treats as a denial, so a payload schema change would have blocked every tool call.
  • Hook launchers failed on Windows when HOME was unset, which blocked every tool call in Copilot SDK sessions and dropped the SessionStart context.
  • Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
  • Bound evaluation definitions, samples, and replies to 1 MiB per file by default (-MaxInputBytes can opt into larger inputs), and provide stable timeout diagnostics. See grading contracts.
  • Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match contains literally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts.
  • Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
  • Set the plugin manifest to the released version in the automated changelog pull request, so a release rollover no longer fails its own manifest version check (CI run 36549550887). Pushes to main that change only CHANGELOG.md or plugin.json are validated without republishing the Customization module.

v6.0.0-preview0003

v6.0.0-preview0003 Pre-release
Pre-release

Choose a tag to compare

@raandree raandree released this 29 Sep 17:09

[v6.0.0-preview0003]

Fixed

  • Block-RemoteMutation now allows the tool call with a warning when the hook payload is not valid JSON, as its message always said. It exited 1, which the Copilot SDK host treats as a denial, so a payload schema change would have blocked every tool call.
  • Hook launchers failed on Windows when HOME was unset, which blocked every tool call in Copilot SDK sessions and dropped the SessionStart context.
  • Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
  • Bound evaluation definitions, samples, and replies to 1 MiB per file by default (-MaxInputBytes can opt into larger inputs), and provide stable timeout diagnostics. See grading contracts.
  • Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match contains literally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts.
  • Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
  • Set the plugin manifest to the released version in the automated changelog pull request, so a release rollover no longer fails its own manifest version check (CI run 36549550887). Pushes to main that change only CHANGELOG.md or plugin.json are validated without republishing the Customization module.

v6.0.0-preview0002

v6.0.0-preview0002 Pre-release
Pre-release

Choose a tag to compare

@raandree raandree released this 29 Sep 12:21

[v6.0.0-preview0002]

Fixed

  • Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
  • Bound evaluation definitions, samples, and replies to 1 MiB per file by default (-MaxInputBytes can opt into larger inputs), and provide stable timeout diagnostics. See grading contracts.
  • Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match contains literally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts.
  • Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
  • Set the plugin manifest to the released version in the automated changelog pull request, so a release rollover no longer fails its own manifest version check (CI run 36549550887). Pushes to main that change only CHANGELOG.md or plugin.json are validated without republishing the Customization module.

v6.0.0-preview0001

v6.0.0-preview0001 Pre-release
Pre-release

Choose a tag to compare

@raandree raandree released this 29 Sep 10:22
314c795

[v6.0.0-preview0001]

Fixed

  • Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
  • Bound evaluation definitions, samples, and replies to 1 MiB per file by default (-MaxInputBytes can opt into larger inputs), and provide stable timeout diagnostics. See grading contracts.
  • Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match contains literally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts.
  • Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.

v5.0.0

Choose a tag to compare

@raandree raandree released this 10 Sep 09:22

[v5.0.0]

Removed

  • The .github/hooks smoke-test probe, which had been failing on every turn since it was committed (2026-09-02). stop-probe.json and Test-HookLoaded.ps1 were scratch: a Stop hook that appended one line to %TEMP%\workspace-hook-probe.log to prove the workspace hook location loads at all. They answered that question on 2026-08-10 and the answer is written into com.github.copilot/hooks/README.md and the changelog entry below — the files themselves had no further job.

    They were not merely idle. The windows override hardcoded D:\Git\CopilotAtelier\.github\hooks\Test-HookLoaded.ps1, the drive the repository sat on when the probe was written, and on Windows that override wins. Every turn on any other machine ended with "The argument … to the -File parameter does not exist". The POSIX command was no better in principle: ./.github/hooks/Test-HookLoaded.ps1 is relative, and the same README says VS Code does not guarantee the working directory, which is why every shipped hook resolves its own path.

    Nothing caught it because nothing looked. The Hook configuration suite in tests/Hooks.Tests.ps1 — which asserts exactly this, that a hook command resolves to a script that exists and carries no shell-interpolated token — is scoped to com.github.copilot/hooks/hooks.json. A second hook file one directory away was outside every gate the repository owns. That suite now enumerates every tracked *.json sitting directly inside a folder named hooks and requires the shipped configuration to be the only one, so the next stray hook file fails the build instead of the chat. The guard was proven by planting one and watching it go red.

Added

  • Add tools/plan-review, an optional local review surface for a Design Concept: it renders the Markdown and its Mermaid diagrams, anchors comments to stable sections, and records a verdict against one specific revision hash. It is opt-in in the strong sense — it is absent from CustomizationDirectory, so the built module and a Gallery install never carry it; no PowerShell source references it; and its Node dependencies are installed explicitly by whoever wants the feature, never by an install, an update, or a validation run.

    A browser verdict is feedback, not sign-off, and the code says so rather than the documentation. An HTTP request proves that something holding the session cookie and the CSRF token posted a content hash; it proves neither identity nor authority to start implementation. Every verdict is therefore persisted with authority: "local-http-feedback" beneath a store header of approvalAuthority: "chat-sign-off-required", the header is repeated on the response and shown above the document, and there is no endpoint that writes a Decision record, triggers a handoff, or runs a command. A --state root resolving inside .memory-bank/decisions is refused at launch. The existing chat sign-off in the Software Architect workflow stays the only thing that authorizes implementation, and the agent body now says that where it points at the tool.

    Revision hashes cover the original document bytes. Comments retain section identity; ambiguous duplicate headings require one exact content match and otherwise stay unanchored. A section key is unique across the whole document, not merely per heading slug: an occurrence ordinal on its own issues risks-2 twice for Risks, Risks, Risks 2, and two sections sharing a key anchor a comment to the wrong heading. Section splitting follows the CommonMark fence rules, so a shorter fence nested inside a longer one cannot expose a fake heading and mis-anchor a comment. Verdict dialogs retain the document and revision they displayed, so a background refresh cannot approve newer content. Requests against stale hashes are refused with 409, and the hash is checked a second time inside the serialized store write — a file edited while the request waits for the lock is refused rather than approved for the bytes it no longer has. The source is read once more after the commit, so a change landing in that last window is reported as superseded instead of being presented as current approval.

    Loopback binding is treated as a reachability reduction, not an authorization boundary. A non-loopback bind address is refused outright, and the allowed authority follows the address actually bound, so ::1 produces [::1]:<port> rather than a hard-coded 127.0.0.1. Every request must carry a Host matching the bound authority; every mutation must additionally carry the exact server Origin, a Sec-Fetch-Site of same-origin or none when the browser sends one, a JSON content type, the per-launch session cookie, and a matching X-CSRF-Token. The session secret is generated per launch and never persisted, so a cookie minted by an earlier server is rejected by the next one even when it reuses the same feedback store. Because cookies are scoped by host and not by port, the cookie name carries a per-launch identifier: opening a second review server in the same browser no longer signs the first one out, and neither server accepts the other's cookie.

    Documents are authorized at launch and addressed on the wire by an opaque sixteen-character identifier, so no request parameter ever names a path. There is no directory listing, no URL fetcher, no shell endpoint, and no generic static handler — vendor assets come from an exact filename allow-list mapped onto node_modules. Every path is realpath-resolved, required to sit inside the declared root, and rejected when any ancestor from the root down is a symbolic link or junction, and the check runs again at read time rather than only at launch, so a link swapped in afterwards still fails.

    Rendering disables raw HTML at the parser instead of filtering it afterwards: markdown-it runs with html: false, and DOMPurify then applies a tag, attribute, and URI allow-list that admits only http, https, and mailto. An image is never fetched — its alternative text is rendered instead, because an image is an implicit external load. Mermaid runs client-side with securityLevel: 'strict' and its SVG is sanitized again before insertion. Responses carry Content-Security-Policy: default-src 'none' with script-src 'self' and no unsafe-eval. The page's own stylesheet, script, and vendor bundles are snapshotted at launch and served from memory, so an asset deleted or swapped afterwards can neither change what the page runs nor leave a request hanging on a broken read.

    Bodies cap at 64 KiB, comment text at 4000 characters, notes at 2000, comments at 200 per document, and documents at 1 MiB. The store is read under a 2 MiB byte bound and fully validated — schema, document identity, every comment field, and the verdict, including its authority, which a stored file can therefore never use to promote itself to sign-off, and every hash, which must be a lowercase SHA-256 digest rather than any bounded string. The write path enforces the same byte bound on the serialized UTF-8 payload, because the count and length bounds do not imply it: 4000 characters of multibyte text cost up to three bytes each, so 200 legal comments could otherwise produce a file the next read refuses. A write that would cross the bound is refused as store-capacity before the temporary file exists, and the stored feedback is left unchanged. A store file that fails any of those checks is reported and left byte-for-byte intact, and the next mutation is refused rather than overwriting somebody's pending review; recovery is a deliberate act by the operator. The exclusive write lock records its owning process: a lock held by a live process is waited on and then refused, a lock is reclaimed only when its named owner is provably gone, the reclaim removes the entries this tool wrote rather than deleting a directory tree it does not own, and a mutation that loses ownership refuses to commit. Server lifetime is bounded by --ttl, and Ctrl+C, the page's Stop server button, and the printed process id all stop it cleanly.

    Add revision-scoped draft recovery, retryable connection errors, an authorized-document selector, and wrapping mobile status text. A draft written against a revision or a section that is no longer current is never re-attached to new content: it is listed under Unsent drafts from an earlier revision with the section and revision it was written on, for explicit discard, and a pending verdict note survives a stale refusal. Switching documents takes a request ticket, so a slow response for one document cannot render under another document's actions. The section outline is a disclosure that starts collapsed on a narrow viewport, and the permanent keyboard tutorial line is gone — the shortcuts remain, named in tooltips and announced to assistive technology. Apply input limits at launch and reload, reject invalid UTF-8, and reject linked feedback roots before reads or writes. Portable Node test commands and desktop/mobile browser regressions cover these boundaries. The ordinary repository gate runs dependency-free Node tests when Node is available and never installs npm dependencies.

    Documented in docs/plan-review.md, with the trust analysis in docs/plan-review-threat-model.md. Rollback is deletion: nothing else in the repository depends on it.

  • Add read-only Get-CopilotAtelierClientAdapter, a thin compatibility adapter that reports how a Custom agent profile is composed for each supported Copilot client and, more importantly, what that client cannot do. The VS Code files under com.github.copilot/agents stay the only source of every shared workflow; the composed body is byte-identical, and only frontmatter is rewritten, so there is no second catalog to drift.

    Discovery is not parity, and the gap is specific. ...

Read more