Repository navigation
Releases: raandree/CopilotAtelier
Release list
v6.1.0-preview0003
[v6.1.0-preview0003]
Added
- Add contributor calibration: every agent pitches answers and questions at your familiarity with each knowledge area (
new,familiar, orexpert;familiaruntil you say otherwise), and innewandfamiliarareas illustrates an abstract finding with a concrete example; every calculated or reconstructed result names its sources and method. Every technical question offers a recommended answer and anot sure, you pickoption; a delegated answer becomes an assumption flagged for expert review, counts only when you write it yourself, and never authorizes an irreversible, destructive, or security-relevant action; a question that authorizes such an action comes without the option. A level changes wording and depth, never warnings, tests, or reviews, and is never written to a repository. See How Much Explanation You Want. - Add
docs/SECURITY-REVIEW.md, the independent security review of contributor calibration with every finding, its severity, and its resolution. - Add the
/simplerand/deeperPrompts, which re-explain the last answer one familiarity level simpler or deeper and keep that level for its knowledge area for the rest of the session.
Changed
- Let
grill-me,software-architect, andgilb-requirements-engineeringacceptnot sure, you pick: the recommended answer, or for a numeric target a level derived fromPast,Record, or a verified benchmark, is recorded as an assumption flagged for expert review instead of blocking the interview. - Document in
agent-evalsthat the Copilot backend's content filter blocks some harmless eval prompts, why an embedded chat transcript makes it worse, and how to keep a comparison fair: situation as system context,FinishReasonlogged per call, retries, and arms compared only on cases complete in both. See harness prerequisites. - Document in
agent-evalsthat compared arms must run at the same time, because backend behavior drifts within hours, and that an LLM judge's disagreement can expose a mislabelled reply, which justifies a label correction only for an objective error. See micro-tests.
Fixed
- Deliver the SessionStart context to Copilot SDK (agent host) and Copilot CLI sessions.
Add-SessionContextwrote it only underhookSpecificOutput, which those hosts ignore; it now also writes the top-leveladditionalContextthey read. - Keep the session clock when a Copilot SDK chat resumes. The runtime reruns the
SessionStarthook withsource: resumewhen it reloads a chat, andAdd-SessionContextrestarted the clock there, so the closing elapsed line and the turn count started over and the injected start time moved; it now keeps a readable clock of the same session. - Keep the
long-running-job-monitorheartbeat state readable while it is rewritten.Start-JobHeartbeat.ps1wrote the state file in place, so-Stop,-TouchStatus, or a wake read at the same moment could find it empty or cut off and fail; it now renames a complete file over it, also on Windows PowerShell 5.1. - Stop the wording of a commit message from choosing the release version.
GitVersion.ymlraised it on words such as "major", "breaking", or "add" anywhere in a message, which is how review text made 6.0.0 a major release; now only the Conventional Commit type in the subject line (feat:,fix:,perf:,!), a line that starts withBREAKING CHANGE:, or a literal+semver:override raises it above the branch's default increment. - Send an authorized push to the user's own terminal. The
PreToolUseblock message,AGENTS.md, and the hooks README told the agent to setCOPILOT_ATELIER_ALLOW_REMOTE=1for the command, but each host starts the hook with its own environment, so a variable set in an agent terminal never reaches the guard. They now also warn never to persist the variable, for example withsetx: every host started afterwards would run without the guard. - Show the
PreToolUseguard's reason to the model in Copilot SDK chats on Windows. That host ran the cross-platformcommandlauncher inside an outer PowerShell that reported the block as 1, so the model saw only "hook errored". A newpowershelllauncher, which the SDK host prefers on Windows, passes exit 2 on, and the guard prints the reason as one JSON object on standard output, which that host merges into the deny on exit 2. The object now carries only the top-levelpermissionDecisionandpermissionDecisionReason: the SDK runtime in VS Code 1.140.0 drops aPreToolUseobject that also carrieshookSpecificOutput, so the model read only "hook exited with code 2". VS Code still reads the reason from standard error.
Security
- Block a push in VS Code Local chats on Windows. VS Code runs a hook's
windowslauncher as the-Commandtext of an outer Windows PowerShell, which reported the guard's exit code 2 as 1, and VS Code treats every exit other than 2 as a warning, so thePreToolUseguard only warned there. The launcher now ends with a statement that passes the inner exit code on. Found by the post-release review of the v6.0.0 hook launcher fix. - Look for the hook scripts under
USERPROFILEbeforeHOME. Since v6.0.0 the launchers triedHOMEfirst, so aHOMEthat a tool such as Git for Windows pointed at a writable tree could run a planted script instead of the deployed guard, and aHOMEon an unreachable network share delayed the guard past its 20-second timeout, which the Copilot SDK host treats as allow. - Scan the raw payload text whenever the
PreToolUseguard cannot walk the payload field by field. A payload that was not valid JSON has been allowed since v6.0.0, and the walk stopped four levels deep, so a push nested more than four levels, or more than the 100 levels Windows PowerShell parses (1,024 in PowerShell 7), went through. The walk now reaches 64 levels, and whatever it cannot reach is scanned as raw text, with JSON escapes decoded and without JSON punctuation, so an argument array such as["git","push"]is caught as well, before the call is allowed. The raw text scan also joins the command-bearing fields the way the walk does, so a command split across fields, such as{"command":"git","args":["push"]}, is caught there too. - Block a push whose arguments come before the command, such as
{"args":["push"],"command":"git"}, the order a serializer that sorts its keys writes. Since v6.0.0 thePreToolUseguard joined the command-bearing fields only in the order they appear, which readpush git; it now also joins each object's executable fields before its argument fields, and for a payload it cannot walk, the whole payload's fields in that order and in reverse. - Block a tool call the
PreToolUseguard cannot inspect within five seconds. Some of its patterns slow down quadratically on one long line ofgitwords: 64 KB of them took 21.5 seconds, past the 20-second hook timeout, which the Copilot SDK host treats as allow. Parsing and walking a payload cannot be interrupted either, and 300,000 small objects kept the guard busy past the timeout too. The whole decision now has a time limit: a payload is parsed only up to 1 MB and walked only up to 20,000 fields and nested objects, the rest is scanned as raw text, a payload over 4 MB is blocked unscanned, and a payload not inspected within the limit is blocked. Ordinary commands still take well under a second.
v6.1.0-preview0002
[v6.1.0-preview0002]
Added
- Add contributor calibration: every agent pitches answers and questions at your familiarity with each knowledge area (
new,familiar, orexpert;familiaruntil you say otherwise), and innewandfamiliarareas illustrates an abstract finding with a concrete example; every calculated or reconstructed result names its sources and method. Every technical question offers a recommended answer and anot sure, you pickoption; a delegated answer becomes an assumption flagged for expert review, counts only when you write it yourself, and never authorizes an irreversible, destructive, or security-relevant action; a question that authorizes such an action comes without the option. A level changes wording and depth, never warnings, tests, or reviews, and is never written to a repository. See How Much Explanation You Want. - Add
docs/SECURITY-REVIEW.md, the independent security review of contributor calibration with every finding, its severity, and its resolution. - Add the
/simplerand/deeperPrompts, which re-explain the last answer one familiarity level simpler or deeper and keep that level for its knowledge area for the rest of the session.
Changed
- Let
grill-me,software-architect, andgilb-requirements-engineeringacceptnot sure, you pick: the recommended answer, or for a numeric target a level derived fromPast,Record, or a verified benchmark, is recorded as an assumption flagged for expert review instead of blocking the interview. - Document in
agent-evalsthat the Copilot backend's content filter blocks some harmless eval prompts, why an embedded chat transcript makes it worse, and how to keep a comparison fair: situation as system context,FinishReasonlogged per call, retries, and arms compared only on cases complete in both. See harness prerequisites. - Document in
agent-evalsthat compared arms must run at the same time, because backend behavior drifts within hours, and that an LLM judge's disagreement can expose a mislabelled reply, which justifies a label correction only for an objective error. See micro-tests.
Fixed
- Deliver the SessionStart context to Copilot SDK (agent host) and Copilot CLI sessions.
Add-SessionContextwrote it only underhookSpecificOutput, which those hosts ignore; it now also writes the top-leveladditionalContextthey read. - Keep the session clock when a Copilot SDK chat resumes. The runtime reruns the
SessionStarthook withsource: resumewhen it reloads a chat, andAdd-SessionContextrestarted the clock there, so the closing elapsed line and the turn count started over and the injected start time moved; it now keeps a readable clock of the same session. - Keep the
long-running-job-monitorheartbeat state readable while it is rewritten.Start-JobHeartbeat.ps1wrote the state file in place, so-Stop,-TouchStatus, or a wake read at the same moment could find it empty or cut off and fail; it now renames a complete file over it, also on Windows PowerShell 5.1. - Stop the wording of a commit message from choosing the release version.
GitVersion.ymlraised it on words such as "major", "breaking", or "add" anywhere in a message, which is how review text made 6.0.0 a major release; now only the Conventional Commit type in the subject line (feat:,fix:,perf:,!), a line that starts withBREAKING CHANGE:, or a literal+semver:override raises it above the branch's default increment. - Send an authorized push to the user's own terminal. The
PreToolUseblock message,AGENTS.md, and the hooks README told the agent to setCOPILOT_ATELIER_ALLOW_REMOTE=1for the command, but each host starts the hook with its own environment, so a variable set in an agent terminal never reaches the guard. They now also warn never to persist the variable, for example withsetx: every host started afterwards would run without the guard. - Pass the
PreToolUseguard's exit code 2 on in Copilot SDK chats on Windows. That host ran the cross-platformcommandlauncher inside an outer PowerShell that reported the block as 1, so the model saw only "hook errored". A newpowershelllauncher, which the SDK host prefers on Windows, passes exit 2 on, and the guard also prints the reason as one JSON object on standard output, which the GitHub hooks reference says that host reads on exit 2. The SDK runtime in VS Code 1.140.0 does not: the call is denied, but the model reads only "hook exited with code 2".
Security
- Block a push in VS Code Local chats on Windows. VS Code runs a hook's
windowslauncher as the-Commandtext of an outer Windows PowerShell, which reported the guard's exit code 2 as 1, and VS Code treats every exit other than 2 as a warning, so thePreToolUseguard only warned there. The launcher now ends with a statement that passes the inner exit code on. Found by the post-release review of the v6.0.0 hook launcher fix. - Look for the hook scripts under
USERPROFILEbeforeHOME. Since v6.0.0 the launchers triedHOMEfirst, so aHOMEthat a tool such as Git for Windows pointed at a writable tree could run a planted script instead of the deployed guard, and aHOMEon an unreachable network share delayed the guard past its 20-second timeout, which the Copilot SDK host treats as allow. - Scan the raw payload text whenever the
PreToolUseguard cannot walk the payload field by field. A payload that was not valid JSON has been allowed since v6.0.0, and the walk stopped four levels deep, so a push nested more than four levels, or more than the 100 levels Windows PowerShell parses (1,024 in PowerShell 7), went through. The walk now reaches 64 levels, and whatever it cannot reach is scanned as raw text, with JSON escapes decoded and without JSON punctuation, so an argument array such as["git","push"]is caught as well, before the call is allowed. The raw text scan also joins the command-bearing fields the way the walk does, so a command split across fields, such as{"command":"git","args":["push"]}, is caught there too. - Block a push whose arguments come before the command, such as
{"args":["push"],"command":"git"}, the order a serializer that sorts its keys writes. Since v6.0.0 thePreToolUseguard joined the command-bearing fields only in the order they appear, which readpush git; it now also joins each object's executable fields before its argument fields, and for a payload it cannot walk, the whole payload's fields in that order and in reverse. - Block a tool call the
PreToolUseguard cannot inspect within five seconds. Some of its patterns slow down quadratically on one long line ofgitwords: 64 KB of them took 21.5 seconds, past the 20-second hook timeout, which the Copilot SDK host treats as allow. Parsing and walking a payload cannot be interrupted either, and 300,000 small objects kept the guard busy past the timeout too. The whole decision now has a time limit: a payload is parsed only up to 1 MB and walked only up to 20,000 fields and nested objects, the rest is scanned as raw text, a payload over 4 MB is blocked unscanned, and a payload not inspected within the limit is blocked. Ordinary commands still take well under a second.
v6.1.0-preview0001
[v6.1.0-preview0001]
Added
- Add contributor calibration: every agent pitches answers and questions at your familiarity with each knowledge area (
new,familiar, orexpert;familiaruntil you say otherwise), and innewandfamiliarareas illustrates an abstract finding with a concrete example; every calculated or reconstructed result names its sources and method. Every technical question offers a recommended answer and anot sure, you pickoption; a delegated answer becomes an assumption flagged for expert review, counts only when you write it yourself, and never authorizes an irreversible, destructive, or security-relevant action; a question that authorizes such an action comes without the option. A level changes wording and depth, never warnings, tests, or reviews, and is never written to a repository. See How Much Explanation You Want. - Add
docs/SECURITY-REVIEW.md, the independent security review of contributor calibration with every finding, its severity, and its resolution. - Add the
/simplerand/deeperPrompts, which re-explain the last answer one familiarity level simpler or deeper and keep that level for its knowledge area for the rest of the session.
Changed
- Let
grill-me,software-architect, andgilb-requirements-engineeringacceptnot sure, you pick: the recommended answer, or for a numeric target a level derived fromPast,Record, or a verified benchmark, is recorded as an assumption flagged for expert review instead of blocking the interview. - Document in
agent-evalsthat the Copilot backend's content filter blocks some harmless eval prompts, why an embedded chat transcript makes it worse, and how to keep a comparison fair: situation as system context,FinishReasonlogged per call, retries, and arms compared only on cases complete in both. See harness prerequisites. - Document in
agent-evalsthat compared arms must run at the same time, because backend behavior drifts within hours, and that an LLM judge's disagreement can expose a mislabelled reply, which justifies a label correction only for an objective error. See micro-tests.
Fixed
- Deliver the SessionStart context to Copilot SDK (agent host) and Copilot CLI sessions.
Add-SessionContextwrote it only underhookSpecificOutput, which those hosts ignore; it now also writes the top-leveladditionalContextthey read. - Keep the
long-running-job-monitorheartbeat state readable while it is rewritten.Start-JobHeartbeat.ps1wrote the state file in place, so-Stop,-TouchStatus, or a wake read at the same moment could find it empty or cut off and fail; it now renames a complete file over it, also on Windows PowerShell 5.1. - Stop the wording of a commit message from choosing the release version.
GitVersion.ymlraised it on words such as "major", "breaking", or "add" anywhere in a message, which is how review text made 6.0.0 a major release; now only the Conventional Commit type in the subject line (feat:,fix:,perf:,!), a line that starts withBREAKING CHANGE:, or a literal+semver:override raises it above the branch's default increment. - Send an authorized push to the user's own terminal. The
PreToolUseblock message,AGENTS.md, and the hooks README told the agent to setCOPILOT_ATELIER_ALLOW_REMOTE=1for the command, but each host starts the hook with its own environment, so a variable set in an agent terminal never reaches the guard. They now also warn never to persist the variable, for example withsetx: every host started afterwards would run without the guard. - Show the reason when the
PreToolUseguard blocks a call in a Copilot SDK chat on Windows. That host ran the cross-platformcommandlauncher inside an outer PowerShell that reported the block as 1, so the model saw only "hook errored". A newpowershelllauncher, which the SDK host prefers on Windows, passes exit 2 on, and the guard also prints the reason as one JSON object on standard output, which that host reads on exit 2.
Security
- Block a push in VS Code Local chats on Windows. VS Code runs a hook's
windowslauncher as the-Commandtext of an outer Windows PowerShell, which reported the guard's exit code 2 as 1, and VS Code treats every exit other than 2 as a warning, so thePreToolUseguard only warned there. The launcher now ends with a statement that passes the inner exit code on. Found by the post-release review of the v6.0.0 hook launcher fix. - Look for the hook scripts under
USERPROFILEbeforeHOME. Since v6.0.0 the launchers triedHOMEfirst, so aHOMEthat a tool such as Git for Windows pointed at a writable tree could run a planted script instead of the deployed guard, and aHOMEon an unreachable network share delayed the guard past its 20-second timeout, which the Copilot SDK host treats as allow. - Scan the raw payload text whenever the
PreToolUseguard cannot walk the payload field by field. A payload that was not valid JSON has been allowed since v6.0.0, and the walk stopped four levels deep, so a push nested more than four levels, or more than the 100 levels Windows PowerShell parses (1,024 in PowerShell 7), went through. The walk now reaches 64 levels, and whatever it cannot reach is scanned as raw text, with JSON escapes decoded and without JSON punctuation, so an argument array such as["git","push"]is caught as well, before the call is allowed. The raw text scan also joins the command-bearing fields the way the walk does, so a command split across fields, such as{"command":"git","args":["push"]}, is caught there too. - Block a push whose arguments come before the command, such as
{"args":["push"],"command":"git"}, the order a serializer that sorts its keys writes. Since v6.0.0 thePreToolUseguard joined the command-bearing fields only in the order they appear, which readpush git; it now also joins each object's executable fields before its argument fields, and for a payload it cannot walk, the whole payload's fields in that order and in reverse. - Block a tool call the
PreToolUseguard cannot inspect within five seconds. Some of its patterns slow down quadratically on one long line ofgitwords: 64 KB of them took 21.5 seconds, past the 20-second hook timeout, which the Copilot SDK host treats as allow. Parsing and walking a payload cannot be interrupted either, and 300,000 small objects kept the guard busy past the timeout too. The whole decision now has a time limit: a payload is parsed only up to 1 MB and walked only up to 20,000 fields and nested objects, the rest is scanned as raw text, a payload over 4 MB is blocked unscanned, and a payload not inspected within the limit is blocked. Ordinary commands still take well under a second.
v6.0.1-preview0001
[v6.0.1-preview0001]
Fixed
- Deliver the SessionStart context to Copilot SDK (agent host) and Copilot CLI sessions.
Add-SessionContextwrote it only underhookSpecificOutput, which those hosts ignore; it now also writes the top-leveladditionalContextthey read. - Keep the
long-running-job-monitorheartbeat state readable while it is rewritten.Start-JobHeartbeat.ps1wrote the state file in place, so-Stop,-TouchStatus, or a wake read at the same moment could find it empty or cut off and fail; it now renames a complete file over it, also on Windows PowerShell 5.1. - Stop the wording of a commit message from choosing the release version.
GitVersion.ymlraised it on words such as "major", "breaking", or "add" anywhere in a message, which is how review text made 6.0.0 a major release; now only the Conventional Commit type in the subject line (feat:,fix:,perf:,!), a line that starts withBREAKING CHANGE:, or a literal+semver:override raises it above the branch's default increment. - Send an authorized push to the user's own terminal. The
PreToolUseblock message,AGENTS.md, and the hooks README told the agent to setCOPILOT_ATELIER_ALLOW_REMOTE=1for the command, but each host starts the hook with its own environment, so a variable set in an agent terminal never reaches the guard. They now also warn never to persist the variable, for example withsetx: every host started afterwards would run without the guard. - Show the reason when the
PreToolUseguard blocks a call in a Copilot SDK chat on Windows. That host ran the cross-platformcommandlauncher inside an outer PowerShell that reported the block as 1, so the model saw only "hook errored". A newpowershelllauncher, which the SDK host prefers on Windows, passes exit 2 on, and the guard also prints the reason as one JSON object on standard output, which that host reads on exit 2.
Security
- Block a push in VS Code Local chats on Windows. VS Code runs a hook's
windowslauncher as the-Commandtext of an outer Windows PowerShell, which reported the guard's exit code 2 as 1, and VS Code treats every exit other than 2 as a warning, so thePreToolUseguard only warned there. The launcher now ends with a statement that passes the inner exit code on. Found by the post-release review of the v6.0.0 hook launcher fix. - Look for the hook scripts under
USERPROFILEbeforeHOME. Since v6.0.0 the launchers triedHOMEfirst, so aHOMEthat a tool such as Git for Windows pointed at a writable tree could run a planted script instead of the deployed guard, and aHOMEon an unreachable network share delayed the guard past its 20-second timeout, which the Copilot SDK host treats as allow. - Scan the raw payload text whenever the
PreToolUseguard cannot walk the payload field by field. A payload that was not valid JSON has been allowed since v6.0.0, and the walk stopped four levels deep, so a push nested more than four levels, or more than the 100 levels Windows PowerShell parses (1,024 in PowerShell 7), went through. The walk now reaches 64 levels, and whatever it cannot reach is scanned as raw text, with JSON escapes decoded and without JSON punctuation, so an argument array such as["git","push"]is caught as well, before the call is allowed. The raw text scan also joins the command-bearing fields the way the walk does, so a command split across fields, such as{"command":"git","args":["push"]}, is caught there too. - Block a push whose arguments come before the command, such as
{"args":["push"],"command":"git"}, the order a serializer that sorts its keys writes. Since v6.0.0 thePreToolUseguard joined the command-bearing fields only in the order they appear, which readpush git; it now also joins each object's executable fields before its argument fields, and for a payload it cannot walk, the whole payload's fields in that order and in reverse. - Block a tool call the
PreToolUseguard cannot inspect within five seconds. Some of its patterns slow down quadratically on one long line ofgitwords: 64 KB of them took 21.5 seconds, past the 20-second hook timeout, which the Copilot SDK host treats as allow. Parsing and walking a payload cannot be interrupted either, and 300,000 small objects kept the guard busy past the timeout too. The whole decision now has a time limit: a payload is parsed only up to 1 MB and walked only up to 20,000 fields and nested objects, the rest is scanned as raw text, a payload over 4 MB is blocked unscanned, and a payload not inspected within the limit is blocked. Ordinary commands still take well under a second.
v6.0.0
[v6.0.0]
Fixed
- Custom agents lost web fetch, search, questions, the browser, and the session tools in VS Code agent-host (Copilot SDK) sessions, because the runtime drops every VS Code tool name it cannot resolve (github/copilot-cli#4594). Every agent now declares the runtime name next to each VS Code name, and the agents that are not contained get a common web, search, question, and session-tool baseline. Contained agents gain only
grep,glob, andask_userfor tools they already had. - The Copilot CLI variant of
software-engineermapped web and search to thewebandsearchaliases, which enable no tool. It now emitsweb_fetch,grep, andglob. Block-RemoteMutationnow allows the tool call with a warning when the hook payload is not valid JSON, as its message always said. It exited 1, which the Copilot SDK host treats as a denial, so a payload schema change would have blocked every tool call.- Hook launchers failed on Windows when
HOMEwas unset, which blocked every tool call in Copilot SDK sessions and dropped the SessionStart context. - Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
- Bound evaluation definitions, samples, and replies to 1 MiB per file by default (
-MaxInputBytescan opt into larger inputs), and provide stable timeout diagnostics. See grading contracts. - Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match
containsliterally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts. - Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
- Set the plugin manifest to the released version in the automated changelog pull request, so a release rollover no longer fails its own manifest version check (CI run 36549550887). Pushes to
mainthat change onlyCHANGELOG.mdorplugin.jsonare validated without republishing the Customization module.
v6.0.0-preview0004
[v6.0.0-preview0004]
Fixed
- Custom agents lost web fetch, search, questions, the browser, and the session tools in VS Code agent-host (Copilot SDK) sessions, because the runtime drops every VS Code tool name it cannot resolve (github/copilot-cli#4594). Every agent now declares the runtime name next to each VS Code name, and the agents that are not contained get a common web, search, question, and session-tool baseline. Contained agents gain only
grep,glob, andask_userfor tools they already had. - The Copilot CLI variant of
software-engineermapped web and search to thewebandsearchaliases, which enable no tool. It now emitsweb_fetch,grep, andglob. Block-RemoteMutationnow allows the tool call with a warning when the hook payload is not valid JSON, as its message always said. It exited 1, which the Copilot SDK host treats as a denial, so a payload schema change would have blocked every tool call.- Hook launchers failed on Windows when
HOMEwas unset, which blocked every tool call in Copilot SDK sessions and dropped the SessionStart context. - Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
- Bound evaluation definitions, samples, and replies to 1 MiB per file by default (
-MaxInputBytescan opt into larger inputs), and provide stable timeout diagnostics. See grading contracts. - Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match
containsliterally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts. - Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
- Set the plugin manifest to the released version in the automated changelog pull request, so a release rollover no longer fails its own manifest version check (CI run 36549550887). Pushes to
mainthat change onlyCHANGELOG.mdorplugin.jsonare validated without republishing the Customization module.
v6.0.0-preview0003
[v6.0.0-preview0003]
Fixed
Block-RemoteMutationnow allows the tool call with a warning when the hook payload is not valid JSON, as its message always said. It exited 1, which the Copilot SDK host treats as a denial, so a payload schema change would have blocked every tool call.- Hook launchers failed on Windows when
HOMEwas unset, which blocked every tool call in Copilot SDK sessions and dropped the SessionStart context. - Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
- Bound evaluation definitions, samples, and replies to 1 MiB per file by default (
-MaxInputBytescan opt into larger inputs), and provide stable timeout diagnostics. See grading contracts. - Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match
containsliterally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts. - Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
- Set the plugin manifest to the released version in the automated changelog pull request, so a release rollover no longer fails its own manifest version check (CI run 36549550887). Pushes to
mainthat change onlyCHANGELOG.mdorplugin.jsonare validated without republishing the Customization module.
v6.0.0-preview0002
[v6.0.0-preview0002]
Fixed
- Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
- Bound evaluation definitions, samples, and replies to 1 MiB per file by default (
-MaxInputBytescan opt into larger inputs), and provide stable timeout diagnostics. See grading contracts. - Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match
containsliterally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts. - Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
- Set the plugin manifest to the released version in the automated changelog pull request, so a release rollover no longer fails its own manifest version check (CI run 36549550887). Pushes to
mainthat change onlyCHANGELOG.mdorplugin.jsonare validated without republishing the Customization module.
v6.0.0-preview0001
[v6.0.0-preview0001]
Fixed
- Preserve unambiguous trigger replies with mixed capitalization or no space after the colon; keep multiline or contradictory replies invalid. Classify query errors structurally and repair the schema-checked sample query IDs.
- Bound evaluation definitions, samples, and replies to 1 MiB per file by default (
-MaxInputBytescan opt into larger inputs), and provide stable timeout diagnostics. See grading contracts. - Harden the bundled evaluation gates against false success: validate case definitions and path-safe IDs, match
containsliterally, require complete numbered samples, bound regex matching, reject malformed trigger replies, and return failing exit codes without dropping incomplete queries from split totals. See grading contracts. - Restore the published v5.0.0 release history and align the static plugin manifest with that release, so post-release builds and plugin update discovery use the recorded version.
v5.0.0
[v5.0.0]
Removed
-
The
.github/hookssmoke-test probe, which had been failing on every turn since it was committed (2026-09-02).stop-probe.jsonandTest-HookLoaded.ps1were scratch: aStophook that appended one line to%TEMP%\workspace-hook-probe.logto prove the workspace hook location loads at all. They answered that question on 2026-08-10 and the answer is written intocom.github.copilot/hooks/README.mdand the changelog entry below — the files themselves had no further job.They were not merely idle. The
windowsoverride hardcodedD:\Git\CopilotAtelier\.github\hooks\Test-HookLoaded.ps1, the drive the repository sat on when the probe was written, and on Windows that override wins. Every turn on any other machine ended with "The argument … to the -File parameter does not exist". The POSIXcommandwas no better in principle:./.github/hooks/Test-HookLoaded.ps1is relative, and the same README says VS Code does not guarantee the working directory, which is why every shipped hook resolves its own path.Nothing caught it because nothing looked. The
Hook configurationsuite intests/Hooks.Tests.ps1— which asserts exactly this, that a hook command resolves to a script that exists and carries no shell-interpolated token — is scoped tocom.github.copilot/hooks/hooks.json. A second hook file one directory away was outside every gate the repository owns. That suite now enumerates every tracked*.jsonsitting directly inside a folder namedhooksand requires the shipped configuration to be the only one, so the next stray hook file fails the build instead of the chat. The guard was proven by planting one and watching it go red.
Added
-
Add
tools/plan-review, an optional local review surface for a Design Concept: it renders the Markdown and its Mermaid diagrams, anchors comments to stable sections, and records a verdict against one specific revision hash. It is opt-in in the strong sense — it is absent fromCustomizationDirectory, so the built module and a Gallery install never carry it; no PowerShell source references it; and its Node dependencies are installed explicitly by whoever wants the feature, never by an install, an update, or a validation run.A browser verdict is feedback, not sign-off, and the code says so rather than the documentation. An HTTP request proves that something holding the session cookie and the CSRF token posted a content hash; it proves neither identity nor authority to start implementation. Every verdict is therefore persisted with
authority: "local-http-feedback"beneath a store header ofapprovalAuthority: "chat-sign-off-required", the header is repeated on the response and shown above the document, and there is no endpoint that writes a Decision record, triggers a handoff, or runs a command. A--stateroot resolving inside.memory-bank/decisionsis refused at launch. The existing chat sign-off in the Software Architect workflow stays the only thing that authorizes implementation, and the agent body now says that where it points at the tool.Revision hashes cover the original document bytes. Comments retain section identity; ambiguous duplicate headings require one exact content match and otherwise stay unanchored. A section key is unique across the whole document, not merely per heading slug: an occurrence ordinal on its own issues
risks-2twice forRisks,Risks,Risks 2, and two sections sharing a key anchor a comment to the wrong heading. Section splitting follows the CommonMark fence rules, so a shorter fence nested inside a longer one cannot expose a fake heading and mis-anchor a comment. Verdict dialogs retain the document and revision they displayed, so a background refresh cannot approve newer content. Requests against stale hashes are refused with409, and the hash is checked a second time inside the serialized store write — a file edited while the request waits for the lock is refused rather than approved for the bytes it no longer has. The source is read once more after the commit, so a change landing in that last window is reported assupersededinstead of being presented as current approval.Loopback binding is treated as a reachability reduction, not an authorization boundary. A non-loopback bind address is refused outright, and the allowed authority follows the address actually bound, so
::1produces[::1]:<port>rather than a hard-coded127.0.0.1. Every request must carry aHostmatching the bound authority; every mutation must additionally carry the exact serverOrigin, aSec-Fetch-Siteofsame-originornonewhen the browser sends one, a JSON content type, the per-launch session cookie, and a matchingX-CSRF-Token. The session secret is generated per launch and never persisted, so a cookie minted by an earlier server is rejected by the next one even when it reuses the same feedback store. Because cookies are scoped by host and not by port, the cookie name carries a per-launch identifier: opening a second review server in the same browser no longer signs the first one out, and neither server accepts the other's cookie.Documents are authorized at launch and addressed on the wire by an opaque sixteen-character identifier, so no request parameter ever names a path. There is no directory listing, no URL fetcher, no shell endpoint, and no generic static handler — vendor assets come from an exact filename allow-list mapped onto
node_modules. Every path is realpath-resolved, required to sit inside the declared root, and rejected when any ancestor from the root down is a symbolic link or junction, and the check runs again at read time rather than only at launch, so a link swapped in afterwards still fails.Rendering disables raw HTML at the parser instead of filtering it afterwards:
markdown-itruns withhtml: false, and DOMPurify then applies a tag, attribute, and URI allow-list that admits onlyhttp,https, andmailto. An image is never fetched — its alternative text is rendered instead, because an image is an implicit external load. Mermaid runs client-side withsecurityLevel: 'strict'and its SVG is sanitized again before insertion. Responses carryContent-Security-Policy: default-src 'none'withscript-src 'self'and nounsafe-eval. The page's own stylesheet, script, and vendor bundles are snapshotted at launch and served from memory, so an asset deleted or swapped afterwards can neither change what the page runs nor leave a request hanging on a broken read.Bodies cap at 64 KiB, comment text at 4000 characters, notes at 2000, comments at 200 per document, and documents at 1 MiB. The store is read under a 2 MiB byte bound and fully validated — schema, document identity, every comment field, and the verdict, including its
authority, which a stored file can therefore never use to promote itself to sign-off, and every hash, which must be a lowercase SHA-256 digest rather than any bounded string. The write path enforces the same byte bound on the serialized UTF-8 payload, because the count and length bounds do not imply it: 4000 characters of multibyte text cost up to three bytes each, so 200 legal comments could otherwise produce a file the next read refuses. A write that would cross the bound is refused asstore-capacitybefore the temporary file exists, and the stored feedback is left unchanged. A store file that fails any of those checks is reported and left byte-for-byte intact, and the next mutation is refused rather than overwriting somebody's pending review; recovery is a deliberate act by the operator. The exclusive write lock records its owning process: a lock held by a live process is waited on and then refused, a lock is reclaimed only when its named owner is provably gone, the reclaim removes the entries this tool wrote rather than deleting a directory tree it does not own, and a mutation that loses ownership refuses to commit. Server lifetime is bounded by--ttl, andCtrl+C, the page's Stop server button, and the printed process id all stop it cleanly.Add revision-scoped draft recovery, retryable connection errors, an authorized-document selector, and wrapping mobile status text. A draft written against a revision or a section that is no longer current is never re-attached to new content: it is listed under Unsent drafts from an earlier revision with the section and revision it was written on, for explicit discard, and a pending verdict note survives a stale refusal. Switching documents takes a request ticket, so a slow response for one document cannot render under another document's actions. The section outline is a disclosure that starts collapsed on a narrow viewport, and the permanent keyboard tutorial line is gone — the shortcuts remain, named in tooltips and announced to assistive technology. Apply input limits at launch and reload, reject invalid UTF-8, and reject linked feedback roots before reads or writes. Portable Node test commands and desktop/mobile browser regressions cover these boundaries. The ordinary repository gate runs dependency-free Node tests when Node is available and never installs npm dependencies.
Documented in
docs/plan-review.md, with the trust analysis indocs/plan-review-threat-model.md. Rollback is deletion: nothing else in the repository depends on it. -
Add read-only
Get-CopilotAtelierClientAdapter, a thin compatibility adapter that reports how a Custom agent profile is composed for each supported Copilot client and, more importantly, what that client cannot do. The VS Code files undercom.github.copilot/agentsstay the only source of every shared workflow; the composed body is byte-identical, and only frontmatter is rewritten, so there is no second catalog to drift.Discovery is not parity, and the gap is specific. ...