Repository navigation
v5.0.0
[v5.0.0]
Removed
-
The
.github/hookssmoke-test probe, which had been failing on every turn since it was committed (2026-09-02).stop-probe.jsonandTest-HookLoaded.ps1were scratch: aStophook that appended one line to%TEMP%\workspace-hook-probe.logto prove the workspace hook location loads at all. They answered that question on 2026-08-10 and the answer is written intocom.github.copilot/hooks/README.mdand the changelog entry below — the files themselves had no further job.They were not merely idle. The
windowsoverride hardcodedD:\Git\CopilotAtelier\.github\hooks\Test-HookLoaded.ps1, the drive the repository sat on when the probe was written, and on Windows that override wins. Every turn on any other machine ended with "The argument … to the -File parameter does not exist". The POSIXcommandwas no better in principle:./.github/hooks/Test-HookLoaded.ps1is relative, and the same README says VS Code does not guarantee the working directory, which is why every shipped hook resolves its own path.Nothing caught it because nothing looked. The
Hook configurationsuite intests/Hooks.Tests.ps1— which asserts exactly this, that a hook command resolves to a script that exists and carries no shell-interpolated token — is scoped tocom.github.copilot/hooks/hooks.json. A second hook file one directory away was outside every gate the repository owns. That suite now enumerates every tracked*.jsonsitting directly inside a folder namedhooksand requires the shipped configuration to be the only one, so the next stray hook file fails the build instead of the chat. The guard was proven by planting one and watching it go red.
Added
-
Add
tools/plan-review, an optional local review surface for a Design Concept: it renders the Markdown and its Mermaid diagrams, anchors comments to stable sections, and records a verdict against one specific revision hash. It is opt-in in the strong sense — it is absent fromCustomizationDirectory, so the built module and a Gallery install never carry it; no PowerShell source references it; and its Node dependencies are installed explicitly by whoever wants the feature, never by an install, an update, or a validation run.A browser verdict is feedback, not sign-off, and the code says so rather than the documentation. An HTTP request proves that something holding the session cookie and the CSRF token posted a content hash; it proves neither identity nor authority to start implementation. Every verdict is therefore persisted with
authority: "local-http-feedback"beneath a store header ofapprovalAuthority: "chat-sign-off-required", the header is repeated on the response and shown above the document, and there is no endpoint that writes a Decision record, triggers a handoff, or runs a command. A--stateroot resolving inside.memory-bank/decisionsis refused at launch. The existing chat sign-off in the Software Architect workflow stays the only thing that authorizes implementation, and the agent body now says that where it points at the tool.Revision hashes cover the original document bytes. Comments retain section identity; ambiguous duplicate headings require one exact content match and otherwise stay unanchored. A section key is unique across the whole document, not merely per heading slug: an occurrence ordinal on its own issues
risks-2twice forRisks,Risks,Risks 2, and two sections sharing a key anchor a comment to the wrong heading. Section splitting follows the CommonMark fence rules, so a shorter fence nested inside a longer one cannot expose a fake heading and mis-anchor a comment. Verdict dialogs retain the document and revision they displayed, so a background refresh cannot approve newer content. Requests against stale hashes are refused with409, and the hash is checked a second time inside the serialized store write — a file edited while the request waits for the lock is refused rather than approved for the bytes it no longer has. The source is read once more after the commit, so a change landing in that last window is reported assupersededinstead of being presented as current approval.Loopback binding is treated as a reachability reduction, not an authorization boundary. A non-loopback bind address is refused outright, and the allowed authority follows the address actually bound, so
::1produces[::1]:<port>rather than a hard-coded127.0.0.1. Every request must carry aHostmatching the bound authority; every mutation must additionally carry the exact serverOrigin, aSec-Fetch-Siteofsame-originornonewhen the browser sends one, a JSON content type, the per-launch session cookie, and a matchingX-CSRF-Token. The session secret is generated per launch and never persisted, so a cookie minted by an earlier server is rejected by the next one even when it reuses the same feedback store. Because cookies are scoped by host and not by port, the cookie name carries a per-launch identifier: opening a second review server in the same browser no longer signs the first one out, and neither server accepts the other's cookie.Documents are authorized at launch and addressed on the wire by an opaque sixteen-character identifier, so no request parameter ever names a path. There is no directory listing, no URL fetcher, no shell endpoint, and no generic static handler — vendor assets come from an exact filename allow-list mapped onto
node_modules. Every path is realpath-resolved, required to sit inside the declared root, and rejected when any ancestor from the root down is a symbolic link or junction, and the check runs again at read time rather than only at launch, so a link swapped in afterwards still fails.Rendering disables raw HTML at the parser instead of filtering it afterwards:
markdown-itruns withhtml: false, and DOMPurify then applies a tag, attribute, and URI allow-list that admits onlyhttp,https, andmailto. An image is never fetched — its alternative text is rendered instead, because an image is an implicit external load. Mermaid runs client-side withsecurityLevel: 'strict'and its SVG is sanitized again before insertion. Responses carryContent-Security-Policy: default-src 'none'withscript-src 'self'and nounsafe-eval. The page's own stylesheet, script, and vendor bundles are snapshotted at launch and served from memory, so an asset deleted or swapped afterwards can neither change what the page runs nor leave a request hanging on a broken read.Bodies cap at 64 KiB, comment text at 4000 characters, notes at 2000, comments at 200 per document, and documents at 1 MiB. The store is read under a 2 MiB byte bound and fully validated — schema, document identity, every comment field, and the verdict, including its
authority, which a stored file can therefore never use to promote itself to sign-off, and every hash, which must be a lowercase SHA-256 digest rather than any bounded string. The write path enforces the same byte bound on the serialized UTF-8 payload, because the count and length bounds do not imply it: 4000 characters of multibyte text cost up to three bytes each, so 200 legal comments could otherwise produce a file the next read refuses. A write that would cross the bound is refused asstore-capacitybefore the temporary file exists, and the stored feedback is left unchanged. A store file that fails any of those checks is reported and left byte-for-byte intact, and the next mutation is refused rather than overwriting somebody's pending review; recovery is a deliberate act by the operator. The exclusive write lock records its owning process: a lock held by a live process is waited on and then refused, a lock is reclaimed only when its named owner is provably gone, the reclaim removes the entries this tool wrote rather than deleting a directory tree it does not own, and a mutation that loses ownership refuses to commit. Server lifetime is bounded by--ttl, andCtrl+C, the page's Stop server button, and the printed process id all stop it cleanly.Add revision-scoped draft recovery, retryable connection errors, an authorized-document selector, and wrapping mobile status text. A draft written against a revision or a section that is no longer current is never re-attached to new content: it is listed under Unsent drafts from an earlier revision with the section and revision it was written on, for explicit discard, and a pending verdict note survives a stale refusal. Switching documents takes a request ticket, so a slow response for one document cannot render under another document's actions. The section outline is a disclosure that starts collapsed on a narrow viewport, and the permanent keyboard tutorial line is gone — the shortcuts remain, named in tooltips and announced to assistive technology. Apply input limits at launch and reload, reject invalid UTF-8, and reject linked feedback roots before reads or writes. Portable Node test commands and desktop/mobile browser regressions cover these boundaries. The ordinary repository gate runs dependency-free Node tests when Node is available and never installs npm dependencies.
Documented in
docs/plan-review.md, with the trust analysis indocs/plan-review-threat-model.md. Rollback is deletion: nothing else in the repository depends on it. -
Add read-only
Get-CopilotAtelierClientAdapter, a thin compatibility adapter that reports how a Custom agent profile is composed for each supported Copilot client and, more importantly, what that client cannot do. The VS Code files undercom.github.copilot/agentsstay the only source of every shared workflow; the composed body is byte-identical, and only frontmatter is rewritten, so there is no second catalog to drift.Discovery is not parity, and the gap is specific. The published custom agents configuration that the Copilot CLI follows documents one model string rather than a priority array, a closed set of tool aliases rather than product-qualified tool identifiers, and no subagent allow-list, handoff, or argument hint — and it ignores an unrecognized tool name, so a profile that loads there can quietly lose the capabilities its own body depends on. The representative profile alone declares dozens of identifiers with no client equivalent. The scope is VS Code Copilot Chat and the Copilot CLI; no other client was checked, and none is claimed.
Four rules are enforced by tests rather than promised in prose. Every mapping is explicit, so an identifier absent from the allow-list is an error rather than a silent drop, and frontmatter is parsed as a strict YAML subset that rejects an unknown top-level field, a duplicate field, a block scalar, an anchor, an alias, a tag, an unterminated list, or an unbalanced quote — a tool list is never partially mapped, and a new safety field cannot disappear without a diagnostic. Nothing is widened to make a workflow run: only a tool that genuinely starts a command may become the shell-execution alias, and the contract names those explicitly, because a product prefix such as
execute/is a namespace rather than proof of execution authority — reading an existing terminal buffer, running a declared VS Code task, and running tests all stay unsupported instead of being traded for a terminal. Every other mapping has to stay inside the capability class of its source identifier. And a restriction that cannot be expressed removes what it guards — the subagent allow-list has no counterpart, so the composed variant loses the delegation tool instead of inheriting unbounded delegation, and the model field is omitted rather than translated into an invented client model identifier. A capability the shared body declares mandatory that cannot be mapped fails the composition instead of shipping without it.A workflow the client cannot run is refused rather than degraded, and the refusal travels inside the file. The shared engineering body offers
review: onandcycle: full, both satisfied in VS Code by dispatching thesecurity-reviewersubagent and by advancing through handoffs, and the Copilot CLI provides neither. The composition therefore prepends an additive client-limitation section naming both modes unavailable and instructing the agent to refuse the request and return it to VS Code Copilot Chat; a required independent review is never quietly downgraded to a written recommendation. The section is composed presentation, not an edit to the shared body, and its boundary says so: explicit begin and end markers, the SHA-256 of the shared body carried in the end marker, and the authoritative body last in the file.Get-CopilotAtelierClientAdapter -RequiredWorkflowthrows instead of returning content when a caller depends on a mode the client cannot honour.The rollout is one profile wide: only
software-engineeris adapted, and any other profile is refused until it has its own passing compatibility test. Neither client is reported as runtime verified, because no client session backs these mappings: both areStructurallyCheckedagainst current documentation. The Copilot CLI is not installed here, and the observation that the source profile loaded in VS Code 1.136.1 is kept as historical source-profile evidence about that file rather than as a property of an arbitrary composed variant.The composed variants are a build artifact written to
output/clientAdapters/<client>/by the newBuild_Client_Adapter_Variantstask, and they are deliberately not deployed — not in the module payload, not in the built module, not written to the canonical target, and not published through the plugin channel. Two profiles for one agent inside~/.copilot/agents, which VS Code and the Copilot CLI both read, would be a duplicate discovery entry rather than a compatibility fix. Tests check the artifact for staleness byte for byte, reject a leftover variant for a profile no longer adapted, and assert the packaging isolation. The build task owns that directory rather than sweeping it, and the bound is three separate properties rather than one marker file. The whole operation — every variant, the manifest schema, its list shape, every entry, duplicate manifest paths, duplicate variant destinations, and every collision — is constructed and validated before the first delete, so a request that is going to be refused leaves the directory exactly as it was found; validating and deleting one entry at a time would already have destroyed a valid first entry by the time a later unsafe one was caught. Every path component is checked by the shared regular-path guard before it is read, deleted, created, written, or enumerated: the output root, the artifact directory, the ownership manifest, each client directory, each generated file, and each destination, because a link in the middle of the path redirects a delete and a write just as effectively as one at either end. And ownership is proved by content, not by a file name:.copilot-atelier-adapter-manifest.jsonrecords the SHA-256 of every file it generated, so a generated file edited in place is refused rather than silently deleted or overwritten, and a file at a destination this build never generated is refused rather than adopted — whatever it contains, because an identical unowned file is the same silent ownership grab as a different one. A names-onlyschema 1manifest cannot prove any of that, so it is refused with the paths it claims and a migration path rather than adopting the hashes of whatever now sits there. A directory without the marker, a reserved build directory such asmoduleorRequiredModules, and a path that is not a direct child of the build output stay refused — a mistypedClientAdapterSubdirectorycan no longer take unrelated build output with it. Authored behavioral evaluation cases are indocs/client-adapter-evals.md; none has been executed, because every one of them needs a model-backed client session. -
Add the
changed-file-validationSkill: an opt-in, bounded validation pass over the files one work batch actually changed. Collection is manual and explicit —Add-ChangedFile.ps1records the paths it is given, deduplicates repeated edits onto one entry with an occurrence count, and normalizes absolute, backslash, and forward-slash forms onto the same project-relative entry. Batches are isolated by session identifier, and the identifier is restricted to letters, digits, period, underscore, and hyphen so it can never name a path. Concurrent collectors merge rather than overwrite: a writer re-reads the store under a bounded lock before it writes, so two sessions collecting at the same moment cannot lose a file.No hook is wired, and the reason is recorded rather than assumed. Of the documented client events,
PostToolUseis the only plausible collector and its edit-tool input contract is not verified for this implementation. The shipped hooks are also mandatory —PreToolUseblocks remote mutation andStopcloses the session clock — so an optional collector inside either one would turn a validation fault into a guard fault and giveStopa way to fail and be retried. The hook configuration is therefore unchanged, still declares exactlyPreToolUse,SessionStart,Stop, andPreCompact, and carries no reference to changed-file collection; a regression asserts that, that no hook script gained the dependency, and that the remote-mutation guard still exits 2 on a push.Invoke-ChangedFileValidation.ps1is the explicit entry point, and it reuses what the repository already has rather than inventing a build system. PowerShell parsing throughParser::ParseFileand PSScriptAnalyzer both run in an owned child worker with a wall clock, because an in-process validator has no wall clock at all and a pathological file or a wedged analyzer would take the session with it; the worker receives its request over standard input, never dot-sources or otherwise runs the file it checks, and is handed an inline analyzer settings hashtable so no projectPSScriptAnalyzerSettings.psd1and no custom rule module is loaded. One validation run executes per session at a time.Markdown is checked by a real markdownlint or not at all.
Markdown.NativeStructureimplements four rules — MD047, unterminated frontmatter, an unterminated code fence, and invalid UTF-8 or a stray control character — and is recorded withcoverage=partial, because.markdownlint.jsoncnever setsdefault: falseand therefore leaves most markdownlint rules enabled and uncovered here.Markdown.Lintis consequently always part of the plan for a markdown file: without a linter it isUnavailableand the entry stays unverified, rather than being verified by a native check that is not markdownlint. Only themarkdownlint-cliinterface is driven, and only after--versionanswers with a version;markdownlint-cli2is reportedUnsupportedInterfacerather than guessed at. The declarative configuration is copied in beside the snapshot and passed explicitly, which is also what disables the linter's nested and ancestor discovery, and a.js,.cjs,.mjs, ormarkdownlint-cli2configuration anywhere from the file's directory up to the project root makes the check unavailable instead of being handed to a linter that would execute it. What the chosen declarative configuration contains is decided by parsing it, not by scanning its bytes: a text scan is not a boundary, because JSON can spellextendswith Unicode escapes. JSON and JSONC are read by a non-executing parser and checked against a conservative rule-map schema, so a dynamic include, a custom rule, a module path at any depth, an unknown key, and an unexpected shape are all refused before the version probe or the linter starts. A format with no trusted parser here — YAML, and JSONC on a host without a comment-tolerant JSON reader, which includes Windows PowerShell 5.1 — is reportedConfigurationFormatUnsupportedrather than copied to an executable unread.A receipt is a claim about exact bytes checked under an exact plan. Validators read an isolated snapshot: the bytes are copied into a generated directory under a generated, metacharacter-free name, and that copy is what is hashed and checked, so the receipt is bound to the byte sequence a validator actually read rather than to a path that may have moved underneath it. Alongside the per-check validator, executable, version, configuration identity, outcome, exit status, and up to twenty located diagnostics, the receipt records a validation plan identity — a SHA-256 over the checks the file is due, resolved from the extension,
-FailOnSeverity, the installed PSScriptAnalyzer version, the SHA-256 of the linter entry point, the bytes of the declarative markdown configuration, and the SHA-256 of the shipped code that performs each check. Hashing the entry point rather than measuring it is what catches a linter replaced in place at the same length with its timestamp preserved, and hashing the shipped worker and helpers is what stops a receipt outliving an edit to the checker itself. That identity is bounded and says so: it does not reach PSScriptAnalyzer's rule implementations beyond their module version, and it does not reach the dependency tree under an npm wrapper, because no supported interface exposes one to a process-free read. Reuse requires all of it to hold: intact shape, results and fields this implementation actually writes, a Boolean change flag by type rather than by coercion, a hash matching the bytes on disk, no change flagged during the run, an identical plan identity, exactly the planned checks with no duplicate and no extra, each carrying the identity the plan names and an exit status that agrees with its result, and a recorded outcome equal to what those checks add up to. A result this implementation cannot produce is treated as unavailable rather than counted towards a pass. Changing-FailOnSeverity, selecting, upgrading, or replacing a linter, upgrading PSScriptAnalyzer, editing the checker, editing.markdownlint.jsonc, and hand-editing the store all invalidate reuse instead of inheriting a pass, andGet-ChangedFileBatch.ps1applies the same rule without starting a process, because nothing in the plan identity needs one. Deleted, renamed, and unsupported files are reported asMissingandUnsupportedand are never verified.Every bound is explicit and every failure is fail-closed. The input-size bound is enforced before any content is read or hashed. External output is drained and capped while the child is still running rather than read to end afterwards, so a talkative or hostile tool cannot grow the retained buffer without limit and neither stream can deadlock on the other; a child that fails to start, one that outruns the bound, and one whose descendants hold the pipe open past the drain grace are reported as
Unavailable,TimedOut, or incomplete, with the exit status left null when it is genuinely unknown, and none of them can produce a pass. A file name is data: because every argument is a name this workflow generated and the working directory is the snapshot directory, no project text reaches a command line at all — which matters on Windows, where a markdownlint entry point is a.cmdshim the command processor re-parses. Path containment is enforced by walking every existing directory from the selected project root down to the file before each read and each write, and again at validation time because a batch collected minutes ago may have gained a link since. A timed-out child is stopped together with the descendants it spawned and nothing else.Nothing outside the batch store is written: no file it validates and no file it was not given is rewritten, the report writes nothing at all, and
Test-CopilotAtelierruns no validator from here. The store lives at.copilot-atelier/changed-file-validation/batches.json, is replaced atomically under the lock, and refuses a store recording another project root, an unsupported schema version, or unparsable JSON. The Skill documents in its own body that it supplements immediate behavior-scoped tests and the full build gate rather than delaying or replacing either.evals/validation-cases.jsoncarries offline cases only, labelledauthoredwithexecutedset to false, because no model-backed sweep has been run for it. The Skill joins theengineeringinstallation profile. -
Add read-only
Get-CopilotAtelierSkillHealth, an on-demand Skill maintenance report
that combines trustworthy usage observations, the evaluation artifacts that already exist, and structural checks, and suggests what a human should investigate, improve, consolidate, or review for retirement without ever changing a Skill, a setting, or an installation. The supported client event contract was established before anything was built: the documented hook events areSessionStart,UserPromptSubmit,PreToolUse,PostToolUse,PreCompact,SubagentStart,SubagentStop, andStop, and not one of them reports that a Skill was selected, loaded, or executed. There is therefore no supported event to observe a Skill activation from, no automatic collection is implemented, capture is disabled by default, and the report names that gap in its output and its help rather than papering over it. No session log, transcript, or cross-workspace history is read.Usage evidence reaches the report only through
-ObservationPath, which takes files or directories the caller selected explicitly. A document is validated against a closed schema —schemaVersion,client,clientVersion,trust,records, and an optionalcoverageblock at the top level andeventId,eventType,skillName,skillSha256,timestampUtc,sessionId,outcomeper record — with bounded file size, record count, and field lengths, a strict UTC timestamp inside a plausible range, and a 64-character SHA-256 content identity.recordsmust be a JSON array, so a scalar or a bare object is refused rather than silently wrapped. An unsupported property, an unsupported schema version, an unsupported event type or outcome label, or an unparsable document is refused rather than partially read, and a rejection names the field and the rule and never echoes the offending value, so no imported payload or secret can travel back out through an error. Prompt text, Skill bodies, tool responses, and free-form notes are not in the schema at all. Nothing is stored, and nothing is uploaded.Provenance survives the import. An event identifier is unique only inside the client and session that minted it, so records are scoped by client, client version, session, and event identifier together, and the same local identifier from two clients no longer collides. Exact copies deduplicate onto one accepted record; records sharing an identity scope that disagree on their payload are reported as a conflict listing the import locator of every variant, and none of them is accepted, because accepting one would mean accepting whichever file happened to be read first. Every accepted record keeps its source path and declaring client, and the accepted set is order-independent. Enumeration is bounded and guarded during the walk: every directory is checked before it is descended into, the file bound is enforced as files are collected rather than after the whole tree has been listed, and an explicitly selected file is guarded against its own parent directory so an ancestor link cannot redirect the read.
The facets stay apart. An observed file read, a Skill activation, a tool execution outcome, demonstrated quality, discoverability, freshness, and description overlap are counted and reported separately, and no facet is collapsed into an invented single health score. A successful load or read is never reported as a passed capability evaluation. Trigger-query sets prove that discovery material was authored, never that it was measured; the command runs no model and adds no second evaluation engine. Every record is bound to the SHA-256 of the body it names, so records written against a different body are counted and labelled separately instead of merged into the current one, and the report exposes the deduplicated event count, the duplicate count, the conflict count, and the observation time window.
Evaluation evidence is read in the shapes the
agent-evalsSkill actually defines rather than in an invented one.evals.jsonis authored input in either the Agent Skills shape or the bundled harness shape and proves only that cases were written.grading.jsonandbenchmark.jsonare run output, and neither carries a Skill identity of its own, so a run describes the current body only when a validated sidecar named after the artifact with its extension replaced —grading.provenance.jsonbesidegrading.json— names this Skill, the SHA-256 of its current body, and a real completion instant no later than the reference instant. A run bound to a different body is labelledDifferentBodyand excluded, and a run whose provenance is missing, malformed, or without a plausible completion instant stays visible asUnboundand is never counted. A run identifier is a quality identity and is consumed only by a graded result that could be counted, so a benchmark arm sharing it costs nothing; copies that agree are counted once asDuplicateRun, and copies that disagree are labelledConflictingRunand none of them is counted. A graded run is its assertion list and the summary is a claim about it: the claim counts only whenpassed,failed, andtotalreconcile exactly with the bounded Boolean verdicts recorded inassertion_results, so a summary with no assertions, a contradicted verdict, an ungraded case, a non-Boolean verdict, and a count that is missing, negative, non-numeric, or outside the supported range are all reported and none becomes a pass. Assertion text and evidence are never read out of an artifact. Every file the report opens is size-bounded, the assertion list is count-bounded, and an oversize artifact is reported as such instead of being read.The supported client event contract is verified rather than asserted. Each enumerated hook event is checked against the deployed authoring Instruction under the inspected content root, the report publishes the resulting verification state and the scope of the check, and the claim is scoped to this implementation — no reliable Skill-activation contract verified for this implementation — instead of asserting that no client anywhere can identify a Skill. A frontmatter fence is likewise not treated as proof of valid metadata: the block is reported as parsed only when every line is a supported top-level key and both
nameanddescriptionresolve to a scalar, and anything else is reported conservatively as unsupported.Missing telemetry means unknown, never zero use, and an imported
SkillFileReadis evidence of a read rather than of an activation. Every record is reported asImported— a document claiming observed trust is recorded as a claim and still read as imported, because no capture this module performs exists. Suggestions are advisory: each cites local evidence and carriesDecision = 'HumanReviewRequired'. ARetirementReviewis never inferred from silence: it is raised only when an import declares explicitly, for one Skill and one body, that activation capture was complete over a window of at least 30 days and 20 sessions that closed within the staleness horizon, and that window recorded no activation. It is never raised for a mandatory Skill and never from a window belonging to a different Skill. -
Add the
reviewed-learning-inboxSkill: an on-demand, project-scoped review queue that turns an explicitly selected local correction into a reviewable suggestion for an existing Skill or Instruction, and never into policy on its own. A candidate carries a stable project-scoped identifier, a sanitized single-line lesson, project-relative evidence locators with their SHA-256 content identity, and separated observations, interpretations, and contradictory evidence. Equivalent lessons deduplicate onto one record and increment an occurrence count, rejections and supersessions are retained so a repeated observation cannot resurrect a decision the user already made, and the confidence field is stored with the note that it is a review label rather than a probability or permission to apply. The store lives at.memory-bank/learning-inbox/candidates.json, outside the routed Memory Bank base, outside every Skill description, and outside the deployed Customization tree, so an unreviewed entry never reaches a trusted context surface.Promotion is append-only, previewed, and hash-gated.
New-LearningPromotionProposal.ps1renders the exact lines a promotion would add together with the destination's current SHA-256 and writes nothing to any Customization;Invoke-LearningPromotion.ps1refuses the change unless the caller passes-Approvewith the SHA-256 of the preview that was actually read. Before the first byte is written it also refuses a proposal from another project, a destination that contradicts the candidate scope, a destination outside an existingSKILL.mdor*.instructions.mdor outside the project, a proposal whose evidence list differs from the authoritative candidate record, and a candidate record, evidence file, destination, or rendered preview that changed after the preview was produced. Every path is checked by walking each existing directory from the selected project root down to the file, so a junction or symbolic link on an intermediate directory cannot redirect a read, a hash, or a write outside the project — the leaf-only check that shipped first could be bypassed that way.The apply step opens one write-exclusive handle, hashes the bytes it is about to append to, and appends through that same handle, so an edit made between the review and the write is refused instead of overwritten and the original byte prefix — including a UTF-8 byte order mark — survives exactly. A destination that is not UTF-8 with or without a byte order mark is refused rather than silently re-encoded, and a failed verification truncates back to the original length. The store is replaced atomically under a bounded, fail-fast per-inbox mutation lock and refuses to overwrite a store that changed after it was read, so two concurrent writers cannot lose a candidate. Repeat application is bound to the block that is actually present rather than to a marker string: a block that matches the candidate record and a recorded promotion reports
AlreadyPromoted, an apply that appended the block but never recorded it reportsReconciliationRequiredand thenReconciledunder the same human approval without touching the destination again, and a forged or edited block is refused with the recovery step named.Text taken from a selected artifact is treated as an untrusted observation at both ends of the pipeline. Intake and promotion apply the same rule: single-line printable prose only, so code fences, shell substitution, markup, control characters, and capability or scope keys such as
tools,model,agents, andapplyToare refused, and hand-editing the store widens nothing. That allow-list is format validation and is documented as such: it is not an injection-proof or secret-redaction boundary, explicitly sanitized input and human review remain necessary, and-Approveplus a preview hash is an approval protocol rather than proof of human identity, which no automatic agent may supply without a current human instruction. OnlySkillandInstructionscopes are promotable at all;NewSkill,NewInstruction, andNewAgentare suggestions for a human and require a written overlap explanation before they can even be recorded.evals/candidate-cases.jsoncarries offline cases only, including one real local correction represented by repository locators and line ranges rather than any transcript, and is labelledauthoredwithexecutedset to false because no model-backed sweep has been run for it. -
Add opt-in installation profiles.
Install-CopilotAtelier,Update-CopilotAtelier, andSetup-CopilotSettings.ps1accept-InstallationProfilewithcomplete(the unchanged no-argument default),engineering,research, anddocument-processing, plus-IncludeSkilland-ExcludeSkillfor adjusting a selection by Skill identifier. Only Skills are selectable: Agents, Instructions, Prompts, and Hooks always deploy in full, andmemory-bank,long-running-job-monitor, andagent-security-reviewstay in every selection because the deployed Instructions and shipped Custom agents load them by name. A selected Skill ships its whole folder, and the Skills it hands part of its workflow to come with it. Unknown identifiers, a Skill both included and excluded, an excluded mandatory Skill, an excluded dependency of a selected Skill, and a cyclic dependency catalog are all rejected before any directory, Discovery link, setting, or Deployment record is touched. A narrowing request is also checked against the payload it narrows: a mandatory Skill the payload does not ship, a selected Skill whose required Skill is absent, and an explicitly selected directory without aSKILL.mdentry point are refused with the dependent and the absent target named, rather than resolving to an incomplete selection. The complete installation keeps deploying a payload exactly as it is and reportsPrerequisiteValidatedasFalseinstead of claiming it validated one. The selection is recorded as an additive optionalSelectionfield in the schema-1 Deployment record, whose shape is validated strictly — arrays rather than scalars, no duplicate or contradictory identifiers — while an absentSelectionstays valid, so records written before profiles existed still read as the complete installation, and an argument-free reinstall or update keeps the recorded selection instead of silently re-expanding it. An inherited selection is re-read and re-resolved once the run holds the local deployment lock, so a concurrent local installer that changed it is followed rather than overwritten from a stale read. Ownership stays with the recordedFileslist alone; a claimedSelectionnever confers it. Switching profiles retires only unchanged Owned files: user-added files stay, and a locally changed file stops the switch with its path named. Profiles apply to the module and clone paths; the native Agent Plugins channel installs the whole package and has no selection mechanism. See choosing what gets installed. -
Add read-only
Get-CopilotAtelierProfile, which resolves every installation profile against a payload and reports the Skills, mandatory Skills, prerequisite-validation status, and counts each one deploys without changing anything.Test-CopilotAteliernow reports the deployedInstallationProfileand fails health when a record excludes a mandatory Skill. -
Add read-only
Get-CopilotAtelierFootprint, an on-demand report of the customization collection's loading footprint with concrete opportunities to reduce unnecessary loading. It reuses the shared deployment directory map and recognizes actual Customization file types: only an explicit broadapplyToand Skill discovery metadata are reported as potential automatic loading, contingent on discovery and client applicability, while missing, malformed, ambiguous, or unsupported scope metadata and unselected Custom agents stay unknown applicability rather than always-loaded, and ancillary documents, scripts, and binary assets are reported as disk footprint rather than model context. Ambiguous frontmatter — an unmatched quote or a duplicate key — is rejected as unsupported rather than trusted. It enforces every mapped root against the selected content root, so a junction at the selected root, at an intermediate namespace folder, or a mapped path that escapes via traversal is reported and never followed, and it labels every byte figure as a file-size estimate rather than measured session loading, proven activation, or duplicate runtime injection. -
Add explicit
-Repairfor modified Owned files and-TargetPathselection through Install, Update, and Setup, with preview support and untracked-file preservation. See repair and recovery. -
Add read-only
Test-CopilotAtelierdeployment diagnostics and hash-awareUninstall-CopilotAtelier, with explicit targets, non-interactive account resolution, and preservation of user content and configuration. See deployment diagnostics and removal. -
Record per-file SHA-256 ownership in the Deployment record and validate paths and metadata before deployment or removal.
-
Bound SessionStart context to 4096 characters by default, configurable from 1024 through 16384 without disabling lifecycle or security hooks.
-
Gate Customization tool bounds, delegation, Prompt overrides, hook timeouts, and remote authorization with adversarial fixtures and a shrink-only MCP baseline. Extend the existing authoring guide with implementation selection and evidence-based pattern promotion.
-
A repository-scoped migration for legacy career, legal, and tax Memory
Bank records (2026-09-04). Thememory-bankSkill now separates planning
from applying: it inventories only direct children of one selected
.memory-bank/, classifies known legacy names, requires explicit decisions
for ambiguous files, saves a metadata-only plan, and previews with-WhatIf.
Apply validates the complete plan before writing, rejects changed sources,
conflicts, path escapes, cross-repository plans, and reparse points, then
performs byte-exact, SHA-256-verified copies without overwriting or deleting
any source. Career Coach, Legal Researcher, and Tax Researcher invoke this
workflow before creating empty namespaced replacements. -
Interactive browser access for every web-capable Custom agent
(2026-09-04). Replace the remaining preview-onlyopenSimpleBrowserentries
with VS Code's built-inbrowsertool while keeping the Contoso profile
browser-free. -
Semantic Custom agent contract tests (2026-09-04). Validate every
handoff target, required delegation surface, DevOps composition contract,
role-specific Memory Bank namespace, browser tool, cross-client README
caveat, and the 30,000-character prompt limit with a shrink-only baseline for
the four existing oversized agents. -
Interactive web-application troubleshooting in the Software Engineer agent (2026-09-04). Replace the preview-only
openSimpleBrowserentry with VS Code's built-inbrowsertool set so the agent can navigate and exercise its product, inspect page content, console errors, and screenshots, verify affected desktop and mobile viewports, fix defects, and repeat the original flow. Browser checks default to agent-opened ephemeral sessions on loopback origins; authenticated state is available only when the user explicitly shares a tab. Session evidence complements rather than replaces repository regression tests, and the Contoso overlay continues to omit browser access. -
A Prompt-led specification completion workflow for any spec-driven
repository (2026-09-02)./complete-specificationsinventories acceptance
criteria, milestone exits, Decision gates, local gaps, test evidence, and an
optional local issue snapshot into a frozen closure matrix. A restricted
controller creates one isolated work item per missing behavior, dispatches one
test-first implementer for each, and sends every result to a fresh read-only
reviewer before integration.Four percentages prevent "implemented" from silently meaning "proven":
implementation, passing unit and integration tests, live verification, and
total specification closure each keep their own numerator and denominator.
Every non-duplicate engineering row stays in the primary denominator. Live
validation defaults to off; enabling it requires a direct containment-profile
digest and a hash-pinned, data-isolated live runner. A pinned append-only
appender hash-chains review and accounting records outside repository
processes' writable roots, and changed build commands cannot run until their
control review passes. Shared and production mutation is prepared as a
supervised procedure and never counted as live evidence. The package has no
web, issue-mutation, or push path, caps work items, Custom agent calls,
concurrency, fix rounds, and run time, and leaves every branch local. -
A session clock, so Post-flight closes with the chat's measured elapsed duration (2026-09-02). The user asked for two more facts at the end of the checklist: when the turn closed, and how long the whole chat had run. The first half already existed —
com.github.copilot/hooks/scripts/Add-SessionContext.ps1injectsSession started at <UTC>and Pre-flight opens every reply with it — which made this look like a formatting change.It is not, because a model has no clock. The opening timestamp is right only because a hook measured it; a closing one composed by the model would drift by the length of the turn, which is the very quantity being reported, and after a compaction the model no longer knows when the session began. The obvious fix is unavailable: VS Code's
UserPromptSubmitsupports the common output format only, with noadditionalContextfield — the same limitation already documented forPreCompact. The events that can inject context areSessionStart, which fires once, andPreToolUse/PostToolUse, which would spend tokens on every tool call of every turn and fold a timing concern into the security guardrail.So the number is measured on disk and read back by the one party that can print it inside the reply.
Add-SessionContextnow also writes the session start to<LocalApplicationData>/CopilotAtelier/sessions/session-<key>.json— on disk, so it survives compaction — and hands the agent the absolute path of a new reader,com.github.copilot/hooks/scripts/Get-SessionElapsed.ps1. The agent runs that reader as the last action of the turn and copies its single line verbatim into the checklist:POST-FLIGHT elapsed: 16m (started 09:15 UTC, measured 09:31 UTC, turn 3).com.github.copilot/rules/postflight.instructions.mdgains a Session clock section forbidding the model from composing, reformatting, or recomputing either number, and telling it to report the duration as unavailable rather than estimate one when the reader is gone.The first attempt printed the line from the
Stophook, and was wrong in a way only a screenshot revealed: VS Code renders a hooksystemMessageas a detached, collapsed Warning from Stop hook box, not as part of the reply. The number was therefore beside the checklist rather than in it, and the user asked again. A hook cannot write inside the model's output and the model cannot read a clock, so the shipped split is the only arrangement that satisfies both.com.github.copilot/hooks/scripts/Write-SessionClose.ps1stays, because the turn counter still has to advance somewhere, but it now reports nothing unless the clock is unreadable — the one case where the agent's own line could not be measured either. Reporting the duration there as well would only have put a second copy in the warning box on every turn.Stopfires once per turn instead of once per tool call and costs no tokens. It emits nodecisionfield: blocking aStoprestarts the agent and bills another turn, which is far more than a timestamp is worth. The clock avoids the temp directory because/tmpis world-writable on Linux, where a predictable name invites another local account to pre-create the path, and avoids.memory-bank/because — unlike a compaction checkpoint — it is machinery rather than knowledge an agent reads, and it has to work in a workspace with no Memory Bank at all. The payload'ssession_idbecomes a path component, so it is stripped to[A-Za-z0-9._-]and capped at 64 characters, with a hash of the working directory as the fallback. The reader is read-only —Stopownsturns, so it reports the turn in progress as one past the closed count — and given no explicit path it picks the newest clock recorded for the current workspace, so a second VS Code window on another folder is never measured here.Deploying the reader exposed a defect the suite had been creating all along. Six
Add-SessionContexttests invoked the hook without-ClockRoot, so every run left real session clocks in the caller's own%LOCALAPPDATA%\CopilotAtelier\sessions— around fifteen of them, including one whose recorded workspace wasC:\demo IGNORE PREVIOUS INSTRUCTIONS, written by the prompt-injection test. That was invisible while only theStophook read the clock, because it looks its own session up by id. The reader searches by workspace, so a clock the tests had written for this repository immediately shadowed the live session and reported a three-minute chat that had been running for an hour and three quarters. Every SessionStart invocation now goes through a helper that pins the clock root toTestDrive, a test asserts the real profile directory gains nothing, and the reader prefers asession-<id>clock over asession-cwd-<hash>fallback — VS Code always supplies a session id, so the hashed name in practice means a test or a non-VS-Code caller.The duration formatter shipped with a bug the tests caught:
[int]1.5rounds in PowerShell, so a 90-minute chat reported2h 30m. It floors explicitly now, andtests/Hooks.Tests.ps1pins five durations that sit where rounding and truncation disagree, alongside the injected reader path, the single-line output contract, the turn-in-progress arithmetic, the workspace preference, the shadowing regression, the clock-root containment, the read-only guarantee, the turn counter, thestop_hook_activecontinuation case, a corrupt clock, an unreadable payload, the absentdecisionfield, and asession_idof../../pwned.
Changed
- Extend the Sampler Skills with version-scoped wiki commit-timeout diagnosis, supported publication-runner rationale, and destination-by-destination recovery after partial publication; retain the real incident as a transcript-graded regression case without claiming behavioral improvement from structural checks or completed requests that did not load the Skills.
- Reject a source tree that overlaps the Canonical target. A clone kept at
~/OneDrive/CopilotAtelier/, the location earlier documentation suggested, now fails before any write; move it aside — for example to~/OneDrive/CopilotAtelier-src/— and reinstall. See repository clone.
Fixed
-
Fix Windows OneDrive detection so a generic
OneDrivevariable or pre-created folder does not select a sync target without account-specific configuration; preserve macOS/Linux discovery and explicit-TargetPathselection. See target selection. -
Preserve the hidden client-adapter ownership manifest in GitHub Actions build artifacts so downstream jobs can verify generated files (CI run #75).
-
Use canonical temporary directories in plan-review filesystem tests on Windows and macOS, with a linked-directory regression, without weakening containment or link rejection (CI run #75).
-
Require the plan-review heading verifier at every document read, and fail the repository gate when a read stops being verified, so the check that keeps a comment anchored to the text the reader actually saw cannot regress unnoticed. An omitted verifier now raises instead of reading the document unchecked. Refuse a request body that nests deeper than the walk bound rather than leaving its deepest keys uninspected.
-
Align local plan-review section anchors with rendered indented ATX and setext
headings, preserve literal trailing hashes, and refuse unsupported structures
before feedback writes, including queued mutations. Report startup failures
and verdict actions without a loaded revision cleanly, and exercise mutation
guard composition with behavioral regressions. See the
plan-review guide. -
Validate payload entries through the shared path guard so Windows Cloud Files placeholders are accepted while redirecting links remain rejected (CI run #72).
-
Fix uninstall failures on dangling Linux and macOS Discovery links and preserve literal POSIX filenames in validation fixtures (CI run #72).
-
Fail deployment health for modified hook files and redirected event commands, including platform overrides and
-Quiet; keep ordinary modified-file warnings distinct. -
Recover interrupted file applies through atomic replacements and per-file Deployment-record checkpoints, including retries with different or older payloads; coordinate local install/removal without claiming a filesystem or cloud-sync transaction.
-
Validate portable path segments consistently before planning and when reading records, reject
.and..explicitly, use target-native filename identity, and diagnose retained capitalized legacy trees without removing them. -
Restore result-serialization type data after successful and failed builds, parse hook JSON with
ConvertFrom-Json, and check built-in and implicit Prompt tool overrides without claiming runtime containment. -
Preserve user-added files and legacy trees during reinstall; reject locally modified files, reparse-point paths, and intervening changes instead of rebuilding deployment directories destructively. Reconcile wanted edits before reinstalling or use explicit
-Repairfor recorded files;-Forcedoes not overwrite them. -
Quote the usage Prompt argument hint so its colon parses as YAML; validate every Prompt with a full YAML parser in the configuration security gate.
-
Pin Pester 5.7.1 and reject unsupported major versions in QA. Bound filesystem references in saved test reports to paths and labels so result export does not traverse live provider and assembly metadata; retain all test counts, failures, and coverage.
-
Repair invalid Custom agent orchestration contracts (2026-09-04). Give
Security Reviewer and Technical Writer an executableresearch-analyst
delegation path, replace the DevOps writer's fictitious inheritance with
explicit composition, and remove the Research Analyst pseudo-handoff that
had no target Custom agent. -
A twenty-minute Pester suite ran five times with
long-running-job-monitorunloaded, because nothing in context said it existed (2026-09-04). InC:\git\RdsFarmManageran agent ran./test.ps1five times and monitored none of them. The cause is a loading mechanism, not a wording problem: Instructions auto-apply byapplyToglob, Skills load only when the model matches theirdescriptionagainst the conversation, andpowershell-execution-safety.instructions.mdwas in context the whole time while the Skill was not. The pointer between them did exist — twice — but never where it would have changed anything: once as "or applylong-running-job-monitorwhen ongoing progress reporting is required", a judgement call attached to the anti-polling bullet, and once under Indefinite processes, a section an agent has already left because a test run is not a daemon.The nuance the earlier 2026-09-01 fix missed is that the run was agent-initiated. The user asked for a code change, never for a test run; "live test", "is it stuck", and "keep me posted" were never typed, so there was no user phrasing for a
descriptionto match. Description matching cannot fire on the agent's own decision to start a long command, which makes the auto-loaded Instruction the only reliable carrier.The Instruction now leads with the launch decision instead of burying it.
Running Tests & BuildsbecomesLong-Running Commands — Detach AND Monitor, Never Direct Executionand opens with a trigger that needs no judgement — anything expected to run past roughly two minutes, and unconditionallyInvoke-Pester,Invoke-Build,build.ps1,test.ps1, any other test or build entry point, any installer, and any deployment entry point — stated to fire on an agent-initiated run with no user request. Detaching and monitoring are stated as one obligation rather than two that can be satisfied separately, because a correctly detached run with noSTARTline, no phase lines, no terminal marker, and no per-turn status line was exactly what happened. The four anti-patterns the session produced are named in two lines each and nowhere else:Select-Object -Lastbuffers the whole stream so every progress check returns the same frozen snapshot; editing source during a verification run scores a stale artifact because build output and Pester discovery are fixed at launch;Tee-Objectoverwrites content while NTFS keepsCreationTime, so file metadata is not elapsed time; and a script running inside the terminal's ownpwshnever appears in a process command line, so liveness cannot be guessed from one. The detach rule itself is unchanged — it was already correct and was already ignored.The Skill keeps the workflow and gains the vocabulary that was actually in play — "run the test suite", "full suite", "Pester run", "Invoke-Pester", "build.ps1", "test.ps1", "verification run", "regression run" — plus "a run the agent starts itself with no user request" in the summary, at 989 characters against the 1000-character soft cap.
DO NOT USE FORgainssampler-build-debugandpester-patternsso the new build-and-test terms cannot buy positives by over-triggering.Measured rather than asserted.
trigger-queries.long-running-job-monitor.jsonholds twelve labelled cases taken from the session itself, and the Skill leaves the uncovered baseline intests/SkillTriggerCoverage.Tests.ps1. Paired arms against the same 47-skill catalogue withclaude-haiku-4.5judging in a fresh context per call: train 5/7 → 6/7, validation 4/5 → 4/5, false positives 0 in both. "Run the full test suite" went 0.33 → 1.00 and the mid-flight unrelated question 0.00 → 0.33. The agent-initiated case stayed at 0/3 in both arms, which is the point rather than a shortfall — it is the measurement that says the Instruction, not the description, has to carry that path. Full result and caveats innotes-evals.mdE11. -
Make remote-mutation hooks resolve deterministically and fail closed
(2026-09-02). Hook commands now use only the exactPLUGIN_ROOTor
~/.copilot/hooks/scriptspath, never a version wildcard, and exit2with
a diagnostic when the security script does not resolve. Missing lifecycle
scripts warn without blocking the agent loop. The command matcher recognizes
git.exe, fully qualified Git executable paths, and GitHub CLI global
repository or hostname options before mutating commands, closing ordinary
bypasses while preserving read-only and local commands. -
A 45-minute live proof ran with
long-running-job-monitorunloaded, and the chat stayed silent for thirty minutes (2026-09-01). An agent launched a live Hyper-V proof in the Vivarium workspace, hand-rolledStart-ProcessplusWaitForExitinstead of the canonical detached launcher, armed no cadence tick, and answered two mid-job turns with no status line. The user had to ask "are you running a task in the background?" and then "didn't we update the skill so the user gets a status update every n minutes?" — a Skill that was never read cannot be followed, so this is three defects inskills/long-running-job-monitor/SKILL.md, not one.The first is a vocabulary gap in the
description, which is the only thing the selector sees. Vivarium's glossary makes proof the canonical term for a live integration run, and theUSE FOR:list carried "live test" and "integration test" but not the word the domain actually uses. It now nameslive proof,proof harness,proof run, andhour-long run; the description stays at 961 characters, under the 1000-character soft cap. The second is a typo in the same list — "log log tail" is now "log tail". Both are one-line fixes that only matter because a description this skill never triggers on is a description that does nothing.The third is structural. Arming the cadence tick was described in the Chat heartbeat section and in a checklist item prefixed "For unattended cadence", so nothing on the launch path itself required it — an agent could follow step 2 to the letter, detach the job correctly, and end the turn with no tick armed. Step 2 now carries the imperative directly: arm in the same turn as the launch, before the turn ends, whenever the job is expected to outrun the cadence interval, and a detached launch with no armed tick is named as the exact failure the Skill exists to prevent. The checklist item is unconditional.
notes-evals.mdgains E10, a trigger-rate eval whose prompt is a live-proof launch in Vivarium's vocabulary that never says "monitor", "heartbeat", or "background", so the Skill has to be selected on the description alone. -
The mandatory disclaimer travelled into two signed submissions to a German tax office (2026-08-31).
com.github.copilot/agents/tax-researcher.agent.mdopened with "include this at the end of every substantive output", and the model did exactly that: an RDG and StBerG notice ended up below the signature block of twoEinspruchsbegründungen, whereskills/german-tax-research/SKILL.mdhad forbidden it since the Skill was written. The defect surfaced only when the taxpayer had already printed and signed both letters.The two rules were both present and contradicted each other. The agent's instruction was unqualified; the Skill's fourth non-negotiable said submissions carry no internal caveats. An unqualified instruction in the agent body beats a rule three sections into a Skill, so the agent is where the fix belongs: the disclaimer now applies to chat replies and internal working papers, and never to a
Schriftsatz,Einspruch,Anlage,Eigenbeleg, orErklärungthat a taxpayer signs. The reason is spelled out rather than asserted — in a letter the taxpayer signs, a notice disclaiming tax advice reads as if an unauthorised third party had drafted it.A rule nobody checks is a rule that fails silently, so both files now carry the check. The agent gains a marker sweep in phase 5 and an anti-pattern for shipping a
SchriftsatzPDF without one. The Skill's non-negotiable 4 names the production failure, lists the search terms —StBerG,RDG,Steuerberatung,intern,Entwurf,Prüfvermerk,TODO— and requires the sweep twice: once against the Markdown and once against the rendered PDF's text layer, because a template or a CSS rule can reintroduce what the source no longer shows. The existingMarker sweepverification item is extended accordingly.
Changed
-
Separate career, legal, and tax records into role-specific Memory Bank
namespaces (2026-09-04). Use.memory-bank/career/,
.memory-bank/legal/, and.memory-bank/tax/; preserve ambiguous legacy
files until the user explicitly assigns and verifies them. -
Document cross-client Custom agent differences and staged sensitive-data
research (2026-09-04). Clarify that plugin discovery does not guarantee
identical model or tool behavior, and preserve local file, OCR, web, and
authenticated-browser workflows by separating private intake, local
transformation, minimized public research, and user-confirmed actions. -
german-tax-researchgains a disclosure economy (2026-08-31). ABegründungaddressed to a tax office had been disclosing which receipts were missing for positions nobody had questioned, explaining at length why items were not claimed, and conceding reductions the office had not proposed. Each sentence was true; together they handed the examiner a worklist he had not written.The new section separates two duties that get conflated.
§ 150 Abs. 2 AOrequires the declared bases of taxation to be complete and true; it does not require a self-assessment of how strong the evidence behind them is.§§ 90, 97 AOoblige cooperation and production — on request, and under theBelegvorhaltepflichtthat request often never comes. One test decides every sentence: does it support an amount that is actually declared?Estimates, deviations from the transmitted return, method changes,
§ 153 AOcorrections, and positions maintained against a contrary document must still be disclosed — silence there is the real risk. What must not be volunteered is the evidentiary weakness of a claimed and consistent position, any reasoning for a position that is not claimed at all, the fact that a figure rests on the taxpayer's own statement where no third-party document could exist, speculation drawn from a bank entry, anticipatory concessions, and promises of documents nobody asked for. Three exceptions keep a non-claimed item in the letter: a cross-year inconsistency the office would otherwise spot, a double-deduction reproach worth forestalling, and a correction against the taxpayer. Two anti-rationalizations, three red flags, aDisclosure sweepverification item, and an anti-pattern make it checkable; thetax-researcheragent gains the matching phase-5 probe and four German anti-patterns.
Changed
-
The Software Engineer agent no longer hands work to
security-revieweron its own judgement (2026-08-28).com.github.copilot/agents/software-engineer.agent.mdgains an explicit independent review switch that isoffby default, so a routine change now ends with the agent's own validation and self-review instead of a subagent dispatch that costs minutes of latency per turn.The old rule read "request an independent review with a subagent for high-risk work" and then listed security or identity boundaries, destructive operations, persistence, concurrency, public APIs, cross-module contracts, and "a large unfamiliar diff". In an agent-customization repository almost every change matches at least one of those, and the
Design and securityrule pointing atagent-security-reviewfor "agents, LLM-backed features, RAG, or MCP servers" matches the rest — so the risk-scaled default behaved as an unconditional handover. The trigger list survives unchanged; what changed is what it triggers.The switch is user-set, not model-set:
review: onrequests one independent review of the finished change,review: autorestores the previous risk-scaled dispatch, andreview: offis the default. Plain language and the existing Run Security Review handoff button both count ason, so the fast path stays available without new syntax to learn.argument-hintadvertises it in the picker.Turning the default off without losing the signal needed one more piece. With the switch off the agent still evaluates the same risk list, but it names the risk instead of reviewing it: the work finishes and the closing line recommends
review: onand states why.com.github.copilot/rules/postflight.instructions.mdgains the matching clause, because the shared Definition of Done gate demanded that independent review "was completed" — a contradiction the model would otherwise have resolved by dispatching anyway. The requirement stands for every other agent; only the deferral path is now named.com.github.copilot/agents/software-engineer-contoso.agent.mdis unaffected by design. Its inlined base body is re-synced, but the overlay pins the switch toonfor security-relevant diffs, new dependencies, new network paths, and first-time repositories, and states that areview: offrequest downgrades nothing there — an overlay that only adds constraints must not inherit a relaxation.tests/SoftwareEngineerAgent.Tests.ps1covers the default, the three switch values, and theargument-hint, so the auto-handover cannot come back silently.
Added
-
elster-form-capture, a Skill for driving the Mein ELSTER web form by machine (2026-08-31). Three full capture runs across two assessment years produced the material:skills/elster-form-capture/SKILL.mdfills a German income tax return field by field while the taxpayer signs in, reviews, and presses Send.The boundary is legal, not technical. Transmission is the taxpayer's declaration of knowledge under
§ 150 Abs. 2 S. 1 AO, so filling fields is assistance and sending is not delegable — Versenden des Formulars is a non-negotiable the Skill never presses. It handles no credentials either: the user authenticates and shares the page.One fact carries the whole Skill. The official ERiC field numbers behind
name="fields[…]"are stable across assessment years; theTeilseiteandZeilenumbers are not. TheAnlage Vwas reorganised for 2023 and renumbered again for 2024 — apportioned costs moved from sub-page 12 to 13, the result and allocation from 17 to 18, and sub-letting left the attachment entirely for a newAnlage V-Sonstige— while not onedata-eru-namechanged. So the Skill addresses fields by Kennzahl, verifies by sub-page heading, and treats a line number from a guide written for another year as a claim to be checked.references/feldkarte-est.mdcarries the harvested numbers forAnlage V,V-Sonstige,N, andVorsorgeaufwand, with the 2023-to-2024 movements tabulated above them.The 29 gotchas are corrections, not advice; each one cost a failed attempt. The three that generalise beyond ELSTER: a
page.goto()discards a select box the server has not yet acknowledged, because thebeforeunloaddialog takes the change with it — three running numbers were set in a loop and only the last survived. The add button of a sub-form shares its id prefix with edit, differing only in a trailing index, so.first()silently overwrote the first foreign country with the second. And a value transferred as eData can itself be the error: an employer reported 0.00 € for statutory health insurance, and the field had to be emptied rather than left at zero.The finding that justifies the whole approach is not the typing. Driving the form mechanically turned out to be the fastest audit of the capture guide that feeds it — it caught a wrong postcode (the insolvent developer's address, not the property's), a stale line reference, and, through a machine comparison of 52 target amounts against the summary page, 1,330 € of deductions that a status table already recorded as captured. Hence the rule that the final check is a comparison and never a reading: values produced by someone else are demonstrably reviewed less carefully than values one typed.
skills/agent-evals/assets/trigger-queries.elster-form-capture.jsoncarries ten positives in both German and English and ten near-miss negatives drawn from the neighbours the Skill must not displace: substantive deductibility andEinspruchdrafting belong togerman-tax-research, login persistence toauthenticated-web-extraction, receipt reading topdf-to-markdownandxlsx-to-markdown, theAnlagenbundle toevidence-package-assembly. "Fill in this PDF form" is included deliberately as the closest false friend. -
An opt-in four-stage development cycle across the agents (2026-08-28).
cycle: fullruns architect → engineer → security reviewer → technical writer as one requested workflow, with the reviewer as the gate: on pass it hands to the writer, on fail it hands back to the engineer. It is off by default and starts only because the user asked for it — never because the work looked like it deserved one.The distinction that makes this safe is the one the same release draws for
review: on: consent at the entry point covers the whole chain. What the previous entry removed was delegation the agent chose; automatic progression inside a cycle the user requested is the opposite thing, so the stages flow without further clicks once the switch is set.Two problems had to be solved before the chain could work at all. The first is close-out: the shared Post-flight makes every substantive turn write the Memory Bank, add a changelog entry, and commit, so four stages would have produced four of each for one change.
com.github.copilot/rules/postflight.instructions.mdnow defers those steps to the final stage, and the three earlier stages verify their own work, refreshactiveContext.md, and hand over. The second is the failure path — reviewer → engineer → reviewer is a loop, not an arrow.The first attempt at bounding that loop was a prose cap: "after two rounds, stop the cycle and report the unresolved findings". It could never fire. A handoff starts the receiving agent with fresh context, so neither side can see, let alone count, the rounds it has already run — the cap was an instruction with no state behind it. Paired with two handoffs that both auto-submitted, it produced a run of fifteen complete
software-engineer↔security-reviewerround trips that only stopped when the session was abandoned. The bound that ships instead is structural: the reviewer's Fix Issues Found handoff setssend: false, so the cycle runs forward on its own but re-entering implementation costs one deliberate click. A ring of auto-submitting handoffs is now a test failure rather than a judgement call.State passes on disk, not through the conversation. The architect writes the signed-off Design Concept to
.memory-bank/decisions/and the engineer reads it from there, because a conversation does not survive a compaction and a subagent never sees one to begin with.Every forward handoff in the cycle sets
send: true, so a transition submits on selection instead of populating the box and waiting for a second confirmation — inside a cycle the user already consented at the entry point, and asking again is the ceremony the switch exists to remove. The reviewer's fail path back to the engineer and every non-cycle handoff keepsend: false. Progression is still surfaced as a handoff rather than an unattended agent switch, which is a platform boundary rather than a design choice: VS Code hands the user to another agent, an agent cannot hand itself.Nobody speaks in switches, so
software-architectcarries a phrase book — "full development cycle", "full workflow", "development cycle", "full SDLC", "full pipeline", "the whole pipeline", "the full agent chain", "all four agents", "design to documentation", "concept to docs", "run the complete workflow" — and, more usefully, a refusal list. "end-to-end" normally means end-to-end tests, and "do it properly", "the whole thing", and "ship it" name nothing; a four-agent cycle is too expensive to start on a guess, so those prompt a question instead. A cycle requested at the engineer rather than the architect hands back upstream, because starting in the middle means there is no signed-off concept to implement.The rules live in the four agent bodies rather than in a Skill, and that is deliberate: this repository already established that a Skill is advisory content while an agent body is mode instruction, which is why
grill-mehad to become thesoftware-architectagent. Close-out ownership has to bind, so it is stated where it binds — but the loop bound is deliberately not prose, because prose is exactly what failed. The only new frontmatter edge issecurity-reviewer→technical-writer, which is what closed the graph; every other leg already existed.tests/DevelopmentCycle.Tests.ps1asserts the stage declarations, the connected chain, the gated fail path, the trigger and refusal vocabulary, the auto-submitting forward handoffs, the escape hatch, that exactly one stage claims close-out, and — by walking the whole handoff graph — that no ring ofsend: trueedges exists in any agent.cycle: offends a running cycle at whichever stage holds the work, and that stage becomes the closer rather than leaving the changelog entry and the commit stranded on a chain nobody will finish. The reviewer additionally has to name any unresolved Blocker on the way out, because a stopped cycle is the easiest place for one to disappear.Both switches are documented where a user will look rather than only in the agent bodies:
README.mdcarries a five-row table from "nothing" tocycle: full, andcom.github.copilot/agents/README.mdexpands it with what each setting actually does. The table leads with the default — doing nothing keeps the work with one agent — because that is the question the previous entry left unanswered. The release-pipeline diagram there gained the technical writer stage it had been missing and the gated fail edge. -
software-engineer-contoso— a corporate overlay on the Software Engineer agent (2026-08-27).com.github.copilot/agents/software-engineer-contoso.agent.mdcarries the wholesoftware-engineer.agent.mdcontract inline and then only tightens it: where the overlay is stricter it wins, where the base is silent the overlay governs, and where both are silent the stricter reading applies. It exists both as a usable agent for regulated work and as the copy-and-rename template forsoftware-engineer-<company>.Inheritance is by inlining, not linking, and the first attempt proved why. Modelled on
devops-training-writer— which states its inheritance fromtraining-writerin prose — the overlay originally opened with a Markdown link to its base and the sentence "read it as part of your operating instructions". Nothing was inherited. VS Code resolves referenced instructions files into the prompt, which is whatchat.includeReferencedInstructionsgoverns and what the documentation means by "reference other files by using Markdown links, for example to reuse instructions files"; an.agent.mdis not an instructions file, so the link is inert and the overlay ran as a bare fragment with the base contract missing. The setting being enabled is not the fix, and the failure is silent — the agent loads, answers, and simply has none of the engineering rules its own text claims to apply.Inlining is also the correct design here rather than a workaround, and the agent's own doctrine is the argument: a rule the model can route around is not a control. An overlay whose base contract depends on the model choosing to open a second file has exactly that weakness, in the one agent least able to afford it. One file now holds the complete envelope, which is also what an auditor needs in a regulated environment. The cost is a duplicate that can drift, so
tests/AgentInheritance.Tests.ps1compares the inlined block byte-for-byte against the base body — dropping its H1 and demoting its H2s, the one documented transformation — and fails the moment the base moves. Line endings are normalised before the comparison because git rewrites them on checkout; without that the test fails on encoding rather than content, which it did on the first run.The containment is in the frontmatter, not only in the prose. The base agent's 45 tools drop to 36:
web/fetch,web/githubRepo,web/githubTextSearch,openSimpleBrowser,github,useMcp,vscode/installExtension,vscode/extensions, andcodeInterpreterare removed, so private-data access and untrusted content cannot combine into the lethal trifecta because the third leg is gone. Prose then closes the three ways an agent reconstructs a removed capability: the terminal (curl,Invoke-WebRequest,ssh, a public-registry install), the user ("switch agents and paste it for me"), and a handoff — which is worth naming explicitly, because a handoff moves the user into another agent's toolset and stops this file binding at that moment.The subagent rule is the one that is easy to get wrong.
agentsis narrowed tosecurity-reviewer, andtechnical-writeris dropped — butsecurity-revieweritself holdsweb/fetch,github, anduseMcp, so delegating to it re-opens the channel the toolset just closed. Removing it was not an option, since the same overlay makes its review mandatory rather than risk-scaled for security-relevant diffs, new dependencies, new network paths, and first-time repositories. The rule that ships instead constrains the dispatch: hand it repository paths, symbol names, and the question, never pasted source, configuration values, hostnames, or data samples, and write every dispatch prompt as if it will leave the boundary.The rest is the control set a regulated employer actually imposes: secrets by reference from the vault with a discovered credential treated as burned (rotate, then scrub — never silently deleted, which hides a leak without revoking it); internal-mirror-only dependencies that need a pinned version, integrity verification, an approved license, and an SBOM entry, all four or none; a "never ship" Blocker list covering hand-rolled crypto, disabled TLS verification, dynamic execution, injection-prone concatenation, wildcard authorization, swallowed security failures, sensitive logging, and — the one usually left implicit — weakening an existing control as a side effect of a feature; synthetic-only test data; separation of duties that ends the agent's entitlements at the local working tree; and a seven-item hard stop whose escalation report names what state was left behind and who must act.
The Memory Bank extension adds
contoso-controls.mdanddata-classification.mdand explicitly refusesthreat-model.md,assessment-log.md, andsecurity-playbooks.md, whichsecurity-reviewerowns — the house rule that an agent does not write another agent's role files is only enforceable if each overlay states which files are not its own.tests/SharedLifecycle.Tests.ps1carries the per-agent baseline so the toolset, the handoff targets, and the Memory Bank section cannot drift without a test naming the change.Inlining has one consequence worth stating outright: the overlay contradicts itself by construction. The inherited block says in its own words that
review: offis the default and that independent review is the engineer's call, while the overlay pins it toreview: onand refuses the downgrade. A precedence clause resolves that — but only for a reader who reaches it, roughly 180 lines later. The reversed defaults are therefore named in the preamble before the inlined block, so the correction arrives ahead of the contradiction rather than after it, andtests/AgentInheritance.Tests.ps1asserts that ordering rather than merely the presence of the precedence language.
Fixed
-
Hook commands no longer carry a
$token for the host to eat (2026-08-27). Every hook died at startup withAn expression was expected after '(', and the only trace was a Warning from Session Start hook balloon — so the never-push block, the Memory Bank probe, and the compaction checkpoint were all silently absent while looking installed. The cause is that the host substitutes$tokens in the command string before the child process parses it."$b = if ($env:PLUGIN_ROOT) { ... } else { Join-Path $env:USERPROFILE '.copilot\hooks' }"reached PowerShell as" = if () { ... } else { Join-Path C:\Users\install '.copilot\hooks' }":$env:USERPROFILEwas resolved by the wrong layer, and$b,$env:PLUGIN_ROOT, and$LASTEXITCODEwere resolved to nothing at all. The previous design assumed the opposite — that VS Code spawns the command with no shell, so each command had to expand its own path — andcom.github.copilot/hooks/README.mdsaid so in as many words.The commands in
com.github.copilot/hooks/hooks.jsonare now written without a single$, which makes them correct under both readings rather than betting on either. Paths come from[Environment]::GetEnvironmentVariable('USERPROFILE')instead of$env:USERPROFILE, and the blocking exit code — which-Commandotherwise flattens to1and would have turned a hard block into a warning — comes fromGet-Variable -Name LASTEXITCODE -ValueOnly. Each candidate path is built with[IO.Path]::Combine('/', <root>, <relative>)so that an unset root yields a drive-rooted path rather than a workspace-relative one; without the leading/, opening an untrusted repository that happened to containcom.github.copilot/hooks/scripts/would have executed its scripts on every tool call.Script resolution also gained the deployment path it was missing. A plugin install materialises at
~/.vscode*/agent-plugins/<host>/<owner>/CopilotAtelier/, which is notPLUGIN_ROOTunless the client chooses to set it — a variable this repository adopted on inference and never confirmed. Each command now probesPLUGIN_ROOT, then the module's~/.copilot/hooks, then the plugin location, and runs the first script that exists, so both supported installs are covered whether or not the client cooperates.tests/Hooks.Tests.ps1grew the guard the old suite could not have had: it models the host's substitution pass over the shipped command, asserts the string survives it unchanged, and then spawns the substituted result and requires exit2. The static "no$" assertion alone would have caught this regression; the executable half is what proves the replacement actually blocks. The first attempt at that test wrapped the command in a second PowerShell and reported1instead of2— an artifact of the wrapper's own exit-code translation, not of the hook — which is itself the reason the substitution model is the one that shipped.
Changed
-
Narrowed the role-file clause in the Pre-flight Instruction so its scope cannot be read backwards (2026-08-27). Step 3 of
com.github.copilot/rules/preflight.instructions.mdclosed on "Create only the active Custom agent's required role files", which parses most naturally as the files that role requires — the agent's entire declared list — and that is the opposite of the intended rule. The authority isskills/memory-bank/SKILL.mdinitialization step 6: an agent's role list is an additive schema, and only the files the current durable workflow actually needs get created. The clause now says exactly that, and adds the negative case the old wording never carried — never scaffold another agent's schema, and never pre-create a declared file this task does not use. The failure mode it prevents is not cosmetic: an eagerly created role file is an empty template that a later turn routes to, reads, and treats as authoritative project knowledge. Both agents that declare a role schema already agreed with the corrected reading, so this aligns the shared contract with them rather than changing behaviour anywhere else. -
Migrated the plugin package to Agent Plugins 1.0 (2026-08-26).
plugin.jsonnow declares"$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json", which is the field VS Code and the GitHub Copilot CLI use to select the format. The 1.0 manifest schema is closed and its component locations are fixed: skills are discovered from a lowercase./skillsat the package root and nowhere else, and theagentsandskillspath fields the legacy Copilot format honoured are no longer manifest fields at all. Declaring the schema without moving anything would therefore not have failed loudly — an unknown top-level field is reported and ignored, so the package would have kept loading while silently contributing zero skills. That trap was already written down as a guard intests/PluginManifest.Tests.ps1, which refused to let the schema be declared while the folder was stillSkills/; this change satisfies the guard rather than deleting it.The repository layout moved to match the format, and the package is now the primary layout rather than a second view of the module payload.
Skills/becameskills/— the portable component location, and the one rename that also fixes a latent bug, because the deployed folder is~/.copilot/skillsand every case-only mismatch was already broken on case-sensitive filesystems. Everything Copilot-specific moved into the client-extension namespace under the names the format gives it:Agents/→com.github.copilot/agents/,Instructions/→com.github.copilot/rules/,Prompts/→com.github.copilot/commands/, andHooks/→com.github.copilot/hooks/with the mandatedhooks.jsonfilename. A client that does not own that namespace ignores it without rejecting the package, so the portable half stays portable. Hooks are a clear gain: the legacy manifest never carried them, so a plugin install previously had no guardrails at all — the never-push block and the Memory Bank probe reached module users only.The
rules/andcommands/file formats are not documented by VS Code or the Copilot CLI, so whether.instructions.mdand.prompt.mdregister from a plugin install is unverified. Moving them anyway is a one-sided bet: if a client rejects them they simply do not load from the plugin, while the module path keeps delivering them to the same deployed locations. There is no state in which it is worse than leaving them at the root, and one in which it is better.Hook commands had to learn two roots. A plugin is installed outside the workspace, so a relative path cannot work, but the same file also ships to
~/.copilot/hooksthrough the module — so each command now resolves$env:PLUGIN_ROOTwhen a plugin host sets it and falls back to the user profile otherwise. Both branches are asserted, and the existing spawn-without-a-shell test clearsPLUGIN_ROOTexplicitly so the user-profile branch is genuinely exercised rather than accidentally taken.The deployment contract is preserved by translation rather than by layout.
Install-CopilotAtelierstill produces~/.copilot/{agents,instructions,skills,prompts,hooks}; because the source and deployed layouts are now deliberately different shapes, the installer carries an explicit deployed-name → source-path map —com.github.copilot/rulesdeploys asinstructions,com.github.copilot/commandsasprompts. Canonical target directories are lowercase to match their discovery links, and because a case-insensitive filesystem quietly reuses the old directory while a case-sensitive one keeps a second copy forever, the installer sweeps the capitalised names from a previous release using a case-sensitive comparison. Renaming the deployed directories to match the namespace was rejected: it would require pinningchat.instructionsFilesLocations, and achat.*FilesLocationssetting replaces the default location map rather than extending it — the same trap that silently disabled every hook location once already.One consequence is recorded rather than hidden. Cross-type relative links are now correct in the package and repository view and wrong in the deployed view, because
rulesandcommandsdeploy asinstructionsandprompts. That direction is deliberate — the package is what a human browses on github.com and what a plugin install materialises — and nothing functional rests on it, since both lifecycle Instructions declareapplyTo: "**"and load regardless. Links into the repository-onlyreference/were already dead in the deployed tree, so this widens an accepted condition rather than introducing a new class of defect.README.mdandAGENTS.mdfollow. The Folder Structure section now shows the deployed tree and the package layout side by side instead of conflating them, and the Agent plugin install path was rewritten: it previously told the reader that "the plugin format does not carry.instructions.mdfiles or hooks, so this path gives you agents and skills only", which the migration makes false in the one direction that matters — hooks are exactly the guardrails a plugin-only user was silently missing. It now also names the two things worth knowing before choosing that path: bundled hooks execute locally and should be reviewed, andrules/commandsregistration is unconfirmed.
Fixed
Install-CopilotAtelier -WhatIfno longer throws (2026-08-27). A dry run aborted with a terminatingItemNotFoundExceptionfromGet-ChildItem: Cannot find path '...\CopilotAtelier' because it does not exist. The legacy-directory sweep enumerated the canonical target unconditionally, but under-WhatIfthe target tree is never created (theNew-Item/Copy-Itemcalls that build it honourShouldProcessand are skipped), so the enumeration hit a path that did not exist. The sweep now runs only when the target is present, which matches how the rest of the function already guards optional paths withTest-Path. Covered by a regression test that asserts-WhatIfneither throws nor creates the canonical target.