Skip to content

Releases: Readtt/devops-studio

DevOps Studio v0.25.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 01:31

Added

  • New models show up on their own. DevOps Studio now reads each connected provider's own model list, so a model released today appears in the model picker without waiting for an app update. Lists refresh every 12 hours and whenever you change a key. Settings → Models shows when they were last checked, says so if a provider couldn't be reached, and has a Check now button. The picker shows each provider's five newest additions; type to search the rest. For OpenRouter, only routes your key can use and that support tools are listed. Free, batch and very small-context routes are left out.
  • New models in the built-in list: Claude Fable 5.1, Claude Opus 5.5 and Claude Sonnet 5.5; GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna; Gemini 3.8 Flash; Grok 4.7 and Grok 4.3; DeepSeek V4.1 Flash; Qwen 3.8 27B on Cerebras; GPT-OSS 120B on Groq; and OpenRouter routes for Claude Opus 5.5, Claude Sonnet 5.5, GPT-6.1 Sol and Grok 4.7.
  • Long runs on Claude Opus 5.5, Sonnet 5.5 and Fable 5.1 keep going after the app trims old tool results. For Anthropic accounts created on or after August 31, these models reject a conversation that has been trimmed. The app now asks Anthropic to skip the affected reasoning instead.

Fixed

  • DeepSeek failed every structured run with "This response_format type is unavailable now". The app asked DeepSeek for a JSON mode it doesn't offer; it now uses the one it does.
  • Gemini 3 Flash was sent a temperature of 0. Google warns this can make Gemini 3 loop or give weaker answers. No Gemini 3 model is sent a temperature now.
  • The Test button on an OpenAI key always said it couldn't fully verify the key. It asked for fewer output tokens than OpenAI allows.
  • Cost readouts used wrong prices for several models. For example, GPT-5.5 output is $30 per million tokens, not $15. Mistral's context windows were also understated: 256K, not 32K or 128K.
  • Models the providers have shut down are gone from the list: DeepSeek Reasoner, DeepSeek V4 Flash (now V4.1 Flash), Grok 4 Fast, Groq's Llama 3.3 70B and R1 Distill 70B, Cerebras's Llama 3.3 70B and Qwen 3 32B, and five OpenRouter routes that were renamed or removed. If one of them was your default, a favorite or a recent pick, it moves to that provider's replacement. The same goes for a saved draft, a Commit Review tab or a Suite Chat thread that used one. Before, it reset to a model you might have no key for. An interrupted run on a retired model no longer offers Resume, since it can only be continued on the model it started on.
  • A tab that fails to open now shows an error inside that tab, with Try again and Close tab, instead of blanking the whole window.
  • The background transcript summarizer no longer picks a reasoning model or a code-completion model when a better choice exists. When every model you have a key for is a reasoning model (Google only, for example), it now gets enough room to think and still write the summary. A summary that ran out of room is thrown away instead of used.
  • The release script now links to the build of the release it just pushed, not the previous one.

Changed

  • The Test button on a provider key now tests with a cheap, stable model from that provider instead of its newest flagship.
  • The AI SDK provider libraries are updated to their latest AI SDK 6 releases.

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.

DevOps Studio v0.24.0

Choose a tag to compare

@github-actions github-actions released this 03 Sep 04:22

Added

  • Commit Review now works forward from what can arrive at the changed code, the way test design does, so it reports handling that was never written and not only changed lines that break something. Each of those findings names the exact input that triggers it.
  • Findings carry Steps to reproduce: numbered steps traced from the code, shown as a disclosure above Evidence on the finding card.
  • A rename sweep runs on every review: when a change renames or re-labels something, the old string is searched across the repos and every leftover (placeholders, tooltips, help text, error messages, tests) is reported.

Changed

  • Severity is rated by what a defect does, not by the shape it takes. Missing handling that loses data or halts a whole batch is no longer a flat "medium", and the severity tooltip on finding cards says the same.
  • Requirements checks stay conservative about intent but now report a concrete, code-visible mismatch at any severity, including low.
  • The verify pass no longer refutes a "handling is missing" finding on thin evidence alone: refuting one requires citing the code that handles the case.

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.

DevOps Studio v0.23.0

Choose a tag to compare

@github-actions github-actions released this 28 Aug 19:02

Added

  • Max output setting for custom OpenAI-compatible endpoints (Settings →
    Models). Sets the output-token limit sent on every request to that endpoint.
    Leave it empty and nothing is sent, exactly as before — but "let the endpoint
    decide" is not free: a proxy in front of Anthropic has to invent a limit, and
    the invented number is often far smaller than the model allows.
  • A warning when a model's answer is cut off mid-write. The generator's
    review pane and Commit Review now say when the answer ran out of room, and
    which part of it was never written. On a custom endpoint the warning links
    straight to the setting that fixes it.

Fixed

  • Bug suggestions went missing on custom endpoints. Test cases are written
    before bugs, so an answer that ran out of room lost the bugs and looked like
    a complete run that simply found none. Commit Review lost findings the same
    way, and a review cut off before it wrote any showed the green "no issues
    found" panel.
  • The Model ID dropdown needed opening twice to fill in. The model-list
    request gave up after 5 seconds. An endpoint that builds its catalogue on the
    first request can take far longer — and because the abandoned request still
    warmed it up, the second open worked. It now waits 30 seconds.
  • The generator can now tell a cut-off answer from an empty one when code
    search is off, instead of reporting both as "the model returned nothing".

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.

DevOps Studio v0.22.2

Choose a tag to compare

@github-actions github-actions released this 21 Aug 15:04

Fixed

  • Resuming an interrupted AI run no longer fails on Claude 5 with "this model does not support assistant message prefill". A resume replays what the run had already read, and that replay could end on the model's own turn — which Claude 5 refuses. Every request now ends with a user turn, on every provider. This affected the Generator, review-pane follow-ups and Commit Review equally, and it hit hardest on exactly the runs worth resuming: the ones stopped just after the model wrote its answer.
  • The native OpenAI GPT-5 models and their OpenRouter equivalents now agree on which settings they accept. They disagreed — the app had decided GPT-5.4 mini rejects a temperature setting on one route while sending it one on the other, and only a rule inside the OpenAI SDK was keeping that off the wire.
  • Mistral Large's context window now reads the same whichever route you reach it through.

Changed

  • When a model refuses a request outright — a setting it dropped, a feature it doesn't have — the error now quotes the provider's own sentence and points you at switching models, instead of offering a Resume that would send the same refused request again and fail the same way.
  • The model list is now checked as part of the build. A new model can't ship without saying whether it takes a temperature setting, the same model reached natively and through OpenRouter can't disagree about that, and every Claude route has to carry its own output limit rather than inheriting whatever the provider SDK happens to assume that month.

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.

DevOps Studio v0.22.1

Choose a tag to compare

@github-actions github-actions released this 19 Aug 22:25

Changed

  • Commit Review findings are now written to be understood in one read. Each finding opens by tying your change to what now goes wrong, explanations are capped at two short paragraphs of plain full sentences — no fragments, no arrow-chain shorthand, no markdown artifacts — and evidence is a short file:line trace of what was actually checked instead of a wall of text.
  • Finding titles state the consequence: declaratively when the failure was fully traced, as an explicit "can ..." when it depends on a situation the review didn't confirm. Maintainability findings state their ongoing cost instead of an invented failure.
  • Generic filler that only sounds helpful is banned from findings, and in a multi-commit review every finding names the commit it concerns.

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.

DevOps Studio v0.22.0

Choose a tag to compare

@github-actions github-actions released this 19 Aug 21:23

Added

  • DevOps Studio now works across several repositories at once. The app used to point at exactly one source folder, so a feature whose code spans repositories was analyzed against a fragment of itself — silently, with nothing to tell you coverage had been lost. Settings now holds a list of source repos, and every surface that reads code reads all of them: the generator, Suite Chat, Commit Review, confidence evaluation, the code viewer and the terminal. Your existing source folder migrates by itself into a one-repo workspace, and at one repo the app looks and behaves exactly as it did.
  • A Source repos block in Settings — add a folder, rename it, remove it, or point "Scan a folder…" at a parent directory and pick up every repo cloned inside it in one go. Each row shows the repo's current branch and which Azure DevOps repository it's bound to. This also closes the dead end where opening a commit review with nothing configured sent you to a Settings tab that had no source-directory control at all.
  • The status bar shows and switches the branch of every repo. Past one repo the folder segment reads "N repos" and the branch segment summarises where they are — the shared branch when they agree, "N branches" when they don't — opening a list you drill into per repo. Repos move independently: a switch running in one can't cancel or steal the toast of one running in another, and the dirty-tree prompt asks about one repo at a time and names it.
  • Commit Review spans the whole workspace. The commit picker merges every repo's history into one timeline, so a change that touches three repositories can finally be reviewed as one change. Every repo is guaranteed a share of the list, so a repo nobody has touched in months stays reachable instead of being crowded out by the busy ones. Which repos the reviewer may read is a separate control from which commits are under review — a commit in one repo often can't be judged without reading another that has no commit in the selection at all.
  • Runs can be narrowed to the repos you care about. The generator, Suite Chat and Commit Review each carry a Repos chip row (shown past one repo): every repo is on by default, and unticking one keeps that run out of it. The generator persists the choice into its checkpoint, so a resumed run keeps the scope it started with instead of quietly widening back to everything.
  • Published code links now deep-link into the right Azure DevOps repository. Each source repo binds to its ADO repo and project, matched from the repo's own origin remote against your organisation's repository list — normalised URL first, then the project and repo parsed out of either URL shape (SSH and legacy visualstudio.com remotes never string-matched the HTTPS one), then a folder-name match that must be unique across the organisation, so a name two projects share is left unbound rather than bound to the wrong one. Every link into a repository living outside the connection's project was previously dead. Settings' repo row is the manual override, with a repository picker grouped by project and a "Detect from remote" action.
  • "Get source code…" is reachable once you already have a repo. It used to hang off the status-bar segment that only appears when no repo is configured, so it vanished the moment a tester had one. It's now in the repo list footer, in the Settings source-repos block, and still in the empty state.
  • The custom-provider Test button now performs a real model call. It fired a bare reachability ping that discarded the response, sent no authorization header, and reported "Reachable — server responded" for any reply at all — so a wrong key, a 404 and a 500 all read as success, and the model ID was never exercised. Test now runs the same one-token generation a real run would, against the base URL, key and model you've drafted, and reports what actually came back. The Model ID field also suggests the models your endpoint lists, while still accepting anything you type.
  • The AI's read-only git access is considerably wider — 31 subcommands instead of 19, so it can answer questions with for-each-ref, show-branch, check-ignore, check-attr and the diff plumbing rather than working around their absence. Write forms remain blocked outright, and ls-remote stays out because it reaches the network, where a credential prompt can hang the call.

Fixed

  • The AI's read gate could be escaped three ways, and each is now closed. A symlink or junction inside a repo (vendor/cache pointing at your home directory) made every file on the machine readable through a path that looked entirely repo-local — the gate checked the literal path and the read followed the link. It now runs on the resolved path, and the read uses that same resolved form so it can't land on a different file than the one just cleared. Separately, files the app refuses to open by name were still readable through the command runner (cat .env, git show HEAD:.env, git log -p credentials.pem), and the source-code exemption that keeps Credentials.cs readable also let .env.ts and .npmrc.js through. Finally, git --git-dir=C:\…\other-repo\.git log pointed the whole command at a different repository, because the guard checking for escapes never looked at a path glued to a flag.
  • Every Azure DevOps deep link built from a published source link was dead. Publishing escaped each / in a path or branch name and nothing ever unescaped it, so the links read back mangled. Two more defects in the same block: unticking "Tag with source branch" — or generating from a non-git or detached-HEAD folder — published links the app itself then refused to read back, and the commit stamp was written empty even though publish had already captured it.
  • Code links were stamped with the wrong repo's branch. A case citing two repositories recorded both links against whichever repo happened to be first. Each link now carries the branch and commit of the repo it actually names, read from that repo's working directory at publish time, and one unreadable folder only costs its own links their stamp.
  • Confidence verdicts read as permanently stale, and every bulk run re-scored the whole suite. A verdict was compared against a single repo's HEAD regardless of what it had read, so past one repo the "skip verdicts that are still fresh" gate never fired — on the app's most expensive path. A verdict now records the repos its own evidence cites and goes stale only when one of those moves. Re-checking that also used to cost a git call per repo per case; a bulk run resolves it once.
  • A suggested fix could be written into the wrong repository. Commit Review's Apply button resolved the patch path against whatever the status bar pointed at rather than the review's own repos.
  • Reopening a saved review bound it to whatever repo was current instead of the repos it was actually run against.
  • A terminal's Quick Prompts injected another repo's base branch into its templates, because the strip read the global source folder rather than the shell's own working directory.
  • Assignee pickers kept offering the previous project's people after a connection change. The roster is per project and was cached for the whole session, so the names on offer weren't assignable on the work items being created.
  • Azure DevOps errors lost the detail that explains them. A rate-limited request reached you as "retry in undefineds" and a server error dropped the excerpt saying what ADO had actually rejected — the error fields were serialized in the wrong case for the frontend reading them.
  • The AI wasted steps on programs that aren't there. Six of the tools its prompts advertised don't exist on a stock Windows PATH, and find and tree resolve to unrelated Microsoft binaries of the same name, so find . -name "*.ts" failed as "FIND: Parameter format not correct". Each attempt burned a step against the run's budget and stayed in the transcript to be re-sent on every later one. The prompts now lead with git and the app's own read tools, and a missing program points at the tool that replaces it. This worked in development and failed in the shipped app, which is why it survived so long.
  • A generated case could be dropped silently when the model followed the new path convention, leaving a run that produced nothing and no explanation.
  • A repo-scope chip could re-include repos you had deselected, when the registry had changed underneath it.
  • Code search hid whole repositories. Search results were concatenated per repo before being truncated, so one repo's eighty matches could bury another's entirely; they're now interleaved.
  • The settings file was rewritten and broadcast on every launch, because the saved repos and the freshly-normalised ones were compared as text and their keys came back in a different order.
  • The empty state mounted two copies of the Get source code wizard, so a request from the Settings window opened one while a click opened the other.
  • The branch picker hid any saved-but-unlisted branch whose name appeared inside its "current branch" sentinel value — a real branch called current vanished from the list rather than being added to it.

Changed

  • Files are addressed as <repo>/<path within repo> throughout — in prompts, in the model's answers, in code links, in the code viewer's header and in tab titles. Tab titles only pick up the prefix when two open viewers collide on the same filename.
  • A bare filename opens the copy in the repo that owns it. The code viewer searches every configured repo instead of silently opening the first one's.
  • There is no fixed tracking-branch option any more. Code links always track the live branch of the repo each link names; a repo with no branch (detached HEAD, or not a git repo) publishes without one rather than claiming a main you never generated from. The per-run "Tag with so...
Read more

DevOps Studio v0.21.1

Choose a tag to compare

@github-actions github-actions released this 11 Aug 23:28

Changed

  • Bug report sections now use standard professional headings — SUMMARY, PRECONDITIONS, STEPS TO REPRODUCE, REQUIRED TOOLS (only when a tool is needed), EXPECTED RESULT, ACTUAL RESULT, TECHNICAL NOTES, and ENVIRONMENT — replacing the casual labels introduced in 0.21.0. PRECONDITIONS remains numbered setup steps, and technical detail remains confined to TECHNICAL NOTES.
  • Generated wording keeps the plain-language rules but is now explicitly held to a professional documentation tone.

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.

DevOps Studio v0.21.0

Choose a tag to compare

@github-actions github-actions released this 11 Aug 22:52

Changed

  • Generated test cases and bug suggestions are now written in plain language a QA tester can follow from the running product alone — no unexplained abbreviations or codes, technical terms explained in plain words on first use, and every exact value and button name kept.
  • Test cases walk through their own setup: preconditions are numbered setup steps at the start of the case (with the exact account, data, and settings), instead of a state the tester has to figure out how to reach.
  • Bug suggestions use a new section layout: WHAT IS BROKEN, SETUP BEFORE YOU START, STEPS TO REPRODUCE (written for a tester with only the running product — no source code, no debugger), TOOLS NEEDED (only when a tool is genuinely required, with plain setup instructions), EXPECTED RESULT, ACTUAL RESULT, NOTES FOR DEVELOPERS (root cause and code references live here), and ENVIRONMENT.
  • Bug titles now lead with the visible problem in plain words, readable by both testers and developers.
  • Suite Chat follows the same plain-language rules and bug layout when it creates or rewrites cases and bugs from chat.

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.

DevOps Studio v0.20.0

Choose a tag to compare

@github-actions github-actions released this 11 Aug 21:00

Added

  • Runs are budgeted by what they cost, not by how many steps they take. A generation, review, or chat run used to stop after a fixed number of reading steps — which cut a cheap run off early and let an expensive one spend far more than you'd expect. Every surface now has a token budget, priced so a cached read costs what a cached read actually costs, and the step count is only a runaway guard. The run header shows what a run has spent against that budget while it works.
  • Long runs no longer run out of room. The app measures the real size of every request instead of estimating it, caps what any single file read or pasted block can contribute, and — once a run gets genuinely long — replaces its oldest file reads with a short note telling the model how to fetch that file again if it still needs it. Recent reads are always kept intact. If that still isn't enough, the middle of the conversation is summarised by the cheapest model you have a key for, so the expensive one keeps working instead of hitting the wall.
  • A run that dies can be picked up where it stopped. Rate limits, dropped connections, a request that didn't fit, or an answer that overran the model's output limit used to lose everything the run had read. Each of those now leaves a resume point: click Resume and the model carries on from its own notes rather than re-reading your codebase. An answer cut off by the output limit retries with more room to write, and a request that was too big comes back smaller than the one that failed. Applies to the generator, its follow-up rounds, and Commit Review.
  • The generator finishes an empty run by itself. The worst failure this app had was a run that read your code for twenty-odd steps and then wrote nothing — all of the cost, none of the output. It now replays what it read with an instruction to stop reading and write the batch, automatically, before any error is shown. You get test cases instead of an error and a button. A run that stopped because it hit its budget still asks first, because that one wanted more reading rather than less.
  • Follow-up rounds remember each other. The review pane showed you every follow-up you'd sent and what each one changed, and none of it reached the model — so a round could cheerfully undo the round before it. That history is now part of the prompt, and a follow-up is metered as a continuation of the run it refines rather than as a fresh one.
  • The Ask remembers what it read. Asking a follow-up question about a draft used to re-read every file from scratch. The conversation now carries the model's own reads with it, and the draft it's shown includes the code references and full repro steps the generator already wrote down — so "explain bug 2" no longer costs sixteen tool calls to recover something the app already knew.

Fixed

  • The default model failed on every run. A stale provider package treated Claude 5 as a model it had never heard of, so it forwarded a sampling parameter Anthropic rejects outright, capped answers at 4,096 tokens, and quietly downgraded structured output. Output limits are now decided per model in the app's own config, where a newly launched model can be described the day it ships instead of waiting for a provider release to catch up.
  • Every turn of a chat re-bought the entire conversation. The thread was rebuilt as prose in the middle of each request — exactly where a prompt cache stops matching — so nothing past the system prompt could ever be a cache hit. Requests are now assembled so each turn extends the previous one, with cache markers where they pay for themselves.
  • A chat turn that stopped without answering showed you its own thinking instead. When the model spent its reading budget and never reached the answer, you saw "I'll dig into the collect code…" presented as the reply. It now finishes from what it already read and gives you the answer. Both the review-pane Ask and Suite Chat.
  • A run that came back empty could hand you a draft the model had abandoned. Anything batch-shaped in the model's mid-run thinking could be picked up and presented as the result — publishable to Azure DevOps, or shown in Commit Review as a verified finding with a patch to apply. Only the model's actual final answer is read now.
  • Resume threw away the reads it exists to preserve. Continuing a failed run handed the model "go read this file again" for most of what it had already read, so it spent the continuation re-reading your code. It carries the full record forward now, and only trims when the run failed because the request was too big to send.
  • Resume was offered when it couldn't possibly work. A first request too large to send had nothing to continue from, but still showed the button — and clicking it sent the identical request, failed on the spot, and charged you for it.
  • A partly-broken answer no longer loses the good half. An answer cut off mid-structure keeps the complete cases and bugs that arrived before the cut instead of dropping the batch, and an answer that restated the format before writing the real one no longer gets mistaken for the example it opened with.
  • The generator ignored a case count you asked for. "Give me 5 cases" produced whatever the model felt like. It also now scales how hard it investigates to the size of what you asked for, rather than always going deep.
  • The broad "where does this live" search under-reported. It stopped counting at a fixed number of matching lines, so a symbol used heavily in one file came back as "1 file matched" when it was used across twenty.
  • The summariser could replace the question you asked with a paraphrase of it, on long chat turns.
  • Turning code search off between two questions in the same chat made every later question fail with a provider error until you turned it back on.
  • A chat answer you stopped halfway was forgotten by the next question, so "keep going from where you stopped" started over.
  • A follow-up round could be told that a previous round had reworked bugs it never touched, whenever you'd unchecked a case in the draft.
  • A failed run and a silent one are now told apart: the error says whether the model hit its output ceiling, ended its turn without writing, or was cut off mid-read — instead of sending everyone to the same JSON-mode setting.
  • The context warning reflects the real size of a request rather than an estimate, so it stops crying wolf on runs with plenty of room left.

Changed

  • Chat answers settle to the answer. While a question is being worked on you see the model's progress and the files it's reading; once it lands, the message shows the answer alone rather than the thinking that preceded it.
  • The run header reads at a glance — model, elapsed time, what the run has spent, and what stopped it.

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.

DevOps Studio v0.19.0

Choose a tag to compare

@github-actions github-actions released this 03 Aug 19:23

Added

  • Requirement-based test suites. Azure DevOps has three kinds of test suite and the app only ever understood one of them. You can now create a suite bound to a user story or PBI — right-click a plan, New suite → Requirement-based, pick the work item — and every case you publish into it links back as "Tested By". That link is the only thing ADO's requirement-coverage reporting reads, so suites created this way show up in coverage instead of sitting outside it. The requirement picker only offers work-item types your project actually treats as requirements, which is process-template dependent and read from your project rather than hardcoded.
  • The generator and Suite Chat now write cases against the requirement. When a suite tracks a work item, its description and acceptance criteria are handed to the model, which is told to prefer one case per criterion and to say so plainly when a criterion is untestable as written rather than inventing scope. Suite Chat answers coverage questions against those same criteria and names the ones nothing covers — "looks well covered" isn't an answer when a criterion is unaddressed. Both surfaces render the requirement from one shared block, so a coverage answer is always judged against the same text the cases were generated from.
  • Confidence evaluation is graded against the requirement too. A case in a requirement-bound suite is now scored knowing the acceptance criteria it was written for, so a step whose expected result contradicts a criterion reads as a real risk instead of an unknown.
  • Query-based suites are recognised and protected. ADO fills these from a work-item query and refuses anything added by hand. Generate, publish, and the Suite Chat create/delete actions are now disabled on them with an explanation of why, rather than letting a full generation run finish and then failing at publish with an opaque server error. Requirement suites likewise refuse renaming and child suites — ADO derives their name from the work item and only nests under static suites.
  • REQ and QUERY badges on suites throughout the app — the plans tree, the command palette, the generator's target chip, and the generation history — so you can tell at a glance what kind of suite you're about to generate into.

Fixed

  • A bug's repro steps could silently lose text. Any < in a work item — an acceptance criterion reading qty < 10 and total > 5, say — swallowed everything up to the next > when the HTML was converted to text, so the model was handed qty 5 and wrote cases for a requirement nobody had written. Script and style blocks also survived as prompt text, unclosed list items ran two criteria together on one line, and the numeric entities Word and Outlook paste in (&#8217;, &#8212;) reached the model undecoded.
  • The app could close silently — no error, no dialog — on text containing non-ASCII characters. Two paths did it: a suite or case whose data Azure DevOps returned in an unexpected shape, and any Azure DevOps error message long enough to be shortened for display. Since error messages routinely quote your own work-item titles and project names, this was reachable by anyone whose ADO content isn't pure ASCII. Both now shorten safely, through one shared helper so a third copy can't reintroduce it.
  • Pasting a screenshot into a work item no longer sends the whole embedded image to the AI as if it were requirement text, and pasting a very large document no longer freezes the window while that text is prepared for a prompt.

Changed

  • Newly created static suites now explicitly inherit the plan's default configurations, matching what the Azure DevOps web UI does when you create a suite there.
  • Work-item text sent to the AI is now explicitly marked as data rather than as instructions. Descriptions and acceptance criteria are editable by anyone with access to your Azure DevOps project, and they reach prompts that can read your source, so the model is told plainly to treat that text as the requirement to test and never as directions addressed to it.
  • Dialogs no longer stretch past their own edges when they contain something long, like a work-item title.

macOS users: the .dmg / .app bundles are unsigned. See docs/install-macos.md for how to bypass Gatekeeper on first launch.