Skip to content

v0.27.0

Choose a tag to compare

@PhenX PhenX released this 07 Sep 17:22
6e21c6d

The biggest release yet: Piwi now explains why a test failed and what to do about it — deterministic clues with a one-line headline, a cross-project failure inbox with bulk triage, one-click reproduce and git-bisect, flaky attempt diffs, page-structure diffs against the last green run, and fix verification that closes the loop — on rebuilt failure pages and with full Playwright 1.63 support.

Full diff: v0.26.1…v0.27.0

✨ Highlights

  • 🧩 Deterministic failure clues — A rules engine explains every failure in one line before the raw error, no AI required.
  • 📥 Failure inbox & bulk triage — Open clusters across projects land in a Home inbox you can filter, snooze, assign, quarantine and triage in bulk — also through MCP.
  • 🔁 Reproduce & git-bisect — Generate a local repro recipe and a git bisect for a failure, run them from the desktop app, and persist the first bad commit on the cluster.
  • 🎞️ Flaky attempt diff — An Attempts tab diffs a flaky test's passing and failing attempts and feeds the flaky classifier.
  • 🧬 Page-structure diff — Sample green ARIA snapshots and diff the page structure against the last green run.
  • ✅ Fix verification — A partial run can verify a fix when it covers the whole cluster; the fixer is notified and can re-run the affected tests in CI.
  • 🧭 Rebuilt failure pages — Execution, cluster, run, project and history pages rebuilt into one column with a situation block, tabbed evidence and shared building blocks.
  • 🎭 Playwright 1.63 — Step subtitles and params, test locks, browser dialogs, .visible() narrowing and contrast, captured end to end.
  • 🔬 Evidence straight from the trace — Console, network and ARIA are derived from the trace without capture fixtures, on one timeline.
  • 📤 Perfetto export — Export runs and executions as Perfetto traces.

Features

Deterministic clues & one-line failures

  • app: add the deterministic failure-clue engine (d78ebfa), load and serve clues per execution (902d5b3), and surface a clue for every failure (72a10a3)
  • ui: surface deterministic clues on the diagnosis tab (d80c4c2)
  • ai: feed clues to diagnosis and spend the auto-diagnose budget on the unclued (b402cd1); base diagnosis staleness on the current evidence (2d73843)
  • app: parse Playwright errors into a structured record and a one-line headline (3dfb5dd), and explain a failure in one line before the raw error (8d1ca43, a7a6a60)
  • app: compute the story, situation, cluster state and next step for failure pages (4293596, 530e18e)
  • app: give every empty evidence card one of three explained states (92bda73)

Failure inbox & triage

  • app: grow the Home open-failures card into a failure inbox (085e33c); the card (f9b6463, 4ec08f5) and its inbox on the UI side (5123e4f)
  • app: list open failure clusters across projects (7b1031b), serve inbox queue data and triage endpoints (ced8b53), and add an inbox-queue filter to the list_open_clusters MCP tool (f9e77b2)
  • app: bulk-triage failing runs and cluster backlogs (a624d1f)
  • app: quarantine a test from its failure pages (c7244fe), and quarantine, link and triage a failure where it appears (9a2d461)
  • app: pin a known issue and surface the owner on a failure cluster (b6540a4)
  • db: add snooze and assignee columns to failure clusters (99dd2e3)
  • app: add a deterministic cluster-title helper (3ae8690); ui: name clusters deterministically wherever the signature was the name (b4b0b57)
  • ui: show and lift snooze on the cluster page, keeping snoozed clusters out of "failing now" (1005aae); add a triage control and additive slots for the cluster header (ba843d2)

"Fixed before" memory

  • app: resolved-cluster memory — "Fixed before" (f972686), added to fix plans (cb71161) and surfaced in UI, demo and docs (7c7bac6)

Reproduce & git bisect

  • app: add pure repro-recipe and git-bisect generators (c764dae), generate a local repro recipe and git bisect for a failure (96dc2e1), and serve them in the fix plan (44e02d6)
  • app: add a Reproduce section to the Fix card (2e0312c)
  • app: run the reproduction and drive the git bisect from the desktop app (8f4715d, 0d0290c, f8fb09b)
  • app: persist a bisected first bad commit on the failure cluster (fe25536)

Flaky attempt diff

  • app: add an Attempts tab that diffs a flaky test's attempts (fa0eda9), a pure attempt-diff builder (897b2bd), and diff failing vs passing attempts (01b6de8)
  • app: serve the attempt diff at /api/test-run-cases/[id]/attempt-diff (d3911a4), and feed live network and attempt-diff signals into the flaky classifier (75584a2)
  • ui: link attempt chips to their sibling executions (01f696d)

Page-structure (ARIA) diff

  • app: structural ARIA snapshot diff (bb61536), sampling green snapshots and diffing the page against the last green (96a74af), serving and ingesting green page samples (aa9452b), and surfacing the page diff in MCP and the docs (87484e9)
  • ui: add the page-diff view with a clue and healing tie-in (e8362f4)

Evidence from the trace

  • app: derive console, network and ARIA from the trace without fixtures (dc4ed66), and get failure evidence without the capture fixtures (43e673e)
  • app: parse and serve aria/screen trace snapshots (5533178), show the 1.63 aria and screen snapshots (e525611); ui: show trace aria/screen snapshots on the execution page (0589070)
  • app: put a failing execution's evidence on one timeline (775c503, a3d56ca), attributing each timeline action to the method it came from (3816cd8)
  • app: report backend-log evidence on the setup ladder (c141a44)

Fix verification & notifications

  • app: let a partial run verify a fix when it covers the whole cluster (ed06fdf), verify fixes from partial runs, notify the fixer and re-run from the dashboard (48a02f8), and re-run a cluster's affected tests in CI from its page (4ae80a0)
  • notifications: close the loop on fix verification (5033f7e) and tell the person who fixed a cluster (846b314)

Playwright 1.63

  • app: capture 1.63 aria JSON, dialogs, .visible() narrowing and contrast (5c8b43b, b4cc0fb)
  • app: display and analyze 1.63 step subtitles and params (e9e4a7f), keep the step target visible and cap params on ingest (773080e), and feed step params into the headline, clues, AI context and attempt diff (444fd53)
  • app: suggest .visible() to narrow strict-mode locator failures (4b88b8c)
  • app: surface test locks end to end (6712a9f) and in the API, MCP and clues (14075b5); db: store test locks per execution and test case (e6db83f); ui: show test locks on the timeline, rows and filters (8b358e1)
  • ui: render the step subtitle distinctly and add a params disclosure (a30d9da)

Baselines & run insights

  • app: compare a run against a chosen baseline in run insights (0bc1799)
  • app: resolve a run's execution from its file, title and retry (cff8dfb)

Sharding & reporter

  • app: lock-aware selection sharding with a split-lock warning (aa4bb8e, 54b6a9a)
  • reporter: append --add-reporter when the config has no Piwi reporter (a904743)
  • reporter: capture 1.63 step subtitles and params (704ce66, 4f866ef), split the step label and back the locator with step params (9e82c40)
  • reporter: capture test locks from the private _locks field (d1d99c6), the contrast browser option (4af2282), and through visible() and no-selector frameLocator() (dc60188)
  • reporter: default aria trace snapshots on Playwright 1.63+ (71a4d27), default screenshot and trace capture via wrapConfig (aeb863e), and sample the ARIA snapshot on passing tests (d6eb620)
  • reporter: print a dashboard link per failed test (d6d1826, 83e2207), and upload attachments that carry a body instead of a path (8460daa)

Export & companion tools

  • app: export runs and executions as Perfetto traces (b54b283, d691eb9); ui: add a Perfetto trace entry to the export menu (1724937)
  • add a companion-tools card and improve streaming error messaging (47e07e4); ui: point the app at companion tools where their job comes up (f8514cb)

Rebuilt failure, run & project pages

  • app: rebuild the execution page as one column with tabbed evidence (7d5f84d); ui: rebuild its header and evidence as tabs (a96312b); fold the fix and history into one column (644ee35, 81ddf15)
  • ui: one situation block at the top of the execution page (faf51b7, f55c906), on one type scale with a "What changed" line (b81f174, 59b05bd)
  • ui: rebuild the failure cluster page as one column (2a49f5a), its header, triage and evidence (561c1d0); fold fix, changes and history into one column (1a36386, b5b5766); a situation block at its top (1d1220d, 3b1b433)
  • ui: show the fix plan on the cluster page (9ebce74) with diagnosis history (1e0fbe0), the diagnosis version history (c2e4b10), unify the diagnosis panel across cluster and execution scope (f5794c3), and fold the fix toolbox open on the story (a978feb, 84ea2ba)
  • app: rebuild the run page header and grouped Tests tab (c07672b, 8aee995), reduce it to Tests, Changes and Timeline tabs (4a7bd07, 194c8ce); ui: add a File + Describe grouping to the Tests tab (3e798b6)
  • ui: rebuild the project page into five tabs (4d6f4f3, e2d126a) with a shared FilterBar and slow-endpoints view (6f5f303); rebuild the test history page around a facts line and TestRow rows (d675ee4)
  • ui: expand a test into its nested step waterfall on the workers timeline (4c2a17c, 867857e); show the error line and cluster badge on failing tree-view rows (e59297e)
  • db: keep an immutable fingerprint source per failure cluster (f8e8a29) and network request start times (4ea3ff3)
Shared UI building blocks, demo seeds & tooling
  • ui: add DetailHeader and tabbed-evidence building blocks (2fb99a8), FixCard and HistoryStrip (3043084), shared TestRow, TestRowGroup and BadgeGroup (ef4cf44), extend TestRow, HistoryStrip and ChartCard for test lists (7ea9653)
  • ui: render every test and execution list as the shared TestRow (a12a112), and the project catalog, flaky and quarantine lists too (1733b58)
  • app: screenshot any route and seed the dev DB in one command (b123511, 65e8771)
  • demo: mirror the trace snapshot endpoints (4283da9), nest web-dashboard steps so the timeline shows step depth (383f711), populate every failure-inbox queue (67c23e0), seed 1.63 step subtitles and params (1aebbd0), green ARIA baselines for the page diff (bdc1778), and test locks on the checkout suite (50f5795)

Bug Fixes

  • ai: hash only the evidence for diagnosis staleness (21040db); support max_completion_tokens in OpenAI-compatible calls (ac09857, d1356e7)
  • app: gate locator healing on a resolution failure and prefer same-environment baselines (1f33ee7, 4aa928a); prefer same-environment, then same-branch baselines for the diffs (4c68883)
  • app: isolate the in-execution page diff to the change that failed (2504556)
  • app: invalidate a cluster's embedding when its exemplar refreshes (30c90e4); refresh a cluster's exemplar as it recurs while keeping the fingerprint source stable (4df19e6, 4936925)
  • app: let a test timeout's pending action outrank the step title (fcd43c5); read the toHaveCount received count from an unbulleted retry line (2662b1b)
  • app: order the latest runs by start_time, not MAX(id) (e0feb7a, d80681d)
  • app: make the attempt diff resolve every attempt, opened from either side (0438396)
  • app: move "Quarantine all affected" to the cluster navbar (186cffa)
  • app: keep sandboxed HTML reports usable (50c415a, 09bac27)
  • app: serve seeded demo evidence media from the file endpoint (96b6599, 2757550); serve video attachments with media types (e944e1e, c5c3fc6) and address video-attachment review feedback (712e50a)
  • ui: render error text with its ANSI colors instead of raw escape codes (697dc41); render slowest-test durations with units (9c58a16); render the failing locator as code and label the environment diff sides (0158b95)
  • ui: surface failure evidence first on the execution page (5703c39, 4a53bbd); move the blocked-tests list into the Diagnosis tab (2e4430b); merge the failure timeline and steps into one table (cf745bd)
  • ui: default the failure cluster page to its latest occurrence (a8cc1a3); show the locator fix on the cluster page only for a locator failure (59f557a); restore the failure-clusters segment on the project failures tab (15bd7d3); retire the last on-screen fix-plan pointer and verdict label (0d15e16); show empty evidence cards expanded so the setup guidance is visible (bf57392)
  • notifications: link alerts to the execution and quote the error head (392a0d5)
  • reporter: always clear a stale green-sample set at run start (87eb92b); keep unsubmitted runs recoverable and surface delivery failures (a1eba73)
  • demo: shift step start times in the seed rebase (5fe9aa4)
  • desktop: stabilize shell E2E test startup timing (2509bf2)
Responsive & mobile polish
  • ui: align shared detail primitives with the smaller mobile gutters (4855652); align the diagnosis rail breakpoint and gate jump chips on rendered sections (cbc6558)
  • ui: fit the project runs table at 1280 px (a6bc24d); fit the steps table and card actions on phone widths (6f1abd0, f8fe525)
  • ui: flatten the evidence tabs so media uses the full card width (a59208b, e6202ed); keep evidence media inside the tab-panel padding (976757e); tidy the evidence tabs and go full-bleed on phones (f366b41); wrap the evidence tab strip on phones (bd1d95a)
  • ui: give folded toolbox summaries a title on hover (162da6f); keep the bisect visible and open folded toolbox sections in specs (0d42099); keep the detail tab strip a flex row so the tab help hint sits beside it (c3b54b6)
  • ui: make detail cards full-bleed below sm (f14a0e1); shrink card and panel gutters below sm (2c9dd7c); shrink stacked mobile gutters on every page (833d562)
  • ui: fix the AI card copy, the missing-cluster state and the demo Share button (67b37c4); repair visual QA defects across the cluster, project and test pages (1db82e9); stop the navbar repeating the breadcrumb and keep the trail readable (753beaf); the 390 px pass, docs rewrite and vocabulary sweep for the failure pages (df5325e); wrap the test history title on phones (ef8c560)