Skip to content

Releases: kimpearce888/supportos

v2.2.1 — The Polish Release: Full-Project Audit, 45 Fixes, Zero New Features

Choose a tag to compare

@kimpearce888 kimpearce888 released this 28 Sep 12:00

v2.2.1 — The Polish Release: Full-Project Audit, 45 Fixes, Zero New Features

After the roadmap completed at v2.2.0, a completely fresh, independent audit re-examined the entire project from a neutral standpoint: four parallel adversarial passes (backend, frontend, consistency/docs, security) that deliberately did not trust the existing test suite, every finding verified and root-caused before a line was changed, plus a human-like browser pass on the real database. It found 45 confirmed defects in existing functionality. This release fixes all of them — and adds no features and removes none.

🔬 Data correctness

  • Report builder organizations metric threw on every run — the SQL referenced a nonexistent conversations.organization_id column; resolved through the conversation's customer now.
  • Sample conversations on grouped reports were always empty — the sample query bound display labels against id group keys (c.mailbox_local_id = 'Support'); samples now bind the raw group key (NULL groups via IS NULL) and respect the aggregate's filters.
  • Customer memory upsert could mislabel its source — AI writes can no longer clobber human-authored entries; human writes relabel honestly.
  • Campaigns with exhausted unknown-outcome recipients stayed sending forever — completion guard tautology fixed; the completion event discloses unknowns.
  • related_ticket_estimate counted knowledge-FTS self-matches, not tickets — now the distinct conversations whose AI analysis cited the document.
  • Conversation detail crashed on analyses lacking optional fields — found by the browser pass on a conversation no automated test had opened.
  • Status writes are atomic; LIKE wildcards are escaped in customer/organization search.

🖥️ Broken or dead UI functionality

  • The composer's typed reply, AI draft and panel state bled across conversations — a draft for ticket A was one click away from ticket B's customer; the detail pane remounts per conversation.
  • A whole family of CSS utilities was referenced but never defined (.col, .between, gap/margin steps, .btn.tiny, table.compact, …) — stacked panels silently laid out as horizontal rows.
  • The report builder no longer dead-ends after switching to a total-only metric.
  • The OAuth "Authorize via browser" flow is completed — /oauth/callback finishes the advertised exchange server-side and verifies the single-use state parameter.
  • The Copilot citation "open" link navigates; the AI Center evaluation toggle reflects immediately; the snooze modal pre-fills local time.
  • Outreach campaigns no longer silently exclude recipients beyond the first 100 (all matching pages load, bounded by the 5,000-recipient snapshot cap, with honest warnings); organizations beyond the first 50 are reachable through a pager.
  • Eight missing empty-state icons; the mentions mark-read button works; timelines render local relative time; a slow translation can no longer land on the wrong message; Escape closes only the topmost dialog; unknown URLs get a 404 page; Cmd/Ctrl+Enter opens the same send confirmation as the button.

🛡️ Reliability & hardening of existing surfaces

  • Stuck running jobs heal via a periodic sweep; the server handles SIGINT/SIGTERM.
  • Campaign creation is bounded (was: 100,000 recipients evaluated in one request).
  • The e2e port collision that broke CI on multi-core runners is gone — and config files now join the typecheck/lint surface (the invalid sequential: true option was invisible to CI).
  • The unanchored screenshots/ gitignore no longer swallows the README gallery (44 images versioned again).
  • DNS-rebinding guard: non-loopback Host headers are refused — a rebinding page is same-origin from the browser's viewpoint, so the CORS allowlist enforced nothing against it.
  • GET endpoints no longer trigger synchronous rebuilds; the legacy /api/ai/memory route enforces the personality red line; known-issue PATCH, release-events and AI integer params validate with zod (422s, not 500/503s).
  • The CLI db scripts load .env like the server (a custom DATABASE_PATH no longer silently migrates the wrong database).

✅ Verification

672/672 tests green (+25 audit regression locks), 452 black-box audit checks (new section O, 0 HIGH / 0 MEDIUM), and a full human-like browser pass on the real 313-conversation database — every fix verified in the running app, all 20 pages plus detail pages walked with zero console errors.

The same project, with the same scope — fixed, synchronized, tested, and documented.

v2.2.0 — The Memory Release: Support Graph, Agent Coaching, Customer Memory, Performance Guards — Roadmap Complete

Choose a tag to compare

@kimpearce888 kimpearce888 released this 28 Sep 09:57

v2.2.0 — The Memory Release (Roadmap Complete)

The memory release — and the closing release of the six-milestone roadmap: SupportOS now understands how everything in a support operation connects, coaches replies before they are sent, remembers customers safely, and stays fast while doing it.

🕸️ Support Graph (plan phase 34)

A relationship layer over the twelve node kinds the plan enumerates — customer, organization, conversation, known issue, issue cluster, incident, knowledge document, agent, campaign, product, custom object, connector row — implemented exactly as the plan prescribes: relational tables, no heavyweight graph database.

  • Derived edges are computed live from the mirror at read time — 26 parameterized SQL branches, each bounded and label-resolving in the same query (no N+1). There is no second copy of any relationship to go stale.
  • Only human-asserted edges are persisted (support_graph_edges, closed five-relation union: related_to / depends_on / blocks / mentions / duplicate_of) — a human judgment is information the database does not already contain.
  • Products become a first-class registry, derived INSERT-OR-IGNORE from the product strings incidents and known issues already carry; a rebuild can only ever add names.
  • Every edge carries provenance (Help Scout mirror / deterministic / AI-derived evidence / human); connector rows honestly have no derived links — the stats page says so.
  • Graph Explorer page: live counts, node search, relation-grouped neighbor lists, human-edge manager; bounded subgraph exploration (depth ≤ 2, ≤ 250 nodes).
  • Three read-only Copilot tools join the registry (22 total): get_graph_neighbors, get_graph_stats, get_customer_memory.

🎓 Pre-Send Agent Coaching (plan phase 35)

The plan's ten checks in the composer — advisory only, ever: no code path blocks, delays or annotates the send.

  • Nine deterministic checks always compute: unanswered customer questions, duplicated questions, explicit timeframe commitments cross-checked against linked ACTIVE incidents, missing acknowledgment when frustration cues exist, excessive wording (preference-aware), insufficient detail, internal information leakage as verbatim 6-gram spans against the conversation's own notes + linked incident/known-issue internals + cited internal-only documents, wrong customer context (foreign conversation numbers, greeting-name mismatch), communication-preference mismatch with human-override precedence.
  • Two local-model checks (unsupported claims, wrong context) under their own ai_runs type agent_coaching — closed verdicts, honest unparseable failure, honest unavailability when the model is down.
  • The full pass/flagged/not-applicable checklist renders with evidence per finding; the last review persists as an audit trail of what the agent was told.

🧠 Customer Support Memory (plan phase 36)

Composed at read time from the tables that already hold each fact — memory can never drift from its sources.

  • Nine sections: known-issue history, previous resolutions with outcome facts, communication preferences (override state), recurring friction patterns (evidence-pinned, "patterns, not judgments"), support-outcome aggregates, campaign history, product & account facts, human-written entries, AI-extracted entries.
  • Every entry carries source, timestamp, confidence, freshness and evidence — "unknown" means no observation, never a guess.
  • Human-written entries are the only persisted rows; AI rows are immutable (they re-derive); human rows edit and delete freely.
  • The personality red line is enforced in code, twice: a closed quarantine pattern list pulls matching stored entries out of usable memory (listed with a reason, purgeable), and human writes matching the list are refused outright with the policy message.

⚡ Performance Polish (plan phase 41)

  • Bounded the remaining unbounded hot paths (gap report, issue-cluster Copilot tool, radar escalated/rating queries, per-customer observations, effectiveness analysis) — bounds disclosed in the data where they bind.
  • The interaction engine's per-GET history backfill is now guarded by cheap COUNT queries.
  • Targeted indexes: customer-ordered conversation history, observation recency, outreach recipient lookups, issue/cluster reverse links.
  • The project's first performance regression guards: a synthetic 2,000-conversation × 400-customer world with CI-safe budgets plus deterministic EXPLAIN QUERY PLAN index assertions.

Verification

  • 647/647 tests green (+53 over v2.1.0): 8 unit + 24 integration + 8 perf-guard + 13 e2e.
  • Black-box audit section N (446 checks, 0 HIGH / 0 MEDIUM): graph/coaching/memory hostile matrices, injection-as-data with tables asserted intact, the red-line refusal surface, AI-availability honesty with no model running.
  • Human-like browser pass on the real upgraded 313-conversation database: centered the graph on INC-001 with derived edges and provenance, asserted a human edge through the UI (DB-verified), watched coaching flag a "within 2 hours" promise against the active incident and catch a verbatim internal-explanation leak with the incident quoted as evidence, saw the honest LM-Studio-down coaching state, wrote a human memory entry through the UI (DB-verified), had a personality-shaped write refused by policy (zero rows stored), and walked all 20 pages with zero console errors.

Fixed

  • Gap-candidate report totals were computed from the (previously unbounded) row set; now exact via dedicated COUNT queries regardless of the display bound.
  • Copilot get_issue_clusters loaded the full cluster table into JS before filtering; the filter now runs inside SQLite with a row bound.

Upgrade path: migration 016 (products registry, support_graph_edges, coaching_reviews, customer_memories.kind, performance indexes) applies automatically on first boot; the graph, memory and coaching surfaces derive from data you already have. The demo seed grew a human graph edge and a human memory entry (empty demo databases only).

v2.1.0 — The Quality Release: Knowledge Gaps, Post-Resolution QA, Effectiveness, Friction, Translation, Advanced Segmentation, Report Builder

Choose a tag to compare

@kimpearce888 kimpearce888 released this 28 Sep 07:48

v2.1.0 — The Quality Release

The quality release: SupportOS learns to look at its own work. The knowledge gap engine (plan phase 26) turns the v1.x documentation-gap detection into a persisted candidate pipeline with human approval — five deterministic kinds (repeated questions with no covering document, repeated questions the existing docs did not solve, conflicting knowledge pairs, missing troubleshooting steps, undocumented new issues), stable dedup keys so rebuilds refresh evidence without duplicating, and human decisions that survive rebuilds untouched; approving marks a candidate and nothing more — nothing ever auto-publishes into the knowledge base. The post-resolution QA pipeline (phase 27) adds the after-close layer that was deliberately missing: per conversation, a deterministic tier that is always computable (back-and-forth after the first reply, repeated 6-word information spans with thread evidence, messages after close, handoffs from the event engine, coarse question-vs-reply counts, first-response and resolution minutes) plus an optional local-model tier (ai_runs type post_resolution_qa, strictly separate from pre-send draft verification) that reports whether the question was answered, whether responses were evidence-supported and whether the right issue was identified — with improvement suggestions that are recommendations for humans, never actions. The historical response-effectiveness report (phase 28) extends the interaction-outcomes layer into observed style → outcome associations (follow-up rate, clarification rate, resolved-after-first, effort, ratings, sample conversations) with the plan's own rule made structural: the notes lead with "OBSERVED ASSOCIATIONS… not causation", small buckets say they are anecdotal, and the audit greps for causal vocabulary. Conversation friction grows six evidence-pinned detections (phase 29): repeated customer explanations (repeated-span matching), repeated agent questions, troubleshooting loops, repeated handoffs (from the M1 event engine), repeated unresolved interactions (customer-level, 90-day same-tag recurrence) and duplicated information requests (request-category + entity-shape matching or the customer saying so directly) — every finding cites thread ids and excerpts, and every detail line says it is a heuristic, a pattern, never a judgment about a person. Local translation (phase 30) arrives with a deterministic language detector (Unicode script ranges first — Han honestly low-confidence — then function-word frequency for Latin scripts, with unknown as a legitimate answer) and LM Studio–only translation with content-hash caching, a system prompt that preserves technical terms/code/URLs/emails verbatim, side-by-side original/translated display in the conversation detail, and no cloud fallback and no automatic sending, ever. Advanced contact segmentation (phase 31) adds seven condition families to the contact-first engine — organization data/properties, ticket custom fields and channel (same-conversation semantics), support-history waiting and previous issues (clusters/known issues), incident exposure, campaign history including not_received exclusions, deterministic support-health aggregates, custom-object links and customer-event timeline conditions — all compiled set-per-node with whitelisted identifiers and bound values; a natural-language suggestion endpoint lets the local model PROPOSE a definition, which the deterministic engine immediately evaluates — the model never selects recipients and nothing saves implicitly (a strict validator rejects any malformed proposal). Outreach (phase 32) inherits every new condition through the existing audience pipeline — snapshots, dedup, explicit review, DNC, duplicate-send protection, timeout reconciliation, audit trail unchanged. The custom report builder (phase 33) closes the release: 20 local metrics × 14 dimensions compiled from closed catalogs only (injection-shaped configs are clean 422s), filters, date ranges, previous-period comparison rendered as differences, bar/table output, saved definitions — and every metric ships its definition and limitations in the response itself; native Help Scout reports stay under their own tab, labeled by origin. Two read-only Copilot tools join the registry (get_knowledge_gaps, get_friction_report). 594/594 tests green (+53 over v2.0.0), the black-box audit grew a v2.1.0 section M (425 checks, 0 HIGH / 0 MEDIUM after fixes — it caught a real safe-deny gap in history_issue and a latent M4 freshness crash that only fires when repeated questions exist), and the human-like browser pass approved a gap candidate through the real UI (DB-verified), recomputed QA deterministically, saw the honest LM-Studio-down error states for both QA and translation, ran the report builder, exercised the new segment conditions, and walked all 19 pages with zero console errors.

Fixed

  • Latent v2.0.0 freshness crash (f.doc_id on fts_knowledge, which has document_id): only fires when repeated questions exist — exactly the scenario the associated-questions feature was built for; surfaced by this release's demo seed adding a repeated question. Column fixed and regression-covered.
  • history_issue segmentation condition fell through to the known-issues link table for unknown issueKind values (audit section M find — matched instead of safe-denying); now a strict closed-vocabulary check.
  • Outreach meta 500 ("" double-quoted SQL literal parsed as an identifier) in the new org-property stats queries — found by this release's own e2e run.
  • QA first_response_minutes now falls back to the first reply thread timestamp when the maintained derived column is absent.
  • Report builder ai_attribute_share bound its parameters in the wrong positional order (SELECT expression params after WHERE params); parameter assembly now follows SQL position order.

v2.0.0 — The Workspace Release: Incidents, Impact, Custom Objects, Connectors, Timeline, Health, Freshness

Choose a tag to compare

@github-actions github-actions released this 28 Sep 05:45

v2.0.0 — The Workspace Release

Incident Workspace · Issue Impact Intelligence · Radar Extensions · Custom Objects · Local Data Connectors · Customer Timeline · Support Health · Knowledge Freshness

The fourth milestone of the 48-phase roadmap (plan phases 18–25). Eight systems, one discipline: derived over stored, closed vocabularies, idempotent by construction, association wording for every correlation, and explicit human control over anything AI-visible or customer-facing.


🚨 Incident / Master-Issue Workspace (plan phase 18)

When something breaks, fifty tickets arrive about one problem. SupportOS now has one place to run the outage.

  • INC-001-coded incidents with status, severity, owner, product/feature, description, internal vs customer-safe explanations, known cause, workaround and resolution.
  • Conversations link to incidents — one issue → many tickets, the plan's explicit requirement. Linked tickets show an active-incident chip in the Inbox so every agent sees the outage context without leaving the ticket.
  • Affected customers and organizations are DERIVED from the linked conversations at read time — never stored, so they can never drift.
  • Engineering references, releases, related known issues/knowledge/campaigns/custom objects, notes and an append-only timeline (dedup-keyed: re-runs can never double-record) complete the record.
  • Declare an incident in one action from a known issue (explanations carried over, its conversations linked) or from a rising cluster (every member conversation linked).
  • incident_update notifications join the closed 15-type union: declared, status/severity changes, conversations linked.

📊 Issue Impact Intelligence (plan phase 19)

The numbers that matter under pressure, through one shared implementation for incidents AND known issues:

  • Affected conversations, distinct customers (a ticket count is never silently used as a customer count), organizations.
  • First/last seen, 7-day growth with direction, trend, affected inboxes, top tags.
  • Products from the local AI attribute layer — honestly unknown until conversations are analyzed.
  • Open/closed distribution and how many customers are waiting right now.
  • Release correlation is a temporal association only: conversations starting within 7 days after a recorded release date, with the non-causal wording in the data itself.

🧭 Issue Radar Extensions (plan phase 20)

Six new deterministic detections, every alert carrying evidence conversation links and association-only wording:

  • Reappearing issues (quiet 30–60d ago, back in the last 30d)
  • Customer concentration (few customers, many tickets)
  • Inbox concentration (one mailbox dominates the cluster)
  • Release-correlation bursts (≥60% of a cluster inside one 7-day window; releases recorded in the window are mentioned as associations)
  • Repeated unresolved patterns (same customer, same cluster, 14+ day spans)
  • Unusual global volume (last 7d vs previous 7d, ≥40% up)

📦 Custom Objects (plan phase 21)

A local extensible object model — Account, Deployment, Subscription, anything you define:

  • Typed field definitions (text, long text, number, date, boolean, select) with immutable types once objects exist.
  • Values are JSON validated at every write by a Zod schema built from your own field definitions — user-defined data never becomes SQL.
  • Relationship edges to customers, organizations, conversations, known issues, incidents and campaigns, with reverse lookups and per-type reporting.
  • Full-text search over titles and property text.

🔌 Local Data Connectors (plan phase 22)

Approved local data sources, refreshed on your schedule:

  • Four kinds: local JSON, CSV, SQLite file, HTTP endpoint.
  • Snapshot semantics: rows keyed by a key column or content hash; vanished rows pruned; failed refreshes mark health=error and never partially overwrite a good snapshot.
  • Fail-closed SSRF guard (the approved plan adjustment): private ranges, loopback, link-local cloud metadata, CGNAT, IPv6 private/loopback, IPv4-mapped tricks, numeric encodings and internal hostnames are refused — validated at configuration time AND re-validated with DNS resolution before every request (rebinding shapes covered).
  • Files live inside the connectors/ folder jail; HTTP responses capped at 10MB with a 10s timeout.
  • Auth material is always redacted (••••••) in every read.
  • The AI sees connector data only where you explicitly allow it — the Local Copilot's connector tool refuses anything else and tells you what IS visible.

🕒 Customer Event Timeline (plan phase 23)

An append-only, dedup-keyed local event log per customer, independent of ticket history:

  • Signups, conversations started/closed, first customer messages, campaign sends/replies, ratings, incident exposure and linked custom object records.
  • Kinds with no observable source (subscription/account/product/integration events) stay honestly absent until a connector or custom object produces them.
  • Organization timelines are the union over member customers; a rebuild action re-derives history idempotently.
  • Migration 014 backfilled observable history (656 events derived live on the real 314-conversation demo database).

🩺 Customer Support Health (plan phase 24)

Operational facts with plain-language definitions and evidence links — waiting, volume, unresolved known issues, incident exposure, response delays, escalation history, an effort proxy — plus attention flags traceable to specific conversations.

There is deliberately no aggregate "score" and no psychological or personal judgments. A number that summarizes a human invites reading it as a judgment; the schema has no such field and the audit asserts its absence.

📚 Knowledge Freshness (plan phase 25)

Lifecycle observability for the local knowledge base:

  • Created / updated / last reviewed / last verified / version / usage (bumped by local searches) / associated recurring questions.
  • Six deterministic flags: stale content, needs review, conflict candidates (title-term overlap), low usage, articles followed by support tickets (temporal/topic association only), articles tied to questions that keep coming back.
  • Review and Verify are human-only timestamps — nothing is edited or published automatically.

🤝 Copilot tool surface

Four new read-only tools join the registry (17 total): search_incidents (derived counts, the distinct-customers note travels with the data), search_custom_objects (type-narrowed FTS, redacted), get_customer_timeline (bounded recent events) and the gated search_connector_data.


🔬 Verification

  • 541/541 tests green (+79 over v1.9.0): 22 unit + 74 integration + 17 e2e — including a full SSRF-guard suite (range matrix, IPv4-mapped IPv6, DNS rebinding shapes, fail-closed semantics), connector snapshot semantics, derived-count invariants, rebuild idempotence, dynamic-schema validation and hostile-input hardening on every new route.
  • Black-box audit extended (section L — 410 checks): incident hostile-creates, XSS/SQL-shaped payloads stored as data with tables asserted intact, SSRF probes across the full private-target matrix, path-jail escapes, auth redaction asserted in responses, timeline rebuild idempotence, the no-score assertion, and a radar honesty sweep that fails on any causal wording. 0 HIGH / 0 MEDIUM after fixes.
  • Human-like browser pass on the real upgraded demo database: incidents declared from a known issue AND a cluster through the UI, severity changed (notification + timeline event verified in the database), a conversation linked by number (which found and fixed a real bug — the flow queried a search parameter the endpoint never had; replaced with an exact-match ?number= filter), the Inbox incident chip verified, customer timeline + support health walked, a custom object type and object created through the designer, a connector created and refreshed with the inferred schema visible, knowledge freshness exercised with the human-only Review action, and all 18 pages walked with zero console errors.

⬆️ Upgrade notes

  • Forward-only migration 014 (idempotent; observable customer-event history backfilled automatically).
  • The connectors/ folder is created on first use and is gitignored — drop your JSON/CSV/SQLite sources there.
  • HTTP connector targets must be public; the SSRF guard is fail-closed by design.
  • Everything in this release is LOCAL SupportOS data — nothing new is written to Help Scout.

v1.9.0 — The Intelligence Release: Local Copilot, AI Attributes, AI Escalation Rules

Choose a tag to compare

@kimpearce888 kimpearce888 released this 28 Sep 03:46

v1.9.0 — The Intelligence Release

Local Copilot · General AI Attribute Layer · AI Escalation Rules

The third milestone of the 48-phase roadmap (plan phases 15–17). Three systems, one discipline: the AI is never the authority — it reads, explains and recommends; deterministic code decides, stores and writes.


🤝 Local Copilot (plan phase 15)

An interactive, read-only assistant inside every conversation's context pane. Ask it what a rep actually asks:

What is this customer asking? What happened in their previous tickets? Have we seen this issue before? What solved the previous cases? What documentation applies? What should I check before replying? Why is this ticket currently considered urgent? Show evidence for that answer.

  • Runs fully locally via LM Studio — no cloud LLM, ever. If AI is disabled or LM Studio is unreachable, the Copilot says so (honest 503s and visible error states) instead of pretending to work.
  • The model never sees SQL. It can only call the allowlisted read-only tool registry — now 13 tools, including 7 new ones: conversation context, customer history, similar conversations, issue clusters, AI analyses, SupportOS metadata, and the AI attribute snapshot.
  • Every tool result is server-validated, bounded and redacted before it reaches the prompt.
  • The tool loop is hard-bounded (max rounds + max tool calls): a tool-looping model gets stopped and must answer from the evidence it already has — honestly.
  • Citations are machine-generated. The source list is built by the server from the tool executions it actually performed — the model cannot fabricate a source that survives. Answers must cite evidence with inline markers.
  • Read-only by construction: there is no write path anywhere in the Copilot. Sessions and messages persist locally; every turn is audited (ai_involvement=true) and recorded in ai_runs with model, prompt version and latency.
  • Deterministic starter questions personalize from local facts (prior tickets, known-issue links, analysis state).

🏷️ General AI Attribute Layer (plan phase 16)

Every conversation gets first-class local attributes — in a closed 14-key catalog:

intent · product · feature · issue · urgency · frustration cues · technical familiarity · customer goal · question count · risk · known issue · issue cluster · response style · escalation signal

  • Two layers. Deterministic slots (urgency, question count, risk, known-issue links, escalation signal…) are computed from observable local facts with zero AI — they always exist. AI slots (intent, product, feature, issue, customer goal, response style) are extracted by LM Studio and stored only with enum-closed values and evidence excerpts whose thread references were actually in the prompt — otherwise they are dropped.
  • Versioned. Every recompute supersedes prior rows (superseded_at) and preserves full history per attribute.
  • Honest unknowns. A key with no current row is unknown — never fabricated, never a guess wearing a confidence badge. Attributes never overwrite Help Scout source data; they live in their own table with their own schema version.
  • Searchable, filterable, reportable.
    • A live inbox filter (aiAttribute / aiAttrOp / aiAttrValue) — compiled through the same viewEngine code path as saved Inbox Views, so a live filter and a saved view can never disagree about what urgency: high means.
    • Saved Inbox Views and Outreach segments can carry ai_attribute conditions.
    • The AI Center → Attributes tab shows a coverage report (known vs unknown per attribute, honest percentages) with a searchable conversation drill-down.
    • A per-ticket snapshot card in the AI sidebar: confidence, source (deterministic/AI), expandable evidence excerpts, honest-unknown list, one-click recompute.

⚡ AI Escalation Rules (plan phase 17)

Automation conditions grow two AI-aware fields:

  • ai_attribute — requires a catalog key (urgency, risk, question_count, …); operator semantics follow the attribute's value type; ordered enums compare by vocabulary position.
  • ai_verification — failed / passed / none, from the latest AI draft verification.

The safety invariant is unchanged: conditions only decide whether a rule matches — every action flows through the existing read / non-destructive / higher-risk approval tiers. "High urgency → review queue" is safe because the attribute decides nothing about the customer; the queue is internal, and the approval tier is unchanged. A missing attribute reads as unknown: it matches equals unknown and never a concrete value.


Verification

  • 462/462 tests green (+54 over v1.8.0): 17 unit (catalog + repository contract + viewEngine compilation), 22 integration (deterministic layer with AI off, tool registry surface, Copilot tool loop with a deterministic fake model, escalation-rule matching), 15 e2e over real HTTP (attribute surface, live filter parity, saved views with attribute conditions, the full Copilot chat loop, the escalation-rule lifecycle, 4xx hardening).
  • Black-box audit section K (383 checks total): injection-shaped attribute keys/values, LIKE-wildcard probes, live-filter parity asserted against the snapshot, hostile copilot bodies, closed-vocabulary automation conditions — final 0 HIGH / 0 MEDIUM. The audit found and fixed real bugs before release: copilot citations that failed to deep-link conversations, and an unknown-session 503 that should have been a 404.
  • Human-like browser pass: the Copilot tab (starter questions, honest offline error, session persistence), the attribute snapshot card (recompute, honest unknowns), the live filter through the real UI (URL-backed, honest notes, parity with the drill-down), the AI Center tabs, the automation rule form with the closed catalog selector — and all 17 pages walked with zero console errors.

Upgrading

Migration 013 (m3_copilot_attributes) applies automatically on first run — it creates ai_attributes, copilot_sessions and copilot_messages. No backfill is forced: a bounded worker job (≤500) computes attributes for previously analyzed tickets, everything else computes lazily on its next analysis or on demand. Pre-1.9.0 history honestly reads as unknown.

Full changelog: CHANGELOG.md

v1.8.0 — The Collaboration Release: Operations Center, Workload & Capacity, Notification Center, Mentions, Side Threads

Choose a tag to compare

@kimpearce888 kimpearce888 released this 28 Sep 01:43

M2 (plan phases 10–14) of the approved 48-phase roadmap.

What's new

🖥️ Operations Center (phase 10)

The whole support operation on one screen — 16 live tiles (unassigned, needs first response, customer waiting, waiting over threshold, urgent, SLA at risk / breached, high customer effort, repeated issue, known issue, AI escalation, issue spike, automation approvals, failed jobs, sync problems, campaign activity), scoped by mailbox, live over SSE. Every conversation tile is a COUNT(*) over one whitelisted parameterized fragment — the same fragment powers its inbox drill-down (?ops=<tileKey>), so a tile number and the list behind it can never disagree (parity-locked by tests and a live audit check).

👥 Team workload & capacity (phase 11)

Per-agent and per-team workload with tiered weighted pressure (each conversation counts once at its highest tier), an explicit configurable capacity model — never inferred from anything about a person — availability from the Help Scout user statuses the mirror already syncs, an honest 7-day average-load approximation, and a suggested assignee that is a read-only recommendation with its reasoning printed next to it. Nothing ever reassigns automatically.

🔔 Notification Center (phase 12)

14 notification types produced by one idempotent sweep — the single producer, single funnel: customer replies, assignments, @mentions, SLA risk/breach, automation approvals, AI escalations, known issues, issue spikes, campaign replies, sync/job failures, customer events. Dedup keys make re-syncs, re-sweeps and crashes structurally incapable of duplicating. Targeting (assignee / mention / broadcast), per-type preferences, unread badge live over SSE, mark read / all-read, source links, retention pruning. The sweep defers its cursor until the first sync settles — history is not news.

💬 Mentions (phase 13)

@agent and @team in internal notes and side-thread messages with exact identity matching only (never prefix/substring — @al never notifies Alex; unknown tokens stay plain text), @autocomplete in the composers, mention highlighting, and a "mentions for me" queue linking back to conversations.

🗂️ Side collaboration threads (phase 14)

Internal-only team discussions attached to a conversation (Support / Engineering / Billing style) — never customer-visible, never synced to Help Scout, by construction. Participants with mention auto-join, resolve/reopen (409 on resolved), existence-validated ids, full audit history.

Quality

  • 408/408 tests (+64 over v1.7.0): mention-parser unit suite, sweep/side-thread/operations integration suites, full-loop e2e (webhook → sweep → SSE → badge → mark read)
  • Black-box audit extended with section J (361 checks, 0 HIGH / 0 MEDIUM after fixes) — hostile scope params, tile/drill parity, hostile capacity models, mention-token identity guessing, XSS payloads, conflict cycles
  • Human-like browser pass: onboarding walked, live SSE badge verified updating without refresh, all flows exercised through the real UI, all 16 pages walked with zero console errors
  • 8 real bugs found and fixed by this release's own testing (each regression-locked): fresh-install notification spam, sweep dropping assignment metadata, a workload crash on a misnamed column, duplicate query-param 500, FK 500s on unknown participant ids, silent non-boolean read flags, capacity key validation, sync_state JSON quoting

Full changelog: https://github.com/kimpearce888/supportos/blob/main/CHANGELOG.md

SupportOS v1.7.0 — The Activity-Intelligence Release

Choose a tag to compare

@kimpearce888 kimpearce888 released this 27 Sep 23:49

What's new

SupportOS stops treating conversations as rows with two timestamps and starts treating them as event histories with a lifecycle.

🕰️ Conversation activity engine

A normalized local event log derived honestly from the Help Scout mirror — messages, notes, lineitem action records, observed changes, local writes — deduplicated with stable keys and honestly sourced. Help Scout exposes no historical change log, so observations are labeled as observations (source: sync, observed: true) and conversations with incomplete thread history report unknown instead of guessing.

🎯 Derived activity fields + deterministic response states

14 indexed activity timestamps per conversation (first response, waiting-since, last status/tag/field change…) power a deterministic response-state machine: Needs First Response / Customer Waiting / Agent Waiting / Recently Responded / Never Responded / Closed / Snoozed / Unknown — one SQL CASE shared by the list, the views and the detail view. No LLM anywhere in the classification.

📅 Date & activity filters (DST-safe)

14 activity fields × 16 date modes in your IANA timezone, with calendar-day and rolling-window semantics labeled distinctly. Verified against the 23h spring-forward day, the 25h fall-back day, Lord Howe's 24.5h day and Kathmandu's 5:45 offset — which caught and fixed a real dayjs timezone-plugin DST bug before release.

💾 Saved Inbox Views

Structured condition trees (AND/OR groups, 21 condition kinds) compiled to parameterized SQL at open time — a view saved with "today" always means the day it's opened. Save-time compile checks, dry-run preview, versioned definitions.

🚩 Priority & custom ticket states

A local SupportOS priority (None→Urgent, optional Help Scout custom-field mapping off by default) and a configurable state layer (New / Investigating / Waiting on Customer / Waiting on Engineering / Ready to Verify / Resolved + your own) with full transition history, per-state lifecycle metrics and bottleneck ranking — layered on Help Scout status, never replacing it.

Verification

  • 344/344 tests (+56): DST boundaries across four zone shapes, response-state SQL/JS equivalence, event derivation + dedup across re-syncs and rebuilds, view compilation for every condition kind (incl. injection-shaped values), saved-view dynamic re-resolution, the v1.4→v1.7 in-place upgrade path
  • Black-box audit extended to 335 checks (0 HIGH / 0 MEDIUM): hostile filters, hostile view definitions, injection-shaped payloads, timeline chronology, rebuild idempotency
  • Human-like browser pass on the real v1.6.0 demo database upgraded in place (314 conversations → 672 derived events): all 14 pages walked with zero console errors

Full changes: CHANGELOG

Installers attach when the desktop workflow finishes. Source: clone and npm run build — or try the 2-minute demo (LOCAL_DEMO_MODE=true, no credentials needed).

SupportOS v1.6.0 — the hardening release (second neutral audit)

Choose a tag to compare

@kimpearce888 kimpearce888 released this 27 Sep 22:27

SupportOS v1.6.0 — the hardening release

A second full neutral audit — three independent adversarial passes (server core, data/sync layer, React client) plus a human-like usage pass — found 2 HIGH + 28 MEDIUM issues. All fixed, each with regression coverage. No new features: v1.6.0 makes everything that already shipped behave the way it already claimed to.

Highlights

HIGH — fixed

  • Outreach send-queue livelock: recipients that exhausted 3 retryable attempts stayed queued — unclaimable but still counted — so sendBatch re-enqueued itself forever and campaigns could never complete. Exhausted rows are swept to failed; Retry failed now resets the attempt budget (it was a silent no-op).
  • Docs embedding churn: every 5-minute incremental sync re-chunked EVERY article, destroying all stored embeddings even when the text was byte-identical — permanent re-embedding, permanently lagging semantic search. Content-hash-gated re-chunking now (migration 010).

Silent-dead features — verified live before/after

  • The entire v1.4.0 webhook-push client UX never fired: the browser's EventSource never subscribed to the server's conversation event — one missing word in a listener list, invisible to tests that asserted the wire event instead of the toast.
  • Write-behind writes never told anyone they landed (found in the human-like browser pass): bulk ops ack "queued", the client invalidates, races the worker, reads stale state. The worker now emits conversation-updated after each completed write — an externally-applied tag appears in an open conversation view within ~2s.
  • Error states on every flaggable query, onError toasts on ~20 silent mutations, Settings forms gate on loaded data (were silently overwriting real config with defaults), 300ms debounce on the per-keystroke audience preview, safeExternalHref allowlist against javascript: hrefs, toast cap, keyboard-operable inbox rows.

Server input hardening — every crash reproduced live before fixing

  • Rate limiter bypass closed: it keyed on spoofable X-Forwarded-For (310/310 requests passed with header rotation); now keyed on the socket address.
  • Six endpoints 500'd on NaN query params; seven routes crashed on missing bodies; numeric name/text 500'd campaign create + simulate-incoming; a 2MB search query crashed FTS5; queue retry/cancel lied with ok:true — all clean 4xx responses now, each with an e2e test.

Data/sync correctness

  • Incremental sync coverage gap: organizations + property definitions only synced during initial sync — post-initial creations never appeared, so segmentation conditions on them could never match.
  • The ISO-vs-datetime('now') date-comparison bug class (12+ sites, up to 24h boundary skew): segmentation date windows, issue trends, reply detection, retention pruning — all normalized to julianday().
  • The satisfaction.ratings webhook never stored ratings on the real provider (read a phantom field); same-millisecond local tag collisions threw UNIQUE; soft-deleted customers never resurrected; FTS ghost rows outlived their threads; backups ran 4x faster than configured and never pruned; failed embedding chunks retried forever; approval jobs were silently completed as no-ops.

Numbers

  • Tests: 259 → 288 (29 new regression tests; every finding with a reproducible failure mode is locked by a named test)
  • 320-check black-box audit script re-run against the fixed build: 0 findings
  • Human-like browser pass: webhook push toast + live list refresh, bulk tag appearing in an open conversation within one worker tick, full reply flow, every page walked with zero console errors

See the CHANGELOG for the complete list and the README for the story and the decision log.


Desktop installers (MSI / NSIS / DMG / AppImage) are built by the tag-triggered workflow and attached below when ready.

⚠️ Upgrading from v1.5.x: migration 010 applies automatically and idempotently on first boot (adds docs_articles.content_hash + embedding attempt counters). Nothing else changes on disk.

SupportOS v1.5.0 — Client Segmentation & Outreach

Choose a tag to compare

@kimpearce888 kimpearce888 released this 27 Sep 09:52

v1.5.0 — the contact-first release

Client Segmentation & Outreach (the full spec), vector search over tickets/threads, business-hours-aware SLA alerts on the Issue Radar, and optional end-to-end encrypted sync for multi-device — built under a fresh independent audit that found and fixed 7 real bugs before shipping. 259/259 tests green (+35).

📣 Client Segmentation & Outreach

  • Contact-first deterministic segment engine over the local mirror — properties answer "which customers?", tags answer "which tickets?", the resolver answers "which customers own those tickets?". The AI may suggest or explain a segment, but it can never decide who gets emailed.
  • Conversation-level tag semantics (ANY / ALL / NONE) resolved before mapping to contacts: "has ALL of timezone, bug" requires ONE conversation carrying both tags — locked by the spec's critical test cases.
  • Why-selected evidence on every row: matched property values, matching tickets with tags/status/dates — the recipient review shows it inline, with a matching-tickets drawer linking into the inbox.
  • Customer property VALUES are now synced (with a raw_json backfill that heals pre-1.5 databases), plus background/age/gender/location on customers.
  • Saved versioned segments (structured condition trees, never SQL); campaign recipients are a static snapshot with the evidence that selected them.
  • One individual Help Scout conversation per customer via POST /v2/conversations — never a shared BCC send; customer identified by id (no accidental duplicate contacts); sends ride the same rate-limited queue as manual replies in small batches.
  • Safe lifecycle: validation before queueing, Do-Not-Contact enforcement, duplicate-send protection, per-recipient states with attempt log + audit trail, timeout → unknown → reconcile-before-retry (never blindly resent), pause/resume/cancel/retry, crash recovery.
  • Personalization with preview ({{first_name}}, {{last_ticket_number}}, …) through the same code path as the send; reply intelligence from the local mirror, honestly labeled.
  • UI: 4-stage wizard (audience → recipient review → compose → explicit final review), campaign monitor with audit events + reports, segments & DNC managers, real-time SSE progress.

🧮 Vector search over tickets/threads

Conversations (subject + customer + tags + thread bodies) are chunked and embedded locally; POST /api/search fuses FTS5 + semantic retrieval with Reciprocal Rank Fusion. Qdrant accelerates when connected; a local cosine scan answers without it. Every hit records which retriever found it.

🚨 Business-hours-aware SLA alerts

GET /api/issues/sla-alerts ages open conversations in business minutes since the last customer message against per-mailbox first-response/resolution targets — breached and at-risk states with per-mailbox rollups, rendered at the top of the Issue Radar. Unconfigured mailboxes are labeled honestly; nothing is guessed.

🔐 Optional end-to-end encrypted sync

.sosync bundles (AES-256-GCM + scrypt): export with a passphrase, move the file however you like, import with integrity + schema checks and an automatic safety backup. No relay server exists by design — a privacy-first product is end-to-end encrypted by construction. Attachments re-download from Help Scout automatically on the other device.

🛡️ Fixed by the fresh independent audit (not the existing test suite)

  • hostile deeply-nested condition trees crashed the whole server (DoS) → depth/node caps, clean 422
  • malformed JSON bodies returned 500 → clean 400
  • POST /api/outreach/dnc with a negative id → FK 500 → validated
  • a mid-batch crash stranded recipients in sending forever → reclaimed at batch start, proven by a crash-recovery test
  • campaign validation/report counts truncated at 1000 recipients → SQL aggregates
  • property backfill expected the wire shape while raw_json stores the normalized shape → both accepted
  • campaign monitor inbox links used the conversation number instead of the local id

Full changelog: https://github.com/kimpearce888/supportos/blob/main/CHANGELOG.md
Docs: README · API integration · Testing


Installers

Download the installer for your platform below (built on native OS runners in CI):

  • Windows: SupportOS_1.5.0_x64_en-US.msi (or the NSIS -setup.exe)
  • macOS: SupportOS_1.5.0_universal.dmg
  • Linux: SupportOS_1.5.0_amd64.AppImage

First run offers a 2-minute demo mode — no Help Scout credentials needed.

SupportOS v1.4.0 — the real-time release

Choose a tag to compare

@kimpearce888 kimpearce888 released this 27 Sep 08:01

SupportOS v1.4.0 — the real-time release

The three follow-up roadmap items are shipped — plus two serious latent bugs found and fixed in the job pipeline underneath the webhook path. 224/224 tests green (+52).

🪝 Incoming webhook push for conversations

  • Register webhooks from the app (Sync Health → Webhook push): POST /api/webhooks/register creates the webhook in Help Scout with your locally configured secret; delete just as easily
  • Conversation changes push in seconds: a convo.* event enqueues a single-conversation sync, and when it lands every connected client receives a real-time conversation-updated SSE event (id, number, subject + an honest reason: webhook)
  • Restart-safe: persisted-but-unprocessed webhook events are drained automatically on boot
  • Demo it end-to-end: POST /api/demo/simulate-webhook drives an event through the exact production path — HMAC-signed self-POST → persist → dedup → job → worker tick → mirror update → SSE

🔎 Semantic docs search (local embeddings + optional Qdrant)

  • Docs mirror articles are chunked and embedded by a background job; vectors are stored locally in SQLite, so semantic search works without Qdrant — and gets ANN speed when Qdrant is running
  • Hybrid retrieval: GET /api/docs/search fuses FTS5 keyword and semantic vector result lists with Reciprocal Rank Fusion; every hit records whether keyword search, semantic search, or both found it, and a mode note explains exactly what ran
  • Honest degradation at every layer (no model → FTS only with setup instructions; Qdrant down → local cosine scan; provider unreachable → FTS only, retried next query)

⏱️ SLA / business-hours reporting per mailbox

  • Per-mailbox schedules: IANA timezone, active weekdays, open window — plus first-response and resolution SLA targets — edited in Settings → Business hours
  • Reports → SLA & business hours: first-response and resolution in wall AND business minutes (median included), met/missed against targets, and live "currently waiting" aging with at-risk counts
  • The business-minutes engine is pure and DST-safe (built on the platform timezone database, unit-tested across a spring-forward transition and a half-hour zone); invalid input returns null instead of a fabricated number

🐛 Two latent job-pipeline bugs fixed (found while wiring the webhook e2e)

  • Every background job was permanently unclaimable at runtime: jobs.run_at was written in ISO-8601 (...T...Z) while the claim loop compares against SQLite datetime('now') (... ...) — 'T' > ' ' lexicographically, so webhook-triggered syncs, attachment downloads, AI jobs and embedding passes all sat queued forever since v1.0.0. Tests stayed green because they called components directly, skipping the claim loop.
  • Job payloads reached the worker as unparsed JSON strings, so payload.remoteId read as undefined and sync jobs "completed" without syncing anything.
  • Both fixes ship with regression tests that enqueue → claim → execute exactly as the worker does.

Full changelog: CHANGELOG.md