Skip to content

v1.0.7: Human Input

Choose a tag to compare

@NiveditJain NiveditJain released this 26 Sep 16:40
· 26 commits to main since this release
2a2fef9

The stable cut of the 1.0.7 line. For the failproofai package itself this is a collector release: failproofaid now records what a person actually typed as its own human_input event on every harness, and it ingests OpenClaw 2026.9.6's compressed transcripts, which it previously failed outright. Enforcement is untouched — nothing under src/ changed since 1.0.6, so upgrading cannot change
what your agents are allowed to do.

The rest of the line ships on its own version lines, both already published: @failproofai/sdk 0.0.1-beta.0, the TypeScript telemetry SDK, and fp-cloud-cli 0.0.1, the first stable release of the Cloud CLI.

npm install -g failproofai

Added

  • The session collector records what a person typed as human_input, beside the model_request it already sent. Every harness writes whatever reached the model into the user role — cron wrappers, injected AGENTS.md and environment context, sub-agent hand-offs and the person's own words alike — so "what did the human say" had no answer short of a per-harness guess downstream. Each source now emits a human_input only when its own record says a person wrote the line, across Claude Code, Codex, OpenCode, pi, Hermes and OpenClaw. The model_request stream is byte-for-byte unchanged, so nothing already shipped re-hashes or duplicates. Measured on a production OpenClaw gateway over 7 days: 7,365 user-role turns, of which 596 were a person. (#839)
  • @failproofai/sdk, the TypeScript telemetry SDK, lands at sdk/typescript. The same 15 events, wire format, spool directory and Evaluator v2 protocol as the Python failproofai-sdk, with instrument() adapters for LangChain.js/LangGraph.js, the Vercel AI SDK, Mastra and LlamaIndex.TS — so a fleet running Node agents and Python agents writes into one pipe. Zero runtime dependencies, dual ESM+CommonJS, Node >= 20.9, versioned independently of this package. (#830)
  • Server-authored evaluations run through a parser and an interpreter, not eval and not node:vm. Every property read goes through one function that checks the actual key at the moment of the read, and the interpreter never constructs a function; a worker_threads sandbox with V8 heap limits and a wall-clock terminate() is the resource bound around it. An evaluation that cannot be sandboxed is refused, never run unbounded. (#830)
  • fp issues gains close, archive, unarchive and clear. close is the second terminal state — a resolved issue reopens when its audit finding recurs, a closed one does not — archive/unarchive toggle a flag orthogonal to state, and clear is the bulk "we changed our agents, give us a fresh board" operation: scoped, previewable with --dry-run, and writing no suppression, so a pattern that survived the agent changes reopens rather than staying hidden. (#815)
  • fp-cloud-cli cuts its first stable release, 0.0.1 — its own PyPI line, which does not move this version. (#815)

Fixes

  • OpenClaw 2026.9.6 sessions were not collected at all. That release stores a large transcript event zstd-compressed in event_zstd and leaves event_json NULL, and the collector failed the whole agent database on the first such row. Compressed rows are now decoded, and a row with neither is skipped. (#839)
  • Claude Code prompts carrying a pasted image were dropped entirely — about 7% of typed prompts on the machine measured. (#839)
  • The TypeScript SDK, after user-style tests on six real agent setups against FailproofAI Cloud, across three rounds: emission order is preserved below the millisecond, SIGTERM no longer leaves open runs showing as running forever, configure() is process-wide so a bundled Next.js route reports the right environment, and everything open at exit closes most-recently-opened-first with the crash's cause named. (#830)
  • The evaluator bin refused every CommonJS evaluator file, because the loader's instanceof check was false across the ESM and CJS copies of the class. It now checks a Symbol.for marker both builds share, with a packaging test running both file kinds through the built bin. (#830)
  • Three fp docs contradictions: what a recurring finding does to a resolved issue (it reopens it — said both ways, four lines apart), what clear's preview count actually guarantees, and what --all-audits covers, which is not all issues. (#815)

Docs

  • A Sentiment page under Trace Agents: what the four feelings and three signals are, how to turn scoring on, which messages are scored (only what a person wrote), and how to read the dashboard. (#839)
  • A TypeScript reference page beside the Python custom-agents one, registered in the English navigation and cross-linked both ways, plus references/typescript.md and references/evaluator.md in the failproofai-sdk skill. (#830)
  • Ending an issue three ways (resolve / close / archive) and clearing a board after an agent change, in the audits guide and the Cloud CLI reference. (#815)

Dependencies

  • 32 of the 36 Dependabot alerts open against the TypeScript SDK's integration fixtures, cleared with minimal overrides entries and without moving a single framework under test — the pinned releases pin their transitives exactly, so an in-range npm update reported every affected fixture already up to date. The two that cannot be fixed this way are documented with why, and SECURITY.md now covers the whole path. (#837)
  • 13 routine dependency bumps (#816-#828), and the bundled skills pinned forward twice.

Full changelog: 1.0.6...1.0.7