Skip to content

Honey v1.3.1: benchmarked output discipline

Latest

Choose a tag to compare

@robertkeus robertkeus released this 03 Aug 15:12
· 29 commits to main since this release
ba137b7

Honey v1.3.1 packages the current multi-tool skill, reproducible benchmark harness, ESON handoff format, and corrected paired per-task reporting.

Benchmark snapshot

  • 29% lower median output across the 23-task mixed suite
  • Up to 70% lower output in focused review workflows
  • 100% objective test pass rate
  • No measurable overall quality loss
  • 100% accessibility pass rate on user-facing tasks
  • 100% lossless recovery in agent handoffs

These results describe different task groups, not one universal saving. Honey does not claim hidden-reasoning reductions or proven dollar savings from this sample.

What is included

  • Rules and installers for Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Windsurf, Cline, OpenClaw, Kiro, Kilo Code, and Hermes Agent
  • Three separate levers for less code, less filler, and compact lossless agent handoffs
  • Safety carve-outs for validation, error handling, auth, accessibility, irreversible work, and explicit requirements
  • ESON tooling and reproducible format benchmarks
  • Public methodology, raw results, and corrections

Honey is an open-source GreenPT product under the MIT license.

Read the method and reproduce the benchmark: https://github.com/Green-PT/honey-for-devs