Skip to content

Releases: zning1994/brainharness-autoresearch

BrainHarness Autoresearch v0.3.0

Choose a tag to compare

@zning1994 zning1994 released this 03 Oct 20:35

从 brainforge-autoresearch 迁移为 brainharness-autoresearch,统一仓库、技能和插件名称。

  • 更新技能入口、元数据和安装命令。
  • 修正 README 中过期的 ClawHub 名称说明。
  • autoresearch.py 及优化流程保持不变。

安装:

/plugin marketplace add zning1994/brainharness
/plugin install brainharness-autoresearch@brainharness

已有 @brainforge 安装不会因 GitHub 更名而自动变为 @brainharness。本次 GitHub Release 与 ClawHub 审核独立,ClawHub 可下载版本以平台状态为准。

验证:插件清单检查及独立配置目录中的实际安装均通过。

完整变更

v0.2.5 — ClawHub republish under brainforge-autoresearch

Choose a tag to compare

@zning1994 zning1994 released this 23 Apr 02:33

Rename-related republish. No functional changes to the optimizer.

Changed

  • SKILL.md `name` field: `autoresearch` → `brainforge-autoresearch`
  • plugin.json and SKILL.md metadata `homepage` → new repo URL
  • Published to ClawHub as `brainforge-autoresearch`; old slug `openclaw-autoresearch` merged in as redirect

Compatibility

Full diff

v0.2.4...v0.2.5

v0.2.4 — Self-optimized SKILL.md

Choose a tag to compare

@zning1994 zning1994 released this 28 Mar 16:51

Self-optimization dogfooding

Two rounds of autoresearch run on its own SKILL.md:

Round Provider Baseline Best Experiments
1 MiniMax M2.7 48.6% 68.1% 15 (3 keep)
2 Claude (improved evals) 80.6% 91.7% 15 (1 keep)

Bug fixes discovered by self-optimization

Both rounds independently found the same documentation bugs:

  • LLM eval field names: pass/fail → pass_description/fail_description (docs didn't match API)
  • contains/not_contains params: value (string) → values (list)
  • Missing fields: LLM eval table now includes required type and name

Other improvements

  • Procedure section: added 3-step summary upfront for agent clarity
  • Step 4/5 merged into actionable "Review results and apply changes"
  • Example restructured: "create eval → run → review" flow
  • Added self-eval.json for future dogfooding runs

Full changelog

v0.2.3 — Configurable HTTP Timeout

Choose a tag to compare

@zning1994 zning1994 released this 25 Mar 05:25

Added

  • --timeout CLI flag to control HTTP timeout per LLM call (default: 180s)

Changed

  • Default timeout: 60s → 180s (large prompts were causing mutation failures on Anthropic API)
  • Timeout stored on provider instance, fully configurable

Usage

# For large style guides or slow APIs
python autoresearch.py --target SKILL.md --evals eval.json --timeout 300

Full changelog: CHANGELOG.md

v0.2.0 — MiniMax Thinking Support + ClawHub Ready

Choose a tag to compare

@zning1994 zning1994 released this 25 Mar 00:45

Fixed

  • MiniMax extended thinking support — models with thinking blocks no longer crash
  • Fallback to thinking content when text block is missing
  • LLM judge max_tokens: 16 → 256
  • Multiple field name inconsistencies (rule/check, pass/pass_description, value/values)
  • test_inputs: accept both strings and {name, input} objects
  • contains rule: support match "all" mode
  • Convergence counter double-increment bug
  • regex uses re.MULTILINE, word_count is CJK-aware

Added

  • --model CLI flag
  • ClawHub metadata in SKILL.md (openclaw.requires)
  • .clawhubignore

Tested

  • brain-search skill: 37.5% → 54.2% pass rate in 5 experiments (MiniMax M2.7)

Full changelog: CHANGELOG.md

v0.1.0 — Initial Release

Choose a tag to compare

@zning1994 zning1994 released this 24 Mar 15:45

Autonomous skill prompt optimizer based on Karpathy's autoresearch methodology.

Highlights

  • Zero-dependency Python script (stdlib only, Python 3.9+)
  • Hybrid eval: rule-based checks + LLM-as-judge
  • Supports MiniMax, OpenAI, Anthropic (auto-detect)
  • Optional live dashboard (Chart.js)
  • Compatible with npx skills add, OpenClaw ClawHub, and standalone use

Install

npx skills add zning1994/openclaw-autoresearch

See CHANGELOG.md for full details.