Skip to content

Releases: xli498/mimo-stable

mimo-stable 1.1.5

Choose a tag to compare

@xli498 xli498 released this 28 Aug 16:33

mimo-stable 1.1.5

  • Added framework-neutral event normalization and policy APIs.
  • Added public API specification, roadmap, packaging checks, and regression coverage.
  • Hardened public API validation and redacted evidence handling.
  • Expanded CI quality coverage to Python 3.13 and 3.14.
  • Normalized Skill frontmatter to the OpenClaw official schema.

Validated locally with pytest (3 passed), fixture benchmark (9/9 passed), package build, isolated API/CLI smoke tests, version checks, and RECORD verification.

v1.1.4

Choose a tag to compare

@xli498 xli498 released this 24 Aug 11:48

Fixed

  • Corrected the package metadata layout so isolated builds and PyPI publishing validate successfully.
  • Includes the changelog cleanup, PyPI links, keywords, classifiers, and English package README metadata.

Verification

  • Wheel build passed.
  • Version consistency check passed.
  • Detector behavior tests passed.
  • Fixture benchmark: 9/9 cases passed.

v1.1.3

Choose a tag to compare

@xli498 xli498 released this 24 Aug 11:43

Fixed

  • Removed stray diff prefixes from the Unreleased changelog entries.
  • Added PyPI project links, keywords, and Python package classifiers.
  • Switched PyPI package metadata to the English README.

Verification

  • Version consistency check passed.
  • Detector behavior tests passed.
  • Fixture benchmark: 9/9 cases passed.

v1.1.2

Choose a tag to compare

@xli498 xli498 released this 24 Aug 11:01

Fixed

  • Corrected the Chinese and English installation instructions to use the published PyPI package.
  • Clarified that mimo-stable provides the mimo-loop-detect CLI and does not expose a mimo_stable Python import API.

Verification

  • Version consistency check passed.
  • Detector behavior tests passed.
  • Fixture benchmark: 9/9 cases passed.
  • Public PyPI installation and CLI smoke test verified after publication.

v1.1.1

Choose a tag to compare

@xli498 xli498 released this 24 Aug 10:46

Highlights

  • Added the English project overview and architecture documentation.
  • Added public integration examples for text detection, tool-call detection, and recovery-policy integration.
  • Added GitHub Actions Trusted Publishing workflow for PyPI.
  • Clarified Python 3.10+ support and release-quality verification boundaries.

Verification

  • Detector behavior tests passed.
  • Fixture benchmark: 9/9 cases passed.
  • Public package installation and CLI smoke test verified after release.

v1.1.0 — 跨模型退化循环检测与工程止损

Choose a tag to compare

@xli498 xli498 released this 24 Aug 05:40

项目定位

这是一个面向国产及其他大语言模型(LLM)的退化循环现象记录、检测与工程侧止损工具。

项目最初来自 MiMo 在特定任务和参数条件下出现重复输出的案例,后续将经验抽象为可复用的检测与恢复决策层。类似现象也可能出现在 GLM 或其他模型中,但不同模型的触发概率、表现形式、根因和有效参数条件不能直接类推。

本项目分享的是:

  • 如何识别输出或工具调用是否进入退化循环;
  • 如何在工程侧限制资源浪费和重复副作用;
  • 如何把案例记录成可复盘、可比较的证据;
  • 如何把检测信号交给上层运行器做保守恢复决策。

本项目不宣称自动修复模型、不宣称彻底解决某个模型的问题,也不提供跨模型故障率结论。

本版本包含

退化循环检测器

scripts/detect_loop.py 支持:

  • 连续相同或高度相似文本块检测;
  • 相同工具及相同参数的连续调用检测;
  • 重复副作用工具调用的提前暂停信号;
  • 显式开启的中文任务语言漂移检测;
  • 默认持续时间门控,减少短时正常重复造成的误报;
  • 显式 --text-mode instant 即时检测模式;
  • JSON 摘要输出,便于接入其他运行器;
  • 参数、输入和 JSON 类型校验;
  • stdin 流式处理,按输出块到达时实时评估,不再等待读完全部输入。

纯决策恢复层

scripts/recovery_policy.py 只负责根据检测摘要输出动作,不执行任何副作用:

  • continue:未检测到循环;
  • pause_and_review:检测到可能重复的副作用工具调用;
  • stop_and_retry_once:首次且符合条件的可重试信号;
  • stop_and_escalate:不再继续自动重试,交给上层或人工处理。

停止生成、切换模型、重试和人工复核仍由调用方负责。检测器和恢复策略不会自行调用工具、切换模型或执行重试。

工程质量与回归

  • 9 个脱敏 fixture benchmark 案例;
  • 连续输出、短时重复、近似文本、参数变化重试、非连续工具调用、工具 key 顺序、副作用重复和语言漂移回归测试;
  • 非法参数和非法 JSON 输入测试;
  • 流式 stdin 行为测试;
  • GitHub Actions 质量检查;
  • pyproject.toml 安装入口;
  • mimo-loop-detect CLI smoke test;
  • 真实项目检查脚本 scripts/test_short.sh 和 scripts/test_long.sh。

快速开始

python3 scripts/detect_loop.py --json --timeout 60 --log fixtures/loop_detected.log

退出码:

  • 0:未检测到循环;
  • 1:检测到循环;
  • 2:参数或输入错误。

将检测摘要交给恢复决策层:

python3 scripts/detect_loop.py \
  --json \
  --timeout 60 \
  --log fixtures/loop_detected.log \
  | python3 scripts/recovery_policy.py --retryable

默认文本检测仍然使用持续时间门控;如果只是需要立即判断重复信号:

python3 scripts/detect_loop.py \
  --json \
  --text-mode instant \
  --log fixtures/repeated_but_short.log

证据边界

当前仓库中的 fixture 和历史日志只用于说明行为和防止回归,不能证明:

  • 某个模型的普遍故障率;
  • 某个参数一定能阻止退化循环;
  • MiMo、GLM 和其他模型具有相同的故障机制;
  • 检测器在所有真实生产日志上都不会误报或漏报。

分享新案例时,建议记录模型精确版本、供应商/端点、时间、任务类型、实际参数、上下文规模、工具调用情况、检测标准、检测延迟和重试结果,并在公开前脱敏。没有充分证据时,使用“观察到”“可能相关”“尚未复现”,不要把相关性写成根因结论。

当前局限

  • 检测器依赖输出块边界;stdin 使用空行切分,其他日志格式需要调用方先适配;
  • 工具调用识别基于常见文本格式,不是所有供应商原生协议的完整解析器;
  • 语言漂移检测目前只对显式 --expect-language zh 开启;
  • 相似度、重复次数和持续时间阈值需要按模型、任务和运行器单独校准;
  • 重复副作用调用可能是合理的幂等重试,生产系统必须结合幂等键、调用结果和重试原因二次判断;
  • 当前为依赖极少的离线工具,不包含真实模型 API,也不代替供应商或框架官方监控能力。

后续方向

  • 收集更多脱敏的跨模型真实案例;
  • 增加不同日志格式的适配器;
  • 在有足够样本后,报告限定条件下的误报、漏报和检测延迟;
  • 完善多模型、多端点和多轮工具调用场景的独立 benchmark。

变更范围

本版本发布在 feat/degenerate-loop-guardrails 分支,对应提交:b169ea8。

它不会修改或宣称修改模型本身,只提供检测信号和保守的工程处置建议。