Repository navigation
Releases: xli498/mimo-stable
Release list
mimo-stable 1.1.5
mimo-stable 1.1.5
- Added framework-neutral event normalization and policy APIs.
- Added public API specification, roadmap, packaging checks, and regression coverage.
- Hardened public API validation and redacted evidence handling.
- Expanded CI quality coverage to Python 3.13 and 3.14.
- Normalized Skill frontmatter to the OpenClaw official schema.
Validated locally with pytest (3 passed), fixture benchmark (9/9 passed), package build, isolated API/CLI smoke tests, version checks, and RECORD verification.
v1.1.4
Fixed
- Corrected the package metadata layout so isolated builds and PyPI publishing validate successfully.
- Includes the changelog cleanup, PyPI links, keywords, classifiers, and English package README metadata.
Verification
- Wheel build passed.
- Version consistency check passed.
- Detector behavior tests passed.
- Fixture benchmark: 9/9 cases passed.
v1.1.3
Fixed
- Removed stray diff prefixes from the Unreleased changelog entries.
- Added PyPI project links, keywords, and Python package classifiers.
- Switched PyPI package metadata to the English README.
Verification
- Version consistency check passed.
- Detector behavior tests passed.
- Fixture benchmark: 9/9 cases passed.
v1.1.2
Fixed
- Corrected the Chinese and English installation instructions to use the published PyPI package.
- Clarified that
mimo-stableprovides themimo-loop-detectCLI and does not expose amimo_stablePython import API.
Verification
- Version consistency check passed.
- Detector behavior tests passed.
- Fixture benchmark: 9/9 cases passed.
- Public PyPI installation and CLI smoke test verified after publication.
v1.1.1
Highlights
- Added the English project overview and architecture documentation.
- Added public integration examples for text detection, tool-call detection, and recovery-policy integration.
- Added GitHub Actions Trusted Publishing workflow for PyPI.
- Clarified Python 3.10+ support and release-quality verification boundaries.
Verification
- Detector behavior tests passed.
- Fixture benchmark: 9/9 cases passed.
- Public package installation and CLI smoke test verified after release.
v1.1.0 — 跨模型退化循环检测与工程止损
项目定位
这是一个面向国产及其他大语言模型(LLM)的退化循环现象记录、检测与工程侧止损工具。
项目最初来自 MiMo 在特定任务和参数条件下出现重复输出的案例,后续将经验抽象为可复用的检测与恢复决策层。类似现象也可能出现在 GLM 或其他模型中,但不同模型的触发概率、表现形式、根因和有效参数条件不能直接类推。
本项目分享的是:
- 如何识别输出或工具调用是否进入退化循环;
- 如何在工程侧限制资源浪费和重复副作用;
- 如何把案例记录成可复盘、可比较的证据;
- 如何把检测信号交给上层运行器做保守恢复决策。
本项目不宣称自动修复模型、不宣称彻底解决某个模型的问题,也不提供跨模型故障率结论。
本版本包含
退化循环检测器
scripts/detect_loop.py 支持:
- 连续相同或高度相似文本块检测;
- 相同工具及相同参数的连续调用检测;
- 重复副作用工具调用的提前暂停信号;
- 显式开启的中文任务语言漂移检测;
- 默认持续时间门控,减少短时正常重复造成的误报;
- 显式
--text-mode instant即时检测模式; - JSON 摘要输出,便于接入其他运行器;
- 参数、输入和 JSON 类型校验;
- stdin 流式处理,按输出块到达时实时评估,不再等待读完全部输入。
纯决策恢复层
scripts/recovery_policy.py 只负责根据检测摘要输出动作,不执行任何副作用:
continue:未检测到循环;pause_and_review:检测到可能重复的副作用工具调用;stop_and_retry_once:首次且符合条件的可重试信号;stop_and_escalate:不再继续自动重试,交给上层或人工处理。
停止生成、切换模型、重试和人工复核仍由调用方负责。检测器和恢复策略不会自行调用工具、切换模型或执行重试。
工程质量与回归
- 9 个脱敏 fixture benchmark 案例;
- 连续输出、短时重复、近似文本、参数变化重试、非连续工具调用、工具 key 顺序、副作用重复和语言漂移回归测试;
- 非法参数和非法 JSON 输入测试;
- 流式 stdin 行为测试;
- GitHub Actions 质量检查;
pyproject.toml安装入口;mimo-loop-detectCLI smoke test;- 真实项目检查脚本
scripts/test_short.sh和scripts/test_long.sh。
快速开始
python3 scripts/detect_loop.py --json --timeout 60 --log fixtures/loop_detected.log退出码:
0:未检测到循环;1:检测到循环;2:参数或输入错误。
将检测摘要交给恢复决策层:
python3 scripts/detect_loop.py \
--json \
--timeout 60 \
--log fixtures/loop_detected.log \
| python3 scripts/recovery_policy.py --retryable默认文本检测仍然使用持续时间门控;如果只是需要立即判断重复信号:
python3 scripts/detect_loop.py \
--json \
--text-mode instant \
--log fixtures/repeated_but_short.log证据边界
当前仓库中的 fixture 和历史日志只用于说明行为和防止回归,不能证明:
- 某个模型的普遍故障率;
- 某个参数一定能阻止退化循环;
- MiMo、GLM 和其他模型具有相同的故障机制;
- 检测器在所有真实生产日志上都不会误报或漏报。
分享新案例时,建议记录模型精确版本、供应商/端点、时间、任务类型、实际参数、上下文规模、工具调用情况、检测标准、检测延迟和重试结果,并在公开前脱敏。没有充分证据时,使用“观察到”“可能相关”“尚未复现”,不要把相关性写成根因结论。
当前局限
- 检测器依赖输出块边界;stdin 使用空行切分,其他日志格式需要调用方先适配;
- 工具调用识别基于常见文本格式,不是所有供应商原生协议的完整解析器;
- 语言漂移检测目前只对显式
--expect-language zh开启; - 相似度、重复次数和持续时间阈值需要按模型、任务和运行器单独校准;
- 重复副作用调用可能是合理的幂等重试,生产系统必须结合幂等键、调用结果和重试原因二次判断;
- 当前为依赖极少的离线工具,不包含真实模型 API,也不代替供应商或框架官方监控能力。
后续方向
- 收集更多脱敏的跨模型真实案例;
- 增加不同日志格式的适配器;
- 在有足够样本后,报告限定条件下的误报、漏报和检测延迟;
- 完善多模型、多端点和多轮工具调用场景的独立 benchmark。
变更范围
本版本发布在 feat/degenerate-loop-guardrails 分支,对应提交:b169ea8。
它不会修改或宣称修改模型本身,只提供检测信号和保守的工程处置建议。