Replies: 4 comments
|
Strong alignment with the verification-first philosophy ? endorsing and cross-linking. Key overlap:
One complement worth adding: after a long task succeeds, extract the lessons into verified rules so the NEXT long task starts smarter ? that's the |
|
这种 superpower 早有了一套规范长任务处理了 |
|
Thanks @zoahdev — the alignment is real, and your Quick mapping of our mechanisms to your three overlaps:
On the complement you suggest (after success, extract lessons into verified rules so the NEXT task starts smarter): we currently have a failure feedback loop (failed traces flow back into templates so the engine assembles each task type better), but it's implicit. Your And yes — happy to be listed in the dsh-ecosystem map. Repo: https://github.com/pridesong/orchestration-skill. If there's a PR/format for the catalog entry, point me at it and I'll submit. |
|
@hytime 完全同意 superpower 在长任务规范上是先行者,这一点没有争议。我们的定位不是替代它,而是回答一个它没覆盖的问题:长任务能不能不依赖贵模型? 差异点其实很具体(可验证,不是口号):
如果 superpower 已经有覆盖这些的方案(尤其廉价模型 + 零散文协议的组合),请指个链接,我认真对比学习;如果没有,那正是这个项目想补齐的空位。 |
Uh oh!
There was an error while loading. Please reload this page.
Show and tell: orchestration-skill — 长任务不该贵到必须用前沿模型 / Long tasks shouldn't cost like they need a frontier model
orchestration-skill 是一个给 DeepSeek Harness 的 Agent Skill:让廉价模型跑出前沿模型的长任务可靠性——把可靠性从模型身上搬进结构里。
核心逻辑
长任务失败不是因为模型在某一步弱,而是不可靠在长链上复利。与其花钱买更强的模型把整条链装进它的上下文,不如把长链拆成廉价模型轻松胜任的小步,让链的可靠性由机械结构承担:
artifacts/<field>.json;断点续跑、零上下文损耗、只重跑失败的那一步validate/check/compare机械拦截幻觉与格式错误,不信任模型自检T3FILE:v1 读取 <file> 并按内容执行);散文无法包裹协议,模型没有漂移空间实测证据(廉价模型全链路)
一个完整医药供应链任务(PVG→EZE 温控空运,1000kg,$11,500 硬预算)在廉价模型上全链路跑通:11 个字段(报价/海关/天气/航线/成本NPV/仪表板/审计/结论)。
降本回退对照实验展示了结构带来的差距:
且当回修经由运行时 mind-decider(读审计证据——固定成本墙 vs 单价墙——再选思维)路由时,精确复现 $12,676:选择是机械可复算的判别决策,不是散文运气。
关键机制
next/parallel/routing展开为排它条件路由表;并行组写同一行 state.csv(多条件 AND)文档
SKILL.md— 协议(给 LLM,schema 驱动契约,派发 prompt pattern 锁死)Description.md— 人类阅读版(理念/架构/工作流)README.md— 卖点定位 + 实测证据English summary: orchestration-skill is an Agent Skill for DeepSeek Harness that makes cheap models deliver frontier-model reliability on long tasks by moving reliability out of the model and into mechanical structure (disk-resident state, mechanical gates at every boundary, multi-brain audit, regex-locked zero-prose T3 dispatch). Evidence: a full pharma supply-chain task (11 fields) ran end-to-end on a cheap model; a cost-reduction rework experiment cut the budget gap by 72% ($12,676 vs $15,730), and the runtime mind-decider reproduced the result exactly. Pure Python stdlib, zero dependencies. Feedback welcome!
All reactions