[Ecosystem] Should deterministic execution boundaries be a first-class Harness capability? #5286
Replies: 1 comment 13 replies
一个真实用户的规则引擎实战:知识与合规之间的那道裂缝你好,读完你在 #5286 的长文,颇有共鸣。我们(一个零编程基础的法律工作者用户 + 一个执行者 AI)在用户侧做了一个规则执行引擎插件(dsh-rule-engine),从 2026 年 8 月中旬至今、用约三周的高强度真实场景验证了你最核心的论断:「Knowledge is not compliance. Model-visible knowledge is not execution authority.」下面按三个层面讲——不只是供案例,更是分享思路并寻求社区建议。 我们在做什么
方案、困难、摇摆(1)路线反复摇摆最初十几个版本多轮重构:意图判定从词表硬编码 → LLM 兜底 → 词表+LLM 双通道;授权语义从「宽泛 any」→「子句动作级授权」→「分点三柱(条件句零授权/显式命名对象锚定/clauseId 隔离)」;执行等级从「规则文本解析」→「等级从动作字段推导」。每次摇摆背后都是真实事故:用户的「请将建议书内容转化到业务通告中」含「建议」被误判成方案、连拦 7 次,最后补上「动作词优先于方案词」。问题还不止于此——词表演进史是一部「盲区→补词→新盲区」的循环史,最终得出「人工枚举永远补不完」的结论,正解是词表只做候选、语言理解主位交给模型。 (2)关键点的僵化与放松我们花大量精力机器化硬拦/时序等级,这部分实测可靠:引擎的「疑问句停机」「授权范围比对」「备份前置」「重启前验证链」等数十次硬拦全部命中执行者真正该被拒绝的动作(每条都可审计查证)。但 自证等级(语义性要求:先查手册再动手、时间词先核对、发现错误先停下、问句回合只答不推进)几乎完全失控——这正好印证你的论断:把规范写进提示词/规则文本,对执行者是「知识」不是「约束」。我们不断「补词表」「补规则」「写手册」,最后承认:答案不是更多文字,而是「判定权移交用户物理动作域(物理确认弹窗)+ 执行边界机器强制」。 (3)通用插件与个人化的耦合痛点执行者在早期把「本机偏好」(作者自己的规则编号→执行器映射)硬编码进插件代码,导致插件对任何其他用户都无法工作;同时测试夹具本机化(针对本机规则写死断言)。更糟的是,执行者以「通用化」为名义继续往上改本机内容——实际上是在旧架构上打补丁,并没有真正做通用化。直到用户严厉指正后,才真正完成:声明式绑定、默认偏好表下沉到配置文件(代码零本机编号)、禁用语义占位、无 AGENTS.md 零加载、发布适用性门禁。但还远远不够。教训:通用插件与个人化偏好必须严格分离——这与你在第 8 节「Harness should preserve reality, not take ownership of user intent」的边界观一致:插件管「现实是否被违反」,不管「用户本来想干什么」。 我们卡住的问题(向社区请教)插件已经 0.5.14 版,但我们很清楚它没做完。已知的八个未完成方向,每条都撞在「模型遵守」这个核心难题上:
执行者犯过的错(有记载的一小部分,按类型+次数)以下全按痕迹可查证(引擎审计 + 手册踩坑记录 100+ 条)。这些都是已被记录在案的;实际同类行为远多于记载。按类型归纳,同类合并:
(这些只是手册记录里的;实际执行者的同类行为远多于记载。如果你需要具体轨迹,我们可以提供脱敏审计日志。) 尾声插件还在开发中,烂摊子远多于成果。写这些不是来共鸣的——是卡在被 #5286 点破的同一个核心矛盾上:「知道」不等于「遵守」,只有执行边界才是可靠的;但执行者连自己都没治好。对规则引擎有什么好的建议(我们的思路如上),请不吝赐教:https://github.com/jilian-dsh/dsh-rule-engine |
Uh oh!
There was an error while loading. Please reload this page.
[Ecosystem] Should deterministic execution boundaries be a first-class Harness capability?
I recently went through a number of DeepSeek Harness community reports together with the trajectories from our own experiments.
At first these looked like unrelated problems:
But after putting them next to each other, I think there is a smaller common question worth discussing:
This is not a proposal for a new Core Runtime.
We built a Plugin prototype to test the idea first:
https://github.com/goatliamia/dsh-runtime
1. Community census: different bugs, recurring execution shapes
A few existing discussions were particularly useful.
Repeated execution without progress
#3171 describes a tool-call failure loop where the model repeatedly tries to repair and execute again.
#3228 shows the more extreme cost side of an unbounded Agent loop.
#3489 is especially interesting: an MCP session has already expired, calls continue returning the same deterministic
-32001failure, and the Agent keeps replanning and calling.These have different root causes.
But at the execution layer they can share the same shape:
Fixing the root cause is still necessary.
For example, an expired MCP session should be repaired in MCP lifecycle handling.
But there is a second question:
I think the answer can be no.
2. We tested this as a no-progress Circuit
We reduced this failure shape into a controlled experiment.
Condition:
Results:
After the Circuit opened, the
Circuit + Deltaarm allowed zero further executions of the known no-progress path.Later, in a separate open-ended creative task, we found another natural instance that had not been designed into the experiment:
That changed our interpretation.
We no longer think this is mainly a “flaky tool retry breaker”.
The more useful primitive appears to be:
The root cause may belong to MCP, a tool implementation, a provider adapter, or something else.
The execution containment does not need to wait for all of those root causes to be fixed.
3. Another community shape: the Harness already knows an action is invalid
A different problem appeared in our own experiments.
The Host deterministically knew:
The Model also discovered the fact through normal probing.
Without an execution Guard:
The Model knew the plugin was required and still eventually attempted to unload it.
With a deterministic pre-execution Guard:
we observed:
The important distinction for us was:
Or more precisely:
If the Harness already knows an invariant, repeating that invariant in the Prompt is a weaker mechanism than enforcing it at the execution boundary.
4. Runtime state changes are a different problem again
#3509 describes failures from long-lived runtimes where composition/readiness can change after initialization.
#941 discusses another related dimension: workspace-scoped configuration, lifecycle, ownership, and runtime state.
These are not no-progress problems.
Here the important shape is:
We tested a minimal lifecycle:
When an action was temporarily invalid, a simple Guard caused repeated verification:
We then tested:
Post-rejection verification decreased:
and request payload decreased:
The same directional result appeared with another model.
This led us to distinguish:
The important part is not the exact API.
It is that a deterministic state transition has different semantics from a static fact.
5. We also found several things that did not work
This mattered just as much as the positive results.
Repeating positive facts
If the Model could already observe a positive runtime fact, injecting it again did not produce a stable improvement.
Provenance as persuasion
Adding:
did not make the Model reliably trust Runtime information more.
Those fields still seem useful for reconciliation, freshness and arbitration.
But:
Injection as enforcement
We explicitly injected that a state had changed from:
and the Model could still execute the invalid action.
Injection-only:
Guard:
So:
This is one reason I do not think the solution is “put more Runtime state into the context”.
6. Persistence also does not imply Exposure
We also tested cross-session state.
A previous Session had already established a deterministic runtime fact.
The next Session could either:
or use persisted state.
Persistence reduced rediscovery, but the more important design result was:
This gave us another primitive that looks surprisingly important:
So the small vocabulary that emerged from the experiments is currently closer to:
These are not proposed as a universal Runtime specification.
They are simply the deterministic responsibilities that survived our experiments so far.
7. Why this matters for cost
We initially measured request payload size.
Later we decoded 152 historical DSH session trajectories and recovered the actual:
That changed our understanding of Runtime cost.
An unnecessary Agent turn is not only another answer.
It can also mean:
Across our historical experiments we observed, depending on the scenario:
These are small, scenario-specific experiments, not general performance claims.
But the mechanism changed how we think about cost:
So the interesting optimization target may not be:
but:
8. We also tested whether this destroys open-ended behavior
A reasonable concern is that stronger execution boundaries could simply make the Agent more conservative.
So we ran an open-ended creative task under three Harness compositions:
The interesting failure was the Off arm:
It completed the artifact while silently violating a Host invariant.
In this tested scene, deterministic enforcement did not reduce the measured successful creative actions.
That does not establish a universal “creativity preservation” result.
But it does show that:
9. Prompt correctness is another reason to separate intent from reality
We also tested correct and incorrect user instructions against different Harness compositions.
One case deliberately asked the Agent to unload a Host-required plugin.
The important result was:
The Harness did not invent a new goal for the user.
It simply refused to make an invalid state transition real.
Another case gave the Agent an incorrect factual assumption:
The action was rejected until the actual state became ready, after which execution succeeded.
This left us with a responsibility boundary that I currently find useful:
The Human can be wrong.
The Model can be wrong.
The Harness does not need to decide what the Human “really meant”.
But when the environment exposes a deterministic invariant, an incorrect belief does not have to become an incorrect world state.
10. Not every problem belongs to this capability
#5106 is a useful counterexample.
A Plugin can fail during discovery / activation and prevent healthy Plugins from loading.
That is a serious ecosystem problem, but a Runtime Plugin cannot repair a Host that fails before the Runtime Plugin itself becomes active.
That responsibility belongs closer to:
This distinction matters.
The point is not:
The process we are trying to follow is:
Sometimes the answer is Runtime.
Sometimes it is MCP lifecycle.
Sometimes it is a Tool.
Sometimes it is the Host.
11. So what capability are we actually asking about?
After these experiments, I think there is a narrower ecosystem question than “Should DSH have a Runtime?”
Something like:
At minimum, the useful semantics we have observed are:
I am deliberately not proposing that all of these need to become Core APIs.
The question is first about responsibility and seam, not API shape.
12. Working prototype and evidence
We built the current experiments and Plugin prototype here:
https://github.com/goatliamia/dsh-runtime
The project is now described as DSH Runtime Capabilities rather than a Universal Runtime.
The repository contains:
The prototype currently includes ideas around:
The repository is meant to make the claim falsifiable.
If a community failure has the same execution shape, we should be able to run it against the capability and see whether it actually helps.
If it fails, that should change the capability or its boundary.
13. What I would especially like feedback on
I am less interested in whether the name “Runtime” is correct than in these questions:
Is deterministic execution containment a distinct Harness responsibility, or should each Tool / MCP / Plugin implement it independently?
Should no-progress be modeled generically at the execution layer, or is that abstraction already too broad?
Where should deterministic state live when it outlives a single Model turn or Session?
Should Runtime state normally stay silent and only expose meaningful transitions?
What real DSH failures would be good adversarial cases for these capabilities?
That last question is especially important to us.
We do not want to keep constructing our own examples indefinitely.
If there are real trajectories where a Guard, Circuit, Delta or persistent runtime state should work — or clearly fails — those are much more valuable than another synthetic benchmark.
Runtime is still a Plugin
One unexpected conclusion from all of this is that I am now less convinced that “Runtime” should become a privileged new Core subsystem.
DSH already has a strong Plugin-oriented architecture.
That gives us another development path:
In that sense, Plugin architecture is not only an extension mechanism.
It can also be an extraction mechanism:
So the proposal here is not:
It is:
All reactions