You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Headroom addresses provider-side context compression with reversibility, token accounting, and benchmark discipline. We are testing the upstream question: before compressing a request, which project instructions, decisions, skills, tools, permissions, history, and file evidence should enter it at all?
Our current result is promising but limited to manually frozen Context Packages: 10 task types, 30 paired A/B trials, 66 real model calls; all 30 pairs used fewer total tokens, with a mean reduction of 50.8%, while irrelevant/forbidden-information leakage fell from 23.3% to 6.7%. We do not claim that this proves an automatic compiler improves task success.
A useful next experiment may be a 2x2 factorial design:
Arm
Context selection
Headroom compression
A
No
No
B
No
Yes
C
Yes
No
D
Yes
Yes
All four arms would use the same OpenCode commit, model, parameters, frozen Chinese coding tasks, equivalent clean workspaces, and independent evaluators. We would retain every actual provider request, tool schema, token count, latency, result, and verification artifact.
The core question is whether selection and compression have separable value: selection should improve context precision and reduce stale/irrelevant knowledge; compression should reduce remaining token cost without reducing task quality.
I would value a critique of this design, especially any interaction effect or leakage path that would make the four arms incomparable. If it survives review, a shared benchmark or adapter could be a concrete first collaboration.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Headroom addresses provider-side context compression with reversibility, token accounting, and benchmark discipline. We are testing the upstream question: before compressing a request, which project instructions, decisions, skills, tools, permissions, history, and file evidence should enter it at all?
Our current result is promising but limited to manually frozen Context Packages: 10 task types, 30 paired A/B trials, 66 real model calls; all 30 pairs used fewer total tokens, with a mean reduction of 50.8%, while irrelevant/forbidden-information leakage fell from 23.3% to 6.7%. We do not claim that this proves an automatic compiler improves task success.
A useful next experiment may be a 2x2 factorial design:
All four arms would use the same OpenCode commit, model, parameters, frozen Chinese coding tasks, equivalent clean workspaces, and independent evaluators. We would retain every actual provider request, tool schema, token count, latency, result, and verification artifact.
The core question is whether selection and compression have separable value: selection should improve context precision and reduce stale/irrelevant knowledge; compression should reduce remaining token cost without reducing task quality.
Project brief and evidence limits: https://feiai2026.github.io/aios-context-compiler/
I would value a critique of this design, especially any interaction effect or leakage path that would make the four arms incomparable. If it survives review, a shared benchmark or adapter could be a concrete first collaboration.
All reactions