The analysis explicitly says that the cost and operational-burden claim is plausible but under-substantiated. The repo and docs describe the SDK as avoiding GPU fine-tune jobs and shifting cost to CPU updates and judge calls, but there is no concrete model showing where the crossover happens.
This ticket should replace the qualitative “cheaper” story with a measurable cost model.
Suggested scope
Add a documented TCO model for episodes evaluated, judge calls, storage writes, and rollout cadence.
Include a break-even example showing when native policy learning is cheaper than a fine-tuning cycle.
Add an operational dashboard or script to estimate cost at different episode volumes.
Clarify that the benefit is conditional and depends on the workload, not universal.
Acceptance criteria
A cost model is published in the repo or docs.
The model includes the main cost drivers and assumptions.
A break-even example is included with explicit inputs and outputs.
The analysis explicitly says that the cost and operational-burden claim is plausible but under-substantiated. The repo and docs describe the SDK as avoiding GPU fine-tune jobs and shifting cost to CPU updates and judge calls, but there is no concrete model showing where the crossover happens.
This ticket should replace the qualitative “cheaper” story with a measurable cost model.
Suggested scope
Add a documented TCO model for episodes evaluated, judge calls, storage writes, and rollout cadence.
Include a break-even example showing when native policy learning is cheaper than a fine-tuning cycle.
Add an operational dashboard or script to estimate cost at different episode volumes.
Clarify that the benefit is conditional and depends on the workload, not universal.
Acceptance criteria
A cost model is published in the repo or docs.
The model includes the main cost drivers and assumptions.
A break-even example is included with explicit inputs and outputs.