Skip to content

Releases: kw2828/OpenJev

Sparse-query control: complete comparison, failed 14/16 rule

Choose a tag to compare

@kw2828 kw2828 released this 23 Sep 03:40

Completed the fixed sparse-query comparison: 24 calibration paths and 360 evaluation paths across five controllers and three sensing settings. All evaluation searches succeeded. Period-two querying cut fully paid controller cost by 53.15% / 40.79% versus always-neural control in the primary settings, but failed two move-quality requirements. The prespecified continuation rule is FAIL, 14/16; no new model or architecture advantage is claimed.

The archive preserves every trajectory, compressed journal, all 114 scientific sources, all engineering attempts and both saved audits. The first audit had a dictionary-versus-scalar timing check bug. A separately frozen corrected verifier agrees with the unchanged data and criteria. Both figure versions and all values remain available. Extract at the repository root; this archive does not include the runtime or eleven inherited model/kernel/prior-evidence dependencies listed in the manifest.

Results and chart · Protocol · Audit correction

Query-value pilot V2: complete, continuation rule failed

Choose a tag to compare

@kw2828 kw2828 released this 23 Sep 03:03

The paired query-value pilot completed all 72 TRAIN paths and all 264 available label panels in 152.92 seconds. Its independently checked scientific continuation rule failed: 7/11 conditions passed. No gate was trained from this recipe and no new architecture advantage is claimed.

The base setting had positive split-half covariance but differing actions spanned only five originating cases. The shifted setting had nonpositive covariance. All anchors, negative returns, censored branches, bootstrap draws, conditions and costs are preserved.

A separate fabricated-event benchmark measured a 5.103x median durable-logging speedup, with byte-identical output and per-event flush/fsync in both implementations. This is a logging result, not a model or end-to-end scientific speedup.

The archive includes complete V2 raw evidence, all logging blocks, original failed and successful engineering records, independent audit outputs, figures and the frozen source closure. The manifest supplies every member's SHA256 and size. Extract at the repository root. Prior model weights, runtime and inherited input evidence are referenced separately; this is not a standalone runtime bundle. Earlier V1 partial records remain unchanged.

See the repository's research/otto-query-advantage-v2-results.md for all eleven conditions, interpretation, and the proposed sparse-query control test. The proposal has not been executed.

Paired query-value pilot: incomplete at the time limit

Choose a tag to compare

@kw2828 kw2828 released this 23 Sep 02:22

The paired value-of-query pilot reached its original 900-second cap after all 72 TRAIN trajectories and 210 of 256 available label panels. One panel was interrupted and 45 were not started. No signal analysis, model training or evaluation was admitted, and no scientific retry or time extension occurred.

The archive preserves all partial records, all 125 source files bound by the plan, engineering attempts, failure review and reporting code. The independent metadata review found that logging occupied 65.36% of measured elapsed time inside the completed panel intervals; this does not establish a future speedup. A separate launch-path mismatch in the unused scientific auditor is documented.

Extract the archive at the repository root. The manifest and SHA256SUMS provide byte-level verification. Original runtime, model weights and inherited input evidence remain dependencies identified by frozen plans and prior releases.

See the attempt report.

Fixed training budget: 80 versus 320 epochs

Choose a tag to compare

@kw2828 kw2828 released this 22 Sep 12:05

The fixed 80-versus-320 epoch comparison completed: six paired fits, identical R64 labels, exact first-80-epoch prefixes, 6,000 updates and 504 fresh full-horizon searches.

Longer training lowered final TRAIN MSE by 11.38% and the chosen-action cost gap against the saved labels by 14.78%. Weighted success rose from 10.19% / 20.42% / 4.52% to 15.34% / 21.57% / 9.52%. Analytic control found all 72 sources. The frozen continuation rule failed: 10/33 passed, including 0/18 competence conditions. No architecture advantage or promoted checkpoint.

The independent saved-record audit agrees on all outcomes and criteria. It checks recorded scores, training prefixes, diagnostic arithmetic, work and costs; it does not regenerate neural scores, gradients or intermediate posterior filtering.

The archive includes all current-phase checkpoints, raw work/evaluation records, fixed TRAIN diagnostics, frozen plans, qualifications, audit, figures and execution sources. SHA256SUMS and manifest.json provide checksums for the archive and its 72 members. Historical inputs and earlier releases remain dependencies specified by the frozen plan; this is not a standalone reconstruction of the entire lineage.

Full results · Protocol · Next proposed diagnostic

Matched teacher-cost learning: complete negative result

Choose a tag to compare

@kw2828 kw2828 released this 22 Sep 10:11

The matched teacher-cost study completed all 558 training states, 32,304 teacher continuations, six fits and 504 fresh evaluation episodes. The continuation-target recipe failed its predeclared rule: 5/33 conditions passed, including 0/18 competence conditions.

Weighted success was 7.75% / 4.11% / 5.83% for continuation targets, versus 20.65% / 8.42% / 9.08% for analytic targets. Analytic control found every source in all 72 cases. No competent learned controller or architecture advantage is established.

The separately frozen V2 saved-record audit agrees on all outcomes and conditions. The archive preserves the original audit failure, its diagnosed summation-order issue, the corrected numerical verifier, and the unchanged empirical data and criteria. The audit checks recorded accounting, random streams, targets and outcomes; it does not regenerate neural scores, gradients or posterior filtering.

Included: all continuation journals, targets, initial and final checkpoints, training records, evaluation traces, plans, original process receipts, both audit attempts, and PNG/SVG comparison figures. The manifest verifies all 614 archive members. Historical dependencies remain identified by the empirical plan and earlier releases.

Archive SHA256: 5a32668d76988eb1917a5530033e7c3d752e151f73e7d1d0c3bff8d90cf83f43.

Full report · Protocol

Teacher-cost cohort: 558 states supported

Choose a tag to compare

@kw2828 kw2828 released this 22 Sep 09:05

The fixed teacher-cost cohort supports all 558 selected states from all 144 original learner-TRAIN trajectories. Initial and late retained states are included; none were repaired, replaced or omitted.

One supervised public replay reconstructed 119,212 updates in 16.85 seconds. Independent saved-record audit agrees with selection, all captured posterior hashes and support predicates, and all 239,828 events. Twenty extractor tests, four execution-failure tests and twelve common-regression tests passed.

No new continuation labels, policy fitting or autonomous evaluation occurred. This is input and objective preparation for a matched learning comparison, not evidence of architecture or gameplay improvement.

The archive contains all new-phase evidence, including the complete replay journal and exact public belief arrays. The manifest provides every file hash and the archive SHA-256. Extract it at the repository root. The auditor additionally requires the pinned historical inputs listed in the plan.

Results and limitations · Prospective protocol · Common regression objective

Teacher target precision: audited 16 versus 64 sample comparison

Choose a tag to compare

@kw2828 kw2828 released this 22 Sep 11:22

Six paired fits and 504 fresh full-horizon searches test 16 versus 64 teacher samples while keeping the model, training states, target scale and optimizer work fixed.

The intervention fails: 7/33 conditions passed, including 0/18 competence conditions. Weighted success changes from 21.43% / 12.79% / 15.11% to 19.78% / 17.89% / 13.06%. The analytic controller finds all 72 sources. No checkpoint is promoted and no architecture advantage is claimed.

Additional supervision costs 96,912 continuations and 1,028.31 seconds. The independent saved-record audit agrees on all outcomes and criteria. It verifies saved evidence without regenerating model scores, gradients or intermediate posterior filtering.

The README includes the all-fit comparison chart. The research report retains all outcomes, 33 checks, complete costs, noise diagnosis and the next proposed test.

The archive contains all newly generated precision and noise evidence, initial/final checkpoints, saved event streams, two figure renders and new source files. Historical inputs remain dependencies identified by the frozen plan and previous releases. Extract paths relative to the OpenJev repository; verify members against manifest.json. This is not a standalone copy of the entire historical dependency chain.

Archive SHA-256: 9cac023b8b7bdeccb6f3531d8883b26dd0fec3a9da5c5ff8b3b0749ba75d099c
Archive size: 712,631,818 bytes; all 630 member hashes verified.

Spatial value learning: all 15 fits and audited scalar results

Choose a tag to compare

@kw2828 kw2828 released this 22 Sep 01:33

The fixed-data comparison completed all fifteen final models across five families and three paired seeds, with 52,800 optimizer updates. All models, predictions and raw records are retained.

The spatial model did not improve mean exposed-validation MSE: it was 2.246% above the matched neighbor-free model and 4.934% above the statistics baseline. Statistics had the lowest validation MSE in every paired seed. Dense128 fit training data best and had the highest validation MSE. No autonomous or novel-architecture advantage is established.

Independent saved-model replay checked 100,470 model-row pairs, matching saved NumPy predictions exactly. Original worker time was 1,513.80 seconds, with a separate 44.99-second audit. These are scalar-fitting and audit costs, not autonomous-controller latency.

The 61,800,541-byte archive contains 337 verified original files, including all fifteen initial/final checkpoints, prediction arrays, complete training journals, audit evidence, plots, 209 frozen scientific sources and applicable licenses. Read RESTORE.md and dependencies.json before restoring; historical data/runtime dependencies remain separate. The strict numerical auditor retains original absolute paths.

Archive SHA-256: 56ba6868ff0359ec6a27d8631085776602b63227c14fd63a0fddff0f72b51662.

Full results · Protocol

OTTO spatial control: completed negative result

Choose a tag to compare

@kw2828 kw2828 released this 22 Sep 06:58

The complete frozen autonomous comparison is a negative result: 32/66 candidate conditions pass; the continuation rule fails. Spatial competence is 0/18, and the other learned heads pass 0/72 descriptive competence checks. Analytic control finds all 72 sources.

All fifteen fixed learned models and the analytic controller completed 1,152 episodes. Spatial weighted success is 40.35%, 27.81% and 16.12% across lambda3/4/5. No learned model meets competence in any setting, and no architecture advantage is established.

The independent saved-output audit replayed 1,736,024 model calls and 27,776,384 branch rows with exact agreement. Its final receipt records 76,316,348 comparisons. Its NumPy neural algebra is shared with the separately qualified deployment code; it independently reconstructs public filtering, branches, choices, weighting, costs and conditions.

Full results, all seeds and cost limitations · Frozen protocol

The two ordered archive parts preserve 329 files and 3,867,523,040 original bytes, including all fifteen final checkpoints, complete trajectories and work journals, the original worker/audit/report, frozen sources, numerical qualification and failed engineering attempts. Read RESTORE.md before extracting; historical TRAIN/VALID caches remain explicitly listed release dependencies. The combined gzip SHA-256 is ee3589a6658894c9418d61fe56a26dfbcdccec92d87e86db765c39638b529e52.

The separately attached PNG/SVG only improve the public chart layout. Their source and receipt are committed in the repository; the original report remains unchanged inside the archive. Neither packaging nor plotting performed model, training or simulator calls.

Learned query gates: no computation savings

Choose a tag to compare

Six small learned gates completed training and 576 fresh autonomous odor-source searches. The frozen continuation rule failed: 30/38 required conditions passed, with 48/62 reported checks passed.

All 18,012 learned gate decisions called the fixed neural planner. The gates inherited its search quality and added overhead, with no query savings or recurrence advantage. All searches found the source, but the changed sensing setting required 2.34 times as many moves under the neural planner and gates as under analytic control.

The independent saved-record audit agrees on 4,299,097 checks. Thirty focused engineering tests passed. The audit validates recorded outcomes, costs, selections and state continuity; it does not independently regenerate the original neural predictions, filtering or gradients.

The archive preserves all current-phase data, initial/final gate checkpoints, parity witnesses, raw journals, frozen plans, process records, audits, figures and execution sources. manifest.json lists member hashes and sizes; SHA256SUMS authenticates the archive and manifest. Inherited OTTO weights, earlier evidence and qualified runtimes remain dependencies identified by the frozen plans. This is not a standalone runtime bundle.

Results and limitations · Evaluation protocol · Prospective next experiment