Skip to content

Releases: andreyivan4enkov/moe-orbit-prefetch

v0.5.7 — open checks, SGD, mixed evidence

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 14:25

What this release is for

Show reviewers what is still unchecked and what the work actually gives, including losses.

New docs

  • docs/OPEN_CHECKS.md
  • docs/WHAT_IT_ACTUALLY_GIVES.md (grounded adjacent uses only)

New live evidence

  • Online SGD baseline + cache-order stand: results/misswait_baselines_v2.md
  • Primary: FAIL vs SGD / frequency / none
  • cache_sens (orbit last): MIXED — beats frequency, still loses to none

No victory spin. Lean HumanEval miss-wait edge vs none remains a separate stand.

v0.5.6 — miss-wait vs classical predictors (honest FAIL)

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 11:59

Summary

  • Live short miss-wait head-to-head: orbit → frequency / LRU / prev_copy / none
  • Verdict: FAIL — orbit miss-wait worse than compared classics on this MacBook stand (wins=0 losses=2 each)
  • Fixed self-deadlock in relieve + non-reentrant store._lock
  • Coalesce takeover if trim races / timed-out loader

Artifacts

  • results/misswait_baselines_v1.md
  • examples/bench_misswait_baselines_v1/

This closes part of the “weak baselines” gap without spinning a loss into a win.

v0.5.5 — broader MacBook suite

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 09:42

What changed

  • added examples/05_gigachat_store_smoke.py
  • expanded docs/LOCAL_VALIDATION_20260729.md
  • verified editable install/import path on the MacBook
  • ignored generated examples/logs/ and examples/reports/

Meaning

This increases the amount of the repository that can be checked locally on the author's MacBook without synthetic scripts or multi-hour reruns.

v0.5.4 — research intent boundary

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 09:11

What changed

  • added docs/RESEARCH_INTENT.md
  • clarified the difference between:
    • objective observation: alternative architecture shows non-noise signal / partial comparability
    • interpretation: different forms of computation/intelligence may exist
  • explicitly labels the second point as hypothesis, not proof from this repo alone

v0.5.3 — exact hardware honesty

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 09:03

What changed

  • exact author lab documented publicly: MacBook Pro 2019 / Intel i9 / 16 GB RAM / Radeon 4 GB
  • added explicit note about strong thermal throttling
  • live results are now framed even more clearly as constrained-machine evidence

Why

This closes the remaining honesty gap around hardware ceilings raised during external review.

v0.5.2 — local live validation + authorship transparency

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 08:54

What changed

  • local no-synthetic validation report added: docs/LOCAL_VALIDATION_20260729.md
  • SparseDeepseekRuntime shared modeled-prefetch state now guarded by _state_lock
  • fixed examples/03_smoke_dynamic_weights_v13.py bootstrap so it runs from the repo clone
  • added docs/AUTHORSHIP.md to explain AI-assisted implementation and maintainer role honestly

What was revalidated locally

  • pytest tests/ -q → PASS
  • examples/02_smoke_expert_slice.py → PASS
  • examples/03_smoke_dynamic_weights_v13.py → PASS after bootstrap fix
  • examples/04_chat_ask.py "Hello" --max-new 16 → PASS
  • GigaChat 10B store smoke → PASS

Limits

  • no synthetic scripts were used for the release verdict
  • GigaChat-20B package-path was not revalidated in the same direct local method in this session
  • no multi-hour lean bench rerun in this release session

v0.5.1 — DeepSeek-family + GigaChat evidence

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 08:23

Scope clarification

Object A targets open DeepSeek-style MoE (DeepSeek-V2-Lite and GigaChat lab).
Not “any neural net”; not Mixtral without a new adapter.

See docs/SUPPORTED_MODELS.md.

New artifacts

  • results/gigachat_v21_orbit_apply.md
  • results/gigachat_v32_lean.md
  • results/gigachat_v34_humaneval.md

v0.5.0 — research prototype (lab-scope honest)

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 08:20

Summary

Industry-style packaging + critical store/prefetch hardening for Object A (MoE orbit prefetch).

Honest lab scope

  • Author compute: MacBook-class laptop (CPU).
  • Published Tier L: lean HumanEval/QuixBugs + residency smokes — not a 100-task GPU suite.
  • See docs/LAB_SCOPE.md, MODEL_CARD.md, docs/EVIDENCE_TIERS.md.

Highlights

  • MODEL_CARD.md, SECURITY.md, CODE_OF_CONDUCT.md, GitHub Actions CI
  • Expert load coalesce, drop_expert, cold-first evict, CUDA-direct guard
  • Prefetch error counters + cancel drains queue
  • Auditor triage: docs/AUDITOR_ISSUES.md

Install

pip install -e ".[runtime]"

v0.4.0 — analysis pack (trajectories + plots)

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 07:53

For reviewers who said the repo looked empty

Start here (no model weights):

  1. analysis/REPORT.md
  2. analysis/figures/hit_curves_vs_baselines.png
  3. analysis/data/orbit_trajectory_200.json (~215KB step records)
  4. docs/ARCHITECTURE.md
  5. docs/MATH.md
  6. `src/moe_orbit_prefetch/sparse_moe_runtime.py` (~900 lines full sparse generate)

```bash
pip install -e ".[analysis]"
python analysis/generate_orbit_trajectory.py
python analysis/plot_orbit_analysis.py
```

Synthetic learnable stream (not live DeepSeek gate): orbit mean hit ~0.30 vs prev ~0.17 / freq ~0.14 / cyclic ~0.08; emergent PASS vs those baselines.

v0.3.0 — full math + Apache-2.0 attribution

Choose a tag to compare

@andreyivan4enkov andreyivan4enkov released this 29 Jul 06:58

Why this release

Answers the request for complete sources + math and an open license with credit when the method is used inside larger systems.

License

  • Apache-2.0 (fully open, no fee)
  • ATTRIBUTION.md: small experiments = Apache only; substantial products = keep NOTICE + credit this repo/method

Math & source map

Still not in git

Model weights (DeepSeek/HF). All first-party code is in the tree for fork/edit.