Skip to content

Examples and evidence

Ginks edited this page Aug 22, 2026 · 2 revisions

Examples & evidence

GenOS is best understood as a collection of testable invariants and research workflows. The repository separates product proofs, experimental examples, benchmarks, and target architecture so that a compelling idea does not become an unsupported claim.

Evidence ladder

Level Required evidence
Implemented Source plus a focused test
Reproduced Exact command passes on a named revision and environment
Measured Harness, warmups, repetitions, raw samples, and machine metadata
Compared Equivalent inputs run through versioned adapters for multiple systems
Externally validated Independent public reproduction

An architecture document is not implementation evidence. A passing correctness test is not a latency benchmark. A deterministic fixture does not prove model quality.

Product proofs to run first

Claim under test Command What it proves
Snapshot, fork, identity/event isolation, diff ./run-demo.sh Core local lifecycle over a deterministic fixture
Safe parallel debugging and promotion ./examples/safe-debugging-demo/run-demo.sh Isolated candidate worlds, shared gate, replay equality, explicit promotion
Event reducer replay cargo test -p genos-store --test replay_tests Stable reconstruction for supported recorded events
Directory-world isolation cargo test -p genos-world --test file_isolation Relative file writes remain isolated between worlds
Isolation limitations stay explicit cargo test -p genos-world --test isolation_boundaries Boundary behavior remains documented and tested

These proofs do not establish deterministic provider inference, network replay, OS-level sandboxing, or enterprise readiness.

Counterfactual execution and selection

Example Question it explores
Counterfactual demo Can sibling agents share a baseline while keeping identity and event streams distinct?
Divergent writes Do branches retain independent memory after a shared snapshot?
Divergent worlds Do parent and sibling directory worlds keep file mutations isolated?
Branch hypotheses Can the reason for each branch remain attached to the outcome?
Counterfactual evaluation Can branch outcomes be scored and compared structurally?
Multi-objective evaluation Can several objective dimensions be preserved?
Pareto selection Which candidates are dominated or non-dominated?
Winner takes branch Can one verified future be promoted explicitly?
Cognitive merge Can selected evidence be reconciled without blindly unioning memory?
Branch evolution Can budget move toward survivors and spawn further exploration?

Genome, phenotype, and heredity

Experiments cover:

  • genome mutation and reversible mutation;
  • genome/phenotype separation;
  • nature-versus-experience cohorts;
  • two-parent breeding and child validation;
  • trait inference, claims, replication, heritability, and promotion;
  • integrated counterfactual cycles over genotype and experience.

The goal is controlled causal study of agent configuration and behavior. These examples do not claim biological equivalence.

Beliefs, memory, and lineage

Focused demos cover belief updates, contradictions, evidence provenance, recursive fork lineage, nearest common ancestors, event correlation, causal chains, experience packets, knowledge synthesis, and failure search.

This surface asks whether an agent system can retain where knowledge came from instead of storing only the latest prose summary.

Reproducibility and storage

Examples include snapshot restoration, snapshot timelines, artifact/snapshot deduplication, model reproducibility classification, dated checkpoint interventions, personal causal replay, and retroactive exploration.

The model-reproducibility work is especially important: it distinguishes exact replay from semantic equivalence, divergence, and unsupported comparisons.

Tools, permissions, and failure

Tool examples exercise successful execution, structured failure, permission boundaries, controlled capabilities, taint, and circuit-breaking behavior. A tool permission model narrows what GenOS exposes; the host still needs a real security boundary for untrusted execution.

End-to-end research scenarios

The repository contains larger scenarios for:

  • unknown-cause bug investigation;
  • adaptive incident search;
  • extreme critical-system refactoring;
  • temporal causal simulation;
  • scientific compression research;
  • security co-evolution;
  • swarm and resilience behavior.

These scenarios may exercise unstable APIs and should be treated as research artifacts.

Published measurements

The safe-debugging benchmark publishes repeated raw samples for its deterministic fixture, including wall time, candidate order, replay verification, merge-gate decisions, model calls, tokens, and model cost.

Because the fixture calls no model, model usage is measured as zero. The benchmark does not compare GenOS against an external framework or establish quality on real provider output.

The dedicated Benchmarks & results page covers the complete B01–B10 portfolio, current resilience and MCP safety results, reproduction commands, and why public scores remain unclaimed.

Missing evidence

Current gaps include:

  • versioned adapters for external agent runtimes;
  • equivalent same-machine cross-framework comparisons;
  • real provider token and cost accounting across branch strategies;
  • model, network, clock, and external-tool replay boundaries;
  • OS-enforced process and network isolation measurements;
  • independent reproduction outside the core project;
  • reviewed benchmark bundles tied to releases.

When a result is absent, report it as missing or unsupported. Do not copy performance numbers from design documents.

Browse the complete examples catalog and proof status.

Clone this wiki locally