Skip to content

Validation and Quality

github-actions[bot] edited this page Oct 8, 2026 · 5 revisions

GraphPaper 0.6.0 · Documentation source · Suggest an edit

Validation and quality

Version 0.6.0 validation

The published v0.6.0 build corresponds to source 00cab75a2cfc13f541573be3bce14f0cc5b97684. Its GitHub-hosted release validation completed successfully before publication.

Check Recorded result
Windows Python regression suite 299 passed
Graph geometry tests 6 passed
Real-HTTP browser workflows All 7 suites passed
Compiled Windows native checks 28 passed, including graph editing/undo, voice reuse, dialogs and save-on-close

Provider and JEV responses were controlled synthetic fixtures; the results test routing, UI and storage rather than live editorial quality. No private project, model account or user computer was used. The release records the archive checksum. The visual tour uses actual captures from this exact release run, with a per-image provenance manifest.

Historical validation

The following versioned reports describe earlier releases, not the current screenshot set or current documentation version.

Version 0.4.0 validation

252 automated Windows tests passed; 1 permission-related test skipped. All 54 real-HTTP browser checks passed, with no JavaScript page errors. The actual compiled Windows application passed 17 native checks, including mode conversion, thesis/force controls, external Graphify, native Save As and pending-edit save before normal exit.

The packaged Python code for the changed modules was compared with the final source as semantic code objects; all UI assets were compared byte-for-byte. The Windows package includes the previously verified public Graphify and official Codex runtimes. No installed user project was changed.

A live Codex test generated argumentative angles from original synthetic evidence and voice samples, then drafted and reviewed a complete short argument. A separate opposing-position test confirmed the editorial policy does not choose which side the author should support. The first outline response was malformed JSON and was rejected; the successful continuation was a manually restarted test using the saved angles. No live JEV call or new user sign-in ceremony was performed. Deterministic tests inspect JEV voice/stance inputs separately.

The tests validate behavior and handoff, not a universal promise of perfect style or zero caveats. Source, browser, native, package and live-test evidence is under validation/v0.4.0. Other-platform results, when available, are the actual hosted CI results rather than inferred Windows coverage. Earlier reports below are historical.

Version 0.3.1 validation

217 automated Windows tests passed (1 platform-permission skip). All 44 browser workflow checks passed, all 10 native-source checks passed, and all 15 checks against the actual compiled executable passed. The compiled-native checks include a real public Graphify worker with synthetic model transport, source provenance, native Save As and save-on-close.

Real Codex-backed extraction was verified separately with synthetic documents and the packaged Graphify worker, using the explicitly selected gpt-6-luna model at low reasoning. It used 1 model request(s), with no duplicate Native pass or API-key fallback. Exact outcomes, dependency versions and platform results are recorded in docs/validation/v0.3.1. This is not a claim of paid API testing for every provider or a macOS desktop build.

The legacy native-test waits were too short on some cold Windows runs. The harness now waits for actual extraction readiness and records export diagnostics while retaining bounded cancellation and shutdown checks. Earlier reports below are historical.

Passing software tests and producing a better article are different questions. This page keeps recorded release checks, live service checks and editorial evaluation separate.

Recorded validation: GraphPaper 0.3.0

The release summary records native Windows validation for the published 0.3.0 build. These are recorded local-build results, not a claim that every current GitHub Actions run passed.

Check Recorded result Evidence
Automated Windows tests 172 passed, 1 symlink-permission skip, 0 failures JUnit report
Base browser workflow 17 checks passed Report
Voice/folder/prose browser workflow 9 checks passed Report
Science/reasoning browser workflow 13 checks passed Report
Native Windows source interface 10 checks passed Report
Actual compiled Windows interface 10 checks passed Report
Compiled backend/assets Passed Report
APA export rendering Four synthetic manuscript pages visually inspected Scope

The 39 browser checks reported no JavaScript page errors. The compiled native test exercises a Windows mouse interaction, bridge initialization, Science protocol editing, Connections, native Save As and persistence of a pending edit before normal process exit. It is not merely a request to a health endpoint.

Live versus controlled tests

Browser/model fixtures are synthetic and labelled. They verify workflow behavior, not prose quality or scientific findings. The release checks did not complete a real user OAuth ceremony or make paid writing/JEV requests.

The separate live metadata check succeeded for PubMed, arXiv, Crossref and Europe PMC. Anonymous Semantic Scholar returned HTTP 429; the failure was recorded rather than turned into zero results. This is a dated connectivity check, not a guarantee of current availability.

The bundled official Codex runtime received a signed-out protocol check. External Graphify execution was not exercised by a live run in the recorded release validation. Mock contract tests cannot establish continued availability of an external endpoint.

APA rendering used the same exporter source with synthetic material in LibreOffice on Linux, not Microsoft Word automation. Visual checks found and corrected blank-page and inherited-font issues. The final manuscript still needs inspection under the target journal’s requirements.

Current CI status

The README’s quality badge is dynamic and links to the actual quality workflow. It should not be replaced with a static “passing” label. A queued, skipped or failed hosted run is distinct from the recorded release results above.

The Windows release workflow requires the compiled native regression before artifact publication. Later CI rebuilds are not allowed to silently replace an existing verified release package. Checksums identify exact assets, not a general claim that all builds are identical.

Evaluate the writing advantage

Compare GraphPaper with a strong direct-writing baseline using the same sources, models, target length and a comparable inference budget. Blind the evaluator to the generation route where practical.

For nonfiction, assess source support, reasoning, useful novelty, counterarguments, structure and author editing time. Count unsupported claims and misleading citations separately from style preferences.

For fiction, assess motivation, agency, canon consistency, scene causality, emotional progression, prose and the ending. Do not reward more generated explanation when it makes the story worse.

For Science, check search completeness claims, actual source access, screening decisions, reported study details, interpretation and references. A model-generated appraisal is not a validated bias assessment or meta-analysis.

Compare direct prompt, outline/review baseline, graph workflow without JEV and graph workflow with JEV. Keep one-source and large-corpus tasks separate. Avoid training a preference model and evaluating it on the same rated examples.

Known limits

Source quotation matching does not prove truth or entailment. Claim review samples important statements rather than exhaustively certifying the manuscript. Graphs can contain plausible but spurious links. Fiction ledgers can become stale after revisions. Science searches are bounded; metadata and abstracts do not count as full-text appraisal.

There is no measured claim that GraphPaper outperforms every direct prompt, no trained global model of an author’s taste, no automatic journal approval and no guarantee that a protected prose edit preserves every nuance.

Historical builds

Earlier results and bug reports remain in the validation directory and versioned release notes. The 0.2.0 native freeze exposed a gap in browser-only/startup testing; 0.2.1 added genuine native-window regression coverage. Historical limitations should not be mistaken for the status of a later tested build.

Clone this wiki locally