Skip to content

v1.0.10 — Improved visualizations, complete asymmetric crypto metrics, ML‑KEM PQC benchmark, more iterations, docs & report tweaks

Choose a tag to compare

@tobrien tobrien released this 01 Jan 01:20
· 8 commits to working since this release
2f8104a

Summary

This release focuses on visualization robustness, completeness of benchmark metrics, improved reporting, and adding a simple ML‑KEM (post‑quantum) micro‑benchmark. The most visible user changes are clearer and properly scaled charts (small differences are now visible), regenerated result pages that include RSA/ECDSA metrics, and documentation that explains handshake metric naming. Internally: benchmark iterations were increased for better statistics, the PQC benchmark tool was added and wired into the benchmark script and Dockerfile, and report/viz generators were hardened to handle missing data.

Why this matters

  • Charts now show small performance differences (0.1%–1%) clearly — no more “collapsed” or empty charts when data is present.
  • Older/incomplete result files that lacked RSA/ECDSA metrics no longer cause broken charts; when data is missing the UI shows a clear message explaining what to run to generate it.
  • Benchmarks now run with more iterations by default so means and stddevs are more reliable (but runs take longer).
  • Basic ML‑KEM (ML‑KEM‑768) microbenchmark is included and can be measured when OpenSSL + tool support exists.

What changed (grouped)

  1. Visualization & charting
  • Small Multiples improvements (scripts/generate-viz.js and scripts/generate-viz-multipage.js):
    • Dynamic Y‑axis auto‑scaling based on actual data range (with padding) instead of a fixed ±10% domain.
    • Increased chart/card heights and margins for better label placement.
    • Prominent percentage labels on bars (one decimal place), color coded (green/red), and a dashed zero reference line.
    • TLS chart width/padding fix to prevent SVG overflow and cramped slope lines (getWidth usage corrected).
    • Added table and explanation areas for block size sensitivity and other helpful UI content.
  • When a group of metrics is entirely missing (e.g., RSA/ECDSA), visualizations now show a descriptive message rather than rendering a zero domain.
  • Multi‑page generator now conditionally generates the Mráz optimization page only when optimized data is present and exposes hasOptimizedData to navigation.

Files of interest: scripts/generate-viz.js, scripts/generate-viz-multipage.js, results/visualizations.html (regenerated), results/* (regenerated pages)

  1. Missing asymmetric-crypto metrics: investigation + results
  • Root cause documented: some committed results were from an earlier run that did not capture RSA/ECDSA metrics.
  • Result JSON files under results/ were updated to include RSA (sign/verify) and ECDSA (sign/verify) metrics so those charts render correctly.
  • Investigation and next‑steps docs added: INVESTIGATION_RESULTS.md and NEXT_STEPS.md describing how to re-run benchmarks and why the data was missing.

Impact: If you see empty RSA/ECDSA charts on GitHub Pages, re-run the full benchmark suite (see Actions below) to populate results and regenerate the pages.

  1. Benchmark iterations and statistics
  • Default iterations increased from 3 → 10 (config/versions.json).
    • Benefit: better statistical confidence for metrics (mean ± stddev).
    • Cost: benchmark runs will take longer (~3× more wall time unless parallelized/CI adjusted).
  • generate-report.js: standard deviation / ± display is now conditional on iterationCount > 1 and report text expanded with more insight into block-size behavior.
  1. PQC (ML‑KEM) microbenchmark support
  • Added a small C benchmark tool: src/mlkem_bench.c.
  • Dockerfile: compiles mlkem_bench into the container so it can be used during benchmark runs.
  • Benchmark script (src/benchmark.sh): detects and runs the custom mlkem_bench when present and parses its output to set metrics.ml_kem_768_ops_sec. If the tool isn’t present or parsing fails the metric is set to 0.
  1. Benchmark script hardening and metric naming
  • src/benchmark.sh: more robust parsing (many metrics now set defensively to 0 if empty), improved warnings when tests return empty values, and uses explicit tls1_3_* and tls1_2_* metric keys alongside deprecated legacy names.
  • README and docs: clarified that legacy metrics handshakes_new_per_sec and handshakes_resume_per_sec are actually TLS 1.3 RSA handshake metrics and are deprecated — explicit metric names are recommended (e.g. tls1_3_rsa_new_cps, tls1_3_rsa_resume_cps).
  1. Docs and developer guidance
  • Many new or updated docs describing the chart improvements, TLS chart width fix, investigation findings, and how to regenerate results and visualizations (docs/SMALL_MULTIPLES_IMPROVEMENTS.md, CHART_IMPROVEMENTS_SUMMARY.md, VIEW_IMPROVEMENTS.md, YOUR_CHART_PREVIEW.md, METRIC_CLARIFICATION_CHANGES.md, INVESTIGATION_RESULTS.md, NEXT_STEPS.md, etc.).
  • scripts/regenerate-from-local.sh updated: it now skips printing a mraz.html message when that file is absent.
  1. Results & generated pages
  • Several generated result files and HTML pages were regenerated and added to the repository under results/ and downloaded-gh-results/ (including index.html, visualizations.html, overview.html, bellingrath.html, schmatz.html, pqc.html, mraz.html when applicable, summary.json, detailed-iterations.json, REPORT.md).
  1. Version bump & repository note
  • package.json version bumped to 1.0.10 and committed.
  • The release includes committed node_modules artifacts (node_modules/.package-lock.json and node_modules/.vite/vitest/results.json) — these look like generated/test artifacts and are typically not committed. Consider removing them from the repo history if they were added unintentionally.

Impact and compatibility

  • User-visible impact

    • Charts will look different (height, label precision, auto scale). This is an improvement for readability but may change snapshots or visual tests.
    • To see RSA/ECDSA charts you may need to regenerate results by re‑running the full benchmark suite if your current results are incomplete.
    • Benchmarks will take longer by default due to iterations increasing from 3 to 10. Adjust CI/workflow timeouts or run modes if needed.
    • New PQC metric ml_kem_768_ops_sec may appear in results; consumers of summary.json should be prepared to see this new key (or 0 when unavailable).
  • Developer impact

    • Visualization code changed substantially in scripts/generate-viz.js and scripts/generate-viz-multipage.js. If you have tests or automation that parse generated HTML, you may need to update them.
    • generate-report.js now conditionally shows stddev and includes expanded narrative text — report parsing should still work but tables may change when iterationCount > 1.
    • Legacy handshake metric names are still written for backward compatibility, but they are marked deprecated. Migrate any downstream analysis to the explicit tls1_3_* / tls1_2_* metric keys where appropriate.

Breaking changes and important considerations

  • No intentional breaking changes to external metric names were made — legacy metrics remain present for backward compatibility. However:
    • New explicit metric keys were added (tls1_3_rsa_, tls1_2_, optimized_*, ml_kem_768_ops_sec). Consumers that expect a fixed set of metrics should verify and accept the new keys.
    • The default iterations increase (3→10) will change the runtime and statistical output (means and stddevs). CI configurations and timeout values should be checked.
  • Node_modules artifacts committed in this release are likely accidental and should be removed. They are not functional changes but can bloat the repo and confuse reviewers.

Recommended actions for users and maintainers

  • If you rely on up‑to‑date RSA/ECDSA data in charts: re‑run the full benchmark suite and regenerate visualizations.
    • Quick commands (from repo root):
      • Run all benchmarks: npm run benchmark # may take ~60 minutes locally depending on environment
      • Or trigger the benchmark workflow in GitHub Actions (preferred for full set): use the repository's benchmark.yml workflow
      • Aggregate & report locally: npm run aggregate:local && npm run report
      • Regenerate visualizations: node scripts/generate-viz.js (or npm run visualize if configured)
  • Review CI/workflow timeouts and resource allocations to account for the iterations change.
  • For downstream tooling that parses results/summary.json, add handling for new metrics (tls1_3_, optimized_, ml_kem_768_ops_sec) and tolerate metrics being zero when not present.
  • Remove node_modules artifacts from version control if they were committed accidentally (recommended):
    • git rm --cached node_modules/.package-lock.json node_modules/.vite/vitest/results.json
    • Add an exception to .gitignore if appropriate and re-commit.

Developer notes (where to review)

  • Visualization logic: scripts/generate-viz.js and scripts/generate-viz-multipage.js
  • Missing data handling and rendering: renderGroupedBarChart / grouped render functions in the above scripts
  • TLS metric naming & deprecation notes: METRIC_CLARIFICATION_CHANGES.md and README.md additions
  • Benchmark script and PQC: src/benchmark.sh and src/mlkem_bench.c; Dockerfile changes compile the mlkem_bench helper
  • Report generation: scripts/generate-report.js
  • Result files regenerated: results/.json and results/.html (inspect these to verify no placeholder/sensitive data was left in exported HTML/JSON)

Changelog highlights (short)

  • Visualizations: auto‑scaling Y axes, one‑decimal percentage labels, taller charts, zero reference line, improved TLS width handling.
  • Results: RSA/ECDSA sign/verify metrics populated into result-*.json so related charts render correctly.
  • Benchmark: iterations increased (3 → 10), report formatting improved, PQC support added via mlkem_bench.c and Dockerfile compilation.
  • Docs: many explanatory docs and investigation notes added to explain missing data and how to re‑run benchmarks.
  • Version: package.json bumped to 1.0.10; node_modules generated artifacts were committed (review/remove if accidental).

If anything in these notes needs to be expanded (commands, locations of specific lines to review, or help removing accidental artifacts), check the files mentioned above or ask for targeted guidance.