Skip to content

Minor Release v6.10.0

Choose a tag to compare

@ruck314 ruck314 released this 13 Apr 23:26
· 260 commits to main since this release
aa4fcaa

Pull Requests Since v6.9.0

Unlabeled

  1. #1162 - Add native C++ ctest suite
  2. #1172 - Add floating point variable types: Float16, Float8, BFloat16, TensorFloat32, Float6, Float4 (ESROGUE-730)
  3. #1164 - Add versioned release and pre-release doc publishing workflows
  4. #1173 - Add SW batcher (CombinerV1/V2) for ESROGUE-516
  5. #1174 - Add SrpV3Emulation module for software-only SRPv3 CI testing (ESROGUE-493)
  6. #1166 - Fix silent fixed-point overflow (ESROGUE-735)
  7. #1168 - fix(memory): prevent false verify failures under concurrent access (ESROGUE-733)
  8. #1167 - Catch GeneralError during init read/write in start() (ESROGUE-734)
  9. #1176 - test: harden UDP packetizer integration test against flakiness
  10. #1171 - Add memBase validation for devices with remote variables (ESROGUE-732)
  11. #1169 - Fix Variable.set/get/post rejecting numpy integer index types (ESROGUE-724)
  12. #1165 - Add path to support Qt6 in the future via qtpy abstraction layer (ESROGUE-688)
  13. #1170 - Fix unpicklable Boost.Python exceptions crashing ZMQ server (ESROGUE-723)
  14. #1177 - fix: modernize super() calls in AxiStreamDmaMon

Pull Request Details

Add native C++ ctest suite

Author: Benjamin Reese bengineerd@users.noreply.github.com
Date: Mon Apr 13 13:45:10 2026 -0700
Pull: #1162 (8686 additions, 159 deletions, 30 files changed)
Branch: slaclab/cpp-tests
Issues: #1162

Notes:

Description

This branch adds the first native C++ regression suite under tests/cpp/ and wires it into the main CMake/CTest flow. The new coverage stays intentionally narrow and fast. It implements deterministic native tests for version helpers, memory bit helpers, stream frame/pool behavior, and stream iterator behavior, plus a Python-enabled API smoke test managed by CTest.

This change also replaces the old standalone tests/api_test sample executable with a CTest-managed smoke test, updates the test documentation to describe how the native suite is built and run from the standard build/ tree, and updates CI so the native test labels run as part of the existing build jobs.

Details

The native test system is added behind ROGUE_BUILD_TESTS in the top-level CMake build and uses an in-tree single-header doctest-compatible harness plus a shared support target so individual test files stay focused on behavior. The suite is organized under tests/cpp/ by subsystem and uses CTest labels (cpp, cpp-core, no-python, requires-python, smoke) to keep local runs and CI selection predictable.

Two product bugs were exposed while building the deterministic native coverage and are fixed here:

  1. src/rogue/interfaces/memory/Master.cpp
    The copyBits, setBits, and anyBits helpers all use a do { ... } while (...) structure after initializing rem = size. When size == 0, the old code still executed the loop body once before checking termination. That made zero-length operations unsafe: a no-op request could still calculate masks, read source bits, inspect destination bits, or write into the destination buffer even though no bits were requested. The fix adds an immediate zero-length return at the top of each helper, preserving the expected semantics: copyBits and setBits become true no-ops, and anyBits returns false without touching memory.

  2. src/rogue/interfaces/stream/Pool.cpp
    The pool destructor previously freed dataQ_.front() inside while (!dataQ_.empty()) but never removed the front element from the queue. In practice that means the destructor repeatedly frees the same pointer while the queue never advances, which can lead to double-free behavior or an infinite drain loop during teardown. The fix pops the queue entry after each free() so the destructor actually walks and releases every pooled allocation exactly once.


Add versioned release and pre-release doc publishing workflows

Author: Benjamin Reese bengineerd@users.noreply.github.com
Date: Mon Apr 13 15:51:39 2026 -0700
Pull: #1164 (1397 additions, 24 deletions, 12 files changed)
Branch: slaclab/doc-hist
Issues: #1164

Notes:

Description

Add versioned documentation publishing on GitHub Pages so release docs can be retained per tag, expose a latest alias for the newest release, and publish development docs from pre-release.

Details

This change:

  • splits docs publishing out of the main CI workflow into dedicated release and pre-release workflows
  • adds a docs publishing helper script that writes versioned docs trees, versions.json, a versions index page, and the latest alias
  • adds a docs version selector and version-status banners to the Sphinx UI
  • adds repo-local implementation and rollout planning documents under docs/plans/

Related

  • release staging workflow supports manual validation under /staging-docs/ before production release publishing is exercised

Add path to support Qt6 in the future via qtpy abstraction layer (ESROGUE-688)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Wed Apr 8 10:44:30 2026 -0700
Pull: #1165 (14 additions, 12 deletions, 4 files changed)
Branch: slaclab/ESROGUE-688
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-688

Notes:

Summary

  • Replace direct PyQt5 imports with qtpy in rogue_plugin.py and time_plotter.py, matching PyDM's own Qt backend abstraction
  • Replace deprecated QFontMetrics.width() with horizontalAdvance() for Qt6 compatibility
  • Add qtpy>=2.0 and pyqtgraph>=0.13 to conda.yml; add qtpy to pip_requirements.txt

Users can now choose their Qt backend (PyQt5 or PySide6) via the QT_API environment variable. Existing PyQt5 users see zero behavior change since qtpy defaults to PyQt5 when QT_API is not set.

Currently, PyDM has limited support for PyQt5 and PySide6, and does not yet support PyQt6. This will allow users to easily migrate to PyQt6 once PyDM adds support for it.

Test plan

  • CI passes on ESROGUE-688 branch (all 4 jobs green)
  • grep -rn "PyQt5\|pyqtSlot" python/ tests/ --include="*.py" returns zero matches
  • flake8 clean on python/ and tests/
  • Full test suite passes under PyQt5 backend
  • Full test suite passes under PySide6 backend

JIRA: ESROGUE-688


Fix silent fixed-point overflow (ESROGUE-735)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Wed Apr 8 10:28:09 2026 -0700
Pull: #1166 (375 additions, 20 deletions, 9 files changed)
Branch: slaclab/ESROGUE-735
Issues: #1166
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-735

Notes:

Summary

  • Fixed-point variables (Fixed/UFixed) now raise a GeneralError when set to a value outside the representable range, instead of silently clipping or storing an incorrect value
  • Added minValue()/maxValue() to Fixed and UFixed Python model classes so the existing C++ range check in Block::setFixed() is activated
  • Replaced the silent positive-edge-case clipping (fPoint -= 1) in C++ with an explicit overflow check that throws with the variable name, attempted value, and valid range

Test plan

  • Existing 103 core tests pass
  • New test_fixed_point_overflow_raises_error covers: positive overflow, negative overflow, unsigned overflow, negative-into-unsigned, and boundary values
  • Verified with Fixed(16, 14) (the etaQ/etaI register scenario, range +-2.0) that overflow now produces a clear error message

Catch GeneralError during init read/write in start() (ESROGUE-734)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Mon Apr 13 12:49:44 2026 -0700
Pull: #1167 (211 additions, 2 deletions, 2 files changed)
Branch: slaclab/ESROGUE-734
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-734

Notes:

Summary

  • Wraps self._read() and self._write() calls in Root.start() with try/except rogue.GeneralError
  • A transaction timeout during initial read/write (e.g. PCIe bus contention when starting multiple SMuRF carriers simultaneously) previously propagated as an unhandled GeneralError and crashed Root.start()
  • The fix catches the exception in start() and logs it via pr.logException(), allowing startup to continue
  • _read() and _write() themselves still raise rogue.GeneralError on timeout — only the start() call site swallows it

Test plan

  • 4 new integration tests in tests/integration/test_init_read_timeout_integration.py
  • test_init_read_timeout_does_not_crash — root survives timeout during init read
  • test_read_raises_on_timeout — _read() raises rogue.GeneralError on timeout (callers outside start() still see the exception)
  • test_init_read_succeeds_without_timeout — sanity check for normal operation
  • test_concurrent_root_start_with_timeouts — 6 roots start simultaneously, 2 with timeouts, all succeed
  • Existing 184 tests pass with no regressions
  • flake8 clean

fix(memory): prevent false verify failures under concurrent access (ESROGUE-733)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Mon Apr 13 09:49:05 2026 -0700
Pull: #1168 (259 additions, 2 deletions, 3 files changed)
Branch: slaclab/ESROGUE-733
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-733

Notes:

Summary

  • Fix race condition in Block::checkTransaction() where verify comparison used live blockData_ instead of the data actually written to hardware
  • When two variables share a block, a concurrent set() on one variable could modify blockData_ between the write and verify check, causing spurious verify errors
  • Add expectedData_ buffer that snapshots blockData_ at write time; checkTransaction() now compares against the snapshot
  • Add regression tests for interleaved and threaded concurrent set scenarios

Test plan

  • test_verify_survives_interleaved_set — reproduces exact ESROGUE-733 scenario
  • test_verify_write_read_roundtrip — basic write-verify-read roundtrip works
  • test_verify_bulk_write_and_verify — bulk write+verify flow works
  • test_verify_threaded_concurrent_set — concurrent threading with 50 iterations
  • Full test suite passes
  • flake8 and cpplint clean

Fix Variable.set/get/post rejecting numpy integer index types (ESROGUE-724)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Tue Apr 7 15:22:09 2026 -0700
Pull: #1169 (52 additions, 0 deletions, 3 files changed)
Branch: slaclab/ESROGUE-724
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-724

Notes:

Summary

  • Variable._set() and _get() C++ methods (via Boost.Python) require a native Python int for the index argument, but numpy integer types (e.g. np.int32) are not auto-converted, causing a type-mismatch error
  • Added index = int(index) conversion at all 5 Python-to-C++ boundary call sites in RemoteVariable.set/get/post and RemoteCommand.set/get
  • Added regression test test_numpy_integer_index_types exercising np.int32, np.int64, np.uint32, np.uint64, and np.intp index types

Test plan

  • New regression test test_numpy_integer_index_types passes
  • All existing integration tests pass (155/155)
  • flake8 clean on all changed files

Fix unpicklable Boost.Python exceptions crashing ZMQ server (ESROGUE-723)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Tue Apr 7 15:22:33 2026 -0700
Pull: #1170 (25 additions, 1 deletions, 2 files changed)
Branch: slaclab/ESROGUE-723
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-723

Notes:

Summary

  • When a Boost.Python.ArgumentError is raised during a ZMQ remote procedure call, pickle.dumps() fails because Boost.Python exceptions cannot be serialized, crashing the server
  • Added a fallback in ZmqServer._doRequest() that converts unpicklable exceptions to a standard Exception preserving the full type name and message
  • Added regression test simulating an unpicklable exception to prevent future regressions

Test plan

  • New unit test verifies unpicklable exceptions are gracefully handled
  • Verified fix works with real Boost.Python.ArgumentError in local build
  • All existing tests pass (180/180)
  • flake8 linter clean

Resolves ESROGUE-723


Add memBase validation for devices with remote variables (ESROGUE-732)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Mon Apr 13 09:51:21 2026 -0700
Pull: #1171 (133 additions, 0 deletions, 2 files changed)
Branch: slaclab/ESROGUE-732
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-732

Notes:

Summary

  • Add _validateMemBase() to Device that raises DeviceError when a device with RemoteVariables or RemoteCommands is added to Root without a memBase, catching silent runtime timeouts at start time
  • Walk the ancestor chain to allow sub-devices that inherit memBase from a parent to continue working
  • Detect Root via the parent self-cycle (node._parent is node) that Root._rootAttached() creates, which safely handles Root subclasses without circular imports
  • Add 5 unit tests covering all validation cases: RemoteVariable error, RemoteCommand error, local-only pass, inherited memBase pass, Root subclass error

Resolves ESROGUE-732

Test plan

  • test_remote_variable_no_membase_raises — DeviceError raised for RemoteVariable device without memBase
  • test_remote_command_no_membase_raises — DeviceError raised for RemoteCommand device without memBase
  • test_local_only_device_no_membase_passes — Local-only device starts without error
  • test_sub_device_inherits_membase — Child device inherits memBase from parent
  • test_root_subclass_no_membase_raises — DeviceError raised for Root subclass without memBase (no infinite loop)
  • Full existing test suite passes with no regressions

Add floating point variable types: Float16, Float8, BFloat16, TensorFloat32, Float6, Float4 (ESROGUE-730)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Mon Apr 13 09:50:52 2026 -0700
Pull: #1172 (3784 additions, 35 deletions, 26 files changed)
Branch: slaclab/ESROGUE-730
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-730

Notes:

Summary

Extends Rogue's variable model system with six new floating point types for NVIDIA GPU interoperability.

New types

Type Format Bits Value Range NVIDIA Support
pr.Float16 / pr.Float16BE IEEE 754 half-precision, 1s/5e/10m 16 ±65504.0 All
pr.Float8 / pr.Float8BE E4M3, NaN=0x7F, no Inf 8 ±448.0 Hopper, Blackwell
pr.BFloat16 / pr.BFloat16BE 1s/8e/7m, IEEE exponent range 16 ±3.39e38 Ampere, Hopper, Blackwell
pr.TensorFloat32 / pr.TensorFloat32BE 1s/8e/10m, 19 bits in 4 bytes 19 (32 storage) ±3.40e38 Ampere, Hopper, Blackwell
pr.Float6 / pr.Float6BE E3M2, no NaN/Inf 6 ±28.0 Blackwell
pr.Float4 / pr.Float4BE E2M1, no NaN/Inf 4 ±6.0 Blackwell

Implementation

Each type follows the same full-stack pattern:

  • C++ model ID constant in Constants.h
  • Block get/set methods with inline bit-manipulation converters in Block.cpp
  • Variable dispatch in Variable.cpp
  • Python Model class with toBytes/fromBytes/fromString/minValue/maxValue
  • Sphinx API documentation (per-type page + consolidated summary reference)

Special-value handling per spec:

  • Float16: full IEEE 754 NaN and ±Inf
  • Float8: NaN = 0x7F, no infinity (clamps to max finite)
  • BFloat16/TF32: full IEEE 754 NaN and ±Inf
  • Float6/Float4: no NaN or Inf (clamps to max finite on encode)

Documentation

Added docs/src/pyrogue_tree/core/float_types_summary.rst — a consolidated quick-reference page with format details table, special-value handling, NVIDIA architecture support matrix, and usage example.

Test plan

  • pytest tests/core/test_float16.py — all boundary values, round-trip, metadata, remote variable integration
  • pytest tests/core/test_float8.py — all boundary values, round-trip, metadata, remote variable integration
  • pytest tests/core/test_bfloat16.py — all boundary values, round-trip, metadata, remote variable integration
  • pytest tests/core/test_tensorfloat32.py — all boundary values, round-trip, metadata, remote variable integration
  • pytest tests/core/test_float6.py — all 64 bit patterns, boundary values, NaN/Inf clamping, remote variable integration
  • pytest tests/core/test_float4.py — all 16 bit patterns, boundary values, NaN/Inf clamping, remote variable integration
  • Full regression: 207 passed
  • flake8 + cpplint clean

Resolves ESROGUE-730.


Add SW batcher (CombinerV1/V2) for ESROGUE-516

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Mon Apr 13 09:50:09 2026 -0700
Pull: #1173 (1316 additions, 3 deletions, 18 files changed)
Branch: slaclab/ESROGUE-516
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-516

Notes:

Summary

  • Add CombinerV1 and CombinerV2 classes that build batcher v1/v2 super-frames from individual stream frames in software (ESROGUE-516)
  • Fix pre-existing segfault in CoreV1/CoreV2::processFrame() where frame_ member was never assigned, causing beginHeader() null dereference when used by Inverter classes
  • Fix CombinerV1 tail byte 7 to write WIDTH encoding (matching VHDL tData(59 downto 56) := WIDTH_C) instead of zero
  • Add 18 pytest regression tests using the SW batcher to test Splitter and Inverter unbatching
  • Add full documentation (C++ API, Python API, built-in modules guide, logger names)

Zero behavioral changes to the unbatcher

The only existing unbatcher files touched in this PR are CoreV1.cpp and CoreV2.cpp. No changes were made to SplitterV1, SplitterV2, InverterV1, InverterV2, or Data.

The two changes to CoreV1/V2 are:

1. Bug fix: frame_ was never assigned in processFrame()

On the pre-release branch, CoreV1::processFrame() and CoreV2::processFrame() never set frame_ = frame. The member frame_ was only ever cleared in reset() but never populated. This means that after a successful processFrame() call, the beginHeader(), endHeader(), beginTail(), and endTail() accessors — which all dereference frame_ — would crash with a null pointer dereference.

This went unnoticed because SplitterV1/V2 don't call those accessors (they only use core.record(x) which works off the list_ and tails_ vectors populated during parsing). However, InverterV1/V2 do call core.beginHeader() and core.beginTail(), so the Inverter classes were silently relying on the fact that frame (the local parameter) kept the shared_ptr alive for the duration of acceptFrame(). If anyone used CoreV1/V2 directly and called beginHeader()/endHeader() after processFrame() returned, it would have crashed.

The fix adds frame_ = frame; after all header validation passes, before the record-parsing loop. This is a correctness fix, not a behavioral change — it makes CoreV1/V2 work as their API promises.

2. Comment fix: Tail Word 1 bits 31:24 description (CoreV1 only)

Changed "Valid bytes in last field" to "Width (bits 3:0 = width encoding, bits 7:4 = 0)" to match the VHDL firmware where tData(59 downto 56) := WIDTH_C. This is documentation-only; no code behavior change.

Summary: The SplitterV1/V2 code paths are completely unaffected — same parsing logic, same record extraction, same metadata handling. The frame_ fix closes a latent null-deref bug in the CoreV1/V2 API that existing callers happened to avoid.

Test plan

  • flake8 tests/protocols/test_batcher.py passes clean
  • All 18 new batcher tests pass (pytest tests/protocols/test_batcher.py -v)
  • Full test suite passes
  • Build succeeds with cmake .. -DROGUE_INSTALL=local && make
  • CI pipeline passes

Add SrpV3Emulation module for software-only SRPv3 CI testing (ESROGUE-493)

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Mon Apr 13 12:01:45 2026 -0700
Pull: #1174 (912 additions, 0 deletions, 11 files changed)
Branch: slaclab/ESROGUE-493
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-493

Notes:

Summary

  • Adds SrpV3Emulation — a software SRPv3 endpoint that receives SRP request frames, processes them against internal memory emulation, and sends response frames
  • Enables CI regression testing of the SrpV3 client without FPGA/ASIC hardware (ESROGUE-493)
  • Uses a worker thread + queue to avoid deadlock from synchronous frame processing within the SrpV3 transaction lock scope
  • Memory allocated on demand in 4 KiB pages with random initialization (emulates uninitialized SRAM)
  • Validates request size and opcode fields, rejecting malformed frames

Test plan

  • 8 pytest tests in tests/protocols/test_srpv3_emulation.py — all passing
  • Tests cover: basic write/read, multiple registers, overwrite, 64-bit wide register, uninitialized read (consistent random data), multi-device address isolation, zero write, invalid request size rejection
  • C++ and Python linters pass (./scripts/run_linters.sh)
  • CI pipeline passes

test: harden UDP packetizer integration test against flakiness

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Mon Apr 13 10:29:54 2026 -0700
Pull: #1176 (145 additions, 35 deletions, 1 files changed)
Branch: slaclab/ESROGUE-730-flaky-udp-test
Jira: https://jira.slac.stanford.edu/issues/ESROGUE-730-flaky-udp-test

Notes:

Summary

  • Replace wall-clock DRAIN_TIMEOUT / reorder-settle timeouts in tests/integration/test_udp_packetizer_integration.py with progress-based polling (wait_for_progress). As long as the monitored counter keeps advancing, the test keeps waiting — so slow CI runners no longer produce spurious failures. It only bails if the counter stops advancing for STALL_WINDOW = 5.0s.
  • Wrap the body in try/finally with deterministic, ordered teardown: flush the reordering stage, stop client then server RSSI, brief unwind sleep, then drop references in reverse link order. This avoids the worker-crash signature observed in https://github.com/slaclab/rogue/actions/runs/24151542407 (C++ use-after-free during Python teardown when upstream stages outlive downstream slaves).
  • Close the reorder-drain race: wait until every generated frame has been observed by the reordering stage (out_of_order.cnt >= FRAME_COUNT) before setting period = 0, so the flush in the setter cannot miss a frame that gets cached a moment later.
  • Add a small public cnt property on RssiOutOfOrder to expose the observation counter safely.
  • Bump CONNECTION_TIMEOUT to 30s (RSSI open is a binary handshake with no incremental progress to poll).

Test plan

  • Looped test_data_path locally 30x — 30 passed / 0 failed.
  • Full non-perf test suite (python -m pytest -n auto --dist loadfile -m "not perf") run locally — all 4 test_data_path parametrizations pass. (Unrelated pre-existing env failures on softioc/pydm/zmq-port were reproduced on unmodified pre-release and are not from this change.)
  • ./scripts/run_linters.sh — clean.
  • GitHub Actions "Full Build Test / Rogue Tests" green on this PR.

fix: modernize super() calls in AxiStreamDmaMon

Author: Larry Ruckman ruckman@slac.stanford.edu
Date: Mon Apr 13 10:23:57 2026 -0700
Pull: #1177 (5 additions, 5 deletions, 1 files changed)
Branch: slaclab/fix-axi-stream-dma-mon

Notes:

Summary

  • Replace deprecated Python 2 super(self.__class__, self).__init__() with Python 3 super().__init__() in all three AxiStreamDmaMon classes to prevent potential MRO infinite recursion with subclassing
  • Clarify BuffCount descriptions in Rx/Tx to document they are static driver configuration values (no polling needed)

Test plan

  • python3 -m compileall -f passes with zero errors
  • flake8 --count passes with zero violations
  • CI regression tests pass