Skip to content

IFSSIM v0.1.2 — 9-state EKF + Schmidt-Kalman bias-freeze + two-phase mission lifecycle

Choose a tag to compare

@raulmoranguerra raulmoranguerra released this 27 May 06:17
· 1 commit to main since this release
Immutable release. Only release title and notes can be modified.
547b84a

[0.1.2] — 2026-05-26

The biggest autonomy-pipeline release since v0.1.0. Three intersecting
threads: a clean 9-state EKF with Coriolis-correct mechanics
(rewrite + post-rewrite hardening), a two-phase mission lifecycle
that separates SetMission (configure) from RuntimeControl (run),
and a stack of SLAM hardening changes that took mid-lap /odom
position drift from 56 m down to under 2 m on the autocross course
and eliminated the late-corner SLAM yaw-cascade pattern.

Sim-side: documentation overhaul, GPL-3.0-or-later relicense, auto-
pull bags onto the host on session-stop, and a compose-watch-based
inner-loop for Python edits without rebuilding the image.

Changed

  • Breaking — odometry filter rewritten as a 9-state EKF.
    The complementary filter that previously lived in
    pipeline/odometry_filter/ is replaced by a hand-rolled EKF over
    state [x, y, θ, vx, vy, ω, ba_x, ba_y, bg_z]. The predict step
    now carries the Coriolis cross-terms (v̇x = ax + ω·vy,
    v̇y = ay − ω·vx), which the complementary filter could not.
    During steady-state cornering the IMU's body-y axis reads
    centripetal acceleration ω·vx — the old filter integrated this
    directly into vy, so /odom.vy drifted unbounded (4+ m/s after
    68 s of cornering) and the inflated state.speed inside
    PurePursuit was the trackdrive ceiling. The new EKF cancels the
    Coriolis term in v̇y exactly when the model is consistent; the
    gtest regression CoriolisCornering.SteadyVyIsNearZero
    synthesises a 60 s constant-radius turn and asserts
    |vy| < 0.10 m/s (pre-rewrite this fails by ~14×).

    Measurements: motor-RPM is now a Kalman update on vx with
    sigma_rpm = 2 cm/s, so RPM dominates over IMU accel integration
    for vx tracking. Steering kinematic-bicycle
    (ω_pred = (vx/L)·tan δ) is a gated update on ω — applied
    when |residual| < threshold, raises slip_flag and rejects the
    update otherwise (the kinematic-bicycle model is wrong under slip,
    so folding it in would corrupt yaw).

    Breaking surface:

    • /brake_pressure subscription dropped from
      odometry_filter_node. Brake authority pulled α_vx toward
      zero during heavy braking pre-rewrite; the EKF's Kalman gating
      handles wheel-lockup naturally via covariance. Bridge still
      publishes /brake_pressure for replay parity, but no autonomy
      consumer subscribes.
    • /odom_diag/effective_alpha_vx topic retired.
    • OdometryFilter::push_brake() method, Params struct (replaced
      by EkfParams), kAlphaVx*, kBetaVyLeak, kBrakeLockup*,
      kRpmStaleS all removed.

    Preserved: /odom (now with non-trivial covariance populated
    from the EKF P matrix), odom→base_link TF, lifecycle, QoS,
    frame names, and the /odom_diag/yaw_residual_rad_s +
    /odom_diag/slip_flag topics.

Added

  • Auto-pull bag onto host filesystem on session stop (#498,
    closes the manual pull-bag.sh step from #490). When Record bag
    (mcap)
    is ticked, clicking Stop Session now:

    1. Finalises the bag in the docker volume (ifssim_bags, on
      container ext4 — fast).
    2. Streams the bag tarball out of dv_pipeline_stack via the
      Python docker SDK (container.get_archive) and extracts it
      into the host's ./bags/<name>/ (bind-mounted to mc_backend
      as /host_bags/).
    3. Cleans the volume-side copy via container.exec_run(rm -rf …).

    All inside the StopBag callback. No user action required to get
    the bag onto the host filesystem; tools/pull-bag.sh is now the
    manual-recovery path, not the default flow. Session log surfaces
    progress + the final host path. Failure semantics are best-effort:
    if the transfer fails (host disk full, etc.), the bag stays in
    the volume and the user gets a clear log line pointing at the
    manual recovery command.

    Env-gated via IFSSIM_BAG_AUTO_PULL (default 1). Set to 0 to
    disable and keep the manual pull-bag.sh flow.

    Why the Python SDK and not the docker CLI: the Linux docker CLI
    inside mc_backend can't pass a Windows host path to docker cp
    (the first : in C:/Users/... gets parsed as a container name,
    and the daemon doesn't recognise /c/Users/... as a host mount).
    The SDK uses the daemon's HTTP API directly — no argv parsing.
    Tarball extraction uses Python 3.12+'s safe filter="data"
    policy, plus a member-path traversal check for defence in depth.

    One security note: the mc_backend container now mounts
    /var/run/docker.sock (for the SDK) and ./bags:/host_bags
    (for the extraction target). docker.sock means anything that
    compromises mc_backend can root the host. The blast radius for
    mc_backend was already broad (it talks RPC to the sim, runs ROS
    actions, holds the API key), so this doesn't materially change
    the threat model for this dev/sim stack. For a hardened
    deployment, set IFSSIM_BAG_AUTO_PULL=0 AND drop both mounts
    from docker-compose.yml.

    The ./bags:/host_bags bind-mount partially reverses #490's
    retirement of host bind-mounts on the autonomy stack — but it
    only applies to mc_backend (low-traffic) and only at session-stop
    (write-only, not on any hot path), so the 30-90 s startup-time
    cost #490 was targeting doesn't apply here. Recording itself
    still lands in the named volume on container ext4 (fast); this
    bind-mount only takes the finalised tarball at the end.

  • Two-phase mission lifecycle: SetMission → RuntimeControl
    (#499, #518). The single StartMission action that previously
    combined configure + activate is replaced by an explicit two-step
    protocol on mission_control_node:

    • SetMission(mission_id) — prepare phase. Resolves the
      mission via MODE_REGISTRY (single source of truth for
      mission → ordered (node, behavior) in
      pipeline/mode_manager/mode_manager/mode_registry.py),
      calls each autonomy node's new ~/setup service with
      (mode_name, behavior), then drives the configure lifecycle
      transition. Per-node progress streams back as Feedback.stage
      so the operator sees live bring-up state instead of a single
      "starting" wait. Numba JIT (~10-20 s on cone_detection_node)
      lands here, not in RuntimeControl.
    • RuntimeControl — activate + run. Once SetMission returns
      success=true, the client opens RuntimeControl;
      mission_control_node activates the prepared stack and streams
      throttle/steering/emergency/finished feedback at 40 Hz.
      sim_supervisor_node subscribes to the feedback topic and
      relays each frame onto /fsds/control_command for the bridge —
      a clean split from before, where the supervisor was the action
      server.

    New packages: pipeline/node_base/ (Python) and
    pipeline/node_base_cpp/ (C++) provide BaseLifecycleNode with
    the ~/setup plumbing; every managed autonomy node now inherits
    from it. pipeline/bringup/ consolidates the launch files
    (sim_pipeline.launch.py, car_pipeline.launch.py,
    full_pipeline.launch.py) — docker/dv_pipeline_stack/ pipeline.launch.py is now a thin include of these.

  • odometry_filter_node (#499 / #518). The IMU+RPM
    complementary filter that previously lived inside
    sim_supervisor_node was lifted into a standalone C++
    BaseLifecycleNode (pipeline/odometry_filter_node/, backed by
    pipeline/odometry_filter/ for the algorithm library). Same
    /odom contract (100 Hz, IMU+RPM+steering+brake_pressure,
    odom→base_link TF, identical diagnostics on
    /odom_diag/*) — but now part of the standard managed bring-up
    via activate_mode, alongside the other four autonomy nodes.
    Matches the real-car split where the uDV's odometry firmware is
    a separate subsystem from the mission interface.

  • supervisor_cli (#518). Terminal client that mirrors the web
    backend's two-phase flow without needing Mission Control:
    ros2 run sim_supervisor supervisor_cli {set_mission, start_mission, run}.
    Reproduces production lifecycle from a shell — useful for
    developing missions/behaviors and triaging session-start bugs.
    Mission IDs match the registry (1=trackdrive, 2=autocross,
    3=accel, 4=skidpad, 5=scruti, 0=tear down).

  • tools/compose-up-and-watch.sh|ps1 (#518). Wrapper that
    runs docker compose up -d then docker compose watch in the
    foreground, syncing host edits under pipeline/* into named
    src volumes for live Python iteration. Replaces the pre-#490
    bind-mount workflow; tools/refresh-bridge.sh is still the
    go-to for C++ / msg / launch / setup.py changes.

  • DriveController / CompositeDriveController / Stanley in
    pipeline/control/control/controllers/. The control node now
    picks its lateral+longitudinal strategy from the behavior
    string passed via ~/setup — pure_pursuit for trackdrive/accel,
    stanley for autocross/skidpad/scruti — instead of a runtime
    ROS parameter. Composite wraps both into a single
    ActuationCommand-returning interface.

Changed

  • Project relicensed to GPL-3.0-or-later. Top-level LICENSE
    added with the full GPL-3.0 text. All package.xml files under
    pipeline/ and ros2/src/ updated from their previous mix of
    MIT (7 packages), Apache-2.0 (4 packages), and GPLv2 (2 packages —
    fs_msgs + ifssim_bridge inherited from upstream FSDS-Sim) to a
    uniform GPL-3.0-or-later. Motivation: keep the stack on a single
    copyleft-compatible licence so prospective GPL-3-only SLAM /
    perception dependencies can be linked in without per-package
    licence-conflict audits. Apache-2.0 contributions remain
    attributable through git history; the relicense reflects forward
    distribution only.

  • Documentation overhaul (#492). Consolidated the three setup
    entry points (readme.md → QUICKSTART.md → GETTING_STARTED_DOCKER.md)
    into one unified path:

    • New docs/SETUP.md — single first-time-user setup, parallel
      Windows + macOS sections. Covers prereqs, submodule init, sim
      build / download, Docker stack, first Mission Control session,
      troubleshooting. Replaces QUICKSTART.md (Mac-only, stale post-#482)
      and GETTING_STARTED_DOCKER.md (UE 5.4 + bind-mount workflow,
      both wrong post-v0.1.0 / #490).
    • New docs/OPERATING.md — daily ops: refresh-bridge cycle,
      recording and retrieving MCAP bags, switching tracks, common
      failures during a session.
    • docs/autonomy_pipeline.md → docs/AUTONOMY.md (renamed, links
      refreshed for the new doc names).
    • docs/FUNCTIONALITIES.md → docs/REFERENCE.md (renamed so
      first-time users don't mistake the 1000-line reference for setup).
    • New docs/CONTRIBUTING.md — branch flow + CI + release
      process. Extracted from readme.md's "Hacking on IFSSIM" section.
    • readme.md slimmed to a thin front door (45 lines). All
      contributor flow moved to CONTRIBUTING.md.
    • .github/ISSUE_TEMPLATE/*.yml + PULL_REQUEST_TEMPLATE.md
      repointed at the new doc names. Code comments in
      docker/, pipeline/sim_supervisor/, ros2/src/ifssim_bridge/
      also updated.

Added (SLAM + EKF hardening, post-v0.1.0 9-state EKF)

A second wave of autonomy-pipeline fixes landed after the initial EKF
rewrite. Each is a small change but they compose into the headline
"/odom position drift mid-lap dropped from 56 m to under 2 m on the
60 s autocross course, integrated yaw error from +19° mean down to
+1.3° mean, and the late-corner SLAM yaw-cascade pattern is gone."

  • /odom as a SLAM pose prior (#545). cone_graph_slam adds a
    BetweenFactorPose3 on consecutive poses driven by the latest
    /odom delta. Bridges cone-poor windows where the previous IMU-only
    prior gave the optimizer too much freedom and let it snap to
    bad cone factors. Anchors the inter-scan pose so cone DA is
    evaluated against a pose consistent with wheel odometry, not just
    IMU pre-integration.
  • Kinematic-bicycle steering BetweenFactor (#543). A second pose
    prior derived from ω = (vx/L)·tan δ, silent during cornering by
    design (the slip-gate suppresses the factor when the kinematic
    bicycle is wrong) but informative on straights — modestly
    constrains yaw without contributing model error during transients.
  • Cascade-skip recovery with force-accept threshold (#541). When
    the DA-failure cascade detector skips 5+ consecutive scans of cone
    factors, force-accept the next scan as new-territory exploration.
    Prevents permanent IMU-only drift when the car enters genuinely new
    cone geometry (the detector's "everything looks new" signature is
    ambiguous between cascade and exploration).
  • Under-observed landmark filter (#536). /Conos (the
    map-frame cone publication consumed by path_planning_node) now
    excludes landmarks observed fewer than 3 times. Filters single-shot
    perception artifacts and phantom DA-spawn ghosts. Cone count on
    /Conos drops from ~130 to ~100 on a clean autocross run; planner
    sees a more stable map.
  • EKF stationary-calibration invariant restored (#534). The
    3-second post-activate stationary-calibration window had been
    drifting due to a Phase-3↔4 ordering issue surfaced by the
    SetMission/RuntimeControl two-phase split. Restored the
    per-sample stationarity gate and added a low-vx gate on the
    steering correction (skip when vx < 3 m/s, where kinematic-
    bicycle equilibrium hasn't built up).
  • Gyro bias freeze + Joseph-form covariance update + slip-aware
    NHC
    (#555). The 9-state EKF's F[OMEGA, BG_Z] = −1 propagation
    produces non-zero P[BG_Z, OMEGA], so the standard Kalman update
    for any non-gyro observation (RPM, kinematic-bicycle steering,
    NHC) leaks into the gyro bias state. Over a 15 s sustained corner
    bg_z walked ±1.8 °/s, integrating into 27° of mid-lap yaw error.
    Fix is per-correction Schmidt-Kalman partitioning — explicitly
    zero K[i] for every state not legitimately informed by the
    observation, then use Joseph form
    P = (I-K·H)·P·(I-K·H)ᵀ + K·R·Kᵀ for the covariance update so P
    stays symmetric and PSD under the modified gain. NHC is now
    always-on with a slip-aware sigma (tight 0.10 m/s when
    !slip_flag, loose 0.50 m/s when slip_flag) — previously it
    was gated off entirely during slip, letting vy run unbounded
    for half of every autocross lap. Net signal-side result:
    /odom yaw-rate offset on straights drops from +0.526 °/s to
    +0.046 °/s; peak /odom-vs-GT body-frame position drift drops
    from 56 m to 1.4 m; SLAM cone-DA cascade events are gone in the
    cascade-pulse scan.
  • sigma_bg_walk 1e-5 → 1e-4 (#539). One-line tune of the
    gyro-bias random-walk process noise. Loosened to let bg_z track
    in-run drift instead of frozen-bias semantics. Subsumed for
    practical purposes by #555's Schmidt-Kalman partitioning but
    kept for the case where bias really does drift physically.

Added (tooling)

  • tools/refresh-bridge.sh stale-volume guard (#548). The
    development workflow rebuilds the dv_pipeline_stack image and
    recreates the container — but compose-watch's named source
    volumes persist across both, leading to a silent failure where
    source edits were rebuilt but the running container still saw
    the previous content. Fixed in two parts: drop the named source
    volumes after compose build, and verify post-recreate that the
    container's source matches the host (sha256, three canary files
    in odometry_filter + cone_slam). Aborts non-zero if any canary
    doesn't match.
  • tools/replay.sh fix-up (#547). Stale executable names
    (ros2 run slam Cone_Detection etc.) updated to the post-#530-
    revert names (cone_detection cone_detection_node,
    cone_slam slam_node, path_planning path_planning_node,
    control control_node). Added explicit lifecycle transitions
    since mode_manager doesn't run in replay. Topic rename
    /cone_slam/state → /slam/pose (per #382) applied to the
    comparator and recorder topic lists.
  • Lichtblick layouts: integrated yaw drift + per-axis SLAM-vs-GT
    plots
    (#537). Two new userNodes / panels in clean_bag_replay
    and slam_debug for diagnosing the yaw-integration and per-axis
    drift modes addressed by #555.

Reverted

  • slam: two-phase rewrite (#530) reverted in #531. The
    architecture couldn't be validated end-to-end on bag _211619
    before the EKF-stack work landed (/odom drift made the
    rewrite's cone DA cascade indistinguishable from baseline). The
    legacy cone_graph_slam_node entry-point retained on dev as the
    shipping SLAM. The two-phase rewrite branch is preserved under
    origin/feat/slam-two-phase-rewrite for future revival once
    /odom is healthy enough to validate the architecture in
    isolation.

Fixed

  • mc-backend: stream bag tarball to disk in auto-pull; fix UI
    badge
    (#528). Large bags would OOM mc_backend when
    container.get_archive was buffered in-memory. Now streamed to
    a temp file with chunked extraction. UI badge for "Bag pulled
    to host" now reflects the actual extraction result instead of
    going stale on streaming failures.

Planned for v0.1.3

  • Cone-floor clipping fix. Spline-spawned cones currently sit a
    few cm into the asphalt on every track because the spline control
    points are at world Z=+100 but the ground mesh's local-bbox top
    sits a few cm higher. Diagnosis + per-probe evidence in
    docs/cone_floor_clipping_fix.md; fix is a line-trace ground-snap
    in the spline_cones* BP construction script (~1 h editor work).
    Visual + LiDAR-perf impact (bottom rings of close cones get
    occluded by the floor mesh). Deferred from v0.1.0 / v0.1.1 /
    v0.1.2 because the autonomy logic isn't affected. Tracking: #483.
  • Adaptive cone DA / cascade hardening. PR #555's signal-side
    cleanup made /odom an honest pose source, but the late-corner
    cone-DA still cascades when the car re-enters a region with cones
    visible from a different angle (autocross is open-line, not a
    closed loop — there is no loop closure to lean on here). Pose
    drifts ~1 m past the 1.0 m Euclidean DA gate, all observations
    look new, cascade-skip-recovery dumps duplicate landmarks into
    the persistent map. Experimental branches were explored this
    release cycle (Mahalanobis DA, cascade percentage-gate
    tightening, per-scan new-landmark rate cap); none reliably
    helped on the validation bag. Real fix likely combines a tighter
    proximity-veto envelope that scales with recent pose uncertainty
    plus a cap on new-landmark commits per scan.