Skip to content

3.1. Handle exceptions

Nathan zhou edited this page Nov 10, 2025 · 10 revisions

Your custom CAN sensor and the roboRIO driver should include robust exception and fault handling to ensure system reliability and prevent undefined behavior.

Exception handling must exist at all levels — MCU firmware, roboRIO driver, and user program.

At the MCU level, handle sensor read errors, I²C/SPI timeouts, and CAN send failures gracefully. Never let a single fault freeze a task; instead, flag data as invalid and continue running.

At the roboRIO driver level, catch all CAN exceptions and handle missing or invalid frames safely. Use flags like isValid to prevent crashes or stuck notifiers, and display clear diagnostics on Dashboard.

At the user program level, always check data validity before using it in control logic. If data is invalid, hold, zero, or disable outputs to keep the robot safe.

Each layer must protect the next — firmware isolates hardware faults, the driver ensures communication safety, and the user program guarantees safe behavior. Together, this ensures the robot stays stable and recoverable under any exception.

Common types of exception

  1. Sensor Offline from MCU or Sensor Exception

    Your sensor should periodically send a status frame to the roboRIO to report the overall health and operational state of its attached components. This frame should be transmitted at a consistent interval (for example, every 100 ms) and include key diagnostic information such as internal sensor status flags, communication errors, temperature (if available), and a simple “alive” flag to confirm activity. By doing so, the roboRIO can continuously monitor the health of the device and detect faults early — such as abnormal reading, I²C read failures, calibration errors, or signal loss.

  2. RoboRIO (Heartbeat) Offline

    The roboRIO usually boots significantly slower than the external microcontroller (MCU). Therefore, upon startup, the sensor or MCU should not immediately assume that the roboRIO is offline simply because the heartbeat frame has not yet been received. RoboRIO offline is different from CAN bus offline. When the roboRIO is offline, the physical CAN network itself usually remains active and functional — other devices on the bus, such as power distribution modules, motor controllers, or additional sensors, may still be transmitting frames normally. In this case, the MCU’s messages will still receive proper acknowledgments (ACKs) from these active devices, confirming that the bus is electrically and logically healthy. Therefore, the MCU should not interpret the absence of roboRIO heartbeat messages as a full CAN bus failure. Instead, it should distinguish between “roboRIO offline” (no heartbeat or control frames from the RIO, but ACKs still present) and “CAN bus offline” (no ACKs, bus errors, or bus-off condition).

  3. CAN Bus Off

    Bus-off usually indicates that the device has become completely detached or electrically isolated from the CAN bus — either due to wiring issues, transceiver faults, or excessive transmission errors that caused the controller to disable itself for protection. When a device enters this bus-off state, it can no longer transmit or receive valid CAN frames, and most importantly, any message it sends will fail to receive an ACK bit from other nodes. In this situation, CAN drivers such as the ESP32 TWAI (Two-Wire Automotive Interface) automatically detect the condition and raise an alert or exception, such as TWAI_ALERT_BUS_OFF or a transmit failure error. Once this occurs, the device must not attempt to continue sending data; instead, it should enter a recovery procedure — typically calling twai_initiate_recovery() — and wait for a successful TWAI_ALERT_BUS_RECOVERED notification before resuming normal operation. This ensures that the bus is stable and that the node reintegrates cleanly into the network without corrupting ongoing communication. Bus-off detection and recovery handling are essential for maintaining CAN network integrity, especially in environments where cable disconnections, power fluctuations, or noise can temporarily isolate a node from the system.

    On the roboRIO side, entering a bus-off state has a different but equally critical impact. When the CAN controller loses synchronization with the network or fails to transmit messages due to missing acknowledgments, any API call that attempts to send a CAN frame will throw an exception—typically UncleanStatusException with the error code -35007 (HAL: CAN Output Buffer Full) or similar. This can occur inside periodic tasks, Notifiers, or command-based subsystems that continuously publish CAN data. If not properly handled, these exceptions can cause the affected thread or notifier to hang indefinitely, preventing it from recovering even after the CAN bus returns to a healthy state.

    To prevent this, all CAN send operations on the roboRIO should be wrapped in a try-catch block, and exceptions should be logged only once rather than repeatedly flooding the console. The driver should then enter a temporary suspend state, skipping further send attempts for a short cooldown period (for example, 500–1000 ms) before retrying. This ensures that the system remains responsive and avoids halting the entire command scheduler or periodic update loop. Once the CAN bus recovers and normal transmission resumes, the driver can automatically clear the fault flag and restore normal behavior. Implementing this retry and suppression mechanism makes the roboRIO-side code far more robust, preventing unrecoverable hangs in long-running control tasks or periodic background threads after transient CAN failures.

  4. User Driver offline

    User driver offline refers to a condition where the sensor has not received any configuration or update frames from the user driver running on the roboRIO — meaning the subsystem responsible for communicating with the device is unavailable, inactive, disabled, or not yet initialized. This state is distinct from both heartbeat loss and full roboRIO offline conditions. In such cases, the roboRIO may still be powered on and transmitting its standard system heartbeat, but the user-space CAN messages (for example, control commands, configuration packets, or data requests) are missing.

    To prevent this condition from being misinterpreted as a sensor fault, the driver should maintain communication with the sensor at all times, not only when the robot is enabled. Even during the disabled state, the driver should continue sending periodic configuration or “keep-alive” messages so that the sensor remains synchronized and responsive. When the MCU detects the absence of user driver frames for an extended period, it should enter a driver-offline mode—the sensor remains operational but stops acting on driver-dependent commands while still transmitting its own periodic status frames. This allows developers to clearly distinguish between a missing driver and a missing device. Once valid user frames are received again, the MCU can automatically clear the driver-offline flag and resume normal operation. This approach maintains safety, avoids unnecessary timeouts, and ensures that the sensor is always ready for immediate operation when the robot is re-enabled.

What type of CAN failure is this?

To reliably distinguish a CAN bus failure from other communication or software-related issues, the system should follow a structured multi-stage diagnostic process. Each stage answers a progressively higher-level question about the health of the communication link and participating devices:

  1. Is the CAN bus off? Is TX working?

    The first and most fundamental check is to verify the electrical and protocol-level status of the CAN bus itself. If the CAN controller (such as TWAI on the ESP32) reports a bus-off condition, it means the node is no longer synchronized with the network or cannot transmit frames successfully. In this state, any TX attempt will fail and no ACKs will be received, confirming a physical or electrical issue (e.g., disconnected wire, shorted line, or faulty transceiver). The firmware should immediately log a “Bus-Off” fault, pause all outgoing transmissions, and attempt a controlled recovery using the driver’s recovery API. Only once the controller reports “Bus Recovered” should normal operation resume.

  2. Am I getting any CAN frame?

    If the bus is not off and transmission appears normal, the next check is to determine whether any frames at all are being received from other devices. If no frames are detected for a prolonged period (for example, 1–2 seconds), it could indicate that the device is isolated or that all other nodes are powered off. However, if transmissions still receive ACKs, this might simply mean there are no active talkers rather than a failure. A healthy CAN network should show at least some traffic from other devices like PDH, PDP, or motor controllers. Lack of any reception while still getting ACKs suggests the network is quiet but functional; lack of ACKs indicates a deeper bus issue.

  3. Am I getting the roboRIO heartbeat?

    Once basic CAN traffic is confirmed, the next layer of verification is to look for the roboRIO heartbeat frame, which indicates that the control system is active and online. The heartbeat typically updates at a fixed interval (around 10–20 ms), and its absence for more than 500–1000 ms signals that the roboRIO may be rebooting or powered down. However, this does not necessarily mean the CAN bus is offline — only that the control node (roboRIO) is temporarily not broadcasting.

  4. Am I getting user driver commands?

    Finally, the system verifies whether it is receiving user-space driver frames — configuration or control updates sent from the custom FRC robot code on the roboRIO. If these frames are missing while the heartbeat is still present, it indicates that the driver or subsystem code is inactive — for example, the robot program hasn’t been deployed, the subsystem isn’t scheduled, or the robot is disabled. In this situation, the CAN bus and heartbeat remain healthy, but the control logic itself is idle. To prevent this from being misinterpreted as a sensor fault, the driver should maintain communication with the sensor at all times, not just when the robot is enabled. This means continuing to send periodic configuration or “keep-alive” messages even while disabled, ensuring the sensor stays synchronized, avoids unnecessary timeouts, and remains ready for immediate operation when the robot is re-enabledBy evaluating these four levels sequentially, the sensor can precisely identify the cause of communication loss:

  • Level 1 (Bus Off): Physical/electrical CAN fault.
  • Level 2 (No Frames): Network silent or isolated node.
  • Level 3 (No Heartbeat): RoboRIO offline or rebooting.
  • Level 4 (No User Commands): Driver inactive or not called fast enough.

This layered diagnostic approach ensures accurate fault isolation, enabling clear visual or log-based feedback for each condition (e.g., different LED patterns for bus-off, RIO offline, or driver inactive) and supporting automatic recovery behavior once communication resumes.

What to do during an exception?

The most important action during any exception is to ensure that the robot remains safe under all conditions. Inform the user program of the presence of a fault is a must have for all CAN sensor driver. The system must never allow invalid or corrupted data to propagate into control logic, as this could result in unpredictable behavior or unsafe motion. When an exception occurs—such as sensor disconnection, CAN bus-off, or data timeout—the firmware should immediately set a clear “isValid = false” or “faultActive = true” flag in its next CAN status frame, allowing the roboRIO to recognize that the data cannot be trusted. On the roboRIO side, the user program should then treat all sensor readings marked as invalid as untrustworthy. Depending on the use case, it can take one of several safe actions:

  • Reset values to zero (for example, current, velocity, or torque).
  • Hold the last known good value temporarily if a smooth transition is needed.
  • Disable outputs or commands if the sensor is critical for feedback control.
  • Alert the driver/operator via Dashboard or driver station log.

Equally important is ensuring that the software never crashes or hangs when exceptions occur. All critical sections—especially CAN transmissions, I2C reads, or data parsing routines—should be wrapped in try-catch blocks or equivalent error-handling logic. The system must continue running background tasks like LED indicators, fault reporting, or retry timers, even when one component fails.

In addition to keeping the robot safe, the firmware should aim to recover from faults as quickly as possible. This involves automatically reinitializing failed interfaces, reattempting CAN recovery after bus-off, or resuming data acquisition once communication is restored. The recovery should be handled gracefully, without requiring a full reboot of the device or the robot code. Once a fault clears and normal data flow resumes, the isValid flag can be restored, and normal operation should continue seamlessly.

This approach—detect, report, protect, and recover—ensures that even during exceptions, the robot behaves predictably, remains safe, and restores functionality autonomously as soon as the fault condition ends.

Advanced Fault Indication System

In many cases, a simple “sensor good” or “sensor bad” flag is not enough to represent the full range of possible issues a device might experience. Since CAN does not provide traditional serial logging, a more advanced fault indication system may be required to communicate detailed diagnostic information to the roboRIO driver and user program.

For example, a magnetic encoder could encounter multiple fault conditions — such as sensor offline, magnet too weak, magnet too strong, or magnet misaligned. Each of these represents a different failure mode and may require different responses from the robot software. To handle this, a structured fault protocol should be implemented, where each fault type is assigned a specific bit or code within the CAN status frame. This allows the driver to decode and display precise fault information, improving both debugging and runtime safety.

In addition to real-time fault flags, the system should support sticky faults whenever necessary, which record whether a specific fault has occurred at any time since boot, even if it later cleared. Sticky faults are especially useful for diagnosing intermittent issues that may not persist long enough to be observed during normal operation. They should be stored and tracked on the MCU side, since CAN packets can be lost and the sensor cannot rely on the roboRIO to maintain persistent fault states. These sticky faults should also be clearable via CAN command from the driver, allowing operators or code to reset the history once the fault has been reviewed.

This kind of structured, multi-level fault reporting system gives both firmware and robot software more insight into device health and helps isolate subtle or transient issues that would otherwise be invisible over the standard CAN telemetry interface.

Use a Watchdog for Sensor Failure Recovery

A watchdog can serve as an effective safeguard against sensor-level failures. In this case, the watchdog’s role is to detect when the sensor itself becomes unresponsive, stuck, or returns invalid data for an extended period — for example, when an I²C device stops responding, a timing loop freezes, or a peripheral driver fails silently.

The roboRIO can implement a software watchdog that monitors incoming sensor data fields (such as voltage, color, distance, or temperature). If the values remain unchanged for longer than expected or the isValid flag stays false for several cycles, the driver can issue a remote reboot or reset command to the sensor through a dedicated CAN API ID. This allows the robot to recover the peripheral automatically without requiring a manual power cycle or full robot reboot.

Alternatively on the MCU side, a local watchdog FreeRTOS task should perform similar supervision internally. It can periodically check whether the sensor reading task is still updating data, whether I²C transactions are succeeding, and whether status flags look healthy. If the MCU detects that the sensor is no longer producing valid readings or its internal driver has locked up, it can safely reset the affected subsystem or even perform a full software restart if recovery fails.

However, this must be implemented carefully. Rebooting the MCU or resetting the sensor too frequently can cause unnecessary downtime, confuse diagnostics, and occasionally disrupt CAN bus timing.

Clone this wiki locally