Skip to content

ZMF1 Hardware Testing

Dan Kalowsky edited this page Oct 6, 2026 · 15 revisions

Hardware Testing

← Agenda · ZMF#1 · 2026-10-06 · 1:45 PM - 3:30 PM

Key Speakers, Leads, and Required Attendees

  • Dan Kalowsky, Qualcomm (Lead Organizer & Moderator)
  • Parthiban, Linumiz (co-organizer)

Decision or Outcome Sought

Decide whether Zephyr wants hardware in the loop testing in CI, and if yes, agree on board tiers, where boards are hosted and how results reach PRs.

By the end of this session, participants should have:

  • Agreed on:
    • How tiered board support will be implemented
    • Where boards used for HIL will be housed
    • Where HIL tests will be utilized in the codebase
    • If D1 is No, how does the Zephyr project leverage the information from hardware vendors already collected for making PR and release determinations?
  • Identified:
    • How shall the project handle board failures that are not included in the HIL?
    • What additional features shall be investigated?
      • Remote debugging
      • Snapshot of firmware+logs to replicate failures
      • Remote flashing
      • Support for twister fixtures
      • LLM support
    • What challenges exist around long term usage/running?
      • Board failures
        • Flash wear
      • Point of Contact
      • Priority of fixing on-site issues
      • Telling infra failures apart from real test failures
      • Keeping boards alive for an LTS lifetime
      • Tool and probe firmware versions across labs
      • Long term uptime testing (Is device alive after 90 days of uptime? Still functioning?)
      • ???
    • What is the effort required to implement?
    • [Open issue requiring follow-up]
  • Assigned:
    • [Owner and next action]

Problem Statement

At over 1000 officially supported boards, the Zephyr project has grown to the point beyond any developer being able to test all board configurations for a change. Currently, the Zephyr CI system depends heavily upon emulators (QEMU, FVP, etc) which have served well. Yet these solutions are proving not always suitable for testing entire classes of boards, devices, and code changes that Zephyr now encompasses. Pull Requests that refactor code or add new features are passing our CI, only to be (sometimes weeks) later reverted due to the discovery of an edge case on a silicon platform. As the Zephyr project continues to grow, ensuring the stability of the code and signaling more confidence to PR submitters and Release team members on changes will be essential.

Example of the issue:

Current state, as far as I can tell:

Scope

In scope

  • Which twister fixtures would benefit most from hardware testing.
  • Which sections of Zephyr code would benefit most from hardware testing.
  • Updating current infrastructure and designs to support boards.

Out of scope

  • Addressing PR and code review short falls. There are opportunities to improve both, which are open for discussion in the Process Working Group.
  • Simplifying processes, such as SDK updates, release, PR notifications, etc. There are opportunities to improve here that should be brought before the Process Working Group.
  • Budgets - from emulation resources to infrastructure. While important, let us first focus on our ideal scenario then work on reducing to what we can accomplish.
  • Negative commentary on vendors participation.
  • Removing the use of twister.

Questions Requiring Agreement

Decisions likely needing TSC approval: D2 (revisits the 2022 TSC vote), D4 (tree and test.yml changes), D5 and D6 (project hosted infra).

D1: Hardware Based Testing

Decision required: Does the project still desire hardware in the loop testing in CI/CD?

Some background:

Possible outcomes:

  • Yes
  • No

D2: Tier Revival

Decision required: Does the project still desire to use the previously proposed hardware tier system in CI/CD?

Some Background:

  • The Process Working Group began exploring a priority tiered hardware system in 2021 (https://github.com/zephyrproject-rtos/zephyr/issues/38566), with the goal of enabling Tier 1 boards as a hardware run requirement.
  • The Process Working Group concept was brought to a TSC vote in September of 2022 (https://lists.zephyrproject.org/g/tsc/message/1464) and passed on a vote. Unfortunately there appears to be limited further efforts.
    • Last mention in Testing WG was 05/08/2025 for a "light Hardware-in-the-loop" testing, which does not include many details.

Possible outcomes:

  • Yes
  • No

D3: Tier Expectations

Decision required: What roles do each tier encompass?

Outcome Matrix: (based upon previous Process Working Group discussions)

Tier PR Nightly / Weekly RC and Release No support?
Tier 1 X X X
Tier 2 X X
Tier 3 X
Tier 4 X

Note: the D4 examples use tier_0 for native_sim, which is not in this matrix. Do we need a Tier 0 for emulators/native_sim?

D4: Identification of Tiers in Tree

Decision Required: How would board tier inclusion be shown in the code?

How do we define a discoverable means to which boards are available on a per tier basis? Does this need to be something that is integratable with twister, or should board tier distinction live outside and feed into driving twister?

By adding a new YAML file in tests? For example tests/tier_support.yaml

tier_board_support:
  tier_0:
    - native_sim
    - native_sim/native/64
  tier_1:
    - arduino_uno_r4@minima/r7fa4m1ab3cfm
  tier_2:
    - arduino_uno_r4@wifi/r7fa4m1ab3cfm

By adding a new mapping to the boards/**/boards.yml?

For example:

  name: rpi_pico2
  full_name: Raspberry Pi Pico 2
  vendor: raspberrypi
+  tier: 2
  socs:
    - name: rp2350a
      variants:
        - name: w
          cpucluster: m33_0
          variants:
            - name: mcuboot
              cpucluster: m33_0
        - name: mcuboot
          cpucluster: m33_0

Other options.. ?

Decision Required: Do tests need to be associated to a tier? If yes, how?

By adding a mapping to tests/**/test.yml? For example:

common:
  tags:
    - filesystem
    - littlefs
+    - tier_2
  platform_allow:
    - nrf52840dk/nrf52840
    - native_sim
    - native_sim/native/64
    - mr_canhubk3

By adding a board mappings to tests/**/test.yml? For example:

common:
  tags:
    - filesystem
    - littlefs
+  tier:
+    platform_allow_0:
+      - native_sim
+      - native_sim/native/64
+    platform_allow_1:
+      - nrf52840dk/nrf52840
+    platform_allow_2:
+      - mr_canhubk3
common:
  tags:
    - filesystem
    - littlefs
+  platform_allow:
+    tier_0:
+      - native_sim
+      - native_sim/native/64
+    tier_1:
+      - nrf52840dk/nrf52840
+    tier_2:
+      - mr_canhubk3

Other methods?

D5: Board Hosting

Decision required: Where shall boards for the HIL CI/CD be hosted?

Where the test boards are hosted introduces several challenges. For example, ensuring uptime of scheduler host and workers, regular upgrade/replacement of boards, and consistent network and power speeds to hosts all need to be considered.

What about boards that need to have rework to support HIL testing?

Usage in various scenarios also introduces engagement requirements. For example, board support testing for an LTS must encompass the lifetime of the LTS.

Outcomes Matrix:

Location Tier 1 Tier 2 Tier 3 No
Project Member
Vendors
Zephyr Project
Individuals

D6: Scheduler Framework

Decision required: Does the project care to have a scheduler framework for CI/CD?

In broad strokes, the process for submitting to a HIL process will likely take on the form of:

  1. PR submission
  2. Zephyr CI generates a test plan for the PR and platform through scripts/ci/test_plan_v2.py.
  3. Twister builds the test artifacts for the board.
  4. The twister built artifacts are submitted for testing to a test scheduler.
  5. Test Scheduler distributes the job to a worker with the correct devices attached to it.
  6. Test worker runs twister with the pre-built artifacts.
  7. The PR is gated until the HIL test is completed and results collected.

Three options exist for following this pattern.

Option 1 - the Zephyr project decides upon a framework and hosts a scheduler that distributes jobs to worker clients connected to it. An alphabetical list of (currently known) possible frameworks:

  • LabGrid
  • Latchport
  • Linaro LAVA

There are talks at this conference (example: https://osselceu2026.sched.com/event/2RaYr/demystifying-board-farms-build-a-mini-desktop-board-farm-with-standard-tools-and-hardware-francesco-cervigni-neoncomputing?iframe=no&w=&sidebar=yes&bg=no) giving more detail about all of these options. Please attend one if you’re curious about details.

Advantages:

  • Clear way to document and ramp up new workers as needed.
  • The Zephyr project can easily disable/enable boards when needed.
  • Clear performance monitoring knobs.
  • Upgrades to workflow are easier.

Disadvantages:

  • Sites with already established HIL infrastructure may not want to convert/support any alternative framework.
  • Another service that Zephyr infrastructure must host.
  • May require vendors to have two HIL systems, external public and internal private.
  • Zephyr CI may still be in the path and very busy.

Option 2 - The Zephyr project stays framework agnostic and only focuses on the resulting twister run outputs.

Golioth has shown it is possible 3 years ago:

Advantages:

  • Already established HIL infrastructures can be used provided all accept a GitHub API based submission.
  • The Zephyr CI is already good at parsing twister output.
  • Can reduce the Zephyr CI load by having build on downstream consumers.

Disadvantages:

  • Limited to no insight into infrastructure failures.
  • Will require coordination on GitHub access tokens to submit work.
  • Difficult to update workflows.

Option 3 - Hybrid. The Zephyr project defines a common result format and submission API, and provides a reference scheduler setup that labs may use or ignore.

Advantages:

  • Existing HIL setups keep working, they only need to report in the common format.
  • New labs have a ready setup to start with.
  • Unifies the HIL reporting format already seen in PRs.

Disadvantages:

  • Two paths to maintain, the reference setup and the API.
  • Still limited insight into infrastructure failures for labs not using the reference setup.

Proposed Session Agenda

Not too strict agenda, but we need to manage the discussion and time box it, i.e. for a 2 hour session

  • Present the problem / proposal (~20-30 minutes)
  • Discuss (1 hours)
  • Closing (30 minutes)

Pre-Session Preparation

Session owner

Before the event, the session owner should:

  • Complete Sections above.
  • Provide links to relevant issues, pull requests, and documentation.
  • Identify two or three representative examples.
  • Describe realistic options rather than only the preferred solution.
  • Highlight decisions that may require TSC approval.

Participants

Participants should:

  • Read the document before the event.
  • Add missing evidence or examples.
  • Comment on the proposed scope.
  • Identify constraints that may have been overlooked.
  • Indicate strong objections before the session where possible.