Repository navigation
ZMF5 Safety and Code Quality
← Agenda · ZMF#5 · 2026-10-06 · Part I: 9:00 AM - 10:45 AM · Part II: 11:00 AM - 12:45 PM
- Tobias K, Inovex (Zephyr Safety Architect, lead host)
- Kate Stewart, Linux Foundation (co-host)
- Alberto Escolar Piedras, Nordic (co-host)
- Anas Nashif, Intel (co-host)
- Pieter De Gendt (co-host)
- Nicole Pappler, Alektometix (Safety Manager)
- Hanuobu Kurokawa, Renesas
- Alexandre Bailon, Baylibre
Work towards clarity on what our initial configurations will be for the safety scope. Understanding of the end-to-end analysis flow - clear understanding of what is in upstream, what will the safety artifacts consist of
Materials shown during the session Anas' slides
By the end of this session, participants should have:
-
Next Steps with Assignee
- [TBD] Enable Clang23 in CI in addition to gcc to take advantage of build checking emerging
- [TBD] Clang 19 --> Clang 23 port
- [Tobias] Submit RFC based on scope defined to enable/disable features possible in Zephyr
- [Tobias] Document Assumptions of Use associated with scope defined
- [Safety WG] Work with maintainers of code in enabled scope to sign off on functional requirements explicitly
- [TBD] reassess Kernel - Architecture APIs with goal of minimize surface to simplify certification and architecture specific code.
- [Safety WG to drive, but include call for participation from others] Pragmatically audit current set of coding guidelines and determine set we'll work with; then enable CI to add in these rules, and then increase over time.
-
Questions to answer
- Where in the project to store integration assumptions of use.
- Recommendation on how all the tools to be used in safety analysis can be qualified.
-
Identified:
- [Open issue requiring follow-up]
-
Assigned:
- [Owner and next action]
Avoid outcomes such as “discuss the issue” or “raise awareness.” Describe what should be different after the session.
In order to achieve safety certification the projects needs to demonstrate what's called systematic capability according to IEC 61508. In essence this means that the projects needs to show that required processes such as requirements capture, design verification or safety analysis are implemented and executed. To this end our Safety Chair Stephane Parenti (@parphane), Wind River has recently published https://github.com/zephyrproject-rtos/zephyr/pull/120739 to outline the comprehensive set of required processes.
As per the current project's work practices there are (unsurprisingly) a couple of open gaps and questions how we should implement these processes in the least friction-causing way.
The open problems that we want to discuss in this session:
- Anas to show current status.
- Mechanisms being used
- Translation to SPDX use with build SBOMs.
- Detection
- Bug Fixes
Companies need to include more than (just) the kernel in their products.
Guidelines would help to explain how the safety process could/should be rolled out to other areas of the Zephyr project:
- Kernel
- Drivers
- Network stacks
- Subsystems
- FS
- Settings
- Arch / Processor
- Libraries
- Rust, etc.
- Compiler
Q1: How to extend to Drivers, SoC Code, Network Stacks, USB etc.
- Getting contributions of analysis - back to extend our scope?
- Maintainers sign off on requirements to be "follow our methodology" & tests contributed to prove it out.
- Paying others to extend beyond and contribute to upstream?
Q2: When companies are willing to perform such activities downstream, how could they upstream and what acceptance criterias have to be met?
Do we have what we need?
Many product makers need to continue fixing/supporting their SW over much longer than any release (LTS) lifecycle. With product lifetimes > 10-20 years, both backporting and jumping to a new release must be much simpler.
Many product vendors also have a need to start developing on very new SoCs (so the HW is supported for long enough). (This new HW will/can not be supported in the previous LTS) => It must be easier for a vendor to start working from main/a non-LTS release and update Zephyr versions as they develop their product, and up to an LTS which properly supports that SoC.
What is the right way to find, document and verify the Kconfig configuration sub-space which is safe.
Background: While Kconfig is one of the key strengths of Zephyr, it also spans a huge space of possible configurations that are impossible to verfiy one by one. The Kernel alone is configured by approximately 200 Symbols (CONFIG_USERSPACE, CONFIG_COOP_ENABLED, ...).
Q1: Are the current Kconfig default values the correct ones for safe use?
Q2: What Kconfig symbols can be altered to remain in the safe configuration sub-space?
Q3: What mechanism could be implemented to detect when a non-safe configuration is built?
Q4: Do we have to add additional scenarios to the kernel tests to better sample the safe configuration space during verification?
Q5: Can we build on top of the hardening yaml effort from https://github.com/zephyrproject-rtos/zephyr/pull/116673 (related to Q1/Q2)
In order to allow product-makers the correct integration of Zephyr into their safety-critical products, the project is required to document the assumptions that the Zephyr code expects from its runtime-environment.
Assumptions of use can be put into three categories:
- What pre-requisites have to be met?
- How must Zephyr be configured itself for safe use (see problem 1)
- What must the user know to correctly make use of the Zephyr API
This leads to the following question:
Q1: Where and how do we store Zephyr's assumptions of use?
- One example, where things can go wrong today: Semaphores must not be initialized more than once.
Is the reqmgmt repository a possible home for this? https://github.com/zephyrproject-rtos/reqmgmt
A safety analysis is a risk based approach to systematically identify and remedy safety risks associated with the usage of a SW or HW component. For this analysis typically safety engineers would need to work together with the subject matter experts (maintainers or code owners in our case). For instance freedom of interference is always one of the main topics to be considered.
Q1: How well can freedom of interference be realized in today's design - Isolation (time, memory) of certain component (driver, subsystem, kernel)
Q2: To what extent can https://github.com/zephyrproject-rtos/zephyr/pull/118015 help? Q3: How can/should maintainers be included in the safety analysis?
To demonstrate systematic capability we also have to qualify and validate the tools that a required to build Software based on Zephyr. As tools the following qualify
- Compilers and Linkers
- Scripts that generate code during build time (e.g. edt_lib.py)
- Scripts that help create the build configuration of the Software (e.g. Kconfig)
- Build system files (e.g. CMakeLists.txt)
Check this link for the complete "set of tools" that we depend on.
Q1: What measures exist today to uniquely identify the tools we are using?
- Name, Version, Vendor, Usage Manual, known bugs, (Normal SBOM + Requirements + Test SPec & Results) (cf. ISO 26262)
- Level of dependency
Q2: What measures exist today to validate that our tools work as expected?
- Specification of what the tool is supposed to do? Inputs and Outputs
- Test suites
Q3: Is our current description of the build process accurate?
- https://docs.zephyrproject.org/latest/build/index.html
- Uniprocessors
- Supervisor Mode
- Simple Scheduling
- User Space API
- Drivers
- Subsystems
- HW Architecture / SoC
- Compilesrs
- Libraries
- Logging
Frame the session around a limited number of explicit questions.
Decision required: [Specific question participants must answer]
Possible outcomes:
- [Option]
- [Option]
- [Option]
Decision required: [Specific question]
Possible outcomes:
- [Option]
- [Option]
- [Option]
Repeat for no more than five or six primary decisions.
Not too strict agenda, but we need to manage the discussion and time box it, i.e. for a 2 hour session
- Present the problem / proposal (~20-30 minutes)
- Discuss (1 hours)
- Closing (30 minutes)
Before the event, the session owner should:
- Complete Sections above.
- Provide links to relevant issues, pull requests, and documentation.
- Identify two or three representative examples.
- Describe realistic options rather than only the preferred solution.
- Highlight decisions that may require TSC approval.
Participants should:
- Read the document before the event.
- Add missing evidence or examples.
- Comment on the proposed scope.
- Identify constraints that may have been overlooked.
- Indicate strong objections before the session where possible.