Repository navigation
RFC: Support for OCI backends, especially the HaloGen, ROCmFPX, and NathanW containers from kyuz0 (Donato) #3585
Replies: 20 comments 54 replies
|
BTW rocmfpx is our number 1 most upvoted feature request, with 15 votes: #2089 |
|
I love this idea! so many great projects |
|
@jeremyfowers are we proposing OCI over nono or whatever the other one was that @abn had suggested before? Is there a reason for that also? I feel like expecting Podman / docker might be a hard pill to swallow. Perhaps maybe we do nono as fallback if no Podman / docker? |
+1 on the standard container interface. |
|
this fork of llamacpp is the combined effort of some of the popular strix-halo llamacpp devs: NathanW, Laurent (AgentionAI), rocmfpx, and others: https://github.com/halo-box/strix-llama.cpp it supersedes them all (though not everything from the other forks is merged yet) because this is all the fork devs essentially combining all the forks into one central strix halo-focused llamacpp fork so, please add support for it! |
|
Loving the interest in broader backend support but for users already running lemonade in a container, how is this expected to work? Passing in the docker socket? |
|
How will this impact Lemonade on Windows? WSL OCI container support seems to be a future thing and the referenced toolbox containers aren't windows compatible as far as I know. |
|
@jeremyfowers the RFC doesn't say, and I think it needs to, about Windows support for custom backends like this. Podman and Docker means WSL which is a whole new topic We can also just call it out in this RFC as Linux only feature |
So I don't really want to pigeon-hole us that OCI containers are only as engines/backends. We could totally become an orchestrator for containers too. Think like managing a Hermes OCI container or Claw OCI container. We then become the glue that "links" the OCI containers for backends to OCI containers for frontends. I don't think it necessarily needs to be done immediately as it's a big scope creep from the original RFC, but I want to make sure this design doesn't preclude it and make it more difficult to accomplish. |
I don't see anything in this design that talks about the security of these containers. Here's the areas that I think need to be discussed.
|
|
@kenvandine can snaps start and stop OCI containers? What runtime can be used? Or this is going to be exclusively to native packages? |
|
Disclosure: this comment was written by Nick's AI assistant. Nick (halogen's author) asked me to review the RFC and the thread and reply with comments; the facts below come from halogen's repositories and release records, and Nick reviewed and approved this before posting. Commitments in it ("will", "happy to") are his. Thanks for the RFC. On the ordering: the license is the gate, and Nick is moving it. A few notes that may make halogen the easy integration rather than the last one, since it already ships as a container and most of what @superm1 listed is how it works today. What the images already are
One launch requirement that has to be declarable ROCm's HSA runtime, inside a container, needs The container interface standard: halogen can be the first image to carry it Proposal: the declaration lives on the image as OCI labels, so Lemonade reads it before starting anything, e.g. with the same models list served on an endpoint once running, for the auto-update idea in the maintenance plan. That covers @iswaryaalex's least-privilege launch and the "read supported models from the container" idea in one place. Nick is happy to implement it first and adjust to whatever shape the maintainers settle on. Source of truth A strong preference: Lemonade pulls the vendor's signed image by digest and @kyuz0 vets it, rather than rebuilding it inside a toolbox. Every published halogen number is per image; a rebuild with a different ROCm is a different engine (ROCm 10 alone moves prefill about +3% and is not bitwise with 7.14). Vetting a pinned digest is the stronger guarantee, it matches the pinning requirement above, and it scales to the other engines in this RFC. License Apache 2.0, both engines, weeks rather than months. Nick is also offering to help on the DS4 conversion PR, since that builds the plumbing this rides on. Linux-only is fine; halogen targets gfx1151 by construction. |
|
First, I want to let everyone know this project, and especially this RFC, is important to me. Me showing up with an AI generated post is a bad first impression. I will be more sensitive to the preference of the maintainers if they do not want AI generated comments in the future. The way I work now will never be the same and I will be respectful of the rules. This is the human version of Nick, not Son of Anton. I have a couple thoughts. I am a massive fan of simplicity and removing technical barriers for users. Donato's toolboxes made it very easy for me to get an environment setup on my Strix Halo when I first started out, so if someone could try out halogen simply by running Halogen (flash) can be picky about memory as it pins 68GB in a continuous block, then eats another 35 in KV cache. It's a memory hog at shipped defaults, regardless of what Also, Donato's toolboxes as a single trusted source makes sense, so please disregard the suggestion above... and agree on the container registry or interface long term. Regarding security, an issue came up when Halogen was first published because I had 0 trust and a docker-compose file that was too permissive. Well that and compiled binaries. I'm not an expert in security but containers are inherently more secure and principle of least privilege should apply. I can only attest that --ipc=host is needed to run the Halogen container in podman on Fedora. If there's a more secure way to access to the host, please let me know. |
|
In regards to testing the backends itself (not just the integration with lemonade) is it enough to rely only on external verification? Given the hardware we have in the dev lab we could either enable testing of the container backends itself or if it help we can onboard @kyuz0 to use the resources in our dev lab for CI. |
|
Late to the party, my thoughts:
|
|
Okay, I'm pretty late to the party given that I've poked @jeremyfowers to move on with this, but here's how I imagine this works. Since with containers, you can have a standard HTTP port where the API lives and you can redirect it to any port you want, this will not be a problem. As for the protocol, we're probably aiming for OpenAI-compatible where applicable and proposing our own protocol where there isn't an industry standard (3D-gen probably?). The only thing this leaves is a format for the containers to describe their configuration. I assume this means each container is shipped with some sort of metadata that describes the input variables, preferrably with some standarization so that things with similar meanings for different engines (such as backend choice, eg. Vulkan vs ROCm) are passed on the same way. And in line with the security discussion, I do think it should be maximum isolation, i.e. the container has no access to the parent filesystem except for mounting the single folder that contains the model images (HF cache?). |
|
I support this proposal very strongly. Please add https://github.com/gufo-org/gufo to the list (near the top). Leading performance and open source. Also please ensure compatibility with distributed inference setup (tp=2 w/ RDMA) |
|
this proposal would make my life much easier |
|
Can we document an exit plan for experimental engines? For example what gets an engine promoted, demoted or removed. |
|
Onwards to implementation! |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Proposal size
Major feature (spans multiple well-scoped PRs)
Updates
User story
Engine innovation is happening fast on Strix Halo, with the emergence of llama.cpp forks like DwarfStar4, ROCmFPX, NathanW, along with new engines like Halogen. Users want to try out new optimizations before they land in mainline core engines like llamacpp. Devs also need a way to do apples-to-apples benchmarking and agentic tests across engines/backends.
The core proposal here is to enable OCI containers on Linux to serve as Lemonade engines/backends, and then to leverage this capability to support new experimental engines in a way that is both agile and secure. Future work may include defining a standard interface for containers that makes these Lemonade integrations even easier in the future.
User stories enabled:
lemonade install halogen:rocmlemonade pull Qwen3.8-27B-ROCmFP4lemonade bench Qwen3.8-27B-GGUF Qwen3.8-27B-ROCmFP4lemonade launch pi -m MODEL --agent-args="-p TEST PROMPT"across multiple core and experimental backends.High-level design
Container source
There are two major problems with adding any new experimental backend to Lemonade today:
This proposal attempts to address both by leveraging @kyuz0's work. He already maintains a Docker Hub with curated containers for popular experimental backends, including a lot of backends that Lemonade doesn't support yet. He is also a security professional and pulling a backend from him is a lot more credible than getting it from a relatively unknown GitHub.
OCI containers as backends
All of our WrappedServers today assume that they are launching native processes. New features:
New Engines
I am proposing to add the following new engines. The first PR will introduce both an engine and the general OCI support, while subsequent PRs should just incremental add 1 engine per PR.
My proposed order:
Breaking changes
Support for our native DS4 binaries is replaced by the container, so we will need to auto-migrate people (help them delete their native installs).
Maintenance plan
This work will include automated GitHub actions that pull new container versions and curated models from @kyuz0. He has already set up actions that automatically update his containers when the upstream backends have releases. Our CI will validate that container-lemonade integration still works, so we are expecting minimal manual maintenance.
In the future, it would be great to define a container interface standard that would make maintenance even easier. For example, if containers support an endpoint that lists their supported/suggested models, then Lemonade could read that to auto-update a supported models list.
Risks
The security risks of supporting experimental projects should be mitigated by containers and by @kyuz0's expertise.
All reactions