Skip to content

fix(backends): pin grpcio-tools to installed grpcio in runProtogen (protobuf gencode mismatch)#10735

Merged
mudler merged 1 commit into
masterfrom
fix/protogen-grpcio-tools-pin
Jul 8, 2026
Merged

fix(backends): pin grpcio-tools to installed grpcio in runProtogen (protobuf gencode mismatch)#10735
mudler merged 1 commit into
masterfrom
fix/protogen-grpcio-tools-pin

Conversation

@localai-bot

Copy link
Copy Markdown
Collaborator

Problem

The vLLM backend (all accelerations) fails to start with:

google.protobuf.runtime_version.VersionError: Detected incompatible Protobuf
Gencode/Runtime versions when loading backend.proto: gencode 7.35.0 runtime 6.33.6.
Runtime version cannot be older than the linked gencode version.

The backend crashes on import backend_pb2 before it can serve, surfacing to the user as failed to load model with internal loader: grpc service not ready.

Reported as a ROCm / R9700 gfx1201 failure (#10718), but it is not GPU-specific — the process never reaches any GPU code. It affects every vLLM variant (and any Python backend that caps the protobuf runtime below the latest grpcio-tools gencode).

Root cause

runProtogen (backend/python/common/libbackend.sh) runs after the backend's requirements are installed, and installs grpcio-tools unpinned:

uv pip install grpcio-tools   # pulls the newest release

The protoc bundled in the newest grpcio-tools stamps backend_pb2.py with Protobuf gencode 7.35.0. vLLM's dependencies pin the protobuf runtime down to 6.33.6. Protobuf enforces runtime >= gencode at import, so the generated module refuses to load.

Fix

Pin grpcio-tools to the grpcio version the backend already installed — the two are released in lockstep, so the bundled protoc's gencode stays in step with the protobuf runtime. Falls back to unpinned when grpcio isn't present (behaviour unchanged for those backends).

grpcio_version="$(python -c 'import importlib.metadata as m; print(m.version("grpcio"))' 2>/dev/null || true)"
[ -n "$grpcio_version" ] && grpcio_tools_spec="grpcio-tools==${grpcio_version}"

Testing

  • bash -n clean; shellcheck -S warning reports nothing new on the changed lines (pre-existing findings elsewhere untouched).
  • Fallback path verified (grpcio absent → unpinned, as before).
  • Full runtime validation requires a backend image build (ROCm/vLLM) — CI territory; cannot be reproduced locally.

Closes #10718

🤖 Generated with Claude Code

runProtogen installed grpcio-tools unpinned, so the protoc it bundles
stamped backend_pb2.py with the newest Protobuf gencode (7.35.0). When a
backend caps the protobuf runtime lower -- vLLM pins protobuf to 6.33.6 --
the import-time guarantee runtime >= gencode fails:

  google.protobuf.runtime_version.VersionError: Detected incompatible
  Protobuf Gencode/Runtime versions ... gencode 7.35.0 runtime 6.33.6

The backend crashes on `import backend_pb2` before it can serve, which
surfaces to the user as "grpc service not ready". It was mis-reported as a
ROCm/gfx1201 failure in #10718 but is not GPU-specific and affects every
vLLM variant (and any backend that caps protobuf below the latest gencode).

Pin grpcio-tools to the grpcio version the backend already installed --
they release in lockstep -- so the generated gencode stays in step with
the protobuf runtime. Falls back to unpinned when grpcio isn't present.

Closes #10718

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
@mudler
mudler enabled auto-merge (squash) July 7, 2026 21:29
@mudler
mudler disabled auto-merge July 8, 2026 13:24
@mudler
mudler merged commit d5d659b into master Jul 8, 2026
65 of 66 checks passed
@mudler
mudler deleted the fix/protogen-grpcio-tools-pin branch July 8, 2026 13:24
@aTosser

aTosser commented Jul 18, 2026

Copy link
Copy Markdown

#10718 is not resolved. There exists still the protobuf gencode/runtime mismatch error in both rocm-vllm and rocm-vllm-development still. grpcio-tools is not pinned to 1.78.0 for some reason. This patch seems ineffective.

@Jpc-Automation

Jpc-Automation commented Jul 22, 2026

Copy link
Copy Markdown

#10718 is not resolved. There exists still the protobuf gencode/runtime mismatch error in both rocm-vllm and rocm-vllm-development still. grpcio-tools is not pinned to 1.78.0 for some reason. This patch seems ineffective.

I can confirm the issue still persist in exactly the same way.
#10718
#10642
#10940

Is there a way to set the grpc version without compiling the whole project/backend?

@mudler should we reopen the issue?

best regards

@localai-bot

Copy link
Copy Markdown
Collaborator Author

@Jpc-Automation thanks for the ping, and for spotting that grpcio-tools was not landing on 1.78.0 — that observation was the thread that unravelled it.

Answering the direct question first: no need to reopen. There is no defect left in master. The two backends you named are failing for two different reasons:

  • rocm-vllm-development (tracks master): already fixed by fix(vllm): generate protobuf 6 compatible stubs #10944. I read backend_pb2.py straight out of the published master-gpu-rocm-hipblas-vllm image (built 2026-07-21) and it stamps gencode 6.31.1, well under the 6.33.6 runtime. It cannot produce the 7.35.0 error. If you are still seeing it, the backend installed under BackendsPath is older than that image — LocalAI installs a backend once and a restart does not re-pull it. Delete the backend and reinstall it.
  • rocm-vllm (tracks latest): genuinely still broken. That image is from 2026-07-14 and predates both fixes, so nothing local will help; it needs the next release build.

On why this PR's pin was ineffective: it pinned grpcio-tools to the installed grpcio version, but grpcio-tools' version tracks grpcio while the gencode its protoc emits tracks protobuf, and the two move independently. Since requirements.txt pins grpcio==1.82.1, the pin resolved to grpcio-tools==1.82.1, which requires protobuf>=7.35.1 and stamps gencode 7.35.0 — precisely the broken generator. #10944 then fixed it by hardcoding 1.78.0.

I have opened #11057 to make that robust rather than lucky: it derives the generator from the installed protobuf runtime (so it cannot rot as protobuf advances), regenerates the stubs after the late vllm install, clears __pycache__ (a stale .pyc can shadow a regenerated stub, since CPython validates by mtime and size and the gencode triple is the same byte width), and fails the build if the generated stub cannot be imported.

Full detail and a stopgap for rocm-vllm in #10940 (comment).

mudler added a commit that referenced this pull request Jul 23, 2026
… regenerate stubs after late installs (#11057)

* fix(backends): choose the protoc generator from the protobuf runtime, and regenerate stubs after late installs

The vLLM backends still crash on startup with

  VersionError: Detected incompatible Protobuf Gencode/Runtime versions when
  loading backend.proto: gencode 7.35.0 runtime 6.33.6

despite #10735 and #10944. Three separate defects kept it alive.

1. runProtogen picked the generator from the installed *grpcio* version.
   grpcio-tools' version tracks grpcio, but the gencode its bundled protoc
   emits tracks *protobuf*, and the two move independently: grpcio-tools
   1.82.1 (the version #10735 pins to, matching grpcio 1.82.1) requires
   protobuf>=7.35.1 and stamps gencode 7.35.0. Pinning to grpcio could
   therefore never constrain the gencode. Constrain the install to the
   protobuf already in the venv instead and let the resolver pick the newest
   compatible grpcio-tools. That both selects a generator the runtime accepts
   and stops protogen from moving the runtime under the backend's other deps.
   This is self-correcting, so the hardcoded GRPCIO_TOOLS_VERSION=1.78.0
   escape hatch from #10944 is no longer needed and is removed.

2. The stubs were generated too early. Most branches of vllm/install.sh (and
   vllm-omni) install vllm *after* installRequirements, and vllm re-resolves
   the protobuf runtime as it lands. Stubs generated against the pre-vllm
   runtime can end up newer than the runtime that finally ships, which is the
   ROCm failure exactly. Regenerate once the dependency set is final.

3. rm -f of the .py sources left __pycache__ behind. CPython validates a .pyc
   against source mtime and size, both of which can be unchanged across a
   regeneration (the gencode triple is the same width whether it reads 7.35.0
   or 6.33.5), so a stale backend_pb2.pyc could shadow the stub just written.

Also fail the build when the generated stub cannot be imported, so a
gencode/runtime mismatch surfaces at image build time instead of reaching
users as an opaque "grpc service not ready".

Verified by driving the real runProtogen through the ROCm install sequence in
a venv harness: before, gencode 7.35.0 against runtime 6.33.6 (reproducing the
reported error verbatim); after, gencode 6.33.5 against runtime 6.33.6 and the
stub imports cleanly.

Closes #10940
Closes #10718

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit]

* fix(backends): regenerate protobuf stubs in the other backends that install after installRequirements

Same defect as the vllm change: installRequirements generates the stubs at the
end of its own run, so any backend that installs further packages afterwards can
have the protobuf runtime moved out from under stubs that were already written.
The gencode stamped into backend_pb2.py then exceeds the runtime that ships and
the backend dies at model load with "grpc service not ready".

fish-speech already had this bug and worked around the symptom: it forces
protobuf>=5.29.0 after installRequirements precisely because "transitive deps
(wandb, tensorboard) may downgrade protobuf to 3.x but our generated
backend_pb2.py requires protobuf 5+". Regenerating after the pin addresses the
cause rather than propping up the runtime to match stale stubs.

Applied to the backends whose post-installRequirements step resolves a
dependency graph and can therefore move protobuf:

  fish-speech             -e . plus an explicit protobuf install
  vibevoice               pip install . (with deps)
  llama-cpp-quantization  gguf / GGUF_PIP_SPEC
  trl                     gguf / GGUF_PIP_SPEC

Deliberately not applied to ace-step and chatterbox (both --no-deps, so the
dependency graph cannot change) or voxcpm (pins setuptools only). gguf does not
depend on protobuf today, but it resolves dependencies, and "this package does
not touch protobuf right now" is exactly the assumption that made the earlier
fix ineffective.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit]

* fix(backends): resolve the protoc generator in a throwaway env so it cannot edit the backend's pinned deps

Installing grpcio-tools into the backend's own venv to generate the stubs also
drags its dependencies in: grpcio-tools 1.82.1 requires grpcio>=1.82.1, so a
backend that pinned grpcio==1.78.1 silently shipped 1.82.1 instead. Caught by
building the llama-cpp-quantization image and reading the versions back out of
the artifact:

  before   grpcio 1.82.1   (requirements.txt pins grpcio==1.78.1)
  after    grpcio 1.78.1   grpcio-tools absent from the venv entirely

Resolve the generator in a throwaway environment instead, still constrained to
the protobuf the backend ships so the gencode stays compatible. The backend's
dependency set is then exactly what its requirements files declared. protoc's
output is plain Python and carries no dependency on the interpreter that
produced it, so generating from a different env is safe; the import check still
runs under the backend's python, since that is the interpreter that has to load
the stubs at model load.

Verified on the rebuilt image: gencode 7.35.0, runtime protobuf 7.35.1, grpcio
back at its pinned 1.78.1, and the shipped stub imports cleanly against 7.35.1.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit]

* fix(backends): bound the protoc generator by BOTH the installed grpcio and protobuf

The generated stubs impose two independent constraints, and every fix so far,
including the previous commit on this branch, satisfied one while violating the
other:

  backend_pb2.py       needs  protobuf runtime >= gencode
  backend_pb2_grpc.py  needs  installed grpcio >= grpcio-tools

Resolving the generator against protobuf alone picked grpcio-tools 1.82.1 for a
backend holding grpcio at 1.78.1, so the gencode was fine but the gRPC stub was
not:

  RuntimeError: The grpc package installed is at version 1.78.1, but the
  generated code in backend_pb2_grpc.py depends on grpcio>=1.82.1.

That is also why installing grpcio-tools into the backend venv appeared to work
earlier: it dragged grpcio up to match, which was load-bearing rather than the
regression it looked like. Isolating the generator removed the accidental fix
and exposed the missing constraint.

Bound grpcio-tools from both sides instead and let the resolver find the newest
version satisfying both. The protobuf ceiling makes it back off to an older
generator when the runtime trails, bounding the gencode; the grpcio ceiling
keeps the _grpc stub loadable. Resolved against the four real runtime pairs
observed in built images:

  grpcio 1.78.1 / protobuf 7.35.1  -> grpcio-tools 1.78.0, gencode 6.31.1  OK
  grpcio 1.78.0 / protobuf 6.33.6  -> grpcio-tools 1.78.0, gencode 6.31.1  OK
  grpcio 1.82.1 / protobuf 6.33.6  -> grpcio-tools 1.81.1, gencode 6.33.5  OK
  grpcio 1.82.1 / protobuf 7.35.1  -> grpcio-tools 1.82.1, gencode 7.35.0  OK

Also restore the import check to cover backend_pb2_grpc as well as backend_pb2.
Narrowing it to backend_pb2 is why the image build passed while CI failed: the
guard could not see the constraint that was actually broken.

Verified by running the CI sequence locally for llama-cpp-quantization, the
backend whose test failed:
  make -C backend/python/llama-cpp-quantization        -> exit 0
  make -C backend/python/llama-cpp-quantization test   -> exit 0, OK

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

VLLM backend failing GRPC rocm hipblas AI PRO R9700 AI TOP 32G gfx1201

4 participants