Repository navigation
Replies: 3 comments 2 replies
|
Spot-on question! Coupling the safety rail's execution domain to the generation LLM client creates a single point of failure. If the LLM API throttles, times out, or gets context-poisoned, your safety gate either fails open or crashes the pipeline. Decoupling safety into an independent, deterministic pre-execution proxy ensures policy enforcement remains isolated from generation failure modes. |
|
I think the distinction here is less about For example, Also agree with your observation about failure behavior. From the current action code I wouldn't describe this as fail-open. If the LLM call raises, the check doesn't turn that into an Using a separate model/client for safety-sensitive checks would make sense when independent failure domains are actually a requirement. It could even be the same provider/model behind a separate endpoint or quota, depending on what failure you're trying to isolate. The current default is convenient, but I wouldn't count an LLM-based rail using the same backend as generation as an independent availability boundary. |
|
@imronreviady thank you, that's the right refinement and my post should have made it. I re-checked against v0.24.1 and the commit I pinned. You're right that it depends on which model is injected. A model entry typed self_check_input replaces llm for that rail, and get_self_check_llm now makes the order explicit: dedicated model, otherwise main. So sharing the generation client is the default for the self-check rails, not a structural property. I also lumped all eight rails together when four already take their own model entries, and I pinned What I think still stands: llm_call is one function used by generation and by all eight rails, and with no task model configured the self-check rails and generation run on the same object. We agree it isn't fail-open at the action level. One question for the maintainers: is falling back to the main model, when no dedicated model is configured, the intended default for a safety rail? If so, it might be worth the docs saying that this puts rail and generation on one provider and quota. @Rehanguards nothing here says the rails fail open, and the code and the reply above show they don't. A deterministic proxy is a different design; it can't make the semantic judgement an LLM-based rail makes, so it complements this question rather than answering it. |
Uh oh!
There was an error while loading. Please reload this page.
Hi,
I've been going through how different guardrail systems get their judgement, across ten agent frameworks and guardrail toolkits, all at pinned commits. This is a design question, not a security report. There's no bypass here and I'm not claiming one. Your SECURITY.md points vulnerabilities at PSIRT and I haven't filed there, because this isn't that.
What I was looking for is whether the component doing the checking gets its answer through the same channel as the thing it's checking. Six of the ten ship a monitor of their own. Two of those six share the channel, and NeMo is one of them, so I wanted to ask rather than assume it was an oversight.
At dc046e4, llm_call in nemoguardrails/actions/llm/utils.py is defined once and called from 10 non-test files. Two of them are your own response generation, actions/llm/generation.py and actions/v2_x/generation.py. Eight are shipped rails:
self_check/input_check/actions.py:67
self_check/output_check/actions.py:74
llama_guard/actions.py:77
content_safety/actions.py:100
hallucination/actions.py:62
topic_safety/actions.py:124
self_check/facts/actions.py:73
patronusai/actions.py:108
Two things follow from that, and I checked both instead of assuming.
First, it's one failure domain rather than two. If llm_call fails, generation and those eight rails go down together. People usually reason about defence in depth as independent layers, but the number of layers here is bigger than the number of things that can fail on their own. I did check whether this turns into a silent allow, and it doesn't. If I make llm_call raise, self_check_input raises too, it doesn't quietly return "safe". That looks deliberate and I think it's the right call, so I'd rather say it up front than leave it implied.
Second, the rail sends the user's text out on the channel it's guarding. self_check_input reads context["user_message"], renders it into the check prompt and hands it to llm_call. I pulled the function out of the pinned commit and ran it against a stub that only records what it receives, and the user's text comes through verbatim. That's unavoidable if the check is an LLM. But it does mean untrusted input reaches the same client and the same provider path your own generation uses.
Two of your rails don't do this, the perplexity based jailbreak detector and the YARA injection detector, so it looks like a property of the LLM-based rails rather than of the library as a whole.
If it's useful, here's what the ones that don't share do. LlamaFirewall's scanners build their own client. Guardrails AI takes the LLM as an injected callable and no validator touches it. LangChain's safety middlewares don't call a model at all. Three different ways, none of which change what a rail actually does.
So my question is whether the sharing is intended, and whether a separate client or endpoint for rails is something you'd consider. I'm writing this work up and I'll report whatever you say accurately, including if the answer is that I've read it wrong. Happy to share the code I used.
Thanks,
Jaswanth
All reactions