Repository navigation
Agent Connectors
Add auditing and shared memory to an agent you already have, by adding one line.
| Status | ✅ Works |
| Verified | 17 September 2026, RTX PRO 4500 Blackwell, live against vLLM 0.29.0 serving Gemma 4 31B: the decorator and the LangChain adapter attached, agent output unaffected |
| Package |
sfab_agents (Python) |
| Needs | the Galahad Python package (it carries libgalahad.so), or GALAHAD_LIBRARY_PATH pointing at your own copy |
pip install galahad-kv # carries the connectors and libgalahad.so
# only if you keep the library elsewhere:
# export GALAHAD_LIBRARY_PATH=/opt/galahad/lib/libgalahad.soimport sfab_agents # raises ImportError if the package is not installed
print("ok")⚠ If that import fails, nothing else on this page will work. Install the
package with pip rather than relying on your current directory.
An AI agent works in steps. Normally, when a step goes wrong, you get the wrong answer at the end and no way to find out where it happened.
This watches each step as it runs. When something fails, it tells you which layer caused it: your application, the model, or the memory. And it lets the agent reuse what it has already read, instead of reading the same documents over and over.
You do not rewrite your agent. You add one line above the function you already wrote.
A support team has a 90 page handbook and ten agents using it.
Without this, every agent rereads the whole handbook before every single question: five hundred readings a day of the same book.
With this, the handbook is read once and shared. Each agent asks its own questions. And when a step fails, you can see exactly which one failed and why.
from sfab_agents.decorator import audit_and_graft
@audit_and_graft(provider="vllm")
def my_agent_step(state):
# your existing code, unchanged
return call_my_model(state["question"])That is the whole integration. The function behaves exactly as before.
from sfab_agents import capability_summary
print(capability_summary())Output from a working install:
sfab-agents on Python 3.11
binding: loaded
galahad: present -- memory reuse available
adapters available: langchain
And on a machine where the Python extension is not reachable:
sfab-agents on Python 3.11
binding: NOT LOADED -- ModuleNotFoundError: No module named 'sfab_py'
nothing is recorded; your agent runs unaffected
galahad: present -- memory reuse available
adapters available: none installed
⚠ Read the second line. "Nothing is recorded; your agent runs unaffected." A broken install does nothing; it does not take your agent down.
audit_and_graft(
provider: str = "unknown", # "vllm", "openai", ...
in_process: bool = False, # True if the model runs in this process
max_logprobs: int | None = None, # how many log probabilities your provider returns
calibrated: bool = False, # True only for a calibrated setup
run_id = None, # group steps into one run
logits_from = None, # where to read logits, if available
topk_from = None, # where to read the top k, if available
halt: bool = False, # stop the loop on a confirmed fault
attributer = None, # optional fault checker, used with halt
drift_from = None, # a function that returns a score for this step
)Every argument is optional. @audit_and_graft() is valid and does the
sensible thing for an unknown provider.
Without logits_from or topk_from, a step is recorded as structure only, and
the capability report says so.
Nothing changes about your return value. The decorator observes; it does not transform.
Not every deployment can be audited to the same depth, and the library tells you which one you are in.
| Tier | Constant | You have | What can be checked |
|---|---|---|---|
| 1 | TIER_1_FULL_LOGITS |
full logits, in process | everything: full divergence detection |
| 2 | TIER_2_LOCAL_TOPK |
local model, top k only | strong |
| 3 | TIER_3_HOSTED_TOPK |
hosted API returning top k | partial |
| 4 | TIER_4_METADATA_ONLY |
hosted API, no probabilities | timing, structure, tool calls |
Most hosted APIs return no probabilities at all, so they land in tier 4.
capability_summary() tells you what your setup supports, and the audit reports
steps_audited / steps_total so you can see what it actually did.
AUDIT_FULL · AUDIT_SURROGATE · AUDIT_NONE
Seven integrations. Import the one you use:
| Framework | Module | Requires |
|---|---|---|
| LangChain | sfab_agents.langchain |
langchain_core |
| LangGraph | sfab_agents.langgraph |
langgraph |
| LlamaIndex | sfab_agents.llamaindex |
llama_index.core |
| CrewAI | sfab_agents.crewai |
crewai |
| AutoGen | sfab_agents.autogen |
autogen_agentchat |
| Letta / MemGPT | sfab_agents.letta |
letta_client |
| plain Python, no framework | sfab_agents.decorator |
nothing extra |
from sfab_agents.langchain import SfabCallbackHandler
handler = SfabCallbackHandler(provider="vllm")
result = llm.invoke("your prompt", config={"callbacks": [handler]})Use one handler per invocation.
✅ Verified live against vLLM 0.29.0 serving Gemma 4 31B: both calls returned, handler attached, agent unaffected.
ImportError: sfab_agents.langchain needs 'langchain_core', which is not installed.
The error names the exact package. Adapters are never skipped silently.
@audit_and_graft()
def broken_step(state):
raise ValueError("my own bug")
# ValueError propagates untouched.The decorator never swallows your error.
If the Python extension is missing, the library is unreachable, or no engine is attached, your agent runs normally and nothing is recorded. The capability summary says so.
The audit never alters what your model returns.
- Entering an audit scope with no engine is refused, naming what to call.
- Capturing nothing is an error, not a silent pass.
- Two concurrent traces are refused, not interleaved.
| ❌ It does not train or fine tune your model | weights are unchanged |
| ❌ It does not make hosted APIs return logits they do not return | tiers 3 and 4 are a real limit |
| ❌ It does not pool failure signals across agents | agents share memory, so work one agent did is reused by another; a failure one agent hit is not automatically applied to the others |
| ❌ It does not run your model | it observes whatever you already run |
⚠ The shared memory benefit is measured: about 3.7× less repeated work at 4 agents, 8× at 10, 23× at 50, on a shared prefix. Speed in seconds depends on your hardware and load; the reliable claim is less repeated work, not a fixed number. See Prefix Sharing.
| Symptom | Cause | Fix |
|---|---|---|
binding: NOT LOADED |
the Python extension is not on the path | check PYTHONPATH; your agent still runs |
galahad: absent |
libgalahad.so not reachable |
install the Galahad package, or set GALAHAD_LIBRARY_PATH=/path/to/libgalahad.so
|
adapters available: none installed |
no framework package present | install the one you use, see Framework adapters |
ImportError: ... needs 'X' |
that adapter's framework is missing | pip install X |
steps_audited much lower than steps_total
|
your tier cannot audit every step | check capability_summary()
|
- Branching and Snapshots: exploring several paths from one memory
- Replay: regression tests for agent runs
- Install