Detect LLM hallucinations and quantify uncertainty in microseconds without secondary NLI cross-encoders.
Traditional epistemic uncertainty estimation in LLMs relies on Semantic Entropy (SE) (Kuhn et al., 2023; Farquhar et al., Nature 2024). While effective, Semantic Entropy requires clustering
This introduces two severe production bottlenecks:
-
Quadratic Cost:
$\binom{K}{2}$ forward passes per query (45 neural evaluations for$K=10$ ). - Serving Latency: Adds $\sim$90 ms of GPU overhead per inference call, making it unusable for high-throughput production serving.
Spanda introduces Exact-Match Normalized Entropy (
Across empirical evaluations spanning two orders of magnitude (1.5B to 120B parameters), Spanda matches or exceeds neural Semantic Entropy on structured reasoning while operating ~90,000$\times$ faster (
As model capacity increases from 1.5B to 27B parameters, internal reasoning coherence causes correct predictions to naturally converge to identical lexical sequences. On mathematical reasoning (GSM8K), exact-match AUROC scales monotonically:
At 7B+ parameters, Spanda achieves the exact same discriminative power as heavy DeBERTa-v3 NLI cross-encoders, rendering the neural clustering step redundant for reasoning.
At the 120B frontier scale on ungrounded factual recall (TriviaQA), the model exhibits Confident Mode Collapse: its parametric memory and RLHF tuning cause it to hallucinate the exact same incorrect answer identically across all
β οΈ Critical Safety Implication: Any system using self-consistency or agreement as a proxy for truth will be systematically deceived by frontier models on ungrounded factual recall. External grounding (RAG) is mandatory in this regime.
| Model Scale | Benchmark | Accuracy | Spanda ( |
Neural SE AUROC | Latency | GPU Req. |
|---|---|---|---|---|---|---|
| Qwen-1.5B | GSM8K | 11.4% | 0.577 | 0.584 | None | |
| Qwen-1.5B | TriviaQA | 32.0% | 0.797 | 0.801 | None | |
| Mistral-7B | GSM8K | 8.2% | 0.706 | 0.705 | None | |
| Mistral-7B | TriviaQA | 45.0% | 0.698 | 0.755 | None | |
| Qwen-27B | GSM8K | 61.2% | 0.889 | --- | None | |
| DeBERTa Baseline | N/A | --- | --- | --- | $\sim$92.4 ms | Required |
Given
The Normalized Shannon Entropy is: $$H_{\text{norm}} = \begin{cases} 0 & \text{if } n = 1 \ \displaystyle\frac{-\sum_{i=1}^n w_i \ln w_i}{\ln K} & \text{if } n > 1 \end{cases}$$
The combined Spanda Risk Score (
-
$R_{sc} = 0$ : Complete consensus (model is confident). -
$R_{sc} \to 1$ : Maximum epistemic divergence (model is guessing / hallucinating).
Spanda is lightweight and requires zero third-party dependencies (pure Python standard library).
pip install spandaOr install from source:
git clone https://github.com/Adarshent/Spnda.git
cd Spnda
pip install -e .from spanda import compute_rsc
# High-consensus query (Model is confident)
samples_confident = ["Paris", "paris.", "Paris", "Paris", "Paris"]
res_conf = compute_rsc(samples_confident)
print(f"R_sc Score: {res_conf['rsc']}") # 0.0
print(f"Dominant Answer: {res_conf['dominant_answer']}") # 'Paris'
# Uncertain / guessing query (Model is hallucinating)
samples_uncertain = ["Berlin", "Rome", "Madrid", "London", "Paris"]
res_unc = compute_rsc(samples_uncertain)
print(f"R_sc Score: {res_unc['rsc']}") # 0.9 (High risk!)from spanda import detect_hallucination
samples = ["42", "42", "24", "17", "99"]
guard = detect_hallucination(samples, threshold=0.35)
if guard["is_uncertain"]:
print(f"π¨ Hallucination Warning (R_sc = {guard['rsc']}). Routing to RAG / Human Review.")
else:
print(f"β
Safe output: {guard['dominant_answer']}")from spanda import batch_compute_rsc
batch = [
["Answer A", "Answer A", "Answer A"],
["Choice 1", "Choice 2", "Choice 3"]
]
results = batch_compute_rsc(batch)
for r in results:
print(r["rsc"], r["dominant_answer"])For mission-critical production pipelines, Spanda provides a 2-Tier Cascaded Guardrail that combines sub-millisecond consensus filtering with context grounding and tool-call safety:
from spanda import CascadedGuardrail
guard = CascadedGuardrail(
uncertainty_threshold=0.3,
grounding_threshold=0.15
)
# 1. RAG Query with Mode Collapse Protection
rag_context = "Documentation: The production cluster runs in us-east-1."
unanimous_hallucination = ["eu-west-3 Paris", "eu-west-3 Paris", "eu-west-3 Paris"]
receipt = guard.evaluate(unanimous_hallucination, context=rag_context)
print(receipt.decision) # 'MODE_COLLAPSE_RISK'
print(receipt.is_safe) # False (Unanimous agreement, but 0% grounded in source!)
print(receipt.tier_executed) # Tier 2
print(receipt.latency_ms) # < 0.05 ms
# 2. Agent Tool Call Argument Verification (e.g. preventing bad 'rm')
tool_calls = [
{"command": "rm -rf /var/cache"},
{"command": "rm -rf /var/log"}, # Conflict detected across parallel paths!
]
agent_receipt = guard.evaluate_tool_calls(tool_calls)
print(agent_receipt.decision) # 'TOOL_ARG_MISMATCH' (Execution blocked!)
# 3. Export SOC2 Audit Receipt
import json
print(json.dumps(receipt.to_dict(), indent=2))| Use Case / Architecture | Recommendation | Rationale |
|---|---|---|
| Math, Code & Structured QA (7Bβ70B) | β Recommended | Coherence Scaling Law ensures exact-match matches neural SE at 0 cost. |
| High-Throughput Production APIs | β Recommended | 90,000x latency reduction without GPU requirements. |
| Free-form Paraphrase QA (<7B) | Small models produce inconsistent surface phrasing. | |
| Ungrounded Facts on Frontier Models (>100B) | β Do Not Use Alone | Subject to Confident Mode Collapse; must combine with retrieval (RAG). |
Run the test suite:
python3 -m unittest discover testsIf you use Spanda in your research or production systems, please cite:
@article{nayak2026spanda,
title={Spanda: Zero-Cost Lexical Entropy Matches Neural Semantic Uncertainty---Until Frontier Models Break It},
author={Nayak, Bhupen},
journal={arXiv preprint},
year={2026},
doi={10.5281/zenodo.22233648},
url={https://doi.org/10.5281/zenodo.22233648}
}This project is licensed under the MIT License - see the LICENSE file for details.