Skip to content

effGen v0.2.9

Choose a tag to compare

@ctrl-gaurav ctrl-gaurav released this 23 May 18:16
· 971 commits to main since this release

effGen v0.2.9 Release — Observability & Reliability

effGen v0.2.9 ships the Observability & Reliability layer — structured JSON logs with automatic secret redaction, OpenTelemetry traces with configurable sampling, Prometheus histograms, SLO burn-rate tracking, per-provider circuit breakers, bulkheads, jittered retries, explicit timeout propagation, a deterministic chaos harness, a Hypothesis fuzz suite covering all 66 built-in tools, a load-testing CLI (effgen loadtest), and six Alertmanager-compatible alert rules. All telemetry is async/non-blocking — a failed export never fails inference. No breaking API changes.


What's Changed

Added

Observability — Structured Logging (effgen/observability/logs.py)

  • StructuredFormatter — emits JSON lines {ts, level, module, event, attributes, trace_id, span_id}. Integrates with OTel span context via _get_otel_ids(). Every attribute dict is filtered through the Redactor before serialisation.
  • get_logger(__name__) helper — log.event("model.call.started", model="llama3.1-8b", cached_tokens=0) event API; module-level structured logger with automatic trace-context injection. Returns a StructuredLogger bound to the calling module name.
  • Migration pass — agent loop (effgen/core/agent.py), model adapters (all 9 providers), router, and tool call sites migrated from ad-hoc print() / logger.info() to the structured logger. No critical path emits unstructured text.

Observability — Secret Redaction (effgen/observability/redact.py)

  • Redactor singleton with nine built-in patterns:

    Pattern Name
    sk-[a-zA-Z0-9]{20,} openai_key
    sk-ant-[a-zA-Z0-9]{20,} anthropic_key
    csk-[a-zA-Z0-9]{20,} cerebras_key
    AIza[0-9A-Za-z_-]{35} google_key
    hf_[a-zA-Z0-9]{20,} hf_key
    gsk_[a-zA-Z0-9]{20,} groq_key
    Bearer [^\s]+ bearer_token
    Slack webhook URLs slack_webhook
    Discord webhook URLs discord_webhook

    Replaces with <REDACTED:openai_key> etc. User-extensible via Redactor.add_pattern(name, regex). Applied at the log encoder — every path is covered without per-call opt-in.

Observability — Metrics: Histograms + SLO (effgen/observability/metrics.py, effgen/observability/slo.py)

  • effgen_model_call_latency_seconds{provider,model,outcome} — Histogram. Buckets: 0.05, 0.1, 0.25, 0.5, 1, 2, 5, 10, 30, 60 s.
  • effgen_tool_call_latency_seconds{tool,outcome} — Histogram. Same bucket spec.
  • effgen_agent_iteration_latency_seconds{preset} — Histogram. Per-iteration wall time.
  • effgen_tokens_total{provider,model,kind} — Counter. kind ∈ {input, output, cached}.
  • SLO(name, target_pct, window_seconds, query) — Defines an SLO with target percentage and rolling window.
  • SLOTracker — record(metric, labels, value) updates the window; burn_rate(name) returns the error-budget consumption ratio against target.
  • /slo FastAPI endpoint — exposes [{name, target_pct, current_pct, burn_rate, window_s}] as JSON.
  • /metrics compatibility — existing endpoint now includes all histogram buckets and _count / _sum in valid Prometheus text format.

Observability — Tracing: Sampling + Span Spec (effgen/observability/tracing.py, effgen/observability/spans.py)

  • Samplers (configurable via ObservabilityConfig.tracing.sampler):

    • AlwaysOn — samples every trace (use only for development / very low traffic).
    • AlwaysOff — disables all traces.
    • ParentBased(TraceIdRatio(p)) — probabilistic head-based sampling; honours parent decision.
    • ParentBased(RateLimited(per_second)) — token-bucket rate-limited sampler; no burst beyond the declared rate.
  • Canonical span-attribute spec (effgen/observability/spans.py) — single source of truth; every attribute name is a typed constant:

    Span name Attributes
    effgen.agent.iteration effgen.agent.preset, effgen.agent.iteration
    effgen.model.call effgen.model.provider, .name, .cached_tokens, .input_tokens, .output_tokens, .reasoning_effort, .thinking_budget, .parts_count
    effgen.tool.call effgen.tool.name, effgen.tool.status
    effgen.router.decision effgen.router.policy, .selected_provider, .considered[]
    effgen.retry.attempt effgen.retry.attempt, effgen.retry.reason
  • Span emission wired into all 9 adapters (OpenAI, Anthropic, Gemini, Cerebras, Groq, Together, Fireworks, Replicate, HFInference), every tool call site, and every router decision. Multimodal effgen.model.parts_count recorded for image/audio/video messages.

Reliability — Timeouts (effgen/reliability/timeouts.py)

  • ReliabilityConfig.default_timeouts — {"model_call": 60, "tool_call": 30, "http": 20}.
  • with_timeout(fn, timeout) / async_timeout(fn, timeout) / apply_timeout(fn, timeout) — sync and async wrappers.
  • audit_no_none_timeouts(source_root) — scans source for timeout=None; fails tests/reliability/test_timeouts.py if any found.
  • httpx propagation — every adapter's httpx.AsyncClient and httpx.Client constructed with timeout=httpx.Timeout(http_timeout).
  • Result: 0 timeout=None violations in the effGen source.

Reliability — Retries (effgen/reliability/retry.py)

  • Retry(max_attempts, base_delay, max_delay, jitter, retryable: Callable[[Exception], bool]) — fully configurable policy dataclass.
  • @retryable(Retry(...)) decorator — works on sync and async callables; respects Retry-After header when the exception carries one; records each attempt in an OTel effgen.retry.attempt span event.
  • is_transient_error(exc) — default retryable predicate: ProviderTransientError, httpx.TransportError, httpx.TimeoutException, HTTP 429 (with Retry-After), HTTP 5xx.
  • Not retried: ModelAuthError, ModelRefusalError, InvalidRequestError, BudgetExceededError.

Reliability — Circuit Breaker (effgen/reliability/circuit.py)

  • CircuitBreaker(name, failure_threshold, recovery_timeout, half_open_probes) — three-state finite-state machine:
    • CLOSED — normal operation; failure counter incremented on each failure.
    • OPEN — trips when failure_count >= failure_threshold; all calls immediately raise CircuitOpenError.
    • HALF_OPEN — after recovery_timeout seconds; allows up to half_open_probes probe calls; resets to CLOSED on success or returns to OPEN on failure.
  • CircuitBreakerRegistry — get(name) → CircuitBreaker singleton registry; wired into ProviderRegistry as get_circuit_breaker(provider).

Reliability — Bulkhead (effgen/reliability/bulkhead.py)

  • Bulkhead(name, max_concurrency, queue_size, queue_timeout) — semaphore-backed concurrency limiter with a bounded FIFO queue.
  • Supports sync (threading.Semaphore) and async (asyncio.Semaphore) variants.
  • Raises BulkheadFullError when queue_size is exhausted and queue_timeout expires.
  • BulkheadRegistry — get(name) → Bulkhead singleton registry; wired into ProviderRegistry as get_bulkhead(provider).

Chaos Harness (effgen/reliability/chaos.py)

  • Chaos(seed) — deterministic fault injector; same seed → same fault sequence across runs.

  • Fault types:

    • NetworkTimeout — raises httpx.ReadTimeout.
    • Http5xx — returns a 500 response.
    • Http429(retry_after) — returns a 429 with Retry-After header.
    • SlowResponse(ms) — time.sleep(ms / 1000) before returning.
    • PartialResponse — returns a truncated/malformed body.
    • MalformedJSON — returns Content-Type: application/json with non-JSON body.
  • registry.with_chaos(Chaos(seed)) — injects chaos as a ProviderRegistry middleware.

  • 4 canonical scenarios in tests/reliability/test_chaos.py:

    Scenario Fault Assert
    A Primary 5xx every 3rd call Router falls over to fallback; agent completes; span tree shows provider switch
    B 429 with Retry-After=2 on primary Retry honours wait; at most 1 extra delay; final answer correct
    C SlowResponse(59_000 ms) + timeout=60s TimeoutError raised cleanly; no hang > timeout+1 s; breaker records failure
    D All providers fail AllProvidersFailed raised; no silent empty string returned

    All 4 scenarios pass deterministically across 10 seeds (273 tests total).

Load-Testing Harness (effgen/tools/loadgen.py, effgen/cli/loadtest.py)

  • LoadGenConfig(concurrency, duration, scenario, provider, model, output_path) — dataclass.
  • Scenarios: fixed (single prompt), synthetic (varied prompts), multi_tool (tool-use heavy prompts).
  • LoadGenRunner.run() → LoadGenReport — async concurrent runner; collects per-request latency, success/failure; computes throughput, p50/p95/p99, error rate.
  • effgen loadtest CLI — --concurrency, --duration, --scenario, --provider, --model, --output. Writes JSON to stdout by default; file path via --output.
  • Live smoke result (mock model, c=10, d=30 s): ~69,260 requests, 0.00% error, p50≈0.3 ms, p95≈4.3 ms, p99≈8.1 ms.

Alerting (effgen/observability/alerting.py, docs/observability/alert_rules.yaml)

  • Six Alertmanager-compatible alert rules (docs/observability/alert_rules.yaml):

    Rule Condition Window
    HighErrorRate error_rate > 5% 10 min
    HighP95Latency p95 > 10 s 5 min
    CostBurnHigh cost > $10/day 1 h
    SLOFastBurn burn_rate > 14.4× 1 h
    SLOSlowBurn burn_rate > 3× 6 h
    CircuitBreakerOpen circuit in OPEN state 5 min
  • AlertWebhook(url).fire(alert) — tries SlackWebhookTool if URL is Slack-like, DiscordWebhookTool if Discord-like, generic httpx.post fallback. fire() is non-raising — delivery failure is logged but never propagated.

  • URL redaction — webhook URL path/token is replaced with *** in all logs (Redactor pattern applied at webhook layer).

Documentation (docs/observability/)

File Contents
overview.md Architecture diagram, quickstart, configuration reference, telemetry flow
metrics.md All metrics with label dimensions, bucket definitions, Grafana query examples
tracing.md Sampler selection guide, span attribute spec, in-memory exporter usage
alerting.md Alertmanager integration, webhook configuration, alert rule reference
loadtest.md Load-testing harness guide with scenarios and JSON report schema

Tests Added

File Tests Coverage
tests/observability/test_logs.py 26 JSON shape, trace_id propagation, module binding
tests/observability/test_redact.py 37 Every built-in pattern on known fixtures; user-extensible pattern; multi-pattern message
tests/observability/test_metrics.py ~20 Histogram recording, Prometheus text format validity
tests/observability/test_slo.py ~15 Rolling-window math, burn-rate formula spot-check
tests/observability/test_tracing.py 65 In-memory span exporter; 3-tool agent span tree; all declared attributes present
tests/reliability/test_timeouts.py 16 Timeout wrappers (sync + async); adapter audit (0 timeout=None violations)
tests/reliability/test_retry.py 33 Jitter distribution, Retry-After honoured, OTel event emission, non-retryable exceptions
tests/reliability/test_circuit.py 36 CLOSED→OPEN, OPEN→HALF_OPEN, HALF_OPEN→CLOSED, HALF_OPEN→OPEN transitions
tests/reliability/test_bulkhead.py 33 Concurrency limit enforcement, queue overflow, sync + async, timeout
tests/reliability/test_chaos.py 273 4 scenarios × 10 seeds; determinism check; middleware wiring; all fault types unit-tested
tests/fuzz/test_tool_fuzz.py 66 × 500 All 66 BaseTool subclasses; no unhandled exceptions; no secret leaks; secret-injection guard
tests/fuzz/test_message_fuzz.py 12 × 500 All 6 ContentPart types; Message role/content/metadata; random sequences
tests/fuzz/test_router_fuzz.py 11 × 500 Random availability + capabilities; valid decision or declared error; no nulls
tests/observability/test_loadgen.py 47 Loadgen library + CLI; JSON report schema; mock smoke; scenario definitions
tests/observability/test_alerting.py 34 Alert rules YAML validity; AlertWebhook fire; non-raising on delivery failure

Source Fixes (Bugs Closed During This Phase)

  • BaseTool.execute() redacts error strings — a secret-in-input leak was identified during fuzz: if a tool's input JSON contained a secret and the tool raised an error, the raw input was included in the ToolError.message. Fixed in effgen/tools/base_tool.py: error messages are now run through Redactor.redact() before being set. Affected 32 of 66 tools (all that include input in error messages).
  • timeout=None cleanup — 3 adapter call sites had implicit None timeouts. Fixed to use ReliabilityConfig.default_timeouts["http"].

Verification Results

Check Result
effgen.__version__ 0.2.9
Every log line in agent.run() JSON + no raw secrets ✓
Redactor covers all 9 built-in patterns 37/37 redact tests pass ✓
/metrics scrape Prometheus-valid histograms ✓
SLO burn-rate math Spot-checked against rolling-window fixtures ✓
OTel span tree (Calculator + WebSearch + 3-tool agent) All declared attributes present in every span ✓
Timeout audit 0 timeout=None violations in source ✓
Breaker CLOSED→OPEN→HALF_OPEN→CLOSED All transitions verified ✓
Bulkhead sync + async Concurrency limits enforced ✓
Chaos Scenario A (5xx fallback) 40 tests × 10 seeds deterministic ✓
Chaos Scenario B (429 + Retry-After) 40 tests × 10 seeds deterministic ✓
Chaos Scenario C (SlowResponse + timeout) 70 tests × 10 seeds, no hang > timeout+1 s ✓
Chaos Scenario D (AllProvidersFailed) 50 tests × 10 seeds, no silent empty string ✓
Fuzz × 500 examples (66 tools) 164/164 pass, 0 unhandled exceptions, 0 secret leaks ✓
Load harness mock smoke (c=10, 30 s) ~69,260 req, 0.00% error, p95≈4.3 ms ✓
Alert rules YAML Syntactically valid; 6 rules; unit test passes ✓
AlertWebhook.fire() Non-raising on delivery failure ✓
Regression suite 835 passed, 0 failed ✓
Wheel build (python -m build) effgen-0.2.9-py3-none-any.whl built cleanly ✓
Wheel smoke python -c "import effgen; assert effgen.__version__ == '0.2.9'" ✓

Upgrading from v0.2.8

No breaking API changes. All observability and reliability modules are new and additive; existing code paths are unaffected.

pip install --upgrade effgen

Observability Quick Start

from effgen.observability import get_logger
from effgen.observability import record_model_call, export_metrics
from effgen.observability.slo import SLOTracker, SLO
from effgen.observability import setup_tracing, ParentBasedSampler, TraceIdRatioSampler

# Structured logger
log = get_logger(__name__)
log.event("agent.started", preset="general", model="llama3.1-8b")
# → {"ts":"2026-05-23T...", "level":"INFO", "module":"...", "event":"agent.started", ...}

# Configure sampling (10% of traces)
setup_tracing(sampler=ParentBasedSampler(TraceIdRatioSampler(0.1)))

# Metrics auto-record on agent/model/tool calls; you can also record directly:
record_model_call(provider="cerebras", model="llama3.1-8b", outcome="ok", latency=0.42)
print(export_metrics())  # Prometheus text format

# SLO tracking
tracker = SLOTracker()
tracker.register(SLO("model_success", target_pct=99.0, window_seconds=3600))
tracker.record("model_success", ok=True)
print(tracker.burn_rate("model_success"))  # 0.0 when all calls succeed

Reliability Quick Start

from effgen.reliability.retry import Retry, retryable
from effgen.reliability.circuit import CircuitBreakerRegistry
from effgen.reliability.bulkhead import BulkheadRegistry
from effgen.reliability.chaos import Chaos

# Retry with jittered backoff
@retryable(Retry(max_attempts=3, base_delay=0.5, max_delay=10.0, jitter=True))
def call_model(prompt: str) -> str:
    ...

# Circuit breaker (per provider, auto-managed)
breakers = CircuitBreakerRegistry()
breaker = breakers.get_or_create("cerebras", failure_threshold=5, recovery_timeout=30)
if breaker.is_call_permitted():  # CLOSED → OPEN → HALF_OPEN
    try:
        result = call_model("hello")
        breaker.on_success()
    except Exception as exc:
        breaker.on_failure(exc)
        raise

# Bulkhead (per provider, auto-managed)
bulkheads = BulkheadRegistry()
bulkhead = bulkheads.get_or_create("cerebras", max_concurrency=10, queue_size=50)
with bulkhead.acquire():
    result = call_model("hello")

# Chaos testing
from effgen.models.registry import ProviderRegistry
registry = ProviderRegistry()
chaos = Chaos(seed=42)
registry.with_chaos(chaos)  # inject deterministic faults

Load-Testing Quick Start

# Default: mock model, 10 concurrent workers, 30 seconds
effgen loadtest

# Custom parameters
effgen loadtest --concurrency 20 --duration 60 --scenario multi_tool

# Against a live provider
effgen loadtest --provider cerebras --model llama3.1-8b --concurrency 5 --duration 15

# Save JSON report
effgen loadtest --output /tmp/loadtest_report.json
cat /tmp/loadtest_report.json | python -m json.tool

Alerting Quick Start

# Copy alert rules to Prometheus
cp docs/observability/alert_rules.yaml /etc/prometheus/rules/effgen.yaml
# Reload Prometheus: curl -X POST http://localhost:9090/-/reload
from effgen.observability.alerting import AlertWebhook, EffGenAlert

webhook = AlertWebhook(url="https://hooks.slack.com/services/...")
alert = EffGenAlert(
    name="HighErrorRate",
    severity="warning",
    summary="effGen error rate exceeded 5%",
    description="error_rate=8.3% over last 10 minutes",
)
webhook.fire(alert)  # non-raising — logs on failure, never throws

Full Changelog: CHANGELOG.md [0.2.9]
News & Highlights: NEWS.md
Observability Overview: docs/observability/overview.md
Metrics Reference: docs/observability/metrics.md
Tracing Guide: docs/observability/tracing.md
Alerting Guide: docs/observability/alerting.md
Load Test Guide: docs/observability/loadtest.md
Alert Rules YAML: docs/observability/alert_rules.yaml