Skip to content

Releases: HelixObs/client-python

v0.5.3

Choose a tag to compare

@chitrangpatel chitrangpatel released this 23 Jun 20:50

Fix: at-fork handlers to prevent gRPC deadlock in child processes

Pipelines that use os.fork() (directly or via multiprocessing / cadcutils) could deadlock permanently when fork happened while a BatchSpanProcessor or BatchLogRecordProcessor background thread held an internal gRPC mutex. The child inherited the locked mutex but the thread was dead, causing any subsequent gRPC call in the child to wait forever.

What changed

Added helixobs/_fork.py with os.register_at_fork() handlers:

  • before fork: force_flush(timeout_millis=500) on all OTel providers — drains the export queue and releases locks before fork copies memory state
  • after fork in child: shutdown() on all providers — stops background threads in the child so it never tries to export telemetry

Both configure_logging(otlp=True) and Instrument.__init__ automatically register their providers. No-op on Windows where os.register_at_fork is unavailable.

v0.5.2 — TLS auto-detection in setup()

Choose a tag to compare

@chitrangpatel chitrangpatel released this 18 Jun 00:17

What's changed

setup() now auto-detects TLS from the endpoint port

When insecure is not passed explicitly, setup() infers it from the endpoint:

  • Port 443 → insecure=False (TLS)
  • Any other port → insecure=True (plaintext)

Explicit insecure=True/False still takes precedence.

This means pipelines only need to update their endpoint env var — no code changes required:

# Before (plaintext to old VM):
HERALD_ENDPOINT=206.12.91.148:4317

# After (TLS to new cluster, insecure auto-detected):
HERALD_ENDPOINT=herald.chimefrb.xyz:443

No API changes — fully backward compatible.

v0.5.1 — TLS log exporter

Choose a tag to compare

@chitrangpatel chitrangpatel released this 17 Jun 23:22

What's changed

Fix: OTLP log exporter now auto-detects TLS from endpoint scheme

configure_logging(otlp=True) previously hardcoded insecure=True on the OTLPLogExporter, making it impossible to ship logs to a TLS endpoint. The exporter now derives TLS mode from OTEL_EXPORTER_OTLP_ENDPOINT:

  • https://collector.example.com:443 → TLS (insecure=False)
  • http://localhost:4317 or bare host:port → plaintext (insecure=True)

This matches the behaviour of the trace exporter and is required when using setup() with a TLS herald endpoint (e.g. herald.chimefrb.xyz:443).

No API changes — fully backward compatible.

v0.5.0 — Deferred entity ID, tel.trace()

Choose a tag to compare

@chitrangpatel chitrangpatel released this 29 May 13:09

New features

Deferred entity ID on operate()

entity_id is now optional on tel.operate(). Open the trace immediately — before the entity is known — and call token.set_entity_id(id) once it is discovered mid-operation. All logs emitted before the call share the same otelTraceID, making them reachable from the Entity Inspector via the operation's trace link.

with tel.operate("stage-replication") as op:
    replicas = await fetch_replicas(storage_id)   # inside the trace
    if not replicas:
        return                                     # no entity → passthrough trace
    op.set_entity_id(dataset_name)                # now linked to the entity
    await deposit_work(replicas)

If the span closes without entity_id, a WARNING is logged and the span is forwarded as a plain OTel trace — no entity_operations row is written.

tel.trace() — plain OTel spans

New method for infrastructure work (HTTP handlers, loops, daemons) where log correlation by trace ID is useful but no entity is being tracked. No need to import the raw OTel tracer separately.

with tel.trace("handle-request", attributes={"method": "POST"}):
    with tel.child_span("validate"):
        ...
    with tel.child_span("write-db"):
        ...

child_span() calls inside a tel.trace() block automatically inherit the trace context. The herald forwards these spans unchanged.

token.set_entity_id(id)

New convenience method on Token. Equivalent to token.set_attribute("helix.entity.id", id) — both suppress the missing-entity-id warning.

Upgrading

All changes are backwards-compatible. Existing tel.operate(operation, entity_id=...) calls are unaffected.

v0.4.2

Choose a tag to compare

@chitrangpatel chitrangpatel released this 21 May 00:42

Bug fixes

  • Logging: Log lines from external libraries (e.g. numpy, scipy) no longer produce a broken GitHub permalink pointing at the instrument's repo. The src field now contains the package-relative path (e.g. numpy/core/fromnumeric.py#L42) instead.

Other changes

  • Added unit tests for _TokenProvider auth and setup() credential forwarding
  • Coverage badge updated to use shields.io live rendering from Coveralls
  • Documentation site link added to README
  • Alloy pipeline snippet and helix_entity_id structured metadata note corrected in docs

v0.4.1 — Auth docs cleanup

Choose a tag to compare

@chitrangpatel chitrangpatel released this 20 May 01:22

Changes

  • User guide authentication section is now instrument-agnostic — no instrument-specific references or URLs
  • Documents the two credential patterns clearly:
    • Static string — for long-lived registration secrets
    • Callable — for short-lived tokens that need refreshing before the 24h HelixObs JWT renewal

No code changes. Upgrade from v0.4.0 only needed if you rely on the user guide.

v0.4.0 — Gateway authentication

Choose a tag to compare

@chitrangpatel chitrangpatel released this 20 May 01:09

What's new

Gateway authentication support

Instrument now accepts two optional parameters for authenticating with the HelixObs gateway:

tel = CHIMEInstrument(
    service_name="chime-frb-pipeline",
    instrument_id="CHIMEFRB",
    endpoint="206-12-91-148.cloud.computecanada.ca:4317",
    credential=os.environ["CHIMEFRB_ACCESS_TOKEN"],
    auth_endpoint="https://206-12-91-148.cloud.computecanada.ca/auth/token",
)
  • credential — registration secret or existing instrument JWT (e.g. CHIMEFRB_ACCESS_TOKEN)
  • auth_endpoint — URL of the gateway POST /auth/token endpoint

The client exchanges the credential for a short-lived HelixObs JWT at startup and attaches it to every OTLP export. Token refresh is automatic (1 hour before expiry) and thread-safe.

Phase 1 (current — plaintext gRPC)

Pass insecure=True (the default). The JWT is embedded as a static header at exporter creation. Valid for 24 hours; restart the process to refresh after expiry.

Phase 2 (future — TLS gRPC)

Pass insecure=False. A gRPC AuthMetadataPlugin refreshes the token per-RPC without recreating the channel.

Upgrading

No breaking changes. The new parameters are optional — existing code without credential/auth_endpoint continues to work unchanged. Auth enforcement on the gateway side is gated by the JWT_SECRET env var (empty = disabled).

v0.3.5

Choose a tag to compare

@chitrangpatel chitrangpatel released this 13 May 13:05

What's changed

  • Fix: log format args were being swallowed — the log record factory was appending src= to record.msg before %-style arguments were resolved, causing OTel SDK warning/error messages to appear as raw %s placeholders in Loki logs. Now calls record.getMessage() first so the full formatted message is preserved.

v0.3.4

Choose a tag to compare

@chitrangpatel chitrangpatel released this 13 May 00:57

What's changed

  • Python 3.8 compatibility — added from __future__ import annotations to instrument.py; the X | None union syntax now works on Python 3.8 and 3.9
  • process_name documented — setup() API table now includes the process_name parameter with a link to the naming convention section
  • Process naming convention — new USER_GUIDE section explaining the InstrumentID/pipeline/stage hierarchy, its rationale (no cross-instrument clashes, Loki regex group filtering), and example Loki queries like {helix_process_name=~"CHIME/l4-pipeline/.*"}

v0.3.3

Choose a tag to compare

@chitrangpatel chitrangpatel released this 12 May 21:54

What's new

Token.add_error(metadata=None)

New method for recording recoverable failures without ending the span.

  • token.error() — fatal: records helix.error event, sets ERROR status, ends the span
  • token.add_error() — non-fatal: records helix.error event, sets ERROR status, keeps the span open

Use add_error() when a sub-step fails but the operation should continue:

with tel.operate("write-header", entity_id=event_id) as token:
    try:
        with tel.child_span("header_analysis.dump_header"):
            dump_header(event)
    except Exception as e:
        logger.error(f"dump_header failed: {e}")
        token.add_error({"stage": "dump_header", "message": str(e)})

User guide updates

  • Documents add_error() in the Token methods reference
  • New "Recoverable failures" pattern in section 5
  • Documents that subprocess output (os.system(), etc.) is not captured by logging