Releases: HelixObs/client-python
Release list
v0.5.3
Fix: at-fork handlers to prevent gRPC deadlock in child processes
Pipelines that use os.fork() (directly or via multiprocessing / cadcutils) could deadlock permanently when fork happened while a BatchSpanProcessor or BatchLogRecordProcessor background thread held an internal gRPC mutex. The child inherited the locked mutex but the thread was dead, causing any subsequent gRPC call in the child to wait forever.
What changed
Added helixobs/_fork.py with os.register_at_fork() handlers:
- before fork:
force_flush(timeout_millis=500)on all OTel providers — drains the export queue and releases locks before fork copies memory state - after fork in child:
shutdown()on all providers — stops background threads in the child so it never tries to export telemetry
Both configure_logging(otlp=True) and Instrument.__init__ automatically register their providers. No-op on Windows where os.register_at_fork is unavailable.
v0.5.2 — TLS auto-detection in setup()
What's changed
setup() now auto-detects TLS from the endpoint port
When insecure is not passed explicitly, setup() infers it from the endpoint:
- Port
443→insecure=False(TLS) - Any other port →
insecure=True(plaintext)
Explicit insecure=True/False still takes precedence.
This means pipelines only need to update their endpoint env var — no code changes required:
# Before (plaintext to old VM):
HERALD_ENDPOINT=206.12.91.148:4317
# After (TLS to new cluster, insecure auto-detected):
HERALD_ENDPOINT=herald.chimefrb.xyz:443No API changes — fully backward compatible.
v0.5.1 — TLS log exporter
What's changed
Fix: OTLP log exporter now auto-detects TLS from endpoint scheme
configure_logging(otlp=True) previously hardcoded insecure=True on the OTLPLogExporter, making it impossible to ship logs to a TLS endpoint. The exporter now derives TLS mode from OTEL_EXPORTER_OTLP_ENDPOINT:
https://collector.example.com:443→ TLS (insecure=False)http://localhost:4317or barehost:port→ plaintext (insecure=True)
This matches the behaviour of the trace exporter and is required when using setup() with a TLS herald endpoint (e.g. herald.chimefrb.xyz:443).
No API changes — fully backward compatible.
v0.5.0 — Deferred entity ID, tel.trace()
New features
Deferred entity ID on operate()
entity_id is now optional on tel.operate(). Open the trace immediately — before the entity is known — and call token.set_entity_id(id) once it is discovered mid-operation. All logs emitted before the call share the same otelTraceID, making them reachable from the Entity Inspector via the operation's trace link.
with tel.operate("stage-replication") as op:
replicas = await fetch_replicas(storage_id) # inside the trace
if not replicas:
return # no entity → passthrough trace
op.set_entity_id(dataset_name) # now linked to the entity
await deposit_work(replicas)If the span closes without entity_id, a WARNING is logged and the span is forwarded as a plain OTel trace — no entity_operations row is written.
tel.trace() — plain OTel spans
New method for infrastructure work (HTTP handlers, loops, daemons) where log correlation by trace ID is useful but no entity is being tracked. No need to import the raw OTel tracer separately.
with tel.trace("handle-request", attributes={"method": "POST"}):
with tel.child_span("validate"):
...
with tel.child_span("write-db"):
...child_span() calls inside a tel.trace() block automatically inherit the trace context. The herald forwards these spans unchanged.
token.set_entity_id(id)
New convenience method on Token. Equivalent to token.set_attribute("helix.entity.id", id) — both suppress the missing-entity-id warning.
Upgrading
All changes are backwards-compatible. Existing tel.operate(operation, entity_id=...) calls are unaffected.
v0.4.2
Bug fixes
- Logging: Log lines from external libraries (e.g.
numpy,scipy) no longer produce a broken GitHub permalink pointing at the instrument's repo. Thesrcfield now contains the package-relative path (e.g.numpy/core/fromnumeric.py#L42) instead.
Other changes
- Added unit tests for
_TokenProviderauth andsetup()credential forwarding - Coverage badge updated to use shields.io live rendering from Coveralls
- Documentation site link added to README
- Alloy pipeline snippet and
helix_entity_idstructured metadata note corrected in docs
v0.4.1 — Auth docs cleanup
Changes
- User guide authentication section is now instrument-agnostic — no instrument-specific references or URLs
- Documents the two credential patterns clearly:
- Static string — for long-lived registration secrets
- Callable — for short-lived tokens that need refreshing before the 24h HelixObs JWT renewal
No code changes. Upgrade from v0.4.0 only needed if you rely on the user guide.
v0.4.0 — Gateway authentication
What's new
Gateway authentication support
Instrument now accepts two optional parameters for authenticating with the HelixObs gateway:
tel = CHIMEInstrument(
service_name="chime-frb-pipeline",
instrument_id="CHIMEFRB",
endpoint="206-12-91-148.cloud.computecanada.ca:4317",
credential=os.environ["CHIMEFRB_ACCESS_TOKEN"],
auth_endpoint="https://206-12-91-148.cloud.computecanada.ca/auth/token",
)credential— registration secret or existing instrument JWT (e.g.CHIMEFRB_ACCESS_TOKEN)auth_endpoint— URL of the gatewayPOST /auth/tokenendpoint
The client exchanges the credential for a short-lived HelixObs JWT at startup and attaches it to every OTLP export. Token refresh is automatic (1 hour before expiry) and thread-safe.
Phase 1 (current — plaintext gRPC)
Pass insecure=True (the default). The JWT is embedded as a static header at exporter creation. Valid for 24 hours; restart the process to refresh after expiry.
Phase 2 (future — TLS gRPC)
Pass insecure=False. A gRPC AuthMetadataPlugin refreshes the token per-RPC without recreating the channel.
Upgrading
No breaking changes. The new parameters are optional — existing code without credential/auth_endpoint continues to work unchanged. Auth enforcement on the gateway side is gated by the JWT_SECRET env var (empty = disabled).
v0.3.5
What's changed
- Fix: log format args were being swallowed — the log record factory was appending
src=torecord.msgbefore%-style arguments were resolved, causing OTel SDK warning/error messages to appear as raw%splaceholders in Loki logs. Now callsrecord.getMessage()first so the full formatted message is preserved.
v0.3.4
What's changed
- Python 3.8 compatibility — added
from __future__ import annotationstoinstrument.py; theX | Noneunion syntax now works on Python 3.8 and 3.9 process_namedocumented —setup()API table now includes theprocess_nameparameter with a link to the naming convention section- Process naming convention — new USER_GUIDE section explaining the
InstrumentID/pipeline/stagehierarchy, its rationale (no cross-instrument clashes, Loki regex group filtering), and example Loki queries like{helix_process_name=~"CHIME/l4-pipeline/.*"}
v0.3.3
What's new
Token.add_error(metadata=None)
New method for recording recoverable failures without ending the span.
token.error()— fatal: recordshelix.errorevent, sets ERROR status, ends the spantoken.add_error()— non-fatal: recordshelix.errorevent, sets ERROR status, keeps the span open
Use add_error() when a sub-step fails but the operation should continue:
with tel.operate("write-header", entity_id=event_id) as token:
try:
with tel.child_span("header_analysis.dump_header"):
dump_header(event)
except Exception as e:
logger.error(f"dump_header failed: {e}")
token.add_error({"stage": "dump_header", "message": str(e)})User guide updates
- Documents
add_error()in the Token methods reference - New "Recoverable failures" pattern in section 5
- Documents that subprocess output (
os.system(), etc.) is not captured by logging