Instrumenting the outbound sync channel (flagd → flag-server) #2044
Replies: 5 comments
|
Just to be perfectly clear, because flagd can be both a sync server and a sync client ... I'm guessing you are looking at adding metrics for the latter? If so, I would add some specific/standard gRPC metrics, but also some more general sync metrics that apply to all sync sources (HTTP, file, gRPC (sync), etc). The stream/connection stuff is gRPC-specific, but staleness (e.g. time since last successful config apply) applies to every source, so I'd rather not make it gRPC-only. We already have |
|
Thanks @toddbaert for the nice suggestions, let me extend the scope and propose a concrete direction.
Universal (all sync providers): Stream-based only (gRPC):
Let me know if you have any further thoughts ! |
|
Hey @toddbaert I’ve implemented the proposed changes in PR #2051, including the client-side sync metrics and gRPC stream lifecycle metrics. I also added a Grafana dashboard snapshot to the PR #2051 to demonstrate how the metrics can be used to detect stale configuration and sync issues. Please let me know if you have any feedback or concerns about the metric naming or the client/server distinction. I’d be happy to incorporate any suggestions. |
|
Sounds good. I will review in the next day or two! |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Context:
Flagd is deployed across multiple servers, and we have encountered issues related to stale flags. During the investigation, we found that there are currently no metrics available to monitor the flagd sync interface or detect issues with the sync channel.
Problem Statement
Today, /metrics exposes inbound evaluation traffic (http_server_) and evaluation outcomes (feature_flag_flagd_), but there are no metrics for the flag-sync-server <-> flagd sync channel. As a result, we currently have no way to reliably alert on conditions such as:
I'm working on a prototype to add the following metrics:
Proposed Metrics:
otelgrpc client stats handler on the sync gRPC client -> gives the standard
rpc.client.*metrics.rpc_client_duration_milliseconds_-> Histogram -> Stream lifetime + _count{rpc_grpc_status_code} = "why streams ended"rpc_client_response_size_bytes_-> Histogram -> Payload size distribution + rate(_sum) = inbound bandwidthflagd-specific sync-lifecycle metrics (otelgrpc records histograms only on RPC end, so for a long-lived
SyncFlagsstream those don't help real-time):flagd_sync_stream_active-> UpDownCounter -> Number of currently open flagd → flag-server sync streams. 1 while the SyncFlags stream is open, 0 when disconnected.flagd_sync_stream_reconnects_total-> Counter -> Total number of successful sync stream reconnections after an initial failureflagd_sync_flag_config_received_total-> Counter -> Total number of flag-configuration payloads received from the flag-server over SyncFlags streams.flagd_sync_flag_config_last_received_timestamp_seconds-> Gauge (seconds) -> Timestamp of the most recent flag-configuration push received from the flag-server.Scope:
The current prototype is limited to the gRPC sync channel.
I'd be interested in feedback on the proposed metrics, particularly whether there are other sync-specific signals that would be useful for monitoring or alerting.
Let me know your thoughts, @toddbaert @beeme1mr.
Thanks !
All reactions