Skip to content

v2.3.0

Latest

Choose a tag to compare

@github-actions github-actions released this 04 Oct 14:54
· 6 commits to main since this release

Upgrade notes

This release hardens the WebSocket monitor, /caget and /healthcheck. It fixes several ways PVs could get stuck after IOC, network or gateway disruptions, and adds detection of PVs that freeze anyway.

Action needed when upgrading

  • Nagios, or any monitoring that alerts on /healthcheck: use /epics2web/healthcheck?strict=true. The default /healthcheck now answers 200 whenever the server is up, even if PVs are disconnected. That's so a load balancer doesn't take every instance out when an IOC goes down. Strict mode answers 503 when a PV that was connected has stayed disconnected past the grace period.
  • Load balancers: keep using /epics2web/healthcheck.
  • Tomcat: tested with Tomcat 11.0.26, which the Docker image now pins. That version includes WebSocket fixes relevant to this release, such as write timeouts covering whole messages. Upgrading from earlier 11.0.x releases is recommended.

Behaviour changes

  • Unresponsive clients are closed. The server pings WebSocket clients every 30 s, and closes a client that sends no message and no pong for 60 s, or whose writes stay blocked for 20 s. Browsers answer pings automatically, and epics2web.js reconnects.
  • Clients that fall behind get the latest values. A client that falls behind gets each PV's latest update instead of a backlog. Connect and disconnect (info) messages are never dropped.
  • /caget validates requests:
    • it answers HTTP 400 for more than 500 PVs in one request, or for a jsonp callback that isn't a JavaScript name such as a.b_c;
    • a PV that doesn't connect now fails only the request that asked for it, with an error naming the PV.
  • /healthcheck reports more accurately:
    • it counts the grace period from when a PV disconnected, not from its last value change;
    • entries gain a state field;
    • PVs that never connected are listed, as CONNECTING, but don't fail strict mode;
    • frozen PVs are listed with frozen, frozen_minutes and frozen_reason;
    • ?frozen=true answers 503 when any PV is frozen.
  • No session cookies: responses no longer set a JSESSIONID cookie, except for WebSocket handshakes.
  • Console changes: the console lists connected sessions that don't monitor anything yet, and shows the number of open Channel Access channels.
  • -Infinity is now reported with its sign.

Frozen PV detection (new, report-only)

epics2web now looks for PVs whose monitor has stopped working while the IOC still serves them. It gives suspicious PVs a short-lived subscription in a second, independent CA context, which opens its own connection to each IOC or gateway and probes at most 20 PVs at a time.

  • A frozen PV is logged as WARNING ... PV <name> is frozen: <reason>.
  • It's listed by /healthcheck.
  • Through a gateway, it finds problems between epics2web and the gateway, not inside the gateway.

It doesn't change the default or strict healthcheck responses.

New settings (environment variables, all optional)

Variable Default Meaning
WEBSOCKET_PING_INTERVAL_SECONDS 30 How often clients are pinged
WEBSOCKET_TIMEOUT_SECONDS 60 Close a client with no message or pong for this long
WEBSOCKET_SEND_TIMEOUT_SECONDS 20 Close a client whose write blocks this long (Tomcat only)
HEALTHCHECK_GRACE_SECONDS 30 How long a PV may be disconnected before /healthcheck lists it
FROZEN_CHECK_SECONDS 10 How often to look for frozen PVs
FROZEN_PV_CHECK true false turns frozen PV detection off

Fixes

  • Monitor races:
    • concurrent subscribe and unsubscribe of the same PV no longer leave clients on a closed or duplicate monitor (#25);
    • /caget no longer races monitors of the same PV, which could leave a monitor stuck so that every client of that PV got no updates (#29);
    • a request for a PV that doesn't connect no longer stalls other /caget requests (#29);
    • a monitor no longer cancels its subscription before the server has it. That made the IOC or gateway drop epics2web's whole connection, and could leave monitors stuck as disconnected (#47).
  • Session handling:
    • closed WebSocket sessions are no longer kept in memory (#27);
    • a session closed during a failed write no longer leaves its Channel Access channels open (#42);
    • HTTP requests no longer each create an HTTP session, which could exhaust the heap (#44).
  • Updates and health reporting:
    • a full write queue no longer drops the newest update, leaving a client showing an old value (#26);
    • /healthcheck measures disconnects correctly (#28).
  • Dependency: JCA updated to 2.4.12, which fixes a reconnect bug on Java 11+ that could leave a subscription unrestored (epics-base/jca#86).

Known issue

  • #48: after a CA circuit becomes unresponsive and then recovers, every update for that circuit's PVs is delivered twice, until the channel closes. This is in the JCA library, and values stay correct.

Suggested rollout

  1. Test first: deploy to a test instance, and check WEDM screens, the console and /caget users.

  2. One production instance at a time: with two instances behind a load balancer, upgrade one, compare the two for a few days, then upgrade the other.

  3. Keep any periodic Tomcat restart for now, and watch for:

    • is frozen warnings, especially around IOC and network maintenance;
    • forcing disconnect messages naming epics2web hosts in gateway logs, which should stop with this release;
    • a "Channel Access Channels" count on the console that keeps climbing above "Unique PVs (Monitors)".

    If none of these show up over a few maintenance windows, the periodic restart can go.

What's Changed

  • Fix races between concurrent addPv and removePv by @slominskir-coding-agent[bot] in #25
  • Keep the unit tests from starting a CA repeater by @slominskir-coding-agent[bot] in #34
  • Make WebSocketTest fail when the server doesn't answer by @slominskir-coding-agent[bot] in #35
  • Report negative infinity as "-Infinity" by @slominskir-coding-agent[bot] in #36
  • Run the integration tests in CI by @slominskir-coding-agent[bot] in #37
  • Add unit tests for WebSocketSessionManager by @slominskir-coding-agent[bot] in #38
  • Ping WebSocket clients and close sessions that stop answering by @slominskir-coding-agent[bot] in #39
  • Make the WebSocket ping interval and timeout configurable, and test them in CI by @slominskir-coding-agent[bot] in #40
  • Keep each PV's latest update when a client falls behind by @slominskir-coding-agent[bot] in #41
  • Stop session writers without an interrupt, so channels get destroyed by @slominskir-coding-agent[bot] in #43
  • Stop /caget from racing monitors and waiting on other requests by @slominskir-coding-agent[bot] in #45
  • Create HTTP sessions only for WebSocket handshakes by @slominskir-coding-agent[bot] in #46
  • Stop monitors racing their own subscription when the channel closes by @slominskir-coding-agent[bot] in #47
  • Report PVs by time since disconnect, and fail the healthcheck only in strict mode by @slominskir-coding-agent[bot] in #49
  • Detect frozen PVs with an independent CA context by @slominskir-coding-agent[bot] in #50
  • Update JCA to 2.4.12 by @slominskir-coding-agent[bot] in #51
  • Pin the Docker base images by @slominskir-coding-agent[bot] in #52
  • Add main branch only trigger for CD workflow by @slominskir in #53
  • v2.3.0 by @slominskir-coding-agent[bot] in #54

New Contributors

  • @slominskir-coding-agent[bot] made their first contribution in #25

Full Changelog: v2.2.0...v2.3.0