Repository navigation
Upgrade notes
This release hardens the WebSocket monitor, /caget and /healthcheck. It fixes several ways PVs could get stuck after IOC, network or gateway disruptions, and adds detection of PVs that freeze anyway.
Action needed when upgrading
- Nagios, or any monitoring that alerts on
/healthcheck: use/epics2web/healthcheck?strict=true. The default/healthchecknow answers 200 whenever the server is up, even if PVs are disconnected. That's so a load balancer doesn't take every instance out when an IOC goes down. Strict mode answers 503 when a PV that was connected has stayed disconnected past the grace period. - Load balancers: keep using
/epics2web/healthcheck. - Tomcat: tested with Tomcat 11.0.26, which the Docker image now pins. That version includes WebSocket fixes relevant to this release, such as write timeouts covering whole messages. Upgrading from earlier 11.0.x releases is recommended.
Behaviour changes
- Unresponsive clients are closed. The server pings WebSocket clients every 30 s, and closes a client that sends no message and no pong for 60 s, or whose writes stay blocked for 20 s. Browsers answer pings automatically, and
epics2web.jsreconnects. - Clients that fall behind get the latest values. A client that falls behind gets each PV's latest update instead of a backlog. Connect and disconnect (
info) messages are never dropped. /cagetvalidates requests:- it answers HTTP 400 for more than 500 PVs in one request, or for a
jsonpcallback that isn't a JavaScript name such asa.b_c; - a PV that doesn't connect now fails only the request that asked for it, with an error naming the PV.
- it answers HTTP 400 for more than 500 PVs in one request, or for a
/healthcheckreports more accurately:- it counts the grace period from when a PV disconnected, not from its last value change;
- entries gain a
statefield; - PVs that never connected are listed, as
CONNECTING, but don't fail strict mode; - frozen PVs are listed with
frozen,frozen_minutesandfrozen_reason; ?frozen=trueanswers 503 when any PV is frozen.
- No session cookies: responses no longer set a
JSESSIONIDcookie, except for WebSocket handshakes. - Console changes: the console lists connected sessions that don't monitor anything yet, and shows the number of open Channel Access channels.
-Infinityis now reported with its sign.
Frozen PV detection (new, report-only)
epics2web now looks for PVs whose monitor has stopped working while the IOC still serves them. It gives suspicious PVs a short-lived subscription in a second, independent CA context, which opens its own connection to each IOC or gateway and probes at most 20 PVs at a time.
- A frozen PV is logged as
WARNING ... PV <name> is frozen: <reason>. - It's listed by
/healthcheck. - Through a gateway, it finds problems between epics2web and the gateway, not inside the gateway.
It doesn't change the default or strict healthcheck responses.
New settings (environment variables, all optional)
| Variable | Default | Meaning |
|---|---|---|
WEBSOCKET_PING_INTERVAL_SECONDS |
30 | How often clients are pinged |
WEBSOCKET_TIMEOUT_SECONDS |
60 | Close a client with no message or pong for this long |
WEBSOCKET_SEND_TIMEOUT_SECONDS |
20 | Close a client whose write blocks this long (Tomcat only) |
HEALTHCHECK_GRACE_SECONDS |
30 | How long a PV may be disconnected before /healthcheck lists it |
FROZEN_CHECK_SECONDS |
10 | How often to look for frozen PVs |
FROZEN_PV_CHECK |
true |
false turns frozen PV detection off |
Fixes
- Monitor races:
- concurrent subscribe and unsubscribe of the same PV no longer leave clients on a closed or duplicate monitor (#25);
/cagetno longer races monitors of the same PV, which could leave a monitor stuck so that every client of that PV got no updates (#29);- a request for a PV that doesn't connect no longer stalls other
/cagetrequests (#29); - a monitor no longer cancels its subscription before the server has it. That made the IOC or gateway drop epics2web's whole connection, and could leave monitors stuck as disconnected (#47).
- Session handling:
- Updates and health reporting:
- Dependency: JCA updated to 2.4.12, which fixes a reconnect bug on Java 11+ that could leave a subscription unrestored (epics-base/jca#86).
Known issue
- #48: after a CA circuit becomes unresponsive and then recovers, every update for that circuit's PVs is delivered twice, until the channel closes. This is in the JCA library, and values stay correct.
Suggested rollout
-
Test first: deploy to a test instance, and check WEDM screens, the console and
/cagetusers. -
One production instance at a time: with two instances behind a load balancer, upgrade one, compare the two for a few days, then upgrade the other.
-
Keep any periodic Tomcat restart for now, and watch for:
is frozenwarnings, especially around IOC and network maintenance;forcing disconnectmessages naming epics2web hosts in gateway logs, which should stop with this release;- a "Channel Access Channels" count on the console that keeps climbing above "Unique PVs (Monitors)".
If none of these show up over a few maintenance windows, the periodic restart can go.
What's Changed
- Fix races between concurrent addPv and removePv by @slominskir-coding-agent[bot] in #25
- Keep the unit tests from starting a CA repeater by @slominskir-coding-agent[bot] in #34
- Make WebSocketTest fail when the server doesn't answer by @slominskir-coding-agent[bot] in #35
- Report negative infinity as "-Infinity" by @slominskir-coding-agent[bot] in #36
- Run the integration tests in CI by @slominskir-coding-agent[bot] in #37
- Add unit tests for WebSocketSessionManager by @slominskir-coding-agent[bot] in #38
- Ping WebSocket clients and close sessions that stop answering by @slominskir-coding-agent[bot] in #39
- Make the WebSocket ping interval and timeout configurable, and test them in CI by @slominskir-coding-agent[bot] in #40
- Keep each PV's latest update when a client falls behind by @slominskir-coding-agent[bot] in #41
- Stop session writers without an interrupt, so channels get destroyed by @slominskir-coding-agent[bot] in #43
- Stop /caget from racing monitors and waiting on other requests by @slominskir-coding-agent[bot] in #45
- Create HTTP sessions only for WebSocket handshakes by @slominskir-coding-agent[bot] in #46
- Stop monitors racing their own subscription when the channel closes by @slominskir-coding-agent[bot] in #47
- Report PVs by time since disconnect, and fail the healthcheck only in strict mode by @slominskir-coding-agent[bot] in #49
- Detect frozen PVs with an independent CA context by @slominskir-coding-agent[bot] in #50
- Update JCA to 2.4.12 by @slominskir-coding-agent[bot] in #51
- Pin the Docker base images by @slominskir-coding-agent[bot] in #52
- Add main branch only trigger for CD workflow by @slominskir in #53
- v2.3.0 by @slominskir-coding-agent[bot] in #54
New Contributors
- @slominskir-coding-agent[bot] made their first contribution in #25
Full Changelog: v2.2.0...v2.3.0