Skip to content

How Polling Works in Depth

hilman2 edited this page Sep 2, 2026 · 2 revisions

How polling works in depth

You do not need this page to use the integration. The README's How polling works is the short version, and Troubleshooting in detail covers the symptoms. This page is for when you want to know why the integration behaves the way it does, or you are about to change it.

The code lives in __init__.py (the coordinator) and api.py (the client wrapper). Both carry long comments at the decisions described here.

One session, held open

The integration opens one Modbus TCP session per config entry and keeps it. Every poll reuses it. This is the opposite of what the integration did until v0.22.0 and of what its predecessor does, both of which disconnected after every poll so a single-slot inverter would be free for other readers in between.

The measurement that changed it, on a KACO Powador 7.8 TL3 at a 30 s interval:

Strategy Result
Reconnect per poll 5 of 6 cycles failed
One session held open 20 of 20 cycles succeeded, 1.6 s each

The device is not flaky. An embedded Modbus stack rebuilding a TCP session every 30 seconds is the thing it copes with worst, and Modbus TCP was designed around a session that stays up.

The session is handed back after each poll in exactly two cases: the Release the Modbus connection between polls option is on, or a second config entry talks to the same host and port. The second case is detected, not configured.

The gateway lock

Every coordinator takes a lock keyed by (host, port) for the whole of a poll. Two entries with different unit IDs behind one IP therefore take turns rather than racing, which is what SolarEdge leader and follower setups and Fronius inverter plus meter need. An entry alone on its host never waits, because nobody else holds its lock.

Writes take the same lock. Before v0.14.0 they did not, and a write could run on the same socket while a poll was mid-read. Since v0.31.0 the transport checks the Modbus transaction id, so such an interleave would end in two timeouts rather than in swapped answers, but the lock is what prevents it in the first place.

The lock is released while the integration sleeps between a failed cycle and its retry, so the other entry on the gateway can poll in the meantime. A write waits up to 120 s for the lock before failing with a readable error; that number is sized for two entries on one slow gateway, where a first connect with a full scan can legitimately hold the lock for a long time.

A poll cycle, step by step

  1. Get the model list. From the live session if there is one. Otherwise connect, then either restore the layout from the cache (below) or walk the model chain.
  2. Compare with last time. A block that was there before and is missing now is noted. If the scan stopped early and no cached layout could stand in, the cycle fails on purpose instead of serving a short list; a short list would drop entities to unknown on a cycle that counted as success, and the stale-data tolerance would never get a chance to hold the last value (#42).
  3. Read the common block once per process, for the device card.
  4. Check the device identity against the stored layout: manufacturer, model and serial number. The firmware version is deliberately not part of it (#49). A changed version is answered with one rescan on the next connect rather than with a failed setup.
  5. Read the nameplate once, from block 120 (WRtg) or 121 (WMax), for the plausibility filter's default ceiling.
  6. Read every polled block. The polled set is what you ticked, plus block 124 whenever the inverter has a battery and blocks 123 and 704 while the export beta is on, intersected with what the device has.
  7. Close the session only if this entry hands the slot back.

On success the failure counters reset, the Repairs issues clear, the layout is saved, and every sensor gets its new value. New blocks that appeared since the last cycle get entities immediately; blocks that vanished keep their entities on the last value.

The layout cache

pysunspec2 only learns a device's blocks by scanning: read the base address, then walk the chain reading an id and a length per block, with a pacing sleep after each one. On a device with 20 blocks that is more than 40 round trips plus 10 s of sleep.

The integration captures the result, base address plus id, address and length per block, and rebuilds the model objects from it on the next connect with three short validating reads instead of a scan. The layout is also written to Home Assistant's storage per config entry, so a restart does not scan either, which is where the scan hurt most: inside the setup timeout, on the slowest devices.

Rules that follow from what a scan can and cannot do:

  • No expiry. A rescan cannot confirm a layout, it can only replace it, and it is the one operation that can silently produce a wrong one. Freshness comes from validating both ends of the chain on every connect, not from a clock.
  • Dropped after a failed cycle. A layout read at addresses that just stopped answering is the one thing not to reuse. The retry connects fresh and scans.
  • Never captured from a partial scan. A truncated chain has a well-formed tail and would validate cleanly while pinning the missing blocks out of existence.

When a scan stops on a Modbus exception partway through the chain (a block the device refuses to serve), the last known good layout is used if it still validates. Without one, the blocks found so far are used for this cycle and the cache stays empty, so the next connect tries again. A scan that stops on a timeout fails the cycle outright: a request went out and nothing came back, so nothing below can tell which block is missing.

Failure handling

A cycle that fails after the integration is running gets one retry: release the lock, flag the client for a rebuild, sleep 5 s, try again. The very first refresh during setup skips the retry so a wrong IP fails fast through Home Assistant's own backoff.

If the retry fails too, the cycle counts as failed, and the entities decide their availability from the count of consecutive failed cycles, not from Home Assistant's last_update_success. Up to 5 consecutive failures the sensors stay available on their last value, which is about three minutes at the default interval. The distinction matters because Home Assistant stops notifying entities after the first consecutive failure; the coordinator notifies them itself when the counter crosses the threshold, and at that moment last_update_success still reads true. An available property that consulted it would never flip.

Error categories and the Repairs panel

Every error from pysunspec2 is translated into one of four categories at the API boundary. The category decides when a Repairs issue appears.

Category Examples Repairs issue after
transport connection refused, socket timeout, CRC error 3 consecutive
protocol no SunSpec base address, chain unreadable 1
device Modbus exception code, value overflow 3 consecutive
transient a one-shot read timeout, an incomplete scan never

The diagnostics download keeps the last errors per category and the current consecutive count per category.

The transport issue has one more gate since v0.28.0 (#52). Most PV inverters power their communication board down at night. If the last successful poll reported an operating state of OFF, SLEEPING, SHUTTING_DOWN or STANDBY, the disconnect that follows is expected and raises nothing; the per-poll retry warning drops to debug as well. The Inverter powers down when idle option forces that decision for devices that vanish without reporting a state, or have none. Protocol and device errors still escalate either way: a device that answers with an error is awake.

A block that was seen once and has been missing from every scan for 600 s wall-clock time raises its own Repairs issue suggesting the device be removed. Wall-clock, not cycles, because a healthy device never rescans and a cycle counter would never advance.

Minimum interval

The options form refuses intervals below 5 s. Not to be gentle on the device: Home Assistant's own coordinator stores an interval of 0 as None and stops polling silently and permanently, with no entity ever going unavailable, and a negative interval schedules a refresh in the past and fires continuously. 5 rather than 1 because the in-cycle retry delay is already 5 s and pysunspec2 sleeps 0.6 s inside every block read, so below that the retry sets the real pace, not the option.

All points of a config entry share the interval. There is no per-point interval; a control loop that needs a fast grid reading is better served by a fast meter than by polling a whole inverter every second.

Writes

A Number, Switch or Select entity writes its point under the gateway lock, then requests a refresh outside the lock so the entity shows what the inverter actually accepted rather than what was sent. The sunspec2.set_export_limit action writes the percentage and the enable flag in one lock hold, because "set the limit and turn it on" is one operation and a poll slipping in between would leave the device with a new percentage and the old enable state.

Where a device has both block 123 and 704, the entities and the action both use 704. Two controls for one physical setting would confuse even if they agreed, and on the one device with evidence they did not.

The embedded pysunspec2

Since v0.30.0 the SunSpec client library is embedded as a fork rather than installed from PyPI. The changes against upstream, and the reason for forking, are recorded in its __init__.py. The one that affects polling: Modbus TCP requests are numbered and a response with another transaction id is dropped as the late answer to an earlier request. Upstream sends id 0 in every frame and takes whatever comes back, which is how a late answer used to land in the wrong register and produce a 10 kW reading on a 9 kW inverter.

Clone this wiki locally