Auto-refreshing/health-checking stale auto/* candidate pools #8551
Replies: 5 comments 1 reply
|
Hey @manchairwang! Good observation — when connections at the front of an OmniRoute already has a 3-layer resilience system that partially addresses this:
What you are describing is a proactive health-check (background pinging to evict stale candidates before they block requests) vs the current reactive approach (evict on failure). That is a meaningful improvement. I have opened #8668 to track this. The existing |
|
Sounds like a solid plan, @manchairwang -- the reset-window/headroom combo strategies should get you most of the way there without a local DB, but building your own health/quota tracker on top is a reasonable belt-and-suspenders approach too. Let us know if you hit anything specific building it out. |
|
Follow-up @manchairwang - I opened #9950 to track refresh and health-check behavior for stale |
|
Follow-up -- this is already tracked and shipped in #9950 (refresh and health-check for auto/* candidate pools). |
|
@manchairwang A correction: my last reply said a refresh and health-check for
There's no periodic refresh or health probe built specifically for |
Uh oh!
There was an error while loading. Please reload this page.
Issue: Auto/* candidate pool has many connections. If some connections in front of the queue failed, the auto/* would keep trying them, which blocks auto/* service.
Suggestion: auto-refresh, or health-checking background, by a set time interval.
All reactions