fix(kibana): fail-fast on permanent errors in readiness check - #173
Conversation
Closes #127. The Kibana readiness check in two places blindly retried 60x5s = 5 minutes regardless of the failure reason. A misconfigured Kibana that can't reach Elasticsearch, has the wrong kibana_system password, or died on a startup error takes the full five minutes to surface a useful error, which is a lot of latency for a config typo. Extend the shell probe with two short-circuits before the HTTP check: if systemctl says the service is dead, exit 2 immediately; if journalctl shows a FATAL line, an 'Unable to connect to Elasticsearch', an authentication failure, or a 'not ready ... fatal' line in the last two minutes, exit 2 immediately. rc=2 is treated as a permanent failure by 'failed_when', so Ansible drops straight into the rescue block and prints the journal instead of waiting out the retries. main.yml already had the systemctl part; add the journal grep to it and give restart_and_verify_kibana.yml the same treatment (it had neither). Both files now handle the same set of exit codes with the same semantics.
|
Warning Review limit reached
Next review available in: 50 minutes You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Follow-up to #173. ansible-lint is right that the journalctl | grep pipe in the Kibana readiness probe should honour \$pipefail — without it, a broken journalctl still lets grep see an empty pipe and exit 1, so we'd fall through to the HTTP check on a machine where journalctl is truly broken. In practice we already had \`|| true\` on the curl call to guard against that, but the linter is enforcing the general rule and the fix is a one-liner.
Closes #127.
The Kibana readiness probe in
roles/kibana/tasks/main.ymlandroles/kibana/tasks/restart_and_verify_kibana.ymlwas doing a straight 60 × 5 s HTTP wait — five minutes of "still not ready?" regardless of whether the failure was transient or terminal. A wrongkibana_systempassword, an unreachable Elasticsearch, or a startup FATAL all took the full five minutes to surface, which is a lot to wait through for a config typo.Extend the shell probe with two short-circuits before the HTTP check: if
systemctl is-activesays the unit is dead, exit 2; ifjournalctl -u kibana --since 2 minutes agoshows aFATALline,Unable to connect to Elasticsearch,Authentication.*failed, orKibana server is not ready.*fatal, also exit 2.failed_when: rc == 2drops Ansible straight into therescueblock, which prints the journal — same rescue path as before, just reached in seconds instead of minutes.main.ymlalready had thesystemctlhalf; I added the journal grep there and gaverestart_and_verify_kibana.ymlthe same treatment (it had neither). Success path is untouched, so no existing scenario regresses.