Skip to content

fix(kibana): fail-fast on permanent errors in readiness check - #173

Merged
Oddly merged 1 commit into
mainfrom
fix/kibana-healthcheck-fail-fast
Aug 4, 2026
Merged

fix(kibana): fail-fast on permanent errors in readiness check#173
Oddly merged 1 commit into
mainfrom
fix/kibana-healthcheck-fail-fast

Conversation

@Oddly

@Oddly Oddly commented Aug 4, 2026

Copy link
Copy Markdown
Owner

Closes #127.

The Kibana readiness probe in roles/kibana/tasks/main.yml and roles/kibana/tasks/restart_and_verify_kibana.yml was doing a straight 60 × 5 s HTTP wait — five minutes of "still not ready?" regardless of whether the failure was transient or terminal. A wrong kibana_system password, an unreachable Elasticsearch, or a startup FATAL all took the full five minutes to surface, which is a lot to wait through for a config typo.

Extend the shell probe with two short-circuits before the HTTP check: if systemctl is-active says the unit is dead, exit 2; if journalctl -u kibana --since 2 minutes ago shows a FATAL line, Unable to connect to Elasticsearch, Authentication.*failed, or Kibana server is not ready.*fatal, also exit 2. failed_when: rc == 2 drops Ansible straight into the rescue block, which prints the journal — same rescue path as before, just reached in seconds instead of minutes.

main.yml already had the systemctl half; I added the journal grep there and gave restart_and_verify_kibana.yml the same treatment (it had neither). Success path is untouched, so no existing scenario regresses.

Closes #127.

The Kibana readiness check in two places blindly retried 60x5s = 5
minutes regardless of the failure reason. A misconfigured Kibana that
can't reach Elasticsearch, has the wrong kibana_system password, or
died on a startup error takes the full five minutes to surface a
useful error, which is a lot of latency for a config typo.

Extend the shell probe with two short-circuits before the HTTP check:
if systemctl says the service is dead, exit 2 immediately; if
journalctl shows a FATAL line, an 'Unable to connect to Elasticsearch',
an authentication failure, or a 'not ready ... fatal' line in the
last two minutes, exit 2 immediately. rc=2 is treated as a permanent
failure by 'failed_when', so Ansible drops straight into the rescue
block and prints the journal instead of waiting out the retries.

main.yml already had the systemctl part; add the journal grep to it
and give restart_and_verify_kibana.yml the same treatment (it had
neither). Both files now handle the same set of exit codes with the
same semantics.
@Oddly Oddly added the ci:run Trigger gated pull request CI label Aug 4, 2026
@coderabbitai

coderabbitai Bot commented Aug 4, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@Oddly, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 50 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: b77bc2c7-3681-4d7f-8f62-c3a79f0b34d1

📥 Commits

Reviewing files that changed from the base of the PR and between e8375ed and 296ed1c.

📒 Files selected for processing (2)
  • roles/kibana/tasks/main.yml
  • roles/kibana/tasks/restart_and_verify_kibana.yml

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Oddly
Oddly enabled auto-merge (squash) August 4, 2026 07:22
@github-actions github-actions Bot removed the ci:run Trigger gated pull request CI label Aug 4, 2026
@Oddly
Oddly merged commit 59911bb into main Aug 4, 2026
9 of 11 checks passed
@Oddly
Oddly deleted the fix/kibana-healthcheck-fail-fast branch August 4, 2026 07:22
Oddly added a commit that referenced this pull request Aug 4, 2026
Follow-up to #173. ansible-lint is right that the journalctl | grep
pipe in the Kibana readiness probe should honour \$pipefail — without
it, a broken journalctl still lets grep see an empty pipe and exit 1,
so we'd fall through to the HTTP check on a machine where journalctl
is truly broken. In practice we already had \`|| true\` on the curl
call to guard against that, but the linter is enforcing the general
rule and the fix is a one-liner.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Kibana health check should fail fast on permanent errors

1 participant