Skip to content

DAOS-19381 placement: tolerate DOWNOUT spare in remap chain - #18737

Closed
wangshilong wants to merge 1 commit into
masterfrom
shilongw/DAOS-19381
Closed

DAOS-19381 placement: tolerate DOWNOUT spare in remap chain#18737
wangshilong wants to merge 1 commit into
masterfrom
shilongw/DAOS-19381

Conversation

@wangshilong

Copy link
Copy Markdown
Contributor

A rebuild resumed by rebuild start can run an older rebuild version against a newer pool map. In that history a failed DOWN/DRAIN shard can select a spare candidate which has already transitioned to DOWNOUT.

That state is valid for placement remap chaining: the selected spare is not a usable final target, but the existing code already updates the failed shard state and re-adds it to the remap list so placement can look for the next candidate. The assertion that rejects a DOWN/DRAIN failed shard remapping through a DOWNOUT spare is therefore too strict and can abort rebuild scans during interactive rebuild stop/start recovery.

Replace the assertion with placement debug logging and keep the existing chained-remap behavior. Add a jump-map placement unit test that drives a DOWN shard through a spare candidate that becomes DOWNOUT and verifies placement continues to a different rebuild target.

Steps for the author:

  • Commit message follows the guidelines.
  • Appropriate Features or Test-tag pragmas were used.
  • Appropriate Functional Test Stages were run.
  • At least two positive code reviews including at least one code owner from each category referenced in the PR.
  • Testing is complete. If necessary, forced-landing label added and a reason added in a comment.

After all prior steps are complete:

  • Gatekeeper requested (daos-gatekeeper added as a reviewer).

A rebuild resumed by rebuild start can run an older rebuild version
against a newer pool map.  In that history a failed DOWN/DRAIN shard can
select a spare candidate which has already transitioned to DOWNOUT.

That state is valid for placement remap chaining: the selected spare is
not a usable final target, but the existing code already updates the
failed shard state and re-adds it to the remap list so placement can look
for the next candidate.  The assertion that rejects a DOWN/DRAIN failed
shard remapping through a DOWNOUT spare is therefore too strict and can
abort rebuild scans during interactive rebuild stop/start recovery.

Replace the assertion with placement debug logging and keep the existing
chained-remap behavior.  Add a jump-map placement unit test that drives a
DOWN shard through a spare candidate that becomes DOWNOUT and verifies
placement continues to a different rebuild target.

Signed-off-by: Wang Shilong <shilong.wang@hpe.com>
@github-actions

Copy link
Copy Markdown

Ticket title is 'Aurora engine crash - determine_valid_spares() Assertion'
Status is 'Open'
https://daosio.atlassian.net/browse/DAOS-19381

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant