Selective mode saving #596

AAnoosheh · 2025-11-21T13:45:05Z

What does this PR do?

Type of change: ? Bug fix

Overview: Filter out KD state from ModelOpt state list when saving. This allows for applying the KD mode after a modelopt checkpoint restore without it complaining that it was already applied previously.

Usage

# Add a code snippet demonstrating how to use this

Testing

Before your PR is "Ready for review"

Make sure you read and follow Contributor guidelines and your commits are signed.
Is this change backward compatible?: Yes/No
Did you write any new necessary tests?: Yes/No
Did you add or update any necessary documentation?: Yes/No
Did you update Changelog?: Yes/No

Additional Information

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

codecov · 2025-11-21T15:24:54Z

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 74.46%. Comparing base (38550b0) to head (d4883ed).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

@@           Coverage Diff           @@
##             main     #596   +/-   ##
=======================================
  Coverage   74.45%   74.46%           
=======================================
  Files         182      182           
  Lines       18250    18255    +5     
=======================================
+ Hits        13588    13593    +5     
  Misses       4662     4662

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:

❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

realAsma

Approving the PR to unblock merging for the release.
Attaching the teacher model design was not a good choice looking back - it is creating lot of headaches. We should just deal with distillation with the loss function or trainer.

We dont need to necessarily need to remove the mtd.convert - but stop maintaining it and use a trainer/loss function based distillation support -

QAT + distillation is a critical piece going forward - It would be great to simplify things - If any design choices turned out problematic - we dont have to keep patching it - we could redirect our energy for a better design.

Cc @kevalmorabia97 @jenchen13 @ChenhanYu

AAnoosheh added 2 commits November 21, 2025 04:03

Skip KD state restore altogether

5a188c0

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

Update test

9a9e4b3

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

AAnoosheh self-assigned this Nov 21, 2025

AAnoosheh requested review from a team as code owners November 21, 2025 13:45

AAnoosheh requested a review from realAsma November 21, 2025 13:45

AAnoosheh changed the title ~~Aanoosheh/selective restore~~ Selective Mode restoration Nov 21, 2025

Update other tests

df1e622

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

AAnoosheh force-pushed the aanoosheh/selective-restore branch from 6f61222 to 929d07a Compare November 21, 2025 15:06

Change to skip during save instead of restore

d4883ed

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>

AAnoosheh force-pushed the aanoosheh/selective-restore branch from 929d07a to d4883ed Compare November 21, 2025 15:10

AAnoosheh changed the title ~~Selective Mode restoration~~ Selective mode saving Nov 21, 2025

realAsma approved these changes Nov 21, 2025

View reviewed changes

AAnoosheh merged commit 01e24fd into main Nov 21, 2025
27 checks passed

AAnoosheh deleted the aanoosheh/selective-restore branch November 21, 2025 17:30

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Selective mode saving #596

Selective mode saving #596

Uh oh!

AAnoosheh commented Nov 21, 2025 •

edited

Loading

Uh oh!

codecov bot commented Nov 21, 2025 •

edited

Loading

Uh oh!

realAsma left a comment •

edited

Loading

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

Selective mode saving #596

Selective mode saving #596

Uh oh!

Conversation

AAnoosheh commented Nov 21, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What does this PR do?

Usage

Testing

Before your PR is "Ready for review"

Additional Information

Uh oh!

codecov bot commented Nov 21, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Codecov Report

Uh oh!

realAsma left a comment • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

AAnoosheh commented Nov 21, 2025 •

edited

Loading

codecov bot commented Nov 21, 2025 •

edited

Loading

realAsma left a comment •

edited

Loading