Skip to content

Enable Strix Halo profiling support - #1

Open
cane122 wants to merge 1 commit into
umarinkovic:mainfrom
cane122:strix-halo-profiling
Open

Enable Strix Halo profiling support#1
cane122 wants to merge 1 commit into
umarinkovic:mainfrom
cane122:strix-halo-profiling

Conversation

@cane122

@cane122 cane122 commented Jun 11, 2026

Copy link
Copy Markdown
  • Add persistent vLLM and Triton cache mounts to reduce compile overhead
  • Fix path handling to support execution from repository root
  • Enable GPU config auto-generation via generate_gpu_yaml.sh
  • Improve orchestrator reliability (absolute paths, python execution)
  • Disable Qwen3-VL on Strix Halo due to OOM during profiling
  • Adjust model configuration for safer memory usage
  • Update README and gitignore for improved usability and clarity

- Add persistent vLLM and Triton cache mounts to reduce compile overhead
- Fix path handling to support execution from repository root
- Enable GPU config auto-generation via generate_gpu_yaml.sh
- Improve orchestrator reliability (absolute paths, python execution)
- Disable Qwen3-VL on Strix Halo due to OOM during profiling
- Adjust model configuration for safer memory usage
- Update README and gitignore for improved usability and clarity
@cane122
cane122 force-pushed the strix-halo-profiling branch from bf56b4e to bd3339b Compare June 11, 2026 11:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant