Repository navigation
Hot Aisle bare metal (experimental)
dstack can now provision Hot Aisle bare metal servers with 8xMI300X in addition to VMs. Bare metal is disabled by default; enable it with bare_metal: true in the backend settings:
projects:
- name: main
backends:
- type: hotaisle
team_handle: hotaisle-team-handle
creds:
type: api_key
api_key: ...
bare_metal: trueAnd the bare metal offers become available:
$ dstack offer --backend hotaisle
# BACKEND RESOURCE TYPE PRICE
1 hotaisle (us-michigan-1) cpu=x86:104 mem=2048GB bm-mi300x-8 $27.12
disk=11
gpu=MI300X:192GB:8Bare metal servers are prepaid for 8 hours. Use a fleet with fixed nodes and delete it when done:
type: fleet
name: hotaisle-bm
nodes: 1
backends: [hotaisle]
instance_types: [bm-mi300x-8]Within the 8 hours, dstack doesn't force-terminate bare metal servers: the instance stays terminating for about 15 minutes, then dstack stops tracking it and the server must be terminated manually in Hot Aisle. Controlled by DSTACK_FF_HOTAISLE_BARE_METAL_NO_FORCE_RELEASE (enabled by default, 0 disables).
What's Changed
- Add Hot Aisle bare metal support by @peterschmidt85 in #4325
- Correctly support NVIDIA B300 in Vast.ai by @jvstme in #4330
- Update Azure GRID driver to 580.178.04 by @peterschmidt85 in #4332
- [Docs] Update vLLM PD example by @un-def in #4336
- Fix pyright with kubernetes 37 by @r4victor in #4347
- Skip placement groups for AWS Capacity Blocks by @r4victor in #4345
- Warn dstack skill about heredocs split across commands items by @r4victor in #4348
- Bound SSHTunnel control commands with timeouts and BatchMode by @r4victor in #4354
- Use Vultr backend credentials when fetching GPU offers by @peterschmidt85 in #4335
- Reconnect gateway replica SSH tunnels after process exit by @peterschmidt85 in #4355
- Fix gateway recovery after high request load by @peterschmidt85 in #4351
Full Changelog: 0.22.2...0.22.3