In version 0.9.17 it says:
Improved memory management for gpt-oss:20b. Increase the chance to run on 32GB system (note: NPU can access <50% of total RAM)
However, when testing on my machine, it fails to run with the following error message:
[XRT] ERROR: Failed to submit command to hw queue (0xc01e0200):
Even after the video memory manager split the DMA buffer, the video
memory manager could not page-in all of the required allocations into
video memory at the same time. The device is unable to continue.
FastflowLM version: 0.9.19
Driver version: 32.0.203.314
Total System RAM: 32GB
GPU UMA buffer size: 512MB
NPU accessible Shared memory size: 15.5GB
In version
0.9.17it says:However, when testing on my machine, it fails to run with the following error message:
FastflowLM version: 0.9.19
Driver version: 32.0.203.314
Total System RAM: 32GB
GPU UMA buffer size: 512MB
NPU accessible
Shared memorysize: 15.5GB