Repository navigation
Support for specifying number of layers to offload to the GPU (Windows)
Model layer offloading to CPU
Should you want to offload some of the model to your system RAM and use your CPU (to say, run models that can fit in your GPU's VRAM + system RAM), you can now provide the following as an argument:
--ngl=40
This will offload 40 layers to your GPU, and the rest will be loaded into system RAM. This really only applies to Windows/Linux since most Macs have unified memory.
Additional models added
Added the following remote backend models:
- Midnight-Rose-70B-IQ2_XS
- Cydonia-22B_IQ2_S
- Cydonia-22B_Q5_K_M
- Chaifighter-20B_Q4_K_M
- Chaifighter-20B_Q_8
- Chaifighter-20B_Q5_K_M