Skip to content

Support for specifying number of layers to offload to the GPU (Windows)

Choose a tag to compare

@Proryanator Proryanator released this 20 Nov 22:13

Model layer offloading to CPU

Should you want to offload some of the model to your system RAM and use your CPU (to say, run models that can fit in your GPU's VRAM + system RAM), you can now provide the following as an argument:

--ngl=40

This will offload 40 layers to your GPU, and the rest will be loaded into system RAM. This really only applies to Windows/Linux since most Macs have unified memory.

Additional models added

Added the following remote backend models:

  • Midnight-Rose-70B-IQ2_XS
  • Cydonia-22B_IQ2_S
  • Cydonia-22B_Q5_K_M
  • Chaifighter-20B_Q4_K_M
  • Chaifighter-20B_Q_8
  • Chaifighter-20B_Q5_K_M