Skip to content

v0.0.1-gptq-llama-cuda

Choose a tag to compare

@1b5d 1b5d released this 23 Apr 21:59
· 20 commits to main since this release

Support Llama based models inference on GPU using GPTQ-for-llama