I built Galahad looking for people to test it on one GPU #29855
corbenicai
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I have build Galahad. It saves the model’s KV cache to disk so it can reuse it after a restart, instead of processing the same context again. It works with llama.cpp, vLLM and SGLang.
I’d love a few people to try it on their own setup. It’s free for non-commercial use on one GPU, on Linux x86-64.
What works? What breaks? Does it help with your workload? All feedback is welcome.
All reactions