Is this still being developed or maintained? #263
Sophist-UK
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
It has been quite a while since this repo has had any significant updates (other than typo fixes) and the AI world is moving almost at the speed of light.
New models implement new inference techniques and e.g. llamacp gets updated to improve performance and handle these new model internals in a cleverer way.
Is this solution still relevant? Does it need any technical updates or is it still perfect doing the job?
Does it only support 70B models in 4GB vRAM or is it more flexible about model and vRAM size?
What models does it now support? Qwen3.5? Minimax2.5?
All reactions