Replies: 1 comment 3 replies
|
Wait this is cool ! But comfyui does the same thing. Isn't it? |
3 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi, i was working on similar idea.
But it was about diffusion models.
https://github.com/Jit-Roy/WeeLLM
It runs large diffusion models with as little as 4 GB of VRAM, without any quantization. It dynamically determines how many layers can fit within the available VRAM and streams the text encoder and transformer layers to the GPU layer by layer, enabling inference on hardware with limited VRAM. It supports both safetensors and GGUF models.
All reactions