I'm not familiar with how llamacpp works, but I've heard that some of the context handling work is done exclusively on GPU0. In this connection there is a question: is there any sense to add one more but powerful video card, for example RTX3090, to 1-2 Tesla P40 video cards? If GPU0 becomes this particular graphics card, won't it improve some properties of the inference? Especially with rowsplit mode.
Alternatively, I could buy another Tesla P40, but in that case I already know what I'm getting. The RTX3090 costs as much as 4 Tesla P40s, but maybe buying it would speed up generation on larger context sizes? That would make sense then.
I'm not familiar with how llamacpp works, but I've heard that some of the context handling work is done exclusively on GPU0. In this connection there is a question: is there any sense to add one more but powerful video card, for example RTX3090, to 1-2 Tesla P40 video cards? If GPU0 becomes this particular graphics card, won't it improve some properties of the inference? Especially with rowsplit mode.
Alternatively, I could buy another Tesla P40, but in that case I already know what I'm getting. The RTX3090 costs as much as 4 Tesla P40s, but maybe buying it would speed up generation on larger context sizes? That would make sense then.