I’ve got a few 4090s that I’m planning on doing this with. Would appreciate even...

andersa · on April 12, 2024

The split is done automatically by the inference engine if you enable tensor parallelism. TensorRT-LLM, vLLM and aphrodite-engine can all do this out of the box. The main thing is just that you need either 4 or 8 GPUs for it to work on current models.

renewiltord · on April 12, 2024

Thank you! Can I run with 2 GPUs or with heterogeneous GPUs that have same RAM? I will try. Just curious if you already have tried.

andersa · on April 12, 2024

2 GPUs works fine too, as long as your model fits. Using different GPUs with same VRAM however, is highly highly sketchy. Sometimes it works, sometimes it doesn't. In any case, it would be limited by the performance of the slower GPU.

renewiltord · on April 12, 2024

All right, thank you. I can run it on 2x 4090 and just put the 3090s in different machine.