Not suggesting a provider but if you're willing to get a Chinamod GPU, go get RTX 3080 20G or RTX 2080 Ti 22G. Get a couple of them and you can probably run Qwen3.8-27B at a reasonable speed.
I've got a RTX 3080 20G for $450 like a year ago. With llama.cpp, bf16 kv cache, kv cache offloaded to RAM, Qwen3.8 27B UD-Q4_K_XL, single RTX 3080 20G, I got 10 tok/s initially and it dropped to 5 tok/s at 50k context. I'm thinking of getting another Chinamod GPU.
Building a dual Chinamod GPU machine would only cost like $1000~$1500. It's not too bad compared with the alternatives!
Or, theres the option of a 32gb v100 (about $600USD on taobao etc), where you can get 1200 of prefill at 80 of decode: https://github.com/geoffwatts/ninfer-v100 - that's _really_ cheap inference, and it's not a modified card - you just need to add a blower or water block.
I've got a RTX 3080 20G for $450 like a year ago. With llama.cpp, bf16 kv cache, kv cache offloaded to RAM, Qwen3.8 27B UD-Q4_K_XL, single RTX 3080 20G, I got 10 tok/s initially and it dropped to 5 tok/s at 50k context. I'm thinking of getting another Chinamod GPU.
Building a dual Chinamod GPU machine would only cost like $1000~$1500. It's not too bad compared with the alternatives!