Have you tried to write a kernel for basic matrix multiplication? Because I have...

woooooo · 2025-09-08T14:08:49 1757340529

Well, CUDA gives you a whole programming language where you have to figure out the optimization for your particular card's cache size and bus width.

I'm saying the API surface of what to offer for LLMs is pretty small. Yeah, optimizing it is hard but it's "one really smart person works for a few weeks" hard, and most of the tiling techniques are public. Speaking of which, thanks for that blog post, off to read it now.

kadushka · 2025-09-08T16:33:04 1757349184

it's "one really smart person works for a few weeks" hard

AMD should hire that one really smart person.

adgjlsfhk1 · 2025-09-08T16:55:27 1757350527

yeah they really should. the primary reason AMD or behind in the GPU space is that they massively under-prioritize software.

astrange · 2025-09-09T06:14:52 1757398492

Not having written one of these (…well I've written an IDCT) I can imagine it getting complicated if there's any known sparsity to take advantage of.

hedgehog · 2025-09-08T18:09:44 1757354984

I assure you from experience that it's more than a smart person for a few weeks.