Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm somewhat dubious about anything talking about low level performance programming at the instruction level that doesn't distinguish between latency and throughput, never mind mention the incredibly out-of-order nature of modern desktop/server class CPU cores.


That's a very important point. For instance on Intel's CPUs multiplication is pipelined - its latency is 3 cycles, but throughput is 1 cycle. Thus completing N multiplication takes 2 + N cycles (in the best case), not 3 * N.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: