Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Agreed. It blew my mind to learn that tiling matrix multiplication is more efficient than untiled, but that you can only show tiled is faster in a hierarchical memory model (for tiled, you have arithmetic intensity that scales with size of your fast memory rather than a fixed constant). Yet in uniform memory model both are O(n^3).


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: