Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide (amd.com)
56 points by matt_d 12 hours ago | hide | past | favorite | 5 comments
 help



As much as I enjoy these articles and for AMD to write more light technical articles, it really feels constrained, even strained, to be unable to cite the equivalent terms from the precursor here (NVIDIA). Another batch of jargon for very similar architectures and programming models... HIP and ROCm have actually made amazing strides in making CUDA developers' porting work easy, and I know playing catchup to a (monopolist) moving target you have no power over is bad... but I feel this is part of the thousand paper cuts.

So when can I buy MI350s for my homelab?

Hopefully the MI350P is available soon, at last a standard PCIe SKU, if a bit too-much previous-generation and gimped compared to the MI350X

MI350 happens to not fit DeepSeek V4 Flash (144GB vs 180+ needed) so it will be dead on arrival.

For production loads? No one will run DS without expert parallelism so it’s not important for model to fit on one gpu.



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: