Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Kimi K3 is in no way going to fit in 96GB RAM * 16 units (1536GB) unless badly quantized and with a small amount of context.


Why not? The model is 2.8T parameters with native MXFP4, which is 1400GB.


I would figure it needs a lot of kV cache and context size which can be much bigger than the model itself




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: