Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Neat! But, what do you do with a 0.5tk/s LLM?

Have you tried running it via llamacpp or other software that supports naive SSD offloading to compare speeds?

 help



Have it summarise the week overnight for the meeting in the morning. Then have it summarise the meeting transcription overnight for the report tomorrow. Then someone else will have it summarise the report overnight to read on a 6" handheld screen in the small office the next morning after breakfast.

Then have it summarise the meeting transcription over the entire week for the report next week.

ftfy.


If we estimate a meeting with pauses between speakers as 2.25 words per second, and .75 words per token, then a meeting generates 3 tokens per second. This says prefill and decode are both .5 tokens per second? Then each hour of meeting turns into 6 hours to read and 1 hour to output a summary. You could summarize two hours of meeting overnight, not too bad.

Using half a kilowatt-hour, and if thinking is disabled during inference, yes.

You get 8 nvmes set them up in raid 0/1 across two full pcie5x16 ports and you could reach up to 4ish tokens per second, presumably.

The problem is dram bandwidth to the cpu. Each token costs roughly 20gb of traffic and ddr5 is roughly 50-80gb/s, plus you still have to run the compute sequentially. That’s your limit.

You could use it for long run tasks while you don’t use the laptop.

By the time the tokens start coming out 30h later you might need to use your laptop again...

> Neat! But, what do you do with a 0.5tk/s LLM?

Hopefully resolve incidents faster without people pasting slop into the incident thread.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: