Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
marcelroed
19 days ago
|
parent
|
context
|
favorite
| on:
GigaToken: ~1000x faster Language model tokenizati...
Time to first token refers to the time until the model outputs one token, which includes the time to process the entire prompt (doing prefill). The GPU time per token is much lower when doing prefill, so the significance of tokenization is higher.
dingdingdang
19 days ago
[–]
Have you done preliminary numbers on replacing tokenizer on, say, llama-server?
marcelroed
18 days ago
|
parent
|
next
[–]
Added numbers here:
https://news.ycombinator.com/item?id=49015014
marcelroed
18 days ago
|
parent
|
prev
[–]
Running the numbers now
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: