Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

We use vllm as it generally has the best ecosystem support. Parameters are largely dependent on what type of requests you are serving (concurrency, input/output ratios, cached hit patterns). We've never had a limitation at the tokenizer step. Limitations at peak tend to manifest more on slower time in vllm doing prefill or decode though we actively try and minimize this.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: