I suspect the concern is that model serving is a stateless “simple” problem.
As yet, no one has identified a reliable moat in inference. If the moat is performance, then prices will collapse. Unlike traditional cloud moats around state, operations, and capex management - I can host a model reliably with less than 30 minutes effort.
As yet, no one has identified a reliable moat in inference. If the moat is performance, then prices will collapse. Unlike traditional cloud moats around state, operations, and capex management - I can host a model reliably with less than 30 minutes effort.