inference.
Low-latency inference at scale.
Serve models on nodes close to your users, autoscaling with traffic. You pay per token, not per idle server.
p50 latency41ms
throughput1,900 tok/s
autoscaletrue
billingper token
Bring your own model or serve from the catalog. Either way, a proof lands for every batch, and you're billed for exactly the compute you touched.