Together AI Launches Reserved Inference Capacity for Frontier Models With 99% Uptime SLA
Together AI is rolling out Provisioned Throughput, a service providing reserved inference capacity for frontier open models. The offering utilizes token-based pricing and includes a 99% uptime service level agreement, allowing customers to secure dedicated compute with guaranteed capacity instead of variable serverless allocations.
Users can access the new tier through MiniMax M3 and GLM-5.2 models, with the company claiming up to a 90% cost reduction compared to OpenAI’s Opus 4.8. The launch extends the provider’s enterprise infrastructure offerings alongside existing pay-as-you-go options.
From the sources (1 posts)
@togethercomputeWe're introducing Provisioned Throughput: reserved inference capacity for frontier open models, with token-based pricing and a 99% uptime SLA. Serverless simplicity, guaranteed capacity, up to 90% lower cost vs. Opus 4.8. Get started with