DeepSeek V4 Pro on Together AI Ranks No. 1 on Artificial Analysis for Output Speed, Latency
DeepSeek V4 Pro on Together AI is now ranked No. 1 on Artificial Analysis for both output speed and latency, according to Together AI.
Together AI said serving V4 well is an inference-systems problem involving KV cache, prefix reuse, kernels and endpoint profiles, and tied the ranking to that systems work.
From the sources (4 posts)
@vipulvedDeepSeek V4 Pro on @togethercompute becomes #1 on both latency and speed.
@tri_daoRT @vipulved: DeepSeek V4 Pro on @togethercompute becomes #1 on both latency and speed.
@togethercomputeRT @vipulved: DeepSeek V4 Pro on @togethercompute becomes #1 on both latency and speed.
@togethercomputeDeepSeek V4 Pro on Together AI is now #1 on Artificial Analysis for both output speed and latency. Serving V4 well is an inference systems problem: KV cache, prefix reuse, kernels, and endpoint profiles. We break down the systems work her