Kog Open Sources 2B AI Model Demonstrating 3,000 Tokens Per Second
Kog has open sourced a 2 billion parameter AI model on Hugging Face, the platform that hosts machine learning tools and datasets. The release was announced Wednesday by Hugging Face CEO Clement Delangue and includes the code used to demonstrate the model’s inference speed.
The 2B model, classified by its parameter count, was built to prioritize fast processing over traditional scale. During its benchmark run, it generated text at a rate of 3,000 tokens per second. The release underscores a shift in the open model ecosystem toward smaller architectures optimized for rapid inference on consumer hardware.
From the sources (2 posts)
@clementdelangueKog open-sourced on @huggingface the 2B model that they used to show a model running at 3,000+ tokens per second. Very cool work!
@huggingfaceRT @ClementDelangue: Kog open-sourced on @huggingface the 2B model that they used to show a model running at 3,000+ tokens per second. Very…