Sakana AI and Nvidia Preview TwELL to Speed Up Large Language Models
AI lab Sakana AI and Nvidia announced a research preview introducing TwELL, a new sparse packing format designed to compress large language models (LLMs) for optimized GPU computing. Developed in collaboration ahead of the International Conference on Machine Learning (ICML) in Seoul, the framework integrates custom CUDA kernels with tiled GPU workloads. According to the companies, the approach demonstrates over 20% speedups for training and inference at billion-parameter model scales, alongside higher savings in peak memory and energy consumption.
The announcement highlights Sakana AI's broader research presence at ICML, where the Tokyo-based lab will present 11 papers covering sparse transformer models, test-time scaling, long-term memory architectures, and masking diffusion methods. TwELL underscores a coordinated industry push to develop hardware-aware compression techniques that reduce compute costs and power draw as AI models grow in size and deployment.
From the sources (3 posts)
@sakanaailabsSakana AI is heading to #ICML2026 in Seoul (July 6–11)! 🐟🇰🇷 Our team will present 11 papers spanning multi-agent coordination, sparse and efficient LLMs, test-time scaling, long-term memory, and agent benchmarks. A thread of everything we'
@sakanaailabs"Sparser, Faster, Lighter Transformer Language Models" will be presented at #ICML2026 Paper: Blog: In collaboration with @NVIDIAAI, we introduce TwELL, a new sparse packing format designed t
@sakanaailabs@NVIDIAAI "UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action Branching" will be presented at #ICML2026 Paper: We introduce UnMaskFork, a test-time scaling framework for Masked Diffusion La