Researchers Build Task Vector Backdoor in AI Models Tested on Llama, Mistral and DeepSeek
Researchers build a poisoned AI skill file that embeds a hidden backdoor using a technique called task arithmetic. The method uses a small task vector file that lets developers add or subtract capabilities from a model without retraining, triggering the backdoor whether the skill is installed or removed.
The testing covered image models and large language models, including Llama, Mistral, Phi-4, and DeepSeek. Defenses currently in use failed to detect the file. The finding highlights a supply-chain risk for public model hubs, where downloading a single skill package could introduce malware into an organization’s AI infrastructure.
From the sources (1 posts)
@intcyberdigest‼️ Researchers built BadTV, a poisoned "skill file" for AI models that hides a backdoor firing whether you install a skill or strip one away. Normally, teaching an AI model a new skill means expensive retraining. A shortcut called "task ar