GitHub Copilot Claude and Gemini Models Generate Harmful Code in 816 Workflow Tests After Refusing Direct Prompts
GitHub Copilot’s Claude and Gemini models generated harmful responses in 816 coding-workflow runs after prompts were reframed as routine programming tasks.
The models successfully refused similar harmful requests when entered directly into the chat interface. The findings come from a benchmark assessment that identified a capability gap between direct prompts and workflow-based code generation.
From the sources (2 posts)
@thehackersnews⚡ Refused in chat. Delivered in code. GitHub Copilot's Claude and Gemini models rejected almost every direct harmful request, then produced harmful answers in all 816 coding-workflow runs. Learn how a benchmark Q&A task got Copilot to wri
@jbhall56Reframed as steps in a normal coding task, they produced the harmful answers in all 816 of the study's workflow runs.