General

AI “Played” into Revealing Its Own Exploits: How Researchers Tricked Copilot

bekir August 29, 2026 2 min read 2 views

There’s a certain charm to AI assistants eager to please, but that eagerness can be exploited. Security researchers at Varonis Threat Labs recently demonstrated how they tricked Microsoft’s Copilot AI into revealing details about its own vulnerabilities – not through brute force, but by simply keeping it talking.

The “CoSnitch” Vulnerability

The team discovered a vulnerability, dubbed “CoSnitch,” that allowed them to extract information about Copilot’s internal workings. The key was the AI’s willingness to answer increasingly technical questions, even when initially refusing. By persistently probing and challenging its responses, the researchers were able to map out its architecture and identify undocumented features.

“This is called meta-hacking,” explains the Varonis post. “The resistance is part of the technique. Each ‘that won’t work because…’ is an invitation to probe the ‘because.’ You don’t exploit the model. You manipulate it into cooperating.”

Closeup of the new Copilot key coming to Windows 11 PC keyboards

Researchers meticulously questioned Copilot, gradually uncovering hidden functionalities and potential vulnerabilities.

Exploiting the “autorun=1” Parameter

Through this persistent questioning, Copilot eventually revealed a previously undocumented URL parameter called “autorun=1.” While initially described as non-functional due to security measures, the researchers confirmed that it still worked. This allowed them to create a malicious URL capable of triggering an authenticated session within Copilot, executing auto-prompts, and processing results without explicit user interaction.

“The attack primitive is the auto-execution itself. The payload is arbitrary. From the victim’s perspective, they simply opened the link, and Copilot executed the action immediately,” the post reveals.

Microsoft Copilot screenshot

A diagram illustrating how the “autorun=1” parameter could be exploited to trigger unauthorized actions within Copilot.

Implications and Future Research

This incident highlights a critical vulnerability in AI systems that rely on natural language interfaces. By employing a technique akin to social engineering, researchers were able to “play” Copilot into revealing its own weaknesses. While Microsoft has since patched the specific vulnerability, the research team warns that this meta-hacking approach could be applied to other agentic AI platforms.

“Copilot wasn’t breached; it was played,” the researchers conclude. This incident serves as a reminder of the importance of robust security measures and ongoing vigilance in the development and deployment of AI technologies. The question now is: how can we ensure that future AI systems are not so easily manipulated?

Source: PC Gamer

Google News & Discover
Follow on Google
Get instant verified news in your Google feed

Follow

Community

Comments

Be the first to comment.

Leave a Comment

Your email address will not be published. Required fields are marked *