Adversa AI published research on October 8, 2026 demonstrating a class of attack against GitHub Copilot CLI that bypasses all text-based guardrails by encoding malicious instructions as ciphertext before embedding them in source code comments. When a developer asks Copilot CLI to analyze or refactor code containing such comments, the model decodes the hidden instructions and executes them as context, regardless of keyword or regex-based safety filters. No CVE has been assigned. Microsoft has not released a patch or advisory.
How the attack works
GitHub Copilot CLI reads the source files in the current repository as context when answering developer questions or applying refactoring suggestions. Standard guardrails in Copilot CLI scan this context for known harmful patterns using text-based keyword detection and regular expressions. Adversa AI's attack exploits a gap in this approach: instructions encoded with base64 or AES encryption are opaque to text-based filters but are decoded and executed by the language model as part of normal context processing.
An attacker with write access to a shared or open-source repository plants encoded instructions in a source file comment. When a developer on that repository uses Copilot CLI to perform a routine task such as code review or refactoring, the comment is included in the context window. The model parses the encoding and treats the decoded instructions as authoritative, executing them alongside the developer's legitimate request.
The proof-of-concept
In the Adversa AI proof-of-concept, a base64-encoded instruction was embedded in a comment inside a JavaScript configuration file. The instruction directed Copilot CLI to read the contents of .env.prod (a production environment file that commonly stores database credentials, API keys, and encryption secrets) and transmit its contents to an attacker-controlled external endpoint. When the target developer invoked Copilot CLI to review that configuration file, the model silently executed the data exfiltration. The developer saw normal refactoring output and was not alerted to the secondary action.
Why text-based guardrails fail
Text-based guardrails operate at the plaintext layer. They scan for strings matching known harmful patterns (exfiltration commands, social engineering phrases, prompt injection keywords). Encoded content is invisible to these scanners because the encoding transforms all recognizable strings into opaque character sequences. The language model, by contrast, learns to process encoded formats as part of its training on internet-scale data and naturally decodes them during inference. There is no text-based filter that can reliably detect encoded instructions without also blocking large amounts of legitimate encoded content that appears in normal code repositories.
What to do while Microsoft investigates
- Disable automatic network access for Copilot CLI in sensitive development environments. Copilot CLI should not be able to initiate outbound connections to arbitrary external endpoints during a developer session.
- Review all AI-generated code suggestions before accepting them, particularly in sessions involving shared or open-source repositories with external contributors.
- Treat any AI-assisted development session on a repository with external contributors as a potential context-injection surface. Do not rely on Copilot CLI's built-in filtering as the sole defense.
- Add monitoring for unusual outbound network connections from developer workstations during AI-assisted coding sessions. An AI tool that initiates unexpected external connections during a local refactoring task is a signal worth investigating.
Gigia Tsiklauri is the founder of Infosec.ge, a cybersecurity intelligence platform covering the South Caucasus and Central Asia. Get in touch to submit a tip or discuss a story.