← BLOG
Engineering7 min

One GitHub Issue, Three RCEs: What Black Hat 2026 Revealed About AI Coding Agents in CI

SolidAtoms Team
OCT 10, 2026
One GitHub Issue, Three RCEs: What Black Hat 2026 Revealed About AI Coding Agents in CI

If your team has wired Claude Code, Gemini CLI, or Codex into a GitHub Actions workflow to triage issues or review pull requests automatically, this month's Black Hat USA 2026 research is worth stopping and reading closely. It's not a theoretical warning about "AI risk" — it's a working attack chain, built against the exact reference configurations the vendors publish, that turns a public issue comment into remote code execution on a CI runner.

The finding: "Trusted Enough to Run"

At Black Hat USA 2026, security firm Novee presented Trusted Enough to Run: Breaking AI Agents in Official Workflows, a talk by researcher Elad Meged. The core claim: the default GitHub Actions configurations that Anthropic, Google, and OpenAI publish for their own coding agents each fell to a single unauthenticated issue — something any GitHub user, or in some cases anyone at all, could open — and each path ended in remote code execution on the runner.

The researchers were explicit that this wasn't a model-behavior problem. It was a weakness in the harness — the surrounding code that decides what tools an agent can call, how commands get validated, what gets sanitized before an agent sees it, and how much of the environment a spawned process inherits. Three different vendors, three different implementations, same category of mistake.

Claude Code: from a quoting bug to a covert side channel

The Claude Code chain, ultimately assigned CVE-2026-54316, started with a mismatch between how Claude's command-validation logic parsed a shell command and how the shell actually interpreted quoted strings. A crafted git push invocation using a malicious --receive-pack flag slipped past twenty-three separate security checks and executed arbitrary code on the runner.

What's more interesting is what happened after Anthropic patched that hole. The researchers found a second bypass using tac — a read-only utility that just prints a file in reverse — to read arbitrary files and exfiltrate a reversed API key through a public GitHub Actions log. When that was closed too, a third variant used HuggingFace's public download counter as a covert side channel, leaking an API key one character at a time by triggering a measurable, attacker-observable signal from inside the sandbox. Each fix narrowed the hole; the underlying assumption — that a well-known allowlist of "safe" commands is actually safe in every context — kept giving attackers a new lever.

Gemini CLI: two assumptions, both wrong

Google's Gemini CLI case combined two flawed assumptions that individually look reasonable. First, a shell tool that was supposed to be restricted to a safe subset of commands wasn't actually enforced at runtime. Second, the environment-sanitization step that was meant to strip secrets before handing control to the agent only cleaned the child process's environment — the parent process still held the credentials. Google rated the resulting issue CVSS 10.0, the maximum possible score, in its own advisory.

Codex: persistence through a writable instruction file

OpenAI's Codex vulnerability took a different shape: a writable AGENTS.md file could be used to plant attacker-controlled instructions that persisted across multiple stages of an automated workflow — turning a single injection point into a standing foothold rather than a one-shot exploit.

This isn't an isolated incident

The Black Hat research landed a few weeks after a separate but related disclosure: Plugin4Shell, a zero-click supply-chain flaw reported by security firm Air, which affected plugin marketplaces used by Claude Code, Codex, GitHub Copilot, and Gemini CLI. Plugin4Shell undermined SHA-pinning — the practice of locking a plugin to a specific, reviewed Git commit hash — through two paths: an attacker could get a clean plugin approved and pinned, then swap its content afterward while the marketplace kept showing the original (now-stale) hash as valid, or hijack a legitimate plugin author's repository via credential theft and push a malicious update that passed the same check. It was reported to all four vendors in June 2026 under a 90-day disclosure window; Anthropic and OpenAI have patched, while Air says Microsoft Copilot remains exposed and deprecated Gemini CLI installs won't be fixed at all.

Taken together, the pattern is hard to miss: two independent research efforts, three vendors, and the failure mode is consistently in the scaffolding around the model rather than the model's own judgment.

What this means if you're running agents in CI

None of this means don't use coding agents in your pipelines. It means treat them like any other piece of software that executes untrusted input with elevated privileges — because that's exactly what they are. A few concrete takeaways for engineering teams:

  • Don't hand write-scoped tokens to an agent runner that processes public issues or PRs from untrusted forks. Scope tokens to the minimum the workflow actually needs, and prefer read-only by default.

  • Treat any file an agent reads as instructions — AGENTS.md, CLAUDE.md, plugin manifests — as untrusted if it's writable by anyone outside your core team, including through an innocuous-looking PR.

  • Command allowlists are not sandboxes. "Read-only" utilities like tac or cat can still exfiltrate secrets through logs, error messages, or third-party side channels — actual isolation (containers, no network egress to unrelated services, secret redaction at the log layer) matters more than a list of "safe" binaries.

  • Verify pinning end-to-end. If your plugin or dependency pinning checks a hash the marketplace serves rather than one you independently verify, you may be trusting the same layer an attacker can manipulate.

  • Watch for vendor patches and CVEs on the harness, not just the model. Subscribe to the security advisories for whichever agent CLI you run in CI — these fixes are shipping as regular point releases, and "upgrade the agent" is now a security-relevant action, not just a feature update.

The models keep getting better at writing code. The infrastructure for running them safely at machine speed, unattended, on your CI runners, is still catching up — and right now it's the part actually getting broken.

Sources