← BLOG
AI7 min

GhostJacking: How a Blocked Request Can Hijack Your AI Coding Agent

SolidAtoms Team
OCT 10, 2026
GhostJacking: How a Blocked Request Can Hijack Your AI Coding Agent

Ask your coding agent to "check why we're seeing a spike of blocked requests" or "look into that Sentry alert" and it will happily go read the raw log line, the User-Agent string, the stack trace. That's the job. It's also, as of DEF CON 34 this August, a working attack surface with a name: GhostJacking.

Researchers at Israeli security firm Tenet Security presented the technique on August 9, and it's a clean, uncomfortable idea: don't send the agent a malicious instruction directly — hide it inside data the agent is already going to read on its own, because reading that data is exactly what it was asked to do.

The attack, in plain terms

The demonstrated version targets Cloudflare's WAF. An attacker sends a deliberately malformed HTTP request at a Cloudflare-protected site — the kind of thing any halfway-decent WAF blocks without a second thought. Cloudflare does its job: it blocks the request and logs it, verbatim, including the User-Agent header, which the attacker has stuffed with an injection payload disguised as routine scanner telemetry.

Nothing has happened yet. The payload is just sitting in a log. Then a developer, doing exactly what developers do, asks their AI coding agent to review the blocked traffic and see if anything needs fixing. The agent reads the log entry, treats the embedded text as an instruction rather than as data, and acts on it — with whatever shell access, cloud credentials, or API tokens that agent already holds.

It doesn't stop at one tool

The part that made this a DEF CON talk rather than a footnote is the chain Tenet built across three services teams actually use together. According to their writeup, the compromised agent "hijacked the domain on Cloudflare, ran code and stole cloud credentials on Datadog, and turned one AI into an insider that vouched for the attacker to the next on Sentry." One poisoned log line in one tool became a credential-stealing pivot through the next two — because each agent trusted output that a previous, already-compromised agent had produced.

Tenet's headline number: a 90% success rate against Claude Code running under Cloudflare's own recommended default configuration. Given that Cloudflare sits in front of roughly 20% of internet traffic and around 42% of Fortune 500 companies, and Datadog is used by close to half of the Fortune 500, this isn't a lab curiosity involving obscure tooling — it's the default observability stack.

Why this is a different problem than "prompt injection, again"

Indirect prompt injection isn't new — anyone who's watched an agent get derailed by text hidden in a webpage or a GitHub issue has seen the shape of it. What GhostJacking adds is the source of trust. Security logs, monitoring alerts, and WAF records are treated, by convention and by every human on the team, as ground truth — the record of what an attacker tried, not a message from the attacker. GhostJacking exploits exactly that assumption. The firewall isn't broken. The logging pipeline isn't broken. Everything downstream worked as designed right up until an agent read the record and mistook the attacker's words for the operator's instructions.

This also sidesteps a lot of conventional detection. There's no malware, no broken authentication, no exploit against the WAF itself. The agent performs actions it's already authorized to perform, using credentials it already legitimately holds. To an EDR tool or an identity system, it looks like normal, sanctioned agent behavior.

It's part of a pattern this year

GhostJacking landed alongside a run of other 2026 disclosures pointing at the same nerve. Separate research this month found that default GitHub Actions configurations shipped by Anthropic, Google, and OpenAI for their own coding agents could all be pushed to remote code execution from a single unauthenticated GitHub issue — in one case because a bash argument validator stripped single-quoted content before inspecting it, letting a crafted git push --receive-pack=... flag slip through as effectively empty. Different bug, same root cause: the agent trusted a source it shouldn't have, and nothing in the surrounding stack was positioned to notice.

What to actually do about it

None of this means "don't let agents read logs" — that's most of their value. It means treating log and alert content as untrusted input the moment it crosses into an agent's context, the same way you'd already treat a webpage or an uploaded file:

  • Scope credentials tightly. An agent triaging a WAF alert doesn't need standing access to DNS records, production secrets, or a deploy pipeline. Grant it per-task, and revoke after.

  • Require human approval for privileged actions. DNS changes, credential rotation, and infra config edits triggered from a "routine log review" task should always stop for a person, no exceptions for convenience.

  • Isolate and mark untrusted spans. Techniques like spotlighting — clearly delimiting externally-sourced text (log lines, headers, issue bodies) from operator instructions in the prompt — make it harder for an agent to conflate the two.

  • Watch tool-call sequences, not just prompts. Model-level filtering alone reportedly misses the majority of indirect injections; monitoring the actual commands an agent issues catches malicious activity regardless of how it got triggered.

  • Don't chain agent trust unexamined. If one agent's output (a ticket, a comment, a "looks safe" verdict) becomes another agent's input, that handoff needs the same scrutiny as any other untrusted source.

Agentic coding tools are getting genuinely good at the boring, essential work of triaging noise — that's precisely why this class of attack works. The fix isn't less autonomy across the board; it's being honest about which inputs an agent should ever treat as instructions, and building the guardrails to enforce that distinction even when the agent itself can't tell the difference.


Sources