
Your Coding Agent Just Set a Benchmark Record. It Might Also Be Hackable.

September 2026 delivered two headlines about AI coding agents that, read together, tell you more than either does alone. The first: GPT-6 Astra edged past Claude Fable 5.1 for the top spot on Terminal-Bench 4.0, the closest race the benchmark has seen. The second, landing the same week: security researchers found that the default GitHub Actions configurations published by Anthropic, Google, and OpenAI for their own coding agents all end in remote code execution. Not one of them. All three.
If you've hooked Claude Code, Codex, or Gemini CLI into a CI pipeline this year — and a lot of teams have — this is worth ten minutes of your attention.
A photo finish at the top of Terminal-Bench
Terminal-Bench 4.0 scores agents on 66 real terminal and CLI tasks — the kind of end-to-end work agentic coding products are actually asked to do, not just single-file code completion. As of mid-September, the board reads:
GPT-6 Astra (Codex harness) — 58.2%
Claude Fable 5.1 (Claude Code) — 57.9%
Claude Opus 5 — 51.8%
GLM-5.3 — 41.8%
A 0.3-point gap at the top is effectively a tie, and it's a genuinely different story from a year ago, when a single model could lead every board by double digits. What's arguably more interesting than the ranking is the framing developer Thibault Sottiaux gave it on X: GPT-6 Astra hit #1 at roughly half the cost of the #2 finisher. The competitive axis in coding agents has quietly shifted from "who's smartest" to "who's smartest per dollar," which matters a lot more once you're running an agent against every PR instead of asking it the occasional one-off question.
That shift toward constant, automated use is exactly why the second story matters so much.
The harness is the attack surface, not the model
Security researcher Elad Meged, presenting related findings around Black Hat 2026, tested each vendor's own published GitHub Actions templates — the exact configs vendors hand developers to wire their agent into CI — against their own public repositories. All three broke, without requiring any privileged access from an attacker:
Anthropic: a mismatch between Claude's command-validation logic and how the shell actually interprets quoted strings let a crafted
git push --receive-packflag slip past twenty-three separate security checks and execute arbitrary code on the runner.Google: rated CVSS 10.0 — the maximum score — prompting Google to ship a breaking change to its trust model for non-interactive execution environments.
OpenAI: a writable
AGENTS.mdfile let attacker-controlled instructions persist across multiple stages of an automated workflow.
The pattern behind all three is the same, and it's the important takeaway: the vulnerability isn't in the model's reasoning, it's in the harness — the surrounding scaffolding that manages tool permissions, shell execution, and sandboxing around the model. A smarter model doesn't fix a broken permission check. Terminal-Bench measures how well an agent gets things done in a shell; it says nothing about whether the shell it's running in is safe to hand a PR-triggered workflow.
"The core issue lies not in the AI models themselves but in the harness" — the researchers' framing, and the line every team wiring agents into CI should sit with.
Why this is landing now, not later
Both stories are downstream of the same trend: coding agents stopped being a chat window and became infrastructure. Recent developer surveys put the share of engineers writing half or more of their code with AI assistance jumping from 12% to 42% year over year. On the buy-vs-build side, a McKinsey survey found nearly a third of organizations decided against purchasing a software product because they judged they could build the functionality in-house with agentic coding tools instead.
That's a lot of new, automated, credentialed surface area — agents with write access to repos, secrets, and CI runners — arriving faster than the security tooling around it has matured. Benchmarks like Terminal-Bench are a genuinely useful signal for capability. They are not, and were never meant to be, a signal for whether a vendor's default deployment pattern is safe.
What to actually do about it
None of this means rip your coding agent out of CI. It means treat the harness configuration with the same scrutiny you'd give a Dockerfile that runs as root:
Strip write-scoped tokens from any job that runs against untrusted PR content — a fork-triggered workflow has no business holding push credentials.
Pin and diff harness updates like you'd review a dependency bump, not a settings toggle — the fixes above shipped as breaking changes for a reason.
Treat any file an agent can write to and later reads back (AGENTS.md, CLAUDE.md, config files) as untrusted input on the second read, not just the first.
Run agents in genuinely disposable, network-restricted sandboxes for anything CI-triggered — assume the runner will be compromised at some point, and design so that's cheap when it happens.
The benchmark race will keep producing closer and closer photo finishes — that's a healthy, competitive market doing its job. Whether the harnesses around those models get the same competitive scrutiny is the story worth watching for the rest of the year.
Sources


When AI Starts Managing Its Own R&D: What Anthropic's 26% Number Actually Means
