← BLOG
Tools7 min

The Coding Agent That Rates Its Own Risk: Inside OpenHands' Safety-First Playbook

SolidAtoms Team
OCT 10, 2026
The Coding Agent That Rates Its Own Risk: Inside OpenHands' Safety-First Playbook

Most of the coding-agent conversation in 2026 is still about capability: which model closes more GitHub issues, which agent tops SWE-bench this month. OpenHands — the open-source, MIT-licensed agent project from All Hands AI — has spent the last year quietly arguing that the harder problem is different: once an agent can run arbitrary shell commands, edit any file, and open pull requests without a human approving each step, the question stops being "how smart is it" and becomes "how do you let it run unsupervised without it wrecking something." Its answer, baked into the rewritten Software Agent SDK it shipped in late 2025, is a genuinely interesting piece of engineering: an agent that scores the danger of its own actions before it takes them.

What OpenHands actually does

OpenHands is an autonomous software engineering agent: point it at an issue or a task description and it writes code, runs the test suite, debugs failures, and submits a pull request end-to-end. In November 2025, the team published "The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents", an arXiv paper describing a from-scratch rewrite of the underlying runtime. That SDK is now the foundation for OpenHands' CLI, its cloud product, and third-party agents built on top of it — and it's what most of the security machinery described below actually lives in.

An agent that grades its own actions

The part worth paying attention to is the default LLMSecurityAnalyzer. Rather than bolting on a separate model to review actions after the fact, OpenHands adds a security_risk field directly to every tool's JSON schema. That means the same model call that decides to run rm -rf build/ or edit a config file also has to self-report how dangerous that action is — LOW, MEDIUM, HIGH, or UNKNOWN — as part of the structured output, with no extra inference call and no separate reviewer model in the loop.

That risk score feeds into a separate confirmation policy layer. A policy like ConfirmRisky() lets a developer decide what happens next: auto-approve everything inside a disposable sandbox, but require a human click-through for anything flagged MEDIUM or HIGH once the same agent is pointed at a production repo. It's a deliberately simple mechanism — the model is judge of its own actions, not an outside critic — and it's an honest bet that a capable model can be trusted to flag its own risky behavior more cheaply and more often than a bolted-on filter would catch it.

Sandboxing underneath the model

The self-scoring is a soft layer; underneath it OpenHands still assumes the model will be wrong sometimes and backs that up with infrastructure. Every agent action runs inside a Docker sandbox with configurable CPU and memory limits per container, executes as a non-root user, and is boxed off from the host filesystem by default. The pitch to self-hosters is explicit in how the project documents it: you can hand the agent real repository access without handing it your machine.

  • Per-container CPU and memory caps, so a runaway agent loop can't take down the host

  • Non-root execution inside the sandbox, closing off a large class of container-escape and privilege-escalation paths

  • A plugin system for extending what tools an agent can call, instead of granting broader default permissions

The benchmark number, with the caveat it deserves

It's worth being precise here rather than repeating a rounder-sounding number that's floating around: OpenHands' own SOTA claim on SWE-bench Verified, published in April 2025, was 60.6%, achieved by running the agent multiple times with Claude 3.7 Sonnet at sampling temperature 1.0, generating several candidate patches per issue, filtering out ones that failed regression and reproduction tests, and using a critic model to pick the best surviving trajectory. That's over a year old at this point, it relied on a heavier multi-sample inference setup rather than one clean pass, and the field has moved since — so treat it as a data point about the technique, not a current leaderboard position. The project itself now publishes an ongoing "OpenHands Index" tracking model performance inside its own harness rather than a single static score, which is a more honest way to present agent benchmarks that shift every time a new frontier model ships.

Why this is the more interesting story right now

OpenHands isn't alone in treating safety scaffolding as a feature rather than an afterthought. Cline's recent releases moved plan-mode switching out of the model's hands entirely — a human now has to explicitly leave plan mode — and hard-block file-editing shell commands while planning. GitHub Copilot Workspace has gone the other direction structurally, splitting a task across separate implementation, testing, and documentation agents that coordinate rather than one agent doing everything with maximal permissions. Different designs, same underlying pressure: as these tools graduate from autocomplete to unsupervised PR-opening, the interesting engineering has quietly shifted from "can the model solve the bug" to "what happens the ten percent of the time it's wrong, at 2am, with shell access."

If you're evaluating agent frameworks for anything beyond a sandboxed demo, the questions worth asking are less about which one wins this week's benchmark and more about the ones OpenHands is answering by default: what does the agent do with a command it isn't sure about, what's actually stopping it from touching things outside its lane, and who signs off before it merges.

Sources