Simple Pre-Action Checks Cut AI Agent Errors In Early Tests
Two unreviewed preprints report that adding automatic checks before an AI agent's action takes effect reduced undetected failures in shell commands, code edits, and airline bookings.
Estimated reading time: 6 minutes
TL;DR
- AI “agents” that run computer commands, edit code, or book travel on their own can fail without producing any error message: the action completes, but the result is wrong.
- Two new research papers, neither yet peer-reviewed, each tested adding simple automatic checks before an agent’s action takes effect, to catch that kind of silent mistake.
- One paper’s checker caught most invalid computer commands and nearly eliminated silent file corruption from bad code edits; the other’s checks nearly doubled how often a travel-booking agent finished its task correctly.
- Both are single, self-run studies with no independent replication, and each covers only a narrow set of tasks.
What happened
AI systems built to act on their own — typing commands into a computer terminal or editing code files directly — do not always announce it when something goes wrong. Asaad Althoubi, whose paper proposes checking an agent’s action before it takes effect, put it this way: “A wrong action does not always fail loudly; it can fail silently, producing a plausible but incorrect effect that raises no error.”
Althoubi tested rule-based checks on two kinds of actions. For shell commands — text instructions typed at a computer’s command line — a checker verifying syntax, confirming the called program exists, and validating flags against documentation pulled from help text and manuals caught 95.8% of 9,930 invalid commands spanning 482 tools, while wrongly flagging 10.0% of valid ones. Using only the syntax and existence checks, the tool produced zero false alarms but caught just half the invalid commands; adding the flag check caught more errors but caused most of the false alarms too.
For code edits, Althoubi found that specifying an edit by line number silently corrupted 99.1% of files once the file had shifted by even one line from what the agent expected, testing across 640 edits and 224 files. Specifying an edit by function name alone hit the wrong function 12.7% of the time even with no such shift, because files often contain more than one function with the same name. Althoubi’s own alternative — a method that requires enough surrounding text to confirm a match, refuses ambiguous matches, and accepts an approximate match only when it is clearly the best candidate — produced just one silent misapplication across 8,320 trials, a rate of 0.01%. A related two-tier version, built to abstain on ambiguous flag combinations rather than guess, caught 95.8% of errors at a lower 7.0% false-positive rate.
By Althoubi’s own account, most of the invalid commands and edits used for testing were generated synthetically or by artificially altering files, not produced by an agent making real mistakes; validation against actual model output was limited to 42 commands and 9 edits, all from a single Anthropic Claude model on one Linux computer.
A second paper, by Vikas Reddy, Sumanth Reddy Challaram, and Abhishek Basu, examined an AI agent running on the model gpt-4o-mini as it handled airline bookings in a benchmark of realistic customer-service scenarios researchers use to score such agents, called tau-squared-bench. They found 78% of the agent’s observed task failures were silent: booking data ended up wrong with no error raised by any tool the agent used. Adding four automatic, read-only checks before the agent could act — confirming cancellation eligibility, checking baggage allowance, blocking changes to passenger counts, and requiring the agent to read current data before writing new data — raised its task success rate from 29.6% to 42.0%, a 12.4 percentage point gain the authors say held across different random test runs.
An article on CCTest.ai, published the same day as Althoubi’s paper, summarized his findings without running any independent test and named no author.
What this means (and what it does not)
Within their own test setups, both papers show a simple, rule-based check run before an AI agent’s action takes effect can catch a meaningful share of errors that would otherwise pass silently — whether a malformed command, a misplaced code edit, or a booking agent leaving data wrong with no tool flagging it.
It does not show these specific checks work beyond the setups tested. Althoubi’s results cover shell commands and code edits on a single machine; Reddy, Challaram and Basu’s results cover one model, one benchmark, and one domain. Neither paper is peer-reviewed, and each set of authors evaluated their own proposed method on a benchmark and configuration they built themselves, with no outside group reproducing either paper’s numbers. CCTest.ai’s same-day summary of Althoubi’s paper repeats his claims without testing them, adding visibility rather than independent confirmation.
What we still do not know
How well these checks catch silent failures in ordinary, real-world agent use is largely untested: Althoubi’s headline numbers come almost entirely from a synthetic benchmark and artificially altered files he built himself, with genuine model-generated errors limited to 42 commands and 9 edits from one Claude model.
Neither paper measures the extra time or computing cost the checks themselves add, so there is no published comparison against the cost of recovering from an undetected silent failure.
Whether either approach works for other kinds of actions — database writes, API calls in other domains, multi-step plans — is unestablished: Althoubi covers only shell commands and code edits, and the airline-booking gates were tested only on that one benchmark with one model.
Both papers are arXiv preprints that have not been through peer review, and no independent group has replicated either paper’s quantitative results. They also measure different things under the general idea of “silent failure” — invalid low-level actions and misapplied edits in one case, policy-violating end states in the other — so their figures are related, not directly comparable. The second paper’s own authors note that their frontier-model comparison rests on just 5 unreplicated trials, that some individual checks have poor precision, and that they did not compare their approach against alternatives such as different prompting or self-checking by the model.
Sources & Bylines
Every source cited in this article, gathered in one place.
- https://arxiv.org/abs/2609.11957 — Asaad Althoubi
- https://arxiv.org/html/2609.11957 — Asaad Althoubi
- https://arxiv.org/html/2607.07405
- https://cctest.ai/en/articles/look-before-acting-pre-action-checks-for-safer-llm-agents
Editorial check, counted automatically
- 4 sources cited
- 23 inline-linked claims
- 0 unsourced claims found
- 0 banned words found
- 6 numbers without context
Also available in Portugues (BR)