OpenAI Introduces Smart Contract Benchmark for AI Agents as AI and Crypto Converge
Highlights
- OpenAI Smart Contract Benchmark tests AI Agents in Crypto for EVM contract bugs.
- EVMbench uses 120 real flaws from 40 audits, including Code4rena and Tempo work.
- GPT-5.3-Codex scored 72.2% in exploit mode, far above GPT-5 at 31.9%.
OpenAI has introduced a new smart contract security benchmark as AI agents gain stronger coding abilities in the crypto sector. Together with Paradigm, OpenAI said the benchmark, called EVMbench, tests how AI systems detect, patch, and exploit serious Ethereum contract bugs. Their effort responds to the growing financial risk, since smart contracts routinely secure over $100 billion in open-source crypto assets.
OpenAI Smart Contract Benchmark Targets Real Audit Vulnerabilities
In their release, OpenAI said EVMbench draws on 120 curated vulnerabilities collected from 40 professional smart contract audits. Notably, most of the issues came from open audit competitions, including Code4rena. OpenAI said the benchmark also includes vulnerability scenarios tied to security auditing work for the Tempo blockchain.
Tempo is described as a purpose-built Layer-1 network designed for high-throughput, low-cost stablecoin payments. Because of that, these scenarios extend the benchmark into payment-focused contract code. The company also said it expects agent-based stablecoin payment activity to grow.
To build the benchmark environments, OpenAI said it adapted existing exploit proof-of-concept tests and deployment scripts when available. However, it said engineers manually wrote missing components when no scripts existed. OpenAI added that it ensured patch tasks remained exploitable while still fixable without breaking compilation.
Detect, Patch, Exploit Modes Test AI Agents Under Pressure
OpenAI said EVMbench evaluates artificial intelligence agents in three modes. That is detect, patch, and exploit. In detect mode, agents audit smart contract repositories and get scored on recall of confirmed vulnerabilities and audit rewards. In patch mode, agents must modify vulnerable contracts while keeping intended functionality intact.
Exploit mode, however, focuses on full end-to-end fund draining attacks in a sandbox blockchain environment. The company said graders verify results using transaction replay and on-chain checks. To support reproducible evaluation, the company said it developed a Rust-based harness to deploy contracts and replay transactions deterministically.
Notably, the exploit tasks run in an isolated local Anvil environment instead of live crypto networks. It also said vulnerabilities used in the benchmark are historical and publicly documented. OpenAI added that the harness restricts unsafe RPC methods to limit abuse.
In exploit testing, OpenAI said GPT-5.3-Codex running via Codex CLI scored 72.2%. However, it said the earlier GPT-5 model scored 31.9%, despite being released just over six months earlier. OpenAI also noted that detect recall and patch success remain below full coverage.
OpenAI Adds New Talent with Agent Hire
While OpenAI pushed EVMbench into public view, it also expanded its agent development team. Notably, they hired Peter Steinberger, founder of the viral open-source AI agent project OpenClaw, previously known as Clawdbot. Sam Altman confirmed on X that Steinberger will join OpenAI to lead work on the “next generation of personal agents.”
Meanwhile, Altman said OpenClaw will transition into a foundation model project supported by OpenAI. The open-source project will continue under that structure, according to the announcement. The hiring drew wide attention as OpenAI increases its focus on autonomous and personal AI agents.
- Goldman Sachs CEO Discloses Bitcoin Stake, Backs Regulatory Push Amid Industry Standoff
- FOMC Minutes Signal Fed Largely Divided Over Rate Cuts, Bitcoin Falls
- BitMine Adds 20,000 ETH As Staked Ethereum Surpasses Half Of Total Supply
- Wells Fargo Predicts Bitcoin Rally on $150 Billion ‘YOLO Trade’ Inflow
- Will Crypto Market Crash as U.S.–Iran War Reportedly Imminent?
- BMNR Stock Outlook: BitMine Price Eyes Rebound Amid ARK Invest, BlackRock, Morgan Stanley Buying
- Why Shiba Inu Price Is Not Rising?
- How XRP Price Will React as Franklin Templeton’s XRPZ ETF Gains Momentum
- Will Sui Price Rally Ahead of Grayscale’s $GSUI ETF Launch Tomorrow?
- Why Pi Network Price Could Skyrocket to $0.20 This Week
- Pi Network Price Beats Bitcoin, Ethereum, XRP as Upgrades and Potential CEX Listing Fuels Demand
















