OpenAI Introduces Smart Contract Benchmark for AI Agents as AI and Crypto Converge

1 hour ago

Journalist

CoinGape comprises an experienced team of native content writers and editors working round the clock to cover news globally and present news as a fact rather than an opinion. CoinGape writers and reporters contributed to this article.

Read full bio

Why Trust CoinGape

CoinGape has covered the cryptocurrency industry since 2017, aiming to provide informative insights to our readers. Our journal analysts bring years of experience in market analysis and blockchain technology to ensure factual accuracy and balanced reporting. By following our Editorial Policy, our writers verify every source, fact-check each story, rely on reputable sources, and attribute quotes and media correctly. We also follow a rigorous Review Methodology when evaluating exchanges and tools. From emerging blockchain projects and coin launches to industry events and technical developments, we cover all facets of the digital asset space with unwavering commitment to timely, relevant information.

OpenAI Smart Contract Benchmark tests AI Agents in crypto

Highlights

OpenAI Smart Contract Benchmark tests AI Agents in Crypto for EVM contract bugs.
EVMbench uses 120 real flaws from 40 audits, including Code4rena and Tempo work.
GPT-5.3-Codex scored 72.2% in exploit mode, far above GPT-5 at 31.9%.

AD Join BC Game Get 840% Bonus

OpenAI has introduced a new smart contract security benchmark as AI agents gain stronger coding abilities in the crypto sector. Together with Paradigm, OpenAI said the benchmark, called EVMbench, tests how AI systems detect, patch, and exploit serious Ethereum contract bugs. Their effort responds to the growing financial risk, since smart contracts routinely secure over $100 billion in open-source crypto assets.

OpenAI Smart Contract Benchmark Targets Real Audit Vulnerabilities

In their release, OpenAI said EVMbench draws on 120 curated vulnerabilities collected from 40 professional smart contract audits. Notably, most of the issues came from open audit competitions, including Code4rena. OpenAI said the benchmark also includes vulnerability scenarios tied to security auditing work for the Tempo blockchain.

Tempo is described as a purpose-built Layer-1 network designed for high-throughput, low-cost stablecoin payments. Because of that, these scenarios extend the benchmark into payment-focused contract code. The company also said it expects agent-based stablecoin payment activity to grow.

To build the benchmark environments, OpenAI said it adapted existing exploit proof-of-concept tests and deployment scripts when available. However, it said engineers manually wrote missing components when no scripts existed. OpenAI added that it ensured patch tasks remained exploitable while still fixable without breaking compilation.

Detect, Patch, Exploit Modes Test AI Agents Under Pressure

OpenAI said EVMbench evaluates artificial intelligence agents in three modes. That is detect, patch, and exploit. In detect mode, agents audit smart contract repositories and get scored on recall of confirmed vulnerabilities and audit rewards. In patch mode, agents must modify vulnerable contracts while keeping intended functionality intact.

Exploit mode, however, focuses on full end-to-end fund draining attacks in a sandbox blockchain environment. The company said graders verify results using transaction replay and on-chain checks. To support reproducible evaluation, the company said it developed a Rust-based harness to deploy contracts and replay transactions deterministically.

Notably, the exploit tasks run in an isolated local Anvil environment instead of live crypto networks. It also said vulnerabilities used in the benchmark are historical and publicly documented. OpenAI added that the harness restricts unsafe RPC methods to limit abuse.

In exploit testing, OpenAI said GPT-5.3-Codex running via Codex CLI scored 72.2%. However, it said the earlier GPT-5 model scored 31.9%, despite being released just over six months earlier. OpenAI also noted that detect recall and patch success remain below full coverage.

OpenAI Adds New Talent with Agent Hire

While OpenAI pushed EVMbench into public view, it also expanded its agent development team. Notably, they hired Peter Steinberger, founder of the viral open-source AI agent project OpenClaw, previously known as Clawdbot. Sam Altman confirmed on X that Steinberger will join OpenAI to lead work on the “next generation of personal agents.”

Meanwhile, Altman said OpenClaw will transition into a foundation model project supported by OpenAI. The open-source project will continue under that structure, according to the announcement. The hiring drew wide attention as OpenAI increases its focus on autonomous and personal AI agents.

Tap here to Add CoinGape as a Trusted Source

Why Trust CoinGape

CoinGape has covered the cryptocurrency industry since 2017, aiming to provide informative insights Read more… to our readers. Our journal analysts bring years of experience in market analysis and blockchain technology to ensure factual accuracy and balanced reporting. By following our Editorial Policy, our writers verify every source, fact-check each story, rely on reputable sources, and attribute quotes and media correctly. We also follow a rigorous Review Methodology when evaluating exchanges and tools. From emerging blockchain projects and coin launches to industry events and technical developments, we cover all facets of the digital asset space with unwavering commitment to timely, relevant information.

Latest
/
Trending

Goldman Sachs CEO Discloses Bitcoin Stake, Backs Regulatory Push Amid Industry Standoff

FOMC Minutes Signal Fed Largely Divided Over Rate Cuts, Bitcoin Falls

BitMine Adds 20,000 ETH As Staked Ethereum Surpasses Half Of Total Supply

Wells Fargo Predicts Bitcoin Rally on $150 Billion ‘YOLO Trade’ Inflow

Will Crypto Market Crash as U.S.–Iran War Reportedly Imminent?

FOMC Minutes Today: Will Bitcoin and Crypto Market Crash After Fed Signals?

Bitcoin vs. Gold: Why Experts Think BTC Will Lag Behind

Here’s Why Crypto Market is Down Today?

Top Crypto Market Events To Watch This Week- Bearish or Bullish?

Here’s Why is Crypto Market is Struggling to Recover (Feb 13)?

Premium Partners

Best Change

Mexc

PrimeXBT

Bitget

Newsletter

Your crypto brief.
Delivered every day.

Insights that move markets
100,000 active subscribers

By signing-up you agree to our Terms and Conditions and Privacy Policy.

RECOMMENDED

price analysis

Best Crypto Presales To Invest In February 2026 – Top Upcoming Presale Tokens

Best Crypto AI Trading Bots In 2026

Top Meme Coins To Buy In February 2026

Best Cloud Mining Platforms For 2026

About Author

Coingapestaff

About Author

Coingapestaff

Investment disclaimer: The content reflects the author’s personal views and current market conditions. Please conduct your own research before investing in cryptocurrencies, as neither the author nor the publication is responsible for any financial losses.

Ad Disclosure: This site may feature sponsored content and affiliate links. All advertisements are clearly labeled, and ad partners have no influence over our editorial content.

OpenAI Introduces Smart Contract Benchmark for AI Agents as AI and Crypto Converge

OpenAI Smart Contract Benchmark Targets Real Audit Vulnerabilities

Detect, Patch, Exploit Modes Test AI Agents Under Pressure

OpenAI Adds New Talent with Agent Hire

Related Articles

AI Coins, Dogecoin Lead Crypto Market Rebound as Elon Musk Lauds Nvidia CEO

Bitcoin, AI Coins Bounce as Nvidia Signs $20B AI Inference Deal with Groq

AI Bubble: Big Short Legend Michael Burry Bets Against AI Giants As NVIDIA And Palantir Stocks Dip

AI Coins Gain as US SEC Crypto Task Force Met Multiple Firms Today

Chamath Palihapitiya Files $250M SPAC Aiming at America’s Strategic Tech Sectors

China Aims to Bridge AI Gap, Challenges US Lead in Crypto and Beyond