Frontier AI Labs Push Into Electronic Hardware Design as Evaluation Benchmarks Evolve
New testing from EEBench shows AI models scoring up to 61.6 percent on electronic circuit design tasks, drawing attention from OpenAI, Anthropic, and xAI.

Frontier artificial intelligence developers are expanding their focus into hardware engineering and printed circuit board design, highlighted by OpenAI demonstrating its GPT-6 Astra model operating within KiCad design software and xAI incorporating dedicated hardware benchmarks into recent model evaluations, according to technical analysis first reported by Hacker News.
The findings stem from testing conducted by EEBench, an evaluation platform built and funded by the team behind hardware description tool atopile. Rather than assessing how AI agents interact with graphical user interfaces in standard computer-aided design software, EEBench requires models to work directly in code using atopile. This framework enables agents to configure components, run SPICE circuit simulations, and evaluate electrical performance deterministically within a unified project workspace.
EEBench assesses models across practical electrical engineering scenarios, such as designing a backup power circuit for a residential utility meter. In that benchmark task, a system must maintain a processor rail above a 3.0-volt brownout threshold for 20 milliseconds following a 5-volt main power outage. The test suite evaluates whether models account for real-world component constraints, including voltage-dependent capacitance drops in ceramic parts, component tolerances, physical package limits, and overall bill-of-materials costs.
Benchmark data logged as of Sept. 1 indicates that Anthropic's Claude Opus 5 currently leads the EEBench V1 suite, scoring 61.6 percent across 13 simulation-backed analog and digital design tasks. xAI's Grok 4.6 ranked second with a score of 57.1 percent, closely followed by Anthropic's Claude Fable 5.1 at 56.4 percent. Older OpenAI systems registered lower scores, with GPT-5.5 reaching 42.3 percent and GPT-5.6 Sol recording 39.4 percent. EEBench has not yet published score data for OpenAI's newly demonstrated GPT-6 Astra.
Major AI labs are increasingly integrating hardware evaluation suites into their development pipelines. xAI included EEBench results in its official Grok 4.6 model card under a section dedicated to engineering acceleration alongside tests for 3D modeling and parametric CAD. According to xAI, Grok 4.6 reached a 60.0 percent score on the benchmark when configured for maximum reasoning effort, benefiting from domain-specific engineering data and reinforcement learning training.
Beyond grading finished designs, automated circuit evaluations are being repurposed to support model post-training. By converting electrical constraints—such as transient responses, ripple voltages, and component cost limits—into deterministic reward signals, developers can provide AI agents with precise feedback loops during training rather than relying on visual schematic inspections.
EEBench V1 currently focuses on circuit synthesis and simulation, leaving physical PCB layout, manufacturing bring-up, and hardware assembly testing for future iterations. While researchers note that current models are not ready to independently design critical systems such as medical devices without human oversight, the benchmarks indicate AI capabilities are steadily advancing across core electrical engineering tasks. Meanwhile, xAI founder Elon Musk has stated that Grok 4.7 is scheduled to launch within weeks following additional training on engineering datasets from SpaceX.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



