How well can an LLM do hardware engineering?

HWE-bench is SWE-bench for hardware. It measures whether a model can do digital engineering — get real work done in design/CAD, modeling/CAE and production/PLM software. Tasks come from real engineering practice, not synthesized from papers. The first release covers CAE.

Leaderboard

HWE-bench

Overall hardware-engineering score · CAE release

SVD Lab
GPT-5.6 Sol
0.90
Kimi K3
0.84
Claude Sonnet 5
0.72
DeepSeek V4 Flash
0.67
Doubao Seed 2.1 Turbo
0.63
GLM-5.3
0.61
GPT-5.6 Luna
0.60
MiniMax M3
0.49
Doubao Seed 2.0 Lite
0.31
00.250.500.751.00

Bars show the overall score. Black whiskers show the 95% interval. 77 tasks across combustion, battery, cfd and packaging; one scored trial per task. All models were run through the same agent scaffold (terminus-2). A model's score moves with how it is driven — the agent scaffold and the API route both.

Run it, or talk to us.

The public task set ships as a Harbor dataset — each task is a container, a prompt, a reference solution and a verifier, runnable against any model with a completions endpoint.

Behind the public set is a much larger private task inventory built for post-training: engineering-simulation environments with a programmatic, non-hackable reward, in families parameterised into as many operating points as an RL run needs. If you are building post-training environments, get in touch.