How well can an LLM do hardware engineering?

HWE-bench is SWE-bench for hardware. It measures whether a model can do digital engineering — get real work done in design/CAD, modeling/CAE and production/PLM software. Tasks come from real engineering practice, not synthesized from papers. The first release covers CAE.

Leaderboard

HWE-bench

Overall hardware-engineering score · CAE release

SVD Lab
GPT-5.6 Sol
0.90
Kimi K3
0.72
DeepSeek V4 Flash
0.67
GPT-5.6 Luna
0.60
MiniMax M3
0.47
Doubao Seed 2.0 Lite
0.31
00.250.500.751.00

Bars show the overall score. Black whiskers show the 95% interval. 77 tasks across combustion, battery, cfd and packaging; one trial per task.

Run it, or talk to us.

The public task set ships as a Harbor dataset — each task is a container, a prompt, a reference solution and a verifier, runnable against any model with a completions endpoint.

Behind the public set is a much larger private task inventory built for post-training: engineering-simulation environments with a programmatic, non-hackable reward, in families parameterised into as many operating points as an RL run needs. If you are building post-training environments, get in touch.