HWE-bench is SWE-bench for hardware. It measures whether a model can do digital engineering — get real work done in design/CAD, modeling/CAE and production/PLM software. Tasks come from real engineering practice, not synthesized from papers. The first release covers CAE.
Overall hardware-engineering score · CAE release
Bars show the overall score. Black whiskers show the 95% interval. 77 tasks across combustion, battery, cfd and packaging; one trial per task.
The public task set ships as a Harbor dataset — each task is a container, a prompt, a reference solution and a verifier, runnable against any model with a completions endpoint.
Behind the public set is a much larger private task inventory built for post-training: engineering-simulation environments with a programmatic, non-hackable reward, in families parameterised into as many operating points as an RL run needs. If you are building post-training environments, get in touch.
The headline is the equal-weighted mean of these four numbers, so a domain where the ordering differs from the headline is worth seeing directly.
Two reference points bracket each run: a reference solution scores 1.0 by construction, and an empty submission has to score below 0.5.