Generation and judgment are different problems
Part 2 ended on this: the judgment rebar requires is presupposed as the API's input. So what if AI takes over that judgment as well? We answer with vendors' own contracts and peer-reviewed benchmarks, not marketing pages.
Published August 10, 2026·Last checked · August 2026
Key points
- Autodesk General Terms §11.1 — “The Offerings are tools … not a substitute for Your professional judgment.” Not an AI clause: it applies to every product and predates AI
- Even a rule-based (non-AI) checker reports “potential” issues rather than confirmed violations, and builds a human decision step into the product
- AECBench (4,800 questions · 9 LLMs) — fluent at recall and comprehension, with a clear decline in table interpretation, complex reasoning and calculation
- SoM-1K (1,065 mechanics-of-materials problems) — best model 56.6%. Models given a written description of the figure beat vision models that saw the figure
- Rebar carries an asymmetric verification cost — a wrong bar looks fine on screen and shows up after the pour
1What Autodesk put in its own terms
Autodesk General Terms (updated 2026-03-30), §11.1 “Offerings are tools.” This is contract language, not an AI marketing page.
“You are solely responsible for … establishing the adequacy of independent procedures for testing the reliability, safety, accuracy, completeness, compliance with applicable legal requirements and industry standards … of any Output, including insights, recommendations, and all items designed with the assistance of the Offerings.”
One more thing matters here: this is not an AI-specific clause. It is a general term covering all Autodesk “Offerings,” and it predates AI. Which makes it not a defence against AI but a statement about what design software is.
2Tools that handle codes and standards land in the same place
| Tool | Capability, per the vendor | Limitation, per the vendor |
|---|---|---|
| UpCodes Copilot (building-code AI) | “an AI-powered chatbot that allows you to ask questions and get an answer specific to your project’s parameters, year, and jurisdiction” | “It may occasionally generate incorrect information.” — the only caution published in the support docs |
| Solibri (rule-based — not AI) | “validating Building Information Models against predefined rules … such as building codes, design guidelines, or project requirements” | Reports results as “potential errors, inconsistencies, or non-compliance issues,” with a human step built into the product (Making Decisions on Checking Results) |
Solibri is not AI — it is a deterministic rule checker. And it still declines to call a result a confirmed violation, and still puts a human judgment step in the product. Even where the rules are unambiguous, the final call stays with a person.
3Measured: how far LLMs get on engineering problems
Impressions aside — here are peer-reviewed benchmarks.
| Study | Scope | Result |
|---|---|---|
| AECBench arXiv:2509.18776 (Sept 2025 · rev. Feb 2026) | AEC domain knowledge across five cognitive levels — recall, comprehension, reasoning, calculation, application. 4,800 questions authored by engineers and expert-reviewed twice, 9 LLMs | “a clear performance decline across five cognitive levels.” Fluent at recall and comprehension, with marked deficits in “interpreting knowledge from tables in building codes, complex reasoning and calculation, and domain document generation” |
| SoM-1K arXiv:2509.21079 (Sept 2025) | 1,065 mechanics-of-materials problems, 8 foundation models | Best model scored 56.6%. LLMs handed an expert's written description of the figure outperformed vision models that looked at the figure — current models misread engineering diagrams visually |
The second result is the telling one: reading drawings is still the weakest link.
In fairness, the counter-evidence: one study reports GPT-4o at 100% on structural analysis problems (arXiv:2504.09754). But it uses 20 questions, and the method is not the LLM calculating — it writes OpenSeesPy Python code and a solver does the maths.
Give it a deterministic solver: 100%. Make it read a diagram and decide: 56.6%.
AI runs calculations whose rules are already encoded. Interpreting the rules to set the values is not there yet.
4And rebar carries one more condition
A wrong dashboard shows on screen. A bad render is visible. Wrong rebar shows up after the concrete is poured.
A hook five degrees short, a development length a hand's width short, a splice sitting where the stress is high — all of it looks perfectly normal on screen. Verification cost is asymmetric here, which makes “mostly right” worth less than it is in other fields.
Sources
- General Terms §11.1 · §14.2 (updated 2026-03-30) — the strongest citation in this seriesAutodesk
- AECBench: A Multi-dimensional Benchmark for LLMs in the AEC Domain (arXiv:2509.18776)arXiv
- SoM-1K: Strength of Materials benchmark (arXiv:2509.21079)arXiv
- Integrating LLMs for Automated Structural Analysis (arXiv:2504.09754) — counter-evidencearXiv
- Copilot overview (support documentation)UpCodes
- Understanding Checking — deciding on checking resultsSolibri
This article is a summary compiled from published material. It is not a legal interpretation and does not constitute advice. Regulation changes frequently, so please check the original documents and the latest notices from the responsible authority before relying on it in practice.
