KP Reddy

A Framework for Evaluating AI in Construction Drawing Review

So many companies are reading drawings, so what?

KP Reddy's avatar
KP Reddy
Jul 16, 2026
∙ Paid

The application of AI vision and multimodal models to construction drawings — architectural sheets, structural details, MEP layouts, shop drawings — is frequently evaluated using a single, undifferentiated criterion: whether the system “reads drawings.” This framing conflates at least four distinct processes — reading, comprehension, understanding, and application — each producing qualitatively different output and carrying a different risk profile. A fifth category, outcomes, is causally downstream of all four but is not determined by any of them in isolation. This paper proposes a five-level framework for evaluating AI drawing-review systems, argues that current benchmarking practice collapses these levels inappropriately, and offers criteria for distinguishing genuine capability from surface performance at each level.

The Five-Level Framework for AI Drawing Review

Q: What problem is this framework trying to solve? A: Vendors and teams often describe AI drawing-review tools with a single claim — “it reads drawings” or “it understands the set.” That single claim hides at least four distinct capabilities, each with a different failure mode. This framework separates them so that claims, evaluations, and deployment decisions can be made about the right level, instead of one level’s performance being used to imply competence at another.

Q: What are the five levels, in plain terms? A:

  • Reading — can it see and decode what’s on the sheet (text, symbols, lines)?

  • Comprehension — can it correctly interpret a single detail or callout on its own?

  • Understanding — can it connect that detail to other sheets, other disciplines, and other phases correctly?

  • Application — can it turn that understanding into the right action, for the right person, in the right form?

  • Outcomes — what actually happened on the project as a result, and who is accountable for it?

Q: Isn’t “reading” basically solved already? A: For clean, CAD-originated PDFs, largely yes. It degrades on scanned sheets, hand-marked redlines, dense overlapping annotations, and non-standard symbol legends. Reading well is necessary but tells you nothing about whether the system understood what it read.

Q: Why isn’t comprehension the same as understanding? A: Comprehension is local — correctly interpreting one detail in isolation. Understanding is systemic — knowing that the wall type on the architectural plan has to match a schedule on another sheet and the life-safety plan’s fire rating, and knowing which of several conflicting versions of a detail currently governs. A system can comprehend every detail correctly and still miss a coordination conflict that only exists across sheets.

Q: Is clash detection the same as “understanding”? A: Not necessarily. A lot of what’s marketed as AI drawing understanding is geometric collision detection on a 3D model — a narrower, more mechanical task. Real cross-discipline understanding also requires knowing which version of a detail governs after an addendum, and recognizing what the set fails to show at all, not just where two objects physically overlap in a model.

Q: Why does “application” need to be broken into function, role, and method? A: Because the same correct understanding supports different actions depending on three independent things:

  • Function — what task is this for (estimation, clash detection, constructibility, takeoff)?

  • Role — whose problem is this (electrical, mechanical, structural, GC, owner’s rep)?

  • Method — how will it be used (a static RFI markup, a 4D sequencing simulation, a field verification)?

A tool validated for one combination of these three shouldn’t be assumed competent at another, even if it’s drawing on the same underlying understanding.

Q: If a tool is accurate at understanding, why would it still fail at application? A: Because application depends on context the tool often doesn’t have — this project’s trade sequencing, this fabrication lead time, whether a bulletin already resolved the discrepancy the tool is flagging. Correct understanding, applied without that context, can still be the wrong action.

Q: Why does the framework treat “outcomes” as separate from “application” instead of the final proof that the system worked? A: Because outcomes are affected by things no drawing-review system controls — field labor quality, material substitutions, weather-driven sequencing. A correct RFI, raised at the right time, can still be followed by a field error unrelated to the review that preceded it. Judging a tool by outcomes alone is noisy; judging it by the level it actually operates at is more reliable.

Q: Where does human sign-off fit into this? A: This is the core of the outcomes level. Historically, an engineer’s stamp, a superintendent’s signature on a submittal log, or a PM’s approval on an RFI was the point where judgment became attributable to a specific person. When an AI system performs application-level judgment — flagging a discrepancy as material, recommending a submittal be held — that judgment needs an equivalent sign-off point before it drives an action in the field. Otherwise the chain from judgment to outcome has a gap where accountability used to sit.

Q: Doesn’t a human reviewing the AI’s output before acting on it just slow things down? A: It adds a step, but that step is what preserves the professional and contractual structure construction already runs on — licensure, stamping, sign-off logs — none of which extend automatically to a system’s output. Removing the sign-off doesn’t remove the judgment being exercised; it just removes the person accountable for it.

Q: What does this mean practically when evaluating or buying an AI drawing-review tool? A: Ask which level the vendor is actually claiming, and test that level specifically — don’t let comprehension-level accuracy stand in for understanding, and don’t let understanding stand in for application. If the claim is at the application level, ask which function, role, and method it was validated for. And check whether the deployment preserves a real sign-off step by an accountable person before the output drives an action, or whether that step has been designed out.

Q: Can this framework apply to anything other than drawings? A: Yes — the same five levels apply to contract text, specifications, and other AEC documents. The mechanics differ (symbol detection versus clause extraction, cross-sheet coordination versus cross-document risk allocation), but the structural gaps between reading, comprehension, understanding, application, and outcomes are the same.

User's avatar

Continue reading this post for free, courtesy of KP Reddy.

Or purchase a paid subscription.
© 2026 KP Reddy · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture