A DEWALT survey of construction professionals published in April found that 86% felt somewhat or very prepared to work with AI. In the same survey, 8% reported using it in their daily work. Most respondents pointed to the same barrier: no formal, job-relevant training. What they had instead was YouTube, a Coursera tab, and whatever a colleague showed them on a Tuesday.
Hold those two numbers next to each other. A workforce that has largely not used these tools on real jobs is confident it knows how. That is the condition under which people get hurt, and it is worth understanding why before you roll a contract review agent out to eleven project managers.
The failure mode is not what people expect
The worry you hear at conferences is that AI will hallucinate a clause and someone will act on it. That happens, and it is manageable, because a wrong answer is a visible answer. The more expensive problem is quieter.
Human factors researchers separate three distinct things that go wrong when people work alongside automation. Automation bias is accepting output without verifying it. Automation-induced complacency is the drop in vigilance that comes from monitoring a system that is usually right. Skill decay is the erosion of a capability through disuse. The literature on this goes back decades to aviation, where the pattern was documented long before anyone used the word AI, and the consistent finding is the uncomfortable part: people generally cannot detect the decline in themselves.
Map that onto a mechanical contractor's office. An estimator runs symbol detection on a panel schedule. It is right the first twelve times. By the twentieth run, the QA pass has quietly compressed from a real check into a scroll. Nobody decided to stop verifying. The habit eroded, and it eroded fastest on exactly the work where the tool performs best, which is also the work where a rare miss is hardest to spot.
This is how a force multiplier becomes a crutch. Not through one dramatic error, but through the gradual transfer of judgment from a person who was accountable for it to a system that is not.
What the tools cannot tell you
Every serious vendor in this space describes its agents as human-in-the-loop. Estimators should run AI takeoff output through their own QA and QC to catch false positives. That instruction is easy to write and hard to operationalize, because it assumes the person doing the checking knows what a false positive looks like.
An estimator with fifteen years on panel schedules knows that a symbol count came back low because the drawing set has a revision cloud nobody incorporated. An estimator with eighteen months does not know that yet, and the agent will not say so. It will return a clean number with no signal that it is working from a stale sheet. The tool cannot flag what it does not know it is missing, and the reviewer cannot catch what they have not yet learned to look for.
Which surfaces the structural problem nobody in the industry has solved. Agents are strongest at counting, measuring, extracting and summarizing. Those are the same tasks junior estimators and assistant PMs have always used to build judgment. Automate them completely and you get faster output this quarter and a thinner bench in five years. The firms that handle this well will be the ones that keep juniors doing some of the work manually, on purpose, at a cost they have chosen to absorb.
What training should look like
Product training is not the same as judgment training, and vendors will happily supply the first. A two-hour session on where the buttons live does not teach anyone when to distrust the answer.
Three things belong in a real program. First, calibration: give people a set of documents where you know the tool is wrong, and let them find the errors themselves. Nothing builds appropriate skepticism faster than watching a confident system return a bad number on a drawing set you understand cold.
Second, a written verification standard per workflow, naming what a person checks before an agent's output moves downstream, with different thresholds for a daily log summary and a contract risk review. Third, clear accountability. Grant Thornton's 2026 AI survey found that firms with stronger governance and clearer lines of accountability were considerably more likely to see measurable results from AI investment, which tracks with how every other control system in construction works.
That last point has a commercial edge. Professional indemnity coverage for AI-generated output is still an unsettled question. If a bid goes out with an error that traces to an unverified agent output, the argument about who owns that error is one you would rather have already resolved in writing.
Some venues have already started shipping the infrastructure for this, with company-specific instruction that teaches agents a firm's own standards and an admin console showing which teams are consuming what. Useful, and not a substitute. The tooling can enforce a policy. It cannot tell you whether your people still know how to do the work.
The question to ask your PMs is not whether they can use the agent. It is whether they could catch it being wrong.