When My AI Told Me I Should Use RSMeans, I Listened
How a Defensible Commercial Repair Cost Estimate Actually Gets Built — Field Measurement, Localized Cost Data, and Where AI Earns Its Place
I saw an ad for a new AI product promising that it could replace huge swaths of my workload. I'm a fan of checking out shiny new things, so I took a peek under the hood. I used the same prompt I run regularly with Claude — who, by the way, helped shape this article.
How effective can [the product] be for Calibre Commercial Inspections?
Very open-ended, deliberately so. I wanted to know what it would promise.
It acknowledged that no AI can replace a human with great eyeballs on site. Yet, it said. Then, in the list of back-office functions it could take off my hands, it described completing a Cost-to-Cure report.
This rang bells.
Claude and I had a discussion about Cost-to-Cure reporting months ago. I had specifically tried to determine how accurate the numbers were by running the same material through ChatGPT and Grok. The numbers came back similar — similar enough to feel like the method was validated. I confirmed with Claude and, paraphrasing the response, was told that all the models were trained on the same data and would carry the same flaws.
What I had confirmed was that the AIs agreed. Not that I had a defensible number.
[Note from Claude, who is helping edit this article and read that paragraph over my shoulder: “Trained on the same data” isn't quite right — the corpora differ. What's defensible is that models built on broadly overlapping public data, using broadly similar methods, produce correlated errors. Agreement between them is evidence of a shared prior, not of accuracy.]
Fine. I'm leaving both versions in, because the correction is the whole article. It caught a wrong technical detail in one pass, and it did not, at any point, offer to price my roof.
Not that AI isn't helpful. It is, enormously. But AI is also vastly oversold as the answer to every difficulty, and a Cost-to-Cure report is a good example of the gap. It arrives as a document, so it looks like paperwork. It isn't. It's an estimate, and an estimate is a position somebody is going to challenge. Heinlein coined a word for understanding a thing so completely you become part of it. Whatever that product was doing, it wasn't that.
What I wanted — needed — was a way to get a great number for my clients and a defensible number for my company.
So the next set of discussions with Claude revolved around finding the data that would do both. RSMeans I knew by reputation, so I started there, and being a contrary sort, my search terms were “RSMeans competitors renovation life cycle.”
That delivered a raft of choices, which I proceeded to muddle through by old-fashioned reading and thinking. This I did not outsource to Claude.
In the end, RSMeans was the organization that floated to the top for everything but price. If you want great tools, you need to spend the money for them.
A couple of studies published recently make the point even better than I can.
GIGO is Still a Thing
In March, a preconstruction software firm handed a 303-page issued-for-construction drawing package to a leading AI model and compared its quantity takeoff line by line against a professional cost estimator working from the same drawings. In May, a construction technology company ran a similar experiment on a road, drainage, and public realm scheme, with a chartered quantity surveyor supplying ground truth and two models competing across 45 line items.
Priced out, the machine-generated quantities came in 51 percent low in the first test and roughly 40 percent low in the second. In the second, 71 percent of individual line items missed by more than 20 percent.
Two details in those results matter more than the headline.
The first is that the cost data was never at fault. Both studies applied the same published unit rates to the machine quantities and to the professional's. The rates were identical. The entire gap was quantity.
The second is what the models were handed: complete, approved, issued-for-construction drawings. A property condition assessment starts with no such thing. The building is 40 years old, the original drawings are missing or describe a floor plan three tenants ago, and the roof has been recovered twice by contractors nobody kept a record of. Whatever those tests measured, they measured under conditions considerably friendlier than a real building offers.
There is one more finding worth pulling out, because it matches what happened on my own desk. In the civils test, two different models made the same mistakes in the same places, which is why the researchers called the errors systematic rather than random.
Running a second model is not a second opinion. It is the same opinion with a different logo on it.
The Quantity Does Not Exist Until Somebody Measures It
On an existing building there is no takeoff to automate, because the number isn't written down anywhere. It has to be produced. That happens on site, with a tape, a wheel, a camera, and a set of eyes that have seen the same failure before.
The field pass produces the quantities — roof area by section, linear feet of failed sealant joint, the count of rooftop units and their nameplate ages, square footage of spalled paving.
But it’s not just the numbers. How are we getting equipment to the roof? Do we have clearances for new electrical? We are back to eyeballs and experience. It’s not just a quote for the new HVAC. It’s the HVAC, the crane, maybe shutting down a street while using the crane (so a safety crew.)
And it produces the judgment calls that move the number further than any measurement does. Is this roof a recover or a tear-off? That question isn't answered by area. It's answered by core samples, by how many layers are already up there, by deck condition, and by what the local code official will accept. The area is identical either way. The cost is nowhere close. Getting that one call wrong outweighs every unit rate in the report.
Which is why quantity error is the dangerous kind. In the civils test the errors did not all run the same direction — some categories overshot on a full-perimeter default, others undershot on a derived volume — and the two partially cancelled at the bottom line. The total came out about 40 percent light rather than displaying the much wider spread underneath. It did not look absurd. It looked like a number somebody had worked out.
A number that is obviously wrong gets caught in review. A number that is wrong underneath a plausible total gets funded. On a capital reserve schedule the consequence is sharper still, because a reserve schedule is read by year, not by total. If the roof line is low and the paving line is high, the 12-year figure can look defensible while year three is badly underfunded. Owners don't spend the total. They spend the year.
The Unit Cost Has to Belong to a Market
A measured quantity is only half a cost. The other half is what that work costs where the building actually sits, and this is the thing a published cost database does that a general-purpose tool cannot.
RSMeans publishes to a national average, built from a composite of major cities and indexed at 100. No city in North America sits at 100. To bridge the gap, the data carries City Cost Index location factors covering roughly 970 United States and Canadian locations, applied separately to material, labor, and equipment, and updated quarterly. A market indexed at 112 runs 12 percent above the national average. A market at 88 runs 12 percent below.
Across Idaho, Eastern Oregon, Eastern Washington, and Western Montana, that adjustment isn't a refinement. It's most of the answer. Lewiston is not the national composite. Neither is Coeur d'Alene, Pendleton, Kennewick, Missoula, or Twin Falls. Four things move in a tertiary market, and they don't move together.
Freight. A membrane roofing order delivered to a small Inland Northwest city doesn't carry the same delivered cost as the same order landing at a distribution point in a major metro.
Labor. Wage rates, union density, and available trade depth vary sharply across a four-state footprint, and the spread between two markets 90 minutes apart is often wider than the spread between either one and the national average.
Mobilization. Where there's no local crane, no local elevator service company, and no local curtain wall crew, the crew drives. Travel, per diem, and equipment transport are real money on a mid-size repair, and they don't appear in a national average at all.
Scarcity. One qualified contractor within 120 miles is a pricing condition, not a line item. It shows up as a premium and as a schedule, and the schedule is frequently the more expensive half.
The index isn't a quote, and it would be dishonest to present it as one. It's a documented, published, quarterly-updated adjustment from a known baseline to a named market. When a single line is large enough to move the deal, the right move is local contractor pricing, and the report should say which lines were priced that way. The difference between a sourced number and a produced one isn't certainty. It's traceability — the reader can see the baseline, see the factor, and see which market it was applied for.
The Framework Is Where the Machine Earns Its Place
The successes in those two studies are as instructive as the misses. Wherever the answer was already a number sitting in a table, the models were close to exact — column counts, base plates, manholes, valve chambers, door assemblies. Concrete volumes computed from a footing schedule landed within 2 percent. Asphalt tonnage from a stated pavement build-up landed within 3 percent. Both research teams reached the same conclusion about the appropriate role: a first-pass index that a professional then corrects.
That's a real job, and on a due diligence report it's a substantial one. A commercial building generates a deficiency list running well past a hundred items across envelope, structure, mechanical, electrical, plumbing, life safety, and site. Every observation has to reach the cost table, carry consistent language, sit in the right priority tier, and land in the correct year of the reserve schedule. Reconciling the observed list against the priced list so nothing walks off between them is exactly the sort of structured, repetitive work a machine does more reliably than a tired human at nine at night.
What it doesn't do is supply a quantity, choose a market, or decide whether the roof is a recover or a tear-off. It wasn't in the building, it can't see the deck, it has no location parameter to set, and it has no way of knowing the index moved last quarter. Ask it to hold the framework and it holds the framework well. Ask it to be the source of the number and it will produce one anyway. That last part is what should make a buyer uneasy.
What to Ask Whoever You Hire
A buyer reading a cost table can't audit the arithmetic. But four questions will reveal how the number was built.
Where did the quantities come from — measured on site, or taken off documents supplied by the seller?
What cost database was used, and which location factor was applied? A number without a named market is a national average wearing a local label.
Which lines were verified against local contractor pricing, and does the report identify them?
Does the estimate assume like-for-like replacement or a code-compliant upgrade? Those are frequently not the same number, and the gap between them is often the negotiation.
An estimate that answers all four is a position that will hold when the seller's engineer pushes back on it. An estimate that answers none of them is a guess with a decimal point.
Three instruments, one number. Measurement supplies the quantity, localized cost data supplies the rate, and the machine holds the structure so nothing observed goes unpriced. Each is good at its own job and unreliable at the other two. A report built properly lets the reader see which did what.
Calibre Commercial Inspections performs Property Condition Assessments, commercial building inspections, and Capital Needs Assessments across Idaho, Eastern Oregon, Eastern Washington, and Western Montana, with opinions of probable cost built from field measurement and RSMeans data localized to the market where the work will actually happen. Contact us to discuss your property

