Datum Technical Brief
One of the most active debates in AI research is whether a single generalist model can handle every domain, or if specific domains need specialized models. For most domains and tasks, generalist models perform well, because the underlying data is text and scale compensates for lack of domain focus. It has become clear, though, that the construction industry is a case where this doesn't hold.
Construction data is heterogeneous and disconnected. RFIs, submittals, drawings, specs, and field reports are stored in different formats and rarely reference each other directly with plain language. Everything relies on the users' context on the project they are working on. Understanding a construction project at a level where AI is truly effective requires tracking dependencies across these documents, not just answering isolated questions about them.
Current AI deployment strategies do not accomplish this. Generalist models can produce fluent and convincing text about construction but fail when modeling how a project actually functions. They lack the structured understanding of sequence, dependency, and process that construction work requires.
Datum addresses this in two parts. First, a data harmonization layer restructures construction documents into a format that reflects project relationships, not just document contents. Second, a custom model is trained and fine-tuned specifically on construction processes, so it understands how projects unfold rather than only recalling facts from training data.
The result is an implementable specialist, cost efficient model with working knowledge of construction, not just linguistic fluency about it.
In practice, Datum connects to Procore (we are working on other integrations such as ACC, Onedrive, and Egnyte), indexes your project files, and answers questions in plain language with citations back to the source document and page.
Procore
36C26126R0019 — SF VAMC
- ▸Addendum 0006 — Statement of Changes.pdf
- ▸1399-CS101 — Civil Site.pdf
- ▸1399-CS501 — Grading Plan.pdf
- ▸RFI-022 — Sound Barrier Response.pdf
To test the model, our team gathered a large set of unstructured construction project documents from open-source government databases. This included RFIs, prime contracts, subcontractor submittals, change orders, field test documents, contract and shop drawings, spec sheets, and other miscellaneous documentation.
We then manually analyzed these documents with experienced construction project managers, and wrote a set of 300 prompt + answer pairs designed to evaluate the models on five key categories shown in the table below.
Spec lookup
exact requirement retrieval
"What galvanizing designation is required for composite metal decking?"
Cross-document
amendment chains, RFI→spec refs
"Which sheets were revised by the addendum and which was added for sound-barrier details?"
Drawing-lite
sheet index + callout xrefs
"Which sheet references detail 12/M-502?"
Workflow entities
CO, deadlines, set-aside
"Who is the contracting officer and their email?"
Abstention/cost
refuses absent facts; token economics
"What is the contract award amount?"
We then fed those prompts, with their respective documents, to 5 different models and scored the answers based on correctness.
The first model we tested was Claude Sonnet 4.6 with only the specific documents that the test question required access to. This is a standard construction-manager workflow — attach what you think is relevant and ask. We also tested Claude Projects (full project library) and Microsoft Copilot (project files through M365). Datum received the same documents through its harmonization pipeline.
On the 300-question ground truth set, Datum scored highest overall at 93.6%. Claude Projects scored 82.3%. Claude with document upload scored 71.9%. Microsoft Copilot scored 69.8%.
Spec lookup
- Claude (upload)
- 87.4%
- Copilot
- 84.6%
- Claude Projects
- 91.2%
- Datum
- 94.2%
Cross-document
- Claude (upload)
- 74.6%
- Copilot
- 71.2%
- Claude Projects
- 86.4%
- Datum
- 95.8%
Drawing-lite
- Claude (upload)
- 59.8%
- Copilot
- 57.4%
- Claude Projects
- 72.6%
- Datum
- 91.4%
Workflow entities
- Claude (upload)
- 61.2%
- Copilot
- 54.8%
- Claude Projects
- 79.8%
- Datum
- 94.6%
Abstention/cost
- Claude (upload)
- 76.5%
- Copilot
- 81.0%
- Claude Projects
- 81.5%
- Datum
- 92.0%
Overall
- Claude (upload)
- 71.9%
- Copilot
- 69.8%
- Claude Projects
- 82.3%
- Datum
- 93.6%
The gap is not vocabulary. Claude and Copilot know construction terms. They do not inherit your project's context. As such, they don't understand construction. Datum does.
To keep in line with our goal of responsible AI deployment, we are not releasing the model to the public yet. Right now, we are working with a smaller group of general contractors and subcontractors to test the model and gather feedback to determine the best way to deploy the model in the construction industry. If you beleive you or your organization could be a good design partner, feel free to fill out the form.
We look forward to seeing what you will build with Datum, but remember: build the future responsibly.
Design-partner release · Projectr Analytics · July 2026