Every filing we read, one click from the original
One square per company-year. Colour is how much artificial-intelligence language that annual report carries. Hover a blue square to read the AI sentence it found; click it to open that 10-K on sec.gov scrolled to that sentence, highlighted. Grey and white squares open their filing at the top, and a company name opens its full record. Nothing on this site is hand-entered, and nothing is more than one click from the document it came from.
Here for the findings instead? Start with the overview, or read every AI passage we found: each one opens its 10-K on sec.gov scrolled to that exact sentence, highlighted.
A blank means the firm filed no 10-K for that fiscal year, usually because it had not yet listed or had been acquired. The right-hand column totals the core AI terms across all of a firm's filings. Shell registrants are excluded by requiring US$10m in assets or revenue and a filing of at least 5,000 words; the filings that fail that screen are still shown, outlined, so nothing is hidden.
Everyone in construction is talking about AI.
Almost nobody says they use it.
We read every mention of artificial intelligence in … annual reports filed with the SEC by … listed US construction and engineering firms between … and …. A 10-K is compulsory and legally actionable, and firms are accountable for what they write in it. Every number on this site links to the filing it came from.
The sudden arrival
Share of operating firms mentioning AI in their annual report. Shaded band is the 95% bootstrap interval.
AI disclosure was flat near zero for eight years, then rose almost vertically in three, in every industry segment at once. That is not how a technology diffuses through an industry.
Opportunity, or liability?
Where AI mentions sit inside the 10-K. Item 1A is where a firm warns investors; Items 1 and 7 are where it describes its own business.
Talking versus doing
Firms mentioning AI, against firms claiming to actually deploy it.
What the AI language actually says
Every AI passage coded by three open-weight language models from three different labs, taking the majority vote. Validated against blind human coding (Cohen's κ = …).
Nothing about the firm explains it
Calendar year alone explains … of the variation in whether a firm discloses AI. Everything about the individual firm, meaning size, profitability, leverage and capital intensity, adds …. Industry segment adds ….
The reason is in the Templates tab: firms are not writing their own sentences.
Look up a contractor
Every firm in the sample, with its complete AI disclosure history. Click any row to see the year-by-year record and jump straight to the filings on sec.gov.
| Firm | Segment | Revenue | AI mentions FY2024–25 |
Intensity per 10k words |
Risk framing | First year | Deploys? |
|---|
Read what they wrote
All … AI passages in the filing panel, with the label three language models agreed on. Search the text, filter by what the passage claims, and open the filing it came from.
The same sentences, at different firms
If firms were independently recognising AI risk, they would describe it in their own words. Frequently they do not. We found … sentence-level templates recurring at two or more unrelated registrants, covering … of all AI passages, including … word-for-word matches between firms.
AECOM, Fluor, Jacobs Solutions and KBR, the four largest engineering and EPC firms in the sample, share four separate AI risk-factor passages word for word. A common external shock explains simultaneous attention to AI. It cannot explain identical sentences. That requires transmission, most plausibly through shared securities counsel.
What a full read of the passages turns up
Every AI passage in the filing panel, read end to end. Three results that only appear at that grain: whose AI the risk language attends to, how eight firms' disclosures moved over a decade, and the six firm-years that bound the sector's disclosure space.
Whose AI the risk disclosures attend to
… of … risk-bearing passages, four of every five, attend to AI in hands other than the reporting firm's.
Eight trajectories, FY2015 to FY2025
AI sentences per fiscal year: capability side of the filing (blue, above each line) and risk side (red, below). Hover any dot for the counts, click it to read that firm-year's passages. Each annotation below is verified verbatim against the raw filing.
The outer range of disclosure positions
Six firm-years that bound the space, quoted verbatim. Between these bounds lies the templated middle documented in the Templates view.
What the AI talk is actually about
Every AI passage was read by three open-weight language models from three different developers and labelled on seven dimensions. Where no two models agreed, the passage is recorded as unresolved rather than forced into a category. These are the majority labels for … passages from … operating firms.
Most named risk is AI in someone else's hands
Risk themes named across all coded passages. Cybersecurity is not AI failing; it is somebody else using AI against the firm.
Which technology gets named
Most AI language names no technology at all.
Whose AI is being described
A firm can disclose AI without ever claiming to touch it.
Slice any dimension
The same coding, cut three ways. Section is the strongest cut: where a sentence sits in the 10-K predicts what it says.
Can these labels be trusted?
Agreement between the three models on each dimension, and against a blind human coding of a random sample. Kappa corrects for agreement by chance, so it is lower than raw agreement and harder to flatter.
How the language changed
The same firms, the same disclosure, four periods. These are the words that separate each period from the others, measured by weighted log-odds with an informative Dirichlet prior. Firm names are stripped first, otherwise one talkative registrant supplies its own name as the most distinctive word of a period.
The grammar flipped to hedging
Hedging words such as may, could and potential, against assertive words such as we use, currently and has. Per 100 words of AI text.
Shorter, and stripped of numbers
Mean passage length, and the share of passages carrying any figure.
When each term entered the vocabulary
First fiscal year each AI term appears anywhere in the panel. Dot size is the number of firms that ever used it.
How this was measured
The filing archive
Registrants in SIC 1500–1799 (Division C, Construction) plus SIC 8711 (engineering services) were identified from the SEC's quarterly Financial Statement Data Sets. We retrieved every Form 10-K for fiscal years 2014–2025 directly from EDGAR (… files, … MB) and froze the archive under a SHA-256 manifest on …, so any later result can be checked against exactly the same text.
Dormant shell registrants, real filers whose entire annual report runs a few hundred words, are excluded by requiring at least US$10m in assets or revenue and a filing of at least 5,000 words. That leaves … firms and … firm-years.
Two measurement details that change the answer
The bare abbreviation AI is matched case-sensitively as a standalone token, because construction filings refer constantly to AIA contract documents, which a naive pattern counts as artificial intelligence.
Splitting a 10-K into its Items is harder than it looks: the table of contents lists every heading before the body starts, and the body is full of cross-references that look identical to headings. Our parser discards a heading followed by another heading within 700 characters, which marks a contents entry, since a real section is followed by thousands of words, then discards grammatical cross-references. It parses 97.0% of filings.
Reading the passages
Counting terms says how much a firm talks about AI, not whether it claims to use it. Every AI passage was therefore coded by three open-weight models from three different laboratories (…) run locally with greedy decoding, taking the majority vote. Different laboratories matter: one model's quirks are indistinguishable from a reading of the text, whereas independently trained models agreeing is evidence about the text. Inter-model agreement was Fleiss' κ = ….
Validation, and what it changed
Machine agreement is a reliability check, not proof of correctness. One author hand-coded a blind random sample of … passages, without seeing the model labels or which 10-K section each passage came from. Agreement was … raw, Cohen's κ = … across six categories and κ = … for the deployment-versus-rest distinction the headline rests on.
The disagreement was systematic and it changed a published number. The models under-call deployment: they read present-tense statements ("we also utilise generative AI tools") as aspirational, and they attribute client-side AI to the filer. So levels on this site come from the hand-coded random sample, at … of passages describing real deployment, while the model coding carries the time series, and every model-based deployment figure is a lower bound.
What this cannot tell you
- It measures what firms say, not what they do. Firms may use AI without disclosing it where it is not material. The gap between saying and doing is the object of study, not a flaw but it is a real limit.
- It covers listed US firms only. The private companies that build most of the built environment never file a 10-K and are invisible here.
- 106 firms is the population, not a sample from it, but it is a small population.
- Application labels were not part of the human validation and should be read as indicative.
Reproducing it
Every figure on this site is generated from public SEC data by code in the project repository. The filing archive is rebuilt from EDGAR by a single script; the analysis re-runs from the frozen archive; this site is baked from the analysis outputs. Nothing is hand-entered.
Every quotation here opens its 10-K on sec.gov at that sentence, highlighted, rather than at the top of a three-hundred-page document. Those 924 links are not assembled in your browser and hoped for: each is cut from the filing's own HTML and then checked against it by replaying the browser's text-matching rules, and one that does not check out is not published. A link that opened the right filing at the wrong paragraph would look exactly like this site inventing a claim, so it is treated as a result to be verified rather than a convenience.