For many institutions, the biggest facilities information problem isn’t a lack of documentation.
It’s the opposite.
Decades of construction drawings, specifications, O&M manuals, project closeouts, studies, reports, and other records have accumulated across paper archives, shared drives, SharePoint libraries, PDFs, image files, and deeply nested folders. The information exists—but finding the right document when someone actually needs it can be remarkably difficult.
Institutions commonly face millions of legacy documents, inconsistent file names and filing structures, missing metadata, non-searchable PDFs, files stored in multiple locations, and institutional knowledge disappearing as experienced staff retire.
That leads to a simple but important principle:
A document you cannot find is knowledge you cannot effectively use.
Artificial intelligence is beginning to change that equation.
archSCAN, LLC‘s recent work with the University of Pittsburgh and University of Massachusetts Lowell shows how AI can help transform massive, inconsistent facilities archives into structured, searchable information—and also demonstrates why humans still need to remain firmly in the loop.
At the University of Pittsburgh, the facilities environment includes 262 buildings, approximately 15 million gross square feet across five campuses, and thousands of construction projects. The source document environment included two drives containing 380 GB of data, 22,410 folders, and 209,293 files.
The challenge was not simply moving those files into a new system, AerieHub® by Aerie Engineering.
Before migration, the files needed to be reviewed, drawings identified, information extracted, metadata standardized, duplicates addressed, and the resulting records organized so people could actually search them.
Doing that manually across hundreds of thousands of files is enormously time-consuming.
This is where AI became valuable.
The workflow developed for the Pittsburgh project illustrates an important lesson about AI: there isn’t one magic button that turns an unstructured archive into a clean database.
Instead, AI works best as part of a sequence of specialized processes.
The first task was identifying which files were actually drawings.
AI classification was used to analyze the collection and distinguish drawings from other document types across formats including DWG, TIF, JPG, and PDF. The process sorted through approximately 210,000 files and an estimated 300,000 drawings.
That alone changes the economics of a large archival project. Rather than having people open files one at a time simply to determine what they are, AI can perform the first pass at scale.
Identifying a drawing is only the beginning.
Facilities drawings contain enormous amounts of useful information—equipment names, room details, notes, system descriptions, dimensions, specifications, title block information, and more.
AI-assisted text recognition can extract printed and handwritten text, making that content searchable. The presentation demonstrates this visually by searching engineering drawings for terms such as “exhaust air duct” and even specific dimensional descriptions.
The result is a fundamental shift: a drawing stops being just an image and starts becoming searchable information.
Generic AI isn’t enough.
The system needs to understand the organization’s data: building names, addresses, project names, project numbers, terminology, title blocks, symbols, formats, and other facilities-specific conventions.
For the Pittsburgh project, standardized source information was supplied to the LLM, and the process was tested on small batches so accuracy could be continuously improved.
This is one of the most important lessons we’ve learned:
The quality of AI output depends heavily on the quality and context of the information you give it.
Once the documents could be interpreted, AI was used to extract metadata.
At the set level, this included information such as campus, building type, building name, project name, project number, issue, set date, and author.
At the individual drawing level, the process extracted information such as discipline, sheet number, sheet title, floor, and sheet date.
That distinction matters.
Searchability isn’t created simply by scanning a document. It comes from transforming the information inside the document into structured metadata that can be filtered, searched, reported, and connected to the facility it describes.
Extracting information is only part of the challenge.
Legacy collections contain duplicates, superseded drawings, inconsistent abbreviations, conflicting naming conventions, different date formats, and incomplete metadata.
In the Pittsburgh collection, approximately 50% of sets were identified as duplicate or superseded. AI-assisted spreadsheet analysis was used to detect duplicate and near-duplicate records, standardize metadata and naming conventions, normalize abbreviations and terminology, validate completeness, and flag exceptions for human review.
This may ultimately be one of AI’s most valuable roles in document management.
Not replacing the document management system.
Not replacing the archivist.
But performing the repetitive comparison, classification, extraction, and normalization work that historically consumed enormous amounts of staff time.
There is another lesson from these projects that deserves just as much attention:
AI sometimes gets things wrong.
Human verification remains essential.
The difference is that people can spend less time manually entering information and more time reviewing and validating what AI has extracted.
In the Pittsburgh workflow, manual indexing had been running at approximately 300 records per day. With AI performing much of the extraction and people concentrating on quality control, throughput increased to approximately 1,000–1,200 records per day—a three- to four-fold improvement.
That is a much more useful way to think about AI adoption.
The question shouldn’t necessarily be:
“How can AI replace this job?”
A better question is:
“Which repetitive parts of this job can AI perform so our people can concentrate on judgment, verification, exceptions, and higher-value work?”
The UMass Lowell project presented a different challenge.
The university was frustrated with its SharePoint document library. The facilities archivist who had created the original organizational system had left years earlier, coded metadata was difficult for others to understand, and individual drawing pages needed to be reconstructed into usable sets.
Once the collection was analyzed, the scale of the problem became clear.
SharePoint contained 415,194 files. Only 14%—57,315 files—had data associated with them. Another 46,139 files were effectively junk, while 286,402 files had no associated metadata at all.
Traditional migration approaches struggle with a collection like this because there isn’t enough reliable metadata to migrate.
AI provides another option: go back to the documents themselves.
The project used AI to identify drawings, analyze folder paths for clues, classify unknown documents, and use computer vision to capture information directly from drawings. The presentation notes that the newer computer-vision approach produced better results than the earlier LLM-based method used on the Pittsburgh project, reducing QC effort and improving accuracy.
That evolution is significant.
We aren’t simply using AI to automate an old process anymore. We’re learning new ways to reconstruct information that organizations thought they had effectively lost.
The end goal isn’t scanning.
It isn’t OCR.
It isn’t metadata extraction.
And it certainly isn’t “using AI” simply because AI is the technology everyone is talking about.
The goal is usable institutional knowledge.
At Pittsburgh, 108,398 drawings were already available in AerieHub at the time of the presentation, with another 50,000 being processed and quality-controlled. Future collections identified for inclusion include O&M manuals, project closeouts, reports and studies, and emergency documents.
Imagine the difference for a facilities team when someone can search by campus, building, floor, discipline, project, drawing number, document type, date, consultant, keyword, or other structured information rather than trying to remember which shared-drive folder a former employee used ten years ago.
That is when digitization becomes digital transformation.
There is also a measurable economic argument.
Based on the experience summarized in our presentation, AI-enabled document management workflows have produced 30–40% reductions in document-processing costs, while certain highly repetitive workflows, such as AP/invoice processing case studies, have demonstrated reductions as high as 80%. At the same time, AI processing costs themselves need to be monitored as adoption and usage grow.
Cost reduction, however, tells only part of the story.
There is also the value of faster access to information, reduced duplication, preserved institutional knowledge, more consistent metadata, improved decision-making, and less staff time spent hunting through folders.
Those benefits can be much harder to quantify—but anyone who has spent hours looking for the “right” version of a 20-year-old drawing understands their value.
These projects reinforce several practical lessons for organizations considering AI-powered document management:
Start with a real problem, not with AI. Identify the repetitive work, inaccessible information, missing metadata, or search problem you are trying to solve.
Standardize wherever possible. AI becomes significantly more useful when it has authoritative building names, project numbers, terminology, and other reference data to work from.
Work in small batches first. Test, review, adjust, and improve before processing hundreds of thousands of documents.
Use different AI tools for different tasks. Classification, text recognition, metadata extraction, duplicate detection, computer vision, and validation are different problems and may require different approaches.
Keep humans in the loop. AI accelerates the work; knowledgeable people establish whether the results can be trusted.
And perhaps most importantly:
Don’t confuse digitization with accessibility.
Scanning 500,000 documents gives you 500,000 digital documents.
Creating structured, standardized, searchable information gives you something much more valuable:
KNOWLEDGE
AI is giving facilities organizations a practical way to unlock that knowledge at a scale that simply wasn’t feasible before.
The technology will continue to mature. The tools will change. Accuracy will improve. Costs will evolve.
But the opportunity is already here.
For universities and other organizations sitting on decades of facilities documentation, the question may no longer be whether AI belongs in document management.
The more useful question is:
How much institutional knowledge is currently sitting in your files—and what could your organization do if people could finally find it?