Walkthrough
How AI reviews an M&A data room
From a raw folder of PDFs, spreadsheets, and decks to cited, checkable answers — the six steps most AI-assisted data room reviews follow.
01Upload and ingestion
The process starts with the data room itself — typically uploaded as a single archive rather than file by file, since transaction rooms commonly run to hundreds of documents across PDF, Word, Excel, PowerPoint, and CSV formats. Each file is parsed according to its type: text is extracted from PDFs and Word documents, spreadsheets are read sheet by sheet and cell by cell, and slide decks are parsed slide by slide. Native, digital files parse cleanly; a scanned, image-only PDF without OCR won't have extractable text, which is a real limitation worth knowing about upfront.
02Indexing
Once parsed, the content is broken into chunks — a page, a slide, a range of cells — and indexed so the system can quickly find which chunks are relevant to a given question. This is what makes it possible to ask "what's the customer concentration in the top five accounts?" and get an answer pulled from the right spreadsheet tab, instead of the system re-reading the entire room from scratch for every question.
03Question answering
With the room indexed, a deal team can ask the kinds of questions an associate would normally chase down manually — about revenue concentration, debt covenants, litigation exposure, key contracts, or employee headcount. The system retrieves the most relevant chunks, drafts an answer grounded in that evidence, and attaches a citation to the exact source. If it can't find supporting evidence, a well-built system says so rather than guessing.
04Request-list mapping
Beyond one-off questions, most deal teams work from a structured due diligence request list. That list can be mapped against the room's evidence automatically — each line item gets checked against what's actually available, rather than a person manually searching the room for every request. See what a due diligence request list is for more on how these lists are typically structured.
05Gap and conflict detection
Two things tend to slow down a manual review the most: noticing when something is simply missing from the room, and noticing when two documents disagree with each other. Both can be checked systematically — every request-list item gets an outcome (evidence found, partially supported, no evidence found, or conflicting), and cross-document contradictions are flagged instead of relying on someone happening to compare both files. See how AI detects conflicting information for a closer look at that specific step.
06Human review
Every answer, gap, and conflict is a proposal for the deal team to confirm, not a conclusion to act on unverified. The team reviews findings, checks citations against source documents where it matters, and decides what's actually significant for the transaction. This is also where the professional judgment that AI doesn't have — interpreting an ambiguous clause, weighing how material a gap really is — gets applied.
Vaultrix follows this exact sequence end to end — see the full AI due diligence page for how citation verification and human review are built into the workflow, or the AI data room page for how it relates to your existing data room.
07Frequently asked questions
How long does it take to ingest a typical M&A data room?
It depends on room size, but ingestion and indexing for a few hundred documents is typically measured in minutes, not days — after which question-answering is close to instant per query.
What file types can be reviewed?
PDF, Word, Excel, PowerPoint, CSV, and plain text are the common formats. Native, digital files work well; scanned image-only PDFs without OCR won't have extractable text.
Does the AI see documents from other deals?
In a well-built system, no — each deal is a separate, isolated workspace, so documents from one data room are never visible when reviewing another.