Can you trust an AI with your data room?

    The question that decides this purchase, and it deserves a mechanical answer rather than reassurance. Here are the four ways it can go wrong, including the one nobody can engineer away.

    Yash Kadam · Last reviewed 6 October 2026

    The four failure modes

    Pointing a language model at confidential documents can fail in four distinct ways. They are not equally serious and they do not have the same fix.

    1. Cross-reader leakage. The model answers reader A using a document only reader B was granted. The serious one.
    2. Cross-tenant leakage. The model answers using another company’s documents entirely. Catastrophic, and the one that should be tested rather than asserted.
    3. Training on your content. Your documents improve a model that then serves someone else.
    4. Confident wrongness. The model answers fluently and incorrectly. No access control helps here.

    1. Cross-reader leakage. Where the limit is applied

    This is the one that matters most in a data room, because different investors legitimately see different things.

    There are two places to apply the access limit, and only one works:

    • After retrieval. Search everything, then filter the output. This is materially weaker and the weakness is subtle: the model has already read the restricted text, and a summary that has absorbed it can leak its substance without ever naming the source. “The founder vesting schedule is not shared with you, but the team section suggests a four-year cliff” is a leak, and output filtering does not catch it.
    • At search time, the document set handed to the model is filtered to what this reader may see before any retrieval happens. Restricted documents are never candidates, so they cannot be quoted, summarised, or confirmed to exist.

    XDrop AI applies it at search time. Concretely: every request resolves the caller’s workspace membership and role before any data is returned, and the document set given to the AI is filtered to what that caller is permitted to read. The practical test is that the same question from two investors with different grants returns different answers, and for one, a refusal rather than a hedge.

    This is the question to put to any vendor, and the answer should be mechanical. If it is “we filter the response”, the boundary is cosmetic.

    2. Cross-tenant leakage. Tested, not asserted

    Every record carries a workspace boundary, and every request resolves the caller’s membership before data is returned. Access across workspaces happens only through deliberate sharing by an authorised user, never by default.

    More usefully: that boundary is covered by an automated regression suite. Anonymous access, privilege escalation, cross-tenant reads, storage-key masking and the legacy grant path. So a regression fails a test rather than leaking data.

    That distinction is the one worth caring about. “We isolate tenants” is a claim every vendor makes. “The isolation breaks a test when it regresses” is a different kind of statement.

    3. Training, the answer should be a flat no

    Your documents are not used to train models. The Privacy Policy names every sub-processor, its location, and its purpose. Published rather than supplied on request, specifically because an investor’s counsel will ask and the answer should not require an email.

    When you evaluate any AI tool, this is a contract question, not a technology question. Ask where it is written down.

    4. Confident wrongness, the one nobody solves

    A language model can produce a fluent, specific, wrong answer. No access control prevents this, and any vendor claiming otherwise is overselling.

    What can be done is make wrongness checkable. Every answer in XDrop AI is cited to the source document and the page it came from, so an investor can open the source and verify rather than trusting the summary. An uncited answer in diligence is worse than no answer, because it cannot be checked and will be relied upon anyway.

    The honest framing: the AI removes the repetitive half of diligence, the questions whose answers are already in the documents. It does not replace reading the documents that decide the round.

    What else is in place

    • Every AI query is audited, who asked, which documents were retrieved, what was answered. Which makes an answer reconstructable later, by you.
    • Passwordless sign-in via one-time codes to email and mobile. No account passwords are stored, so there is no password database to breach.
    • httpOnly session cookies, so the session token cannot be read by client-side scripts.
    • Session management. Active sessions and devices can be viewed and revoked.
    • NDA gating, per boardroom, using your own agreement, with time-limited or revocable access.
    • Hosting and object storage in India (Mumbai), with TLS in transit.
    • Breach notification to affected users and the Data Protection Board as required by the DPDP Act, 2023.

    The full detail (tenant isolation, authentication, encryption, auditing, sub-processors, retention and deletion, incident response, responsible disclosure) is on the security page.

    What we do not claim

    XDrop AI makes no claim of SOC 2, ISO 27001 or any other formal certification. We will claim them when they have been independently audited and awarded, and not before.

    We mention this because the absence is deliberate and you should know how to read it. A vendor listing certifications it does not hold is telling you something about how it treats every other claim on the page.

    Five questions to ask any AI data room vendor

    1. Is the access limit applied at search time or after retrieval? Ask for the mechanism.
    2. Is tenant isolation covered by automated tests, or asserted in a document?
    3. Is every answer cited to a source document and page?
    4. Where is my data stored, and which sub-processors see it? Is that list published?
    5. Which certifications do you actually hold, as opposed to align with?

    If a vendor cannot answer the first question mechanically, the rest does not matter much. Related reading: what is an AI data room and who should see what.

    XDrop AI is a data room for Indian fundraising, with an AI that answers investor questions and cannot read what you have not shared.

    Start free