Where does your AI tool send your data? A due-diligence checklist for EU companies
The fastest way to fail an audit in 2026 is the AI tool someone adopted in 2025. Not because AI is forbidden - it is not - but because nobody asked, before pasting customer contracts into a chat window, the same questions procurement would ask any other processor: where does the data go, who touches it, and what do they do with it?
This checklist is written for the person who has to answer for that decision later: the DPO, the IT manager, the founder who signs the DPA. It works for any AI tool - assistants, meeting transcribers, RAG platforms, coding tools.
The five questions that matter
- Where is the data hosted and processed? Not where the company is headquartered - where the servers are. "EU region available on the enterprise plan" means the plan you are actually buying may process in the US. Get the region for YOUR tier in writing.
- Who are the sub-processors? Most AI products are assembled: a hosting provider, one or more model providers, analytics, support tooling. Each is a data path. A serious vendor publishes the list; a missing list is a finding in itself. Pay attention to where the model inference happens - that is where your prompts and documents actually travel.
- Is your content used for training? The answer must be an unambiguous no for business data, in the contract, not in a blog post. The distinction to look for: retrieval-based tools read your documents at question time to answer; that is not training, and a vendor who explains the difference clearly usually has thought about the rest too.
- What is retained, and for how long? Prompts, uploaded documents, generated answers, logs. Some retention is legitimate (abuse prevention, your own history feature); indefinite and unspecified is not. Ask how deletion works when you offboard - "your workspace is deleted" should include backups, on a stated timeline.
- Is there a real DPA, and does it match reality? A signable data processing agreement naming the sub-processors, the transfer mechanism for any non-EU hop (SCCs, adequacy), breach notification duties, and audit rights. Then the one-minute reality check: does the privacy policy on the website contradict what sales just told you?
Second-order checks that separate vendors
- Data minimisation by design. The best answer to "how do you protect the files" is "we never store them". Tools designed to derive what they need and discard the rest (counts, findings, vectors, scores rather than raw copies) shrink the entire risk surface.
- Access from where, by whom? Support staff jurisdictions matter: EU hosting with unrestricted remote admin access from a non-EU parent company undermines the point.
- An exit that preserves your work. Your prompt libraries, answer libraries, and configurations are an asset. Can you export them in a usable format?
The uncomfortable rule of thumb: if a vendor cannot answer these five questions in one email,
your data is the product research. The vendors with good answers reply the same day, because they
get asked daily and built for it.
Where Olyteck stands on each
We publish our answers because we ask other vendors the same questions. Olyteck's products are hosted in the EU (Scaleway, Paris region) under French jurisdiction, with a published DPA and sub-processor list on the trust center. Customer content is not used to train models. The security products follow "counts findings, never files" - they store results, scores and metadata, not copies of your documents or messages. Ask, our RAG assistant, answers from your documents at question time and cites its sources.
Run the checklist on us too - that is what it is for. Then run it on the AI tools your teams already use, starting with the ones nobody remembers approving. The list of tools that fail question one is usually the real output of the exercise.