Where does your AI tool send your data? A due-diligence checklist for EU companies
Key takeaways
- Adopting an AI tool is a processor decision, so ask five questions before you buy: where the data is hosted and processed, who the sub-processors are, whether your content is used for training, what is retained and for how long, and whether a real DPA exists.
- Hosting claims have to be checked for the tier you are actually buying, because an EU region offered on the enterprise plan does not cover the plan you signed.
- The sub-processor list matters most for model inference, because that is where your prompts and documents actually travel.
- A commitment that business content is not used for training belongs in the contract rather than in a blog post, and a retrieval tool reading your documents at question time is not the same thing as training.
- Olyteck hosts in the EU (Scaleway, Paris region) under French jurisdiction, publishes a DPA and sub-processor list on the trust center, and does not use customer content to train models.
The fastest way to fail an audit in 2026 is the AI tool someone adopted in 2025. Not because AI is forbidden - it is not - but because nobody asked, before pasting customer contracts into a chat window, the same questions procurement would ask any other processor: where does the data go, who touches it, and what do they do with it?
This checklist is written for the person who has to answer for that decision later: the DPO, the IT manager, the founder who signs the DPA. It works for any AI tool - assistants, meeting transcribers, RAG platforms, coding tools. If your team keeps asking whether it is safe to put company documents into AI, this is how you answer it with evidence instead of a shrug.
The five questions that matter #
- Where is the data hosted and processed? Not where the company is headquartered - where the servers are. "EU region available on the enterprise plan" means the plan you are actually buying may process in the US. Get the region for YOUR tier in writing.
- Who are the sub-processors? Most AI products are assembled: a hosting provider, one or more model providers, analytics, support tooling. Each is a data path. A serious vendor publishes the list; a missing list is a finding in itself. Pay attention to where the model inference happens - that is where your prompts and documents actually travel.
- Is your content used for training? The answer must be an unambiguous no for business data, in the contract, not in a blog post. The distinction to look for: retrieval-based tools read your documents at question time to answer; that is not training, and a vendor who explains the difference clearly usually has thought about the rest too.
- What is retained, and for how long? Prompts, uploaded documents, generated answers, logs. Some retention is legitimate (abuse prevention, your own history feature); indefinite and unspecified is not. Ask how deletion works when you offboard - "your workspace is deleted" should include backups, on a stated timeline.
- Is there a real DPA, and does it match reality? A signable data processing agreement naming the sub-processors, the transfer mechanism for any non-EU hop (SCCs, adequacy), breach notification duties, and audit rights. Then the one-minute reality check: does the privacy policy on the website contradict what sales just told you?
Second-order checks that separate vendors #
- Data minimisation by design. The best answer to "how do you protect the files" is "we never store them". Tools designed to derive what they need and discard the rest (counts, findings, vectors, scores rather than raw copies) shrink the entire risk surface.
- Access from where, by whom? Support staff jurisdictions matter: EU hosting with unrestricted remote admin access from a non-EU parent company undermines the point.
- An exit that preserves your work. Your prompt libraries, answer libraries, and configurations are an asset. Can you export them in a usable format?
- Grounded, cited answers as a safety feature. A tool that answers only from the documents you gave it, and cites the source, cannot invent a fact into a contract and cannot wander into data you never provided. Grounding limits leakage in both directions, which makes it a security property, not just a quality one.
The uncomfortable rule of thumb: if a vendor cannot answer these five questions in one email, your data is the product research. The vendors with good answers reply the same day, because they get asked daily and built for it.
Where Olyteck stands on each #
We publish our answers because we ask other vendors the same questions. Olyteck's products are hosted in the EU (Scaleway, Paris region) under French jurisdiction, with a published DPA and sub-processor list on the trust center. Customer content is not used to train models. The security products follow "counts findings, never files" - they store results, scores and metadata, not copies of your documents or messages. Ask, our RAG assistant, answers from your documents at question time and cites its sources.
Run the checklist on us too - that is what it is for. Then run it on the AI tools your teams already use, starting with the ones nobody remembers approving. The list of tools that fail question one is usually the real output of the exercise.
FAQ #
Is it safe to put company documents into an AI tool? #
It can be, but the safety is a property of the specific tool and the contract behind it, not of AI in general. Work through the five questions above: the hosting and processing region for the tier you are actually buying, the sub-processor list, whether your content is used for training, what is retained and for how long, and whether there is a signable DPA. A vendor who cannot answer those in one email has given you an answer anyway.
Does the GDPR forbid using AI tools? #
No. AI is not forbidden; the problem is adopting a tool without asking the questions procurement would ask of any other processor. In practice that means knowing where the data goes, who touches it, what they do with it, and having a data processing agreement that names the sub-processors and the transfer mechanism for any non-EU hop. This checklist is a practical purchasing aid rather than legal advice, so bring your DPO or counsel in on your own situation.
How can I tell whether my data leaves the EU? #
Ask for the processing region that applies to your plan in writing, because "EU region available on the enterprise plan" may not describe the tier you signed. Then read the published sub-processor list and look specifically at where model inference happens, since that is where your prompts and documents actually travel. Check support access as well: EU hosting combined with unrestricted remote admin access from a non-EU parent company undermines the point.
Can an AI vendor train its models on my company data? #
That depends entirely on what the contract says, which is why the answer needs to be an unambiguous no for business data in the agreement itself and not in a blog post. Ask what the commitment covers: prompts, uploaded documents, generated answers and logs. It is worth knowing the distinction too, because a retrieval-based tool that reads your documents at question time to answer is not training on them, and a vendor who explains that difference clearly has usually thought about the rest.
One useful Microsoft 365 email a month
New guides, findings from real tenants, and the occasional checklist. No sales sequence, unsubscribe in one click.