# Where does your AI tool send your data? A due-diligence checklist for EU companies

> Is it safe to put company documents into AI? The five questions to answer before you adopt an AI tool: hosting, sub-processors, training, retention, DPA.

Source: https://olyteck.com/blog/where-does-your-ai-tool-send-your-data
Published: 2026-07-28 | Updated: 2026-07-28
Author: Oleg Garasym, Olyteck (France, EU-hosted)
Product: Olyteck Ask
Topics: gdpr, ai tools, data residency, confidential documents, eu hosting, dpa, sub-processors, sovereignty, due diligence
License: free to quote with attribution to Olyteck and a link to the source URL.

---

The fastest way to fail an audit in 2026 is the AI tool someone adopted in 2025. Not because AI is
forbidden - it is not - but because nobody asked, before pasting customer contracts into a chat
window, the same questions procurement would ask any other processor: where does the data go, who
touches it, and what do they do with it?

This checklist is written for the person who has to answer for that decision later: the DPO, the IT
manager, the founder who signs the DPA. It works for any AI tool - assistants, meeting
transcribers, RAG platforms, coding tools. If your team keeps asking whether it is safe to put
company documents into AI, this is how you answer it with evidence instead of a shrug.

## The five questions that matter

1. **Where is the data hosted and processed?** Not where the company is headquartered - where the servers are. "EU region available on the enterprise plan" means the plan you are actually buying may process in the US. Get the region for YOUR tier in writing.
2. **Who are the sub-processors?** Most AI products are assembled: a hosting provider, one or more model providers, analytics, support tooling. Each is a data path. A serious vendor publishes the list; a missing list is a finding in itself. Pay attention to where the *model inference* happens - that is where your prompts and documents actually travel.
3. **Is your content used for training?** The answer must be an unambiguous no for business data, in the contract, not in a blog post. The distinction to look for: retrieval-based tools read your documents at question time to answer; that is not training, and a vendor who explains the difference clearly usually has thought about the rest too.
4. **What is retained, and for how long?** Prompts, uploaded documents, generated answers, logs. Some retention is legitimate (abuse prevention, your own history feature); indefinite and unspecified is not. Ask how deletion works when you offboard - "your workspace is deleted" should include backups, on a stated timeline.
5. **Is there a real DPA, and does it match reality?** A signable data processing agreement naming the sub-processors, the transfer mechanism for any non-EU hop (SCCs, adequacy), breach notification duties, and audit rights. Then the one-minute reality check: does the privacy policy on the website contradict what sales just told you?

## Second-order checks that separate vendors

- **Data minimisation by design.** The best answer to "how do you protect the files" is "we never store them". Tools designed to derive what they need and discard the rest (counts, findings, vectors, scores rather than raw copies) shrink the entire risk surface.
- **Access from where, by whom?** Support staff jurisdictions matter: EU hosting with unrestricted remote admin access from a non-EU parent company undermines the point.
- **An exit that preserves your work.** Your prompt libraries, answer libraries, and configurations are an asset. Can you export them in a usable format?
- **Grounded, cited answers as a safety feature.** A tool that answers only from the documents you gave it, and cites the source, cannot invent a fact into a contract and cannot wander into data you never provided. Grounding limits leakage in both directions, which makes it a security property, not just a quality one.

> The uncomfortable rule of thumb: if a vendor cannot answer these five questions in one email,
> your data is the product research. The vendors with good answers reply the same day, because they
> get asked daily and built for it.

## Where Olyteck stands on each

We publish our answers because we ask other vendors the same questions. Olyteck's products are
hosted in the EU (Scaleway, Paris region) under French jurisdiction, with a published DPA and
sub-processor list on the trust center. Customer content is not used to train models. The security
products follow "counts findings, never files" - they store results, scores and metadata, not
copies of your documents or messages. Ask, our RAG assistant, answers from your documents at
question time and cites its sources.

Run the checklist on us too - that is what it is for. Then run it on the AI tools your teams
already use, starting with the ones nobody remembers approving. The list of tools that fail
question one is usually the real output of the exercise.

## FAQ

### Is it safe to put company documents into an AI tool?

It can be, but the safety is a property of the specific tool and the contract behind it, not of AI in general. Work through the five questions above: the hosting and processing region for the tier you are actually buying, the sub-processor list, whether your content is used for training, what is retained and for how long, and whether there is a signable DPA. A vendor who cannot answer those in one email has given you an answer anyway.

### Does the GDPR forbid using AI tools?

No. AI is not forbidden; the problem is adopting a tool without asking the questions procurement would ask of any other processor. In practice that means knowing where the data goes, who touches it, what they do with it, and having a data processing agreement that names the sub-processors and the transfer mechanism for any non-EU hop. This checklist is a practical purchasing aid rather than legal advice, so bring your DPO or counsel in on your own situation.

### How can I tell whether my data leaves the EU?

Ask for the processing region that applies to your plan in writing, because "EU region available on the enterprise plan" may not describe the tier you signed. Then read the published sub-processor list and look specifically at where model inference happens, since that is where your prompts and documents actually travel. Check support access as well: EU hosting combined with unrestricted remote admin access from a non-EU parent company undermines the point.

### Can an AI vendor train its models on my company data?

That depends entirely on what the contract says, which is why the answer needs to be an unambiguous no for business data in the agreement itself and not in a blog post. Ask what the commitment covers: prompts, uploaded documents, generated answers and logs. It is worth knowing the distinction too, because a retrieval-based tool that reads your documents at question time to answer is not training on them, and a vendor who explains that difference clearly has usually thought about the rest.

