Audio is among the most sensitive categories of data a company handles. A voice is biometrically identifiable. A meeting recording contains names, opinions and trade secrets. An interview reveals personal views that were never meant for an audience.
And yet companies upload audio into transcription services every day without ever having reviewed how those services handle data. That is not only a legal exposure; it destroys trust, and it carries fines.
This article sets out which GDPR requirements actually apply, where the risk sits with US-based providers, and how to check a vendor before the first file leaves your network.
Why audio deserves more care than most data
Under the GDPR, audio is personal data as soon as a speaker is identifiable, which for speech recordings is effectively always. The human voice is a biometric characteristic: even with no name mentioned, a person can be identified from it.
A transcript captures the spoken content. The underlying audio file carries considerably more:
- Voice profiles that permit biometric identification
- Emotional state, readable from tone and pace
- Background sound that can reveal location
- Content that may include trade secrets, third-party personal data, or special categories under Article 9: health, political opinion, religious belief, trade union membership
Article 9 raises the bar sharply. If an employee mentions their health in passing during a meeting and that meeting is transcribed, you are processing health data, with every consequence that follows.
The everyday version of this problem is mundane: someone searches for "audio to text", finds a free tool, uploads the recording and pastes the result into a document. That the file now sits on US servers, and may be used to improve a model, occurs to nobody in the chain. This is not an edge case. It happens in companies of every size.
What the GDPR requires
Using an external transcription service is processing on your behalf under Article 28. That has concrete consequences.
Article 28: the data processing agreement
Before any personal data goes to a transcription service, a data processing agreement must be in place. It governs subject matter and duration, nature and purpose, the types of data and categories of data subjects, the controller's rights and obligations, the processor's technical and organisational measures, the rules for sub-processors, and deletion or return of the data at the end.
Without that agreement, using the service is unlawful. There is no exception and no threshold below which it stops mattering.
Many free online transcription tools offer no such agreement. Some do not mention the topic in their terms at all. For a company, those tools are simply not usable, regardless of output quality.
Article 44: transfers outside the EEA
Sending personal data outside the European Economic Area requires one of a short list of instruments: an adequacy decision for the destination country, standard contractual clauses plus supplementary measures, binding corporate rules, or explicit consent, which is rarely workable at scale.
For audio, the supplementary-measures bar is high, because the provider needs the audio in cleartext in order to transcribe it. The measure regulators point to first, strong encryption with keys held in the EU, is exactly the one a US transcription provider cannot offer.
Article 17: the right to erasure
A data subject can require deletion. In practice you must be able to show that the audio file is gone, the transcript is deleted or anonymised, the provider has deleted its copies, and nothing remains in backups or training sets.
A provider that uses customer audio for model training makes this impossible to satisfy. Data that has entered a training set cannot be removed from it selectively.
Article 25: data protection by design
Protection has to be built in rather than added later. For vendor selection that means data handling is a selection criterion, not a question you get to after the pilot.
The four checks
Processing location. Where is the audio processed, not only stored? Does the entire chain stay inside the EEA? Is the location documented somewhere you can point an auditor at?
The agreement. Does the provider offer a DPA that contains everything Article 28 requires, with the technical and organisational measures written down? Is it signed before the first upload, rather than after the first incident?
Deletion. When is audio deleted after processing, and when are transcripts? Does it happen automatically or does someone have to remember? Can the provider evidence it?
Sub-processors. Which ones are involved? Is audio handed to third-party AI providers such as OpenAI, Google or AWS? Is any of it used for training? How are you told when a new sub-processor is added?
Add access control to the list: who can reach the files, is data encrypted in transit and at rest, and is access logged.
Where US cloud providers create a problem
Many transcription services run on AWS, Google Cloud or Azure. Even with servers physically in Europe, two legal issues remain.
Schrems II. In July 2020 the Court of Justice invalidated the EU-US Privacy Shield (case C-311/18), finding that US law does not provide adequate protection for EU personal data, mainly because of the reach of US intelligence access. The successor framework from 2023 exists, but several supervisory authorities and practitioners doubt it would survive the same scrutiny.
The CLOUD Act. The 2018 act obliges US companies to give US authorities access to stored data regardless of where it physically sits. A board meeting transcribed through a US service on servers in Frankfurt is legally reachable by US authorities, without a European court order and without notice to anyone affected.
The practical consequence is the part most buyers miss: a large share of transcription vendors call the APIs of OpenAI, Google or Amazon behind the scenes. Your audio reaches those companies, you have no direct contract with them, their handling is outside your control, and a deletion request has to travel through several hops. An EU-registered vendor with a US speech API in the middle has the same transfer problem as a US vendor, only less visibly.
That is the question worth asking in a vendor call, and it is a specific one: where does the GPU that processes my audio physically live?
How DeepScript answers it
Own servers in Germany. All audio is processed on dedicated servers DeepScript operates at Hetzner in Germany, in ISO 27001 certified data centres. Not a virtual machine at a US hyperscaler: physical servers under German jurisdiction. Upload, recognition, transcription and storage all happen there.
No third-party AI. The speech models are ours and run on our infrastructure. No OpenAI Whisper API, no Google Speech-to-Text, no AWS Transcribe. No transfer to US companies, no dependency on a third party's data practices, and no route by which your data could reach someone else's training set.
Deletion you configure. Audio is deleted after processing; the default retention is 30 days and immediate deletion is available. Transcripts stay until you remove them.
A real DPA. Article 28 compliant, listing purpose, technical and organisational measures, deletion periods, obligations on both sides, and the complete sub-processor list, all of them in the EU. It can be signed directly on the website, without a sales call.
A published sub-processor list. If you want to know who can reach your data, the answer is written down rather than described.
What to do with this internally
If you are the data protection officer: inventory which transcription tools are actually in use, including the ones individual departments adopted without asking. Run the four checks against each. Verify a signed DPA exists for every one. Define an approval path for new tools that touch audio. And train people on the single fact that changes behaviour: audio is personal data.
If you run IT: look for shadow usage, publish a short list of approved tools so people have a compliant default, block the rest if you can, and offer an API integration so the compliant path is also the convenient one. Shadow IT is usually a symptom of the approved route being slower.
If you are on the board: fines reach 4% of global annual turnover, a compliant service costs less than one incident, and a written policy on audio processing is cheap to adopt and expensive to lack.
Conclusion
Transcription is a genuinely useful tool. The sensitivity of audio just means the choice of provider deserves more care than the choice of a note-taking app.
The requirements are not vague: sign the agreement, verify the processing location, know the sub-processors, secure the deletion path. A service that fails those four is not an option for a company, however good its transcription quality is.
Our own answers, in detail and with the sub-processor list, are in the Trust Center. If your meetings run on Zoom or Teams, the compliant pattern for those specifically is in this article.