Interviews are the raw material of journalism and qualitative research, and the transcript is what makes them usable: quotable, searchable, codeable. Done badly, transcription eats the day. Done well, it is a background task that costs a few minutes of attention.
This is the whole process, from the recorder to a transcript you can quote from.
Before the interview
The most important decision about transcription is made before the interview, not after. A recording made with a decent microphone in a quiet room comes back nearly error-free from any modern tool. A recording from a phone speaker in a busy café produces nothing but work, whatever software you point at it afterwards.
The microphone hierarchy
Bad: phone microphone in a pocket, laptop microphone at a metre's distance, a conference microphone in a large room with one distant speaker.
Acceptable: phone directly in front of the speaker at no more than 30 cm, a USB headset, the built-in microphone of a good laptop in a quiet room.
Good: a lavalier microphone clipped to the shirt, a portable recorder, a studio USB microphone.
Optimal: one lavalier per speaker on a separate track. This is the gold standard: it separates the speakers physically and it lets the transcription engine do clean diarization instead of guessing.
The room
Most research interviews happen in a living room, an office or a café, and few of those are acoustically kind.
- Smooth walls are bad. Curtains, bookshelves and sofas damp the reflections.
- Air conditioning, fridges and fans are the most common unnoticed killers. Switch them off before you start.
- Cafés work with a lavalier close to the speaker. With a phone on the table, background conversation makes the audio unusable.
- Remote interviews over Zoom, Teams or Meet often produce better audio than a cheap external microphone, because the platforms have good noise suppression built in.
Format and backup
Record lossless (WAV) or high-bitrate MP3, at least 128 kbps. Low-bitrate lossy formats create artefacts that models interpret as words, which is to say as errors.
If the interview matters, record on two devices at once. Losing thirty minutes because the recorder was switched off in a bag is a nightmare a backup phone avoids for free.
During the interview
A few habits make the later transcription dramatically easier.
Introduce every speaker by name at the start. "Today I am speaking with Dr Schmidt, professor of psychology in Heidelberg." That gives the model, and you, a clear anchor.
Spell proper nouns once. It ends up in the transcript, and it can go straight into the custom vocabulary for the engine.
Avoid long silences. Even an "mm, yes" helps the model detect a speaker change. Complete pauses invite speaker confusion.
Do not talk over each other. Crosstalk is the single biggest enemy of diarization. If somebody interrupts, pause briefly and let them finish.
Name the topic transitions. "Let us move to the next point, data protection." Good for the listener, and good for searching the transcript later.
After the interview
Step 1: prepare the file
Move the recording to the computer. If you have two tracks, one per speaker, mix them into a single stereo file: most transcription APIs separate speakers better when they arrive on different stereo channels.
If the file is long, over two hours, do not split it. Modern APIs handle long files without trouble, and splitting cuts the context and makes consistent speaker numbering harder.
Step 2: generate the transcript
Standard German or English with clean audio needs the standard model. Dialect, several speakers or a difficult background is what the premium model is for.
In DeepScript:
- Upload the file.
- Choose the model, standard at €0.18 per hour or premium at €0.27.
- Set the language explicitly. Auto-detection is fine, explicit is safer.
- If you have a list of names or technical terms, create it as custom vocabulary.
- Start the upload. A 60-minute interview takes roughly 5 to 10 minutes to process.
Step 3: first pass
The finished transcript arrives with speaker separation and timestamps. Read the first five minutes and the last third: that tells you how clean it is without reading all of it.
If the engine gets particular words consistently wrong, add them to the custom vocabulary and run it again. That pays off quickly if you conduct many similar interviews.
Step 4: correction
Two strategies. Inline in the editor, with audio sync: click, correct, jump to the next spot. Good for medium-length interviews. Export and edit elsewhere, as TXT, DOCX or SRT, in Word, Google Docs or a journalism tool. Good for very long interviews or when several people correct in parallel.
Rule of thumb for correction time:
- Standard speech, clean audio: about 10 minutes per 60 minutes of recording.
- Dialect-coloured speech, mediocre audio: 30 to 45 minutes per hour.
- Strong dialect, difficult audio: 60 to 90 minutes per hour.
For comparison, fully manual transcription takes 4 to 6 hours per hour of recording. Even on difficult material, machine transcription plus correction saves around 70 per cent of the time.
Worked example
- Length
- 60 min
- Model
- Premium
- Cost
- €0.27
A one-hour interview on the premium tier, including speaker separation and custom vocabulary. Add 10 to 45 minutes of correction depending on audio quality.
Quoting and citing
Two habits that save trouble later.
Keep the timestamp with the quote. A quote in a manuscript should carry the interview identifier and the time, so that a fact-checker or a co-author can verify it against the audio in seconds rather than by listening through the file. This is also what makes a transcript defensible if a source disputes what they said.
Do not silently smooth the wording. Repairing grammar in a direct quote changes what the person said. Mark omissions, keep the rest verbatim, and if you tidy filler words, do it consistently and say so in the method section.
Data protection, briefly
Interviews are personal data, and research interviews frequently contain special categories under Article 9 GDPR: health, political opinion, religious belief. That has three practical consequences: get consent for the recording and for the processing, choose a provider that processes inside the EU and does not train on your material, and define when the audio and the transcript get deleted before you start collecting.
For anything with named individuals in a sensitive context, the detail is in this article.
In short
Three pillars: a good recording, the right provider, systematic correction. Roughly twenty minutes of preparation per pillar saves hours per interview, and the preparation is the part nobody can do for you afterwards.
If your material is dialectal, that is the largest single factor and it has its own article. Three transcriptions are free without a credit card, which is enough to test your recording setup at the free tool.