DeepScript
Accuracy & languages

Transcribing Swiss German, Bavarian and Austrian Reliably

Why most transcription APIs fail on German dialects, how the providers compare on Mundart audio, and the five practical steps that move accuracy from 70 to 95 per cent.

Julian KisselJulian KisselAugust 14, 2026Updated August 14, 20266 min read

Key takeaways

  • Three things decide dialect quality: a provider with DACH tuning, custom vocabulary that is actually maintained, and a usable recording.
  • Generalist US APIs deliver mediocre results on dialectal audio, and the cost lands in review time rather than in the invoice.
  • Proper nouns and dialect words belong in the vocabulary list rather than in the hope that the engine guesses them.
  • Auto-detection is unreliable on Swiss German: the engine sometimes guesses Dutch, so set the language explicitly.

Anyone who has tried to transcribe a Swiss German interview with a standard API knows the result: a text that reads like a bad translation, full of words that were never said. This is not a bug in a particular product. It follows from how these models are built, and it is worth understanding before you blame the vendor.

If your material comes from Switzerland, Austria or southern Germany, this is the single largest accuracy factor you have, larger than the choice between two well-regarded APIs.

Why dialects are hard

A transcription engine is a neural network mapping audio to text units. During training it has heard millions of hours of speech, overwhelmingly in the standard variety of the target language. For German that means newsroom pronunciation and careful Hochdeutsch.

As soon as a speaker departs from it, three things happen at once.

Phonetic shifts. In Bavarian a hard k often softens; in Swabian the final e disappears. The engine knows these phonemes but hears them in the wrong context and writes what its training taught it to expect.

Lexical particularities. Grüezi, Servus, Jause, Bub, Velo: words that do not exist in the standard variety or carry a different meaning there. A model that barely heard them replaces them with the nearest plausible thing, which is usually wrong.

Grammatical variation. Bavarian word order differs from standard German. A model trained on the standard "corrects" the grammar while transcribing, and in doing so changes the meaning rather than the spelling.

The result: on pure Swiss German dialect, word accuracy of even the best US APIs falls to somewhere between 50 and 70 per cent. On moderately coloured German, Munich standard speech for example, the range is 80 to 90 per cent depending on provider.

How the providers handle DACH dialects

These are observations from evaluations on customer material rather than a published benchmark. Treat them as orientation and measure your own recordings before deciding.

OpenAI Whisper. The best-known open model and the basis of many commercial APIs, trained on 680,000 hours of multilingual audio of which German is a noteworthy but not dominant share. On standard German it does well, around 94 per cent. On Swiss German it falls to roughly 60 to 70 per cent, and on strongly Bavarian or Austrian material to 75 to 85. Strength: open source, runs locally. Weakness: no fine-tuning on DACH varieties and no custom vocabulary at all.

Google Speech-to-Text. Google offers a model for German (Switzerland), which is honourable but primarily trained on Swiss standard German, meaning what is spoken on Swiss television, not on Mundart. On a classic dialect interview it reaches its limits quickly. Strength: several regional variants, solid API. Weakness: Switzerland is not Swiss German, and processing is primarily in the US.

AssemblyAI Universal-2. Trained for generalisation rather than for dialect. Solid on standard German, with no explicit DACH focus. Its boost-words feature helps with proper nouns, not with phonetics. Strength: high general accuracy. Weakness: no dialect tuning, US servers.

DeepScript Premium. Built for the DACH market and fine-tuned on Austrian, Swiss and southern German material. On strongly dialectal recordings we typically land 5 to 15 percentage points above generalist US APIs. Strength: DACH tuning, own servers in Germany, custom vocabulary in every tier. Weakness: smaller language coverage outside Europe.

Five steps that move the number

Whichever provider you use, these make the difference between 70 and 95 per cent.

1. Actually use custom vocabulary

Proper nouns, technical terms and dialect words that occur in the material belong in the vocabulary list. For Mundart recordings:

  • Swiss German: Grüezi, Sitzig, Velo, Znacht, Zvieri
  • Bavarian: Grüß Gott, Pfiat di, Servus, Brotzeit, Maß
  • Austrian: Jause, Erdäpfel, Marille, Kassa, Sackerl

Even a list of 30 to 50 words drops the error rate on frequent terms sharply. This is the cheapest intervention available and the one most often skipped.

2. Take the premium model when the dialect is strong

Standard speech does not need it. Clearly dialectal recordings, especially Mundart interviews, do. Going from €0.18 to €0.27 per hour pays for itself on a 60-minute interview the moment it saves half an hour of correction.

3. Audio quality is the biggest lever

Even the best engine struggles with bad audio. Microphone distance of 15 to 30 cm on a headset, 50 to 80 cm on a stand. Watch for background noise: fridges, air conditioning and traffic are the usual killers. At least 64 kbps MP3, better 128, and a WAV file beats a compressed call recording. For multi-speaker recordings, one microphone per speaker if you can, which improves both the transcription and the speaker separation.

4. Set the language explicitly

Auto-detection works well on clear standard speech and is unreliable on Swiss German, where the engine sometimes guesses Dutch or a Scandinavian language. Set the language to German explicitly, or to the regional variant if your provider offers one.

5. Turn on speaker separation

On multi-speaker dialect interviews, diarization does more than improve readability: it improves word accuracy, because the model adapts per speaker instead of averaging across everyone. DeepScript includes it in every tier, so there is no reason to leave it off.

What this means for the vendor choice

If your material is standard German, almost any credible API will do and the decision comes down to price, data residency and API quality. If your material is dialectal, the ranking changes and the gap is large enough that it dominates every other criterion. A provider five points better on your actual audio saves roughly twenty minutes of review per hour transcribed, which is worth more than the entire transcription cost.

The evaluation that answers this takes an afternoon: take ten representative recordings from your real archive, run them through two or three candidates, and measure word error rate and correction time rather than reading the marketing pages.

Our own comparison of the providers, with prices and language coverage, is in this article. Three transcriptions are free without a credit card, which is enough to test a dialect sample at the free tool.

Swiss GermanBavariandialecttranscriptionDACHaccuracy
Accuracy & languages

How accuracy is measured, why vendor numbers aren't comparable, and what happens with dialects.

Related pages

Try it yourself?

Three transcriptions free, no credit card. Data stays in Germany.

Transcribing Swiss German, Bavarian and Austrian Reliably | DeepScript