Speech-to-text and text-to-speech API

Speech AI that understands and speaks Swiss German.

Transcribe Züritüütsch, Bärndütsch and every other dialect into clean Standard German text, and answer in real dialect, not an accent. One API for realtime and batch, built for voice bots, contact centres and call analytics.

No credit card. Billed per second after that.

POST /v1/tts
curl -X POST https://api.suisse-speech.ch/v1/tts \
  -H "X-API-Key: $SUISSE_SPEECH_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Grüezi, schön sind Sie da.",
       "voice": "anna", "language": "de-CH", "dialect": "zurich"}' \
  --output greeting.mp3
< 1 sto the first audio in realtime synthesis
250 msbetween live transcript updates
5 + 14Swiss German dialects: 5 hand-tuned, 14 more regions
12voices, each one speaks every language

Two directions, one API

The same key, the same language codes and the same error format for listening and speaking.

Speech-to-Text

Swiss German speech in, Standard German text out, with word timings and confidence. Stream it live or send whole recordings.

  • Live transcripts over WebSocket, updated every 250 ms
  • Detects German, French, Italian and English on the fly
  • Your own vocabulary: names, products, places
  • Long recordings as jobs, up to 200 MB

Learn more →

Text-to-Speech

Write in Standard German and hear real dialect, with its own words, grammar and sound. Twelve voices, from telephone to studio quality.

  • Zürich, Bern, Basel and Lucerne German, hand-tuned
  • Realtime streaming for voice bots, studio model for prepared audio
  • Swiss numbers, dates and CHF amounts spoken correctly
  • Telephone-ready: G.711 µ-law and A-law at 8 kHz

Learn more →

Why teams choose Suisse Speech

Real dialect, not an accent

Swiss German output uses dialect wording and grammar. It is not Standard German read with a Swiss accent.

Swiss details handled

Phone numbers grouped the Swiss way, per dialect. Dates, times and CHF amounts in their spoken form.

One conversation, four languages

Recognition detects German, French, Italian or English. Synthesis answers with the same language code.

Your audio is not stored

Audio, text and transcripts are processed in memory. We keep only content-free metadata for billing.

Free sandbox for your tests

Sandbox keys return realistic test audio and transcripts, never billed. Made for continuous integration.

Honest about limits

Every refusal names its reason: your rate, your plan, your credit or our capacity. Nothing is queued silently.

Built for Swiss voice products

Voice bots and contact centres

Listen with the realtime transcriber, answer with the realtime voice. The caller hears the reply in their language and, if you like, in their dialect.

Call transcription and analytics

Turn recorded calls into searchable Standard German text with word timings, as a batch job that runs well below real time.

IVR prompts and announcements

Produce prompts, greetings and announcements in studio quality, in the dialect of your region. Phone numbers are read out the Swiss way.

Media, e-learning and accessibility

Voice-over for content in Swiss German, French and Italian, and transcripts of spoken Swiss German for subtitles and archives.

Pay per second, start with an hour

Speech-to-text from CHF 1.35 per hour of audio, text-to-speech from CHF 4.60. No subscription and no minimum per request.

Frequently asked questions

Can the API transcribe Swiss German dialect?

Yes. Swiss German speech from any region is transcribed. Because Swiss German has no standard spelling, the transcript is written in Standard German, which is also what search, analytics and language models handle best.

Does the synthesis speak real dialect or Standard German with an accent?

Real dialect. For Swiss German, the text is rendered into the dialect: its wording, grammar and pronunciation. Zürich, Bern, Basel and Lucerne German are hand-tuned. 14 further regions are available on a best-effort basis.

Which languages are supported?

Swiss German, Swiss High German, Swiss French and Swiss Italian first. German, French, Italian and English with their regional varieties, and about 90 language varieties in total.

How is usage billed?

Per second of audio, with no minimum per request. Realtime (streaming) and batch have separate rates, listed on the pricing page. You top up prepaid credit, and every new account starts with 60 free minutes.

Do you store my audio or transcripts?

No. Audio, text and transcripts are processed in memory and not stored. Results of batch jobs are kept only until you collect them, for 48 hours at most.

Try it on your own audio

Get 60 free minutes of speech-to-text and text-to-speech. No credit card required.