How to transcribe Arabic phone calls accurately (and why most models fail on them)

A phone line throws away most of the sound a speech model listens to. Here is what that does to Arabic transcription, what we measured, and how to set up call transcription that holds up.

nutq teamPublished Updated 4 min read

A headset resting on a desk in a quiet call centre at dusk, with rows of empty workstations and city lights blurred behind it.

If you transcribe Arabic customer calls, you already know the pattern: a vendor's demo on clean audio looks excellent, and the transcripts from your own call centre do not. That gap is not bad luck. To transcribe Arabic phone calls well you are solving a different problem from transcribing podcasts, and it needs to be tested as one.

Why is Arabic phone audio so hard to transcribe?

Two things happen at once on a call.

The line throws sound away. Classic telephone audio is sampled at 8 kHz and passes roughly 300–3,400 Hz. Much of what distinguishes similar consonants lives above that band. Add a compression codec, a mobile connection that drops packets, and a caller on speakerphone in a car, and a model gets a fraction of the signal it was trained on.

The caller speaks dialect. Customer calls in the Gulf are in Gulf Arabic, often with English product names mixed in. A model that leans on Modern Standard Arabic when unsure — see Gulf Arabic vs MSA — leans hardest exactly when the audio is worst.

The two compound. Less signal means more guessing, and a model that guesses in MSA rewrites more of what was said.

How much worse do models get on phone calls?

We tested on the audio that matters: a genuine narrow-band call-in, not telephone audio simulated from studio recordings. On it, nutq made 40% fewer character errors than the leading Arabic vendor.

That test sits inside a wider evaluation of 986 clips — broadcast, spontaneous interviews and telephony — each scored against a human transcript. We report character error rate rather than word error rate because Arabic attaches prefixes and suffixes to words, so a single wrong letter can count as a whole wrong word; characters give a fairer picture of how close the text is.

How do you transcribe recorded Arabic calls?

Send each recording to the transcription endpoint — one request per file:

curl -X POST https://nutq.dev/v1/transcribe \
  -H "Authorization: Bearer $NUTQ_API_KEY" \
  -F "file=@call-2026-09-24-1043.wav" \
  -F "language=ar"

A few practical notes for call centres:

  • Keep the original recording. Do not upsample or "enhance" 8 kHz audio before sending it; that adds nothing the model can use and can add artefacts.
  • Split stereo recordings. If your recorder puts the agent on one channel and the customer on the other, send them separately and you get a clean transcript per speaker.
  • Batch is fast. At roughly 22× realtime, an hour of calls comes back in under three minutes. Files up to one hour and 100 MB go in a single request.
  • Log usage. Each response states exactly what the request cost; failed requests are never billed.

The full request and response are covered in the Arabic speech to text API guide.

How do you transcribe live Arabic calls?

For live calls there is a WebSocket endpoint. You push raw 16-bit PCM and receive a transcript each time the speaker pauses:

wss://worker.nutq.dev/v1/listen?encoding=linear16&sample_rate=8000&channels=2&channel=1
ParameterWhat it does
sample_rate8000 to 48000 — send telephone audio at its real 8 kHz
channelsHow many interleaved channels you are sending
channelWhich channel to transcribe, e.g. only the customer
languageauto by default; the first long turn sets the language for the call

It speaks Deepgram's live protocol, so an existing Deepgram client works by changing the URL. Be clear about one limit before you build on it: results arrive once an utterance has ended — about 700 ms after the speaker stops — not word by word as they talk. The API reference lists every message.

Can the transcription run inside our own network?

For banks, insurers and government, the question is often not accuracy but where the audio goes. nutq deploys the same models inside your network, on your hardware:

Hosted APIOn-premise
Where audio is processednutq's serversYour network
Audio leaving your networkYes, over TLSNone
Pricing$0.15 per hourNo per-minute fee
HardwareNoneOne 24 GB GPU

For regulated work in the UAE, that is usually the whole decision. Talk to us about deployment.

What should you check before choosing a call transcription vendor?

  1. Test on your calls — the real line, the real accents — not the vendor's demo files.
  2. Score against a native speaker's transcript written as spoken, dialect included.
  3. Count character errors, and read the transcripts for dialect words rewritten into MSA.
  4. Check the last seconds of each call and the product names you care about.
  5. Ask where the audio is processed and stored, and whether that can change.

Your $5 of free credit at sign-up covers about 33 hours of audio — enough for a proper test on real calls.

Questions people ask

Why is phone audio harder to transcribe than other recordings?

Traditional telephone audio is sampled at 8 kHz and keeps roughly 300–3,400 Hz. That removes much of the high-frequency detail that separates similar consonants, and adds codec artefacts and line noise. Models trained mostly on clean, wide-band audio have less to go on.

Can I transcribe stereo call recordings where each side is on its own channel?

Yes. For live streams, the channel parameter picks which interleaved channel to transcribe, so you can transcribe the customer and the agent separately. For recorded files, split the channels and send each one.

Can Arabic call transcription run without sending audio to the cloud?

Yes. nutq deploys the same models inside your own network on a single 24 GB GPU. No audio leaves the network and there is no per-minute fee.

How fast is batch transcription of call recordings?

About 22 times faster than realtime: an hour of recorded calls transcribes in under three minutes.

Can I use nutq for a voice agent on the phone?

Yes. The listening side streams the call and returns each turn when the caller pauses, and the speaking side answers in a Gulf Arabic voice. There is a ready configuration for Vapi in the docs.

ShareXLinkedInWhatsApp

Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.

اقرأ بالعربية

Hear it on your own audio.

$5 of credit to start — about 33 hours of transcription. No card.

All articles