# How to transcribe Arabic phone calls accurately (and why most models fail on them)

> A phone line throws away most of the sound a speech model listens to. Here is what that does to Arabic transcription, what we measured, and how to set up call transcription that holds up.

- Source: https://nutq.dev/blog/transcribe-arabic-phone-calls
- Published: 2026-09-24
- Updated: 2026-09-25
- Author: nutq team (https://nutq.dev)

## Short answer

Arabic phone calls are hard to transcribe because telephone audio is narrow-band — it keeps roughly the 300–3,400 Hz range — and callers speak in dialect. Models tuned on clean MSA degrade sharply on it. On a genuine narrow-band call-in, nutq made 40% fewer character errors than the leading Arabic vendor. You can transcribe recorded calls with one API request, stream live calls over a WebSocket, or run the same models on-premise so recordings never leave your network.

If you transcribe Arabic customer calls, you already know the pattern: a vendor's demo on clean audio looks excellent, and the transcripts from your own call centre do not. That gap is not bad luck. To **transcribe Arabic phone calls** well you are solving a different problem from transcribing podcasts, and it needs to be tested as one.

## Why is Arabic phone audio so hard to transcribe?

Two things happen at once on a call.

**The line throws sound away.** Classic telephone audio is sampled at 8 kHz and passes roughly 300–3,400 Hz. Much of what distinguishes similar consonants lives above that band. Add a compression codec, a mobile connection that drops packets, and a caller on speakerphone in a car, and a model gets a fraction of the signal it was trained on.

**The caller speaks dialect.** Customer calls in the Gulf are in Gulf Arabic, often with English product names mixed in. A model that leans on Modern Standard Arabic when unsure — see [Gulf Arabic vs MSA](/blog/gulf-arabic-vs-msa-transcription) — leans hardest exactly when the audio is worst.

The two compound. Less signal means more guessing, and a model that guesses in MSA rewrites more of what was said.

## How much worse do models get on phone calls?

We tested on the audio that matters: a **genuine narrow-band call-in**, not telephone audio simulated from studio recordings. On it, nutq made **40% fewer character errors** than the leading Arabic vendor.

That test sits inside a wider evaluation of **986 clips** — broadcast, spontaneous interviews and telephony — each scored against a human transcript. We report character error rate rather than word error rate because Arabic attaches prefixes and suffixes to words, so a single wrong letter can count as a whole wrong word; characters give a fairer picture of how close the text is.

## How do you transcribe recorded Arabic calls?

Send each recording to the transcription endpoint — one request per file:

```bash
curl -X POST https://nutq.dev/v1/transcribe \
  -H "Authorization: Bearer $NUTQ_API_KEY" \
  -F "file=@call-2026-09-24-1043.wav" \
  -F "language=ar"
```

A few practical notes for call centres:

- **Keep the original recording.** Do not upsample or "enhance" 8 kHz audio before sending it; that adds nothing the model can use and can add artefacts.
- **Split stereo recordings.** If your recorder puts the agent on one channel and the customer on the other, send them separately and you get a clean transcript per speaker.
- **Batch is fast.** At roughly 22× realtime, an hour of calls comes back in under three minutes. Files up to one hour and 100 MB go in a single request.
- **Log `usage`.** Each response states exactly what the request cost; failed requests are never billed.

The full request and response are covered in the [Arabic speech to text API guide](/blog/arabic-speech-to-text-api).

## How do you transcribe live Arabic calls?

For live calls there is a WebSocket endpoint. You push raw 16-bit PCM and receive a transcript each time the speaker pauses:

```text
wss://worker.nutq.dev/v1/listen?encoding=linear16&sample_rate=8000&channels=2&channel=1
```

| Parameter | What it does |
|---|---|
| `sample_rate` | 8000 to 48000 — send telephone audio at its real 8 kHz |
| `channels` | How many interleaved channels you are sending |
| `channel` | Which channel to transcribe, e.g. only the customer |
| `language` | `auto` by default; the first long turn sets the language for the call |

It speaks Deepgram's live protocol, so an existing Deepgram client works by changing the URL. Be clear about one limit before you build on it: results arrive once an utterance has ended — about 700 ms after the speaker stops — not word by word as they talk. The [API reference](/docs) lists every message.

## Can the transcription run inside our own network?

For banks, insurers and government, the question is often not accuracy but where the audio goes. nutq deploys **the same models inside your network**, on your hardware:

| | Hosted API | On-premise |
|---|---|---|
| Where audio is processed | nutq's servers | Your network |
| Audio leaving your network | Yes, over TLS | None |
| Pricing | $0.15 per hour | No per-minute fee |
| Hardware | None | One 24 GB GPU |

For regulated work in the UAE, that is usually the whole decision. [Talk to us about deployment](/docs).

## What should you check before choosing a call transcription vendor?

1. Test on **your** calls — the real line, the real accents — not the vendor's demo files.
2. Score against a native speaker's transcript written **as spoken**, dialect included.
3. Count character errors, and read the transcripts for dialect words rewritten into MSA.
4. Check the last seconds of each call and the product names you care about.
5. Ask where the audio is processed and stored, and whether that can change.

Your $5 of free credit at [sign-up](/sign-up) covers about 33 hours of audio — enough for a proper test on real calls.

## FAQ

### Why is phone audio harder to transcribe than other recordings?

Traditional telephone audio is sampled at 8 kHz and keeps roughly 300–3,400 Hz. That removes much of the high-frequency detail that separates similar consonants, and adds codec artefacts and line noise. Models trained mostly on clean, wide-band audio have less to go on.

### Can I transcribe stereo call recordings where each side is on its own channel?

Yes. For live streams, the channel parameter picks which interleaved channel to transcribe, so you can transcribe the customer and the agent separately. For recorded files, split the channels and send each one.

### Can Arabic call transcription run without sending audio to the cloud?

Yes. nutq deploys the same models inside your own network on a single 24 GB GPU. No audio leaves the network and there is no per-minute fee.

### How fast is batch transcription of call recordings?

About 22 times faster than realtime: an hour of recorded calls transcribes in under three minutes.

### Can I use nutq for a voice agent on the phone?

Yes. The listening side streams the call and returns each turn when the caller pauses, and the speaking side answers in a Gulf Arabic voice. There is a ready configuration for Vapi in the docs.
