Arabic speech to text API: a developer's guide to transcribing Gulf Arabic
One POST request, a file, and Arabic text back — with the dialect intact. What to send, what comes back, what it costs, and where Arabic speech recognition usually goes wrong.
nutq teamPublished Updated 4 min read

Most Arabic speech-to-text APIs will happily take your audio. The problem shows up in the text that comes back: a Gulf speaker says one thing, and the transcript reads like a newsreader said something else. This guide walks through transcribing Arabic with the nutq Arabic speech to text API — the request, the response, the cost — and the places where Arabic transcription usually fails, so you can test for them before you ship.
What does an Arabic speech to text API actually do?
It takes a recording and returns the words in it as text. For Arabic that simple job has three hard parts:
- Dialect. People in Dubai, Riyadh and Amman do not speak Modern Standard Arabic (MSA) on the phone. Many models are trained mostly on MSA and "correct" dialect into it, which changes what was said. We cover this in detail in Gulf Arabic vs MSA transcription.
- Audio quality. Much of the Arabic audio worth transcribing — support calls, sales calls, voice notes — is narrow-band and noisy. Accuracy measured on clean broadcast audio tells you little about a call centre.
- Code-switching. Gulf speech mixes in English words mid-sentence. A model that forces one language per file mangles them.
How do you transcribe Arabic audio with one API call?
Create a key in the dashboard, then post the file:
curl -X POST https://nutq.dev/v1/transcribe \
-H "Authorization: Bearer $NUTQ_API_KEY" \
-F "file=@call.mp3"
There is no SDK to install. The same request from JavaScript:
const body = new FormData()
body.append('file', file)
const res = await fetch('https://nutq.dev/v1/transcribe', {
method: 'POST',
headers: { Authorization: `Bearer ${process.env.NUTQ_API_KEY}` },
body,
})
const { text, usage } = await res.json()
And from Python:
import requests
with open("call.mp3", "rb") as f:
res = requests.post(
"https://nutq.dev/v1/transcribe",
headers={"Authorization": f"Bearer {NUTQ_API_KEY}"},
files={"file": f},
)
print(res.json()["text"])
What comes back?
{
"text": "هلا وغلا فيك، شحالك اليوم؟",
"language": "ar",
"language_source": "detected",
"language_confidence": 1,
"usage": { "audio_seconds": 13.23, "cost_usd": 0.00056 },
"model": "nutq-listen-1"
}
Three fields are worth knowing about:
language_sourcesays how the language was chosen:requested(you named it),detected(you sentauto),corrected(you named one, but the recording was clearly another), orfallback(too short to judge, so Arabic).usageis what this request actually cost — not an estimate. Log it next to the transcript and your billing reconciles itself.modelnames the engine that served the request, so any transcript can be traced to a version.
Which Arabic dialects and languages are supported?
| Speech | Support |
|---|---|
| Gulf Arabic (Emirati, Saudi, Kuwaiti, Qatari…) | Tuned for it; dialect kept as spoken |
| Levantine Arabic | Tuned for it; dialect kept as spoken |
| Modern Standard Arabic | Supported |
| Maghrebi, Yemeni, Sudanese Arabic | Recognised, with lower accuracy |
| 13 other languages | en, fr, es, de, it, pt, nl, pl, el, ja, ko, zh, vi |
If you know the language, pass it as language=ar (or any code above). If the recording is clearly in another language, the API transcribes what was actually spoken and tells you so, rather than forcing the wrong language onto it.
How accurate is it on real-world audio?
We report one test, because it is the one that matches how Arabic audio is actually recorded. On a genuine narrow-band phone call-in, nutq made 40% fewer character errors than the leading Arabic vendor. The wider evaluation scored 986 clips — broadcast, spontaneous interviews and telephony — against human transcripts. The details, and why phone audio breaks most models, are in Transcribing Arabic phone calls.
Speed matters for batch work too: transcription runs about 22× faster than realtime, so an hour-long recording comes back in under three minutes.
What does Arabic transcription cost?
| Price | |
|---|---|
| Speech to text | $0.15 per hour of audio |
| Free credit on sign-up | $5, about 33 hours, no card |
| Failed requests | Not billed |
Files can be up to 100 MB and one hour long; longer recordings are split and stitched for you, so you send one request either way.
How do you test an Arabic speech API before committing?
Do not test on clean samples. Take twenty recordings that look like your production traffic — the real phone line, the real accents — and have a native speaker transcribe them by hand. Then compare each API's output against those transcripts, counting character errors. Look specifically at:
- dialect words turned into their MSA equivalents,
- English words inside Arabic sentences,
- names, numbers and product terms,
- the last seconds of each file, where some pipelines drop audio.
Your $5 of free credit covers far more audio than that test needs.
Can it run on our own servers?
Yes. For banks, government and anyone whose recordings cannot leave the building, the same models run inside your network on a single 24 GB GPU, with no audio sent to us and no per-minute fee. Talk to us about deployment.
Where to go next
- The full API reference covers streaming, errors and OAuth for apps that act on a user's behalf.
- Turning text back into speech? See Arabic text to speech with voice cloning.
- Captioning video? Auto Arabic subtitles shows the no-code route.
Questions people ask
Is there a speech to text API that supports Arabic dialects?
Yes. nutq's speech-to-text model, nutq-listen-1, is tuned for Gulf and Levantine Arabic and writes what was said in dialect rather than converting it to Modern Standard Arabic. Maghrebi, Yemeni and Sudanese Arabic are recognised too, with lower accuracy.
How much does Arabic transcription cost?
$0.15 per hour of audio, billed per second. New accounts get $5 of credit with no card, which covers about 33 hours of transcription. Requests that fail are not billed.
What audio formats and lengths can I send?
Any common audio or video format, up to 100 MB and one hour per file. Long recordings are split and stitched back together for you.
Do I need to tell the API which language the audio is in?
No. The default is auto, which detects the language from the first couple of seconds of speech. You can name one of 14 languages instead, and the response tells you which language was used and how it was decided.
Can I transcribe live audio instead of files?
Yes. A WebSocket endpoint accepts raw PCM and returns a transcript each time a speaker pauses. It uses Deepgram's live protocol, so an existing Deepgram client works by changing the URL.
Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.
اقرأ بالعربيةHear it on your own audio.
$5 of credit to start — about 33 hours of transcription. No card.
Keep reading
All articles
Speech to text3 min read
How to transcribe Arabic voice notes (WhatsApp, Telegram and phone recordings)
In the Gulf, a lot of business happens in voice notes — orders, complaints, instructions. Here is how to turn them into text you can search, forward and act on, in the dialect they were spoken in.

Speech to text3 min read
Real-time Arabic transcription over WebSocket (a drop-in for Deepgram clients)
If you have already written a Deepgram client, you have already written most of a real-time Arabic transcriber. Change the URL, and here is exactly what comes back — and what does not.

Speech to text4 min read
How to transcribe Arabic phone calls accurately (and why most models fail on them)
A phone line throws away most of the sound a speech model listens to. Here is what that does to Arabic transcription, what we measured, and how to set up call transcription that holds up.