Arabic speech to text API: a developer's guide to transcribing Gulf Arabic

One POST request, a file, and Arabic text back — with the dialect intact. What to send, what comes back, what it costs, and where Arabic speech recognition usually goes wrong.

nutq teamPublished Updated 4 min read

A developer's desk by a window in the early morning, with a laptop showing an audio waveform, a small microphone and a cup of Arabic coffee.

Most Arabic speech-to-text APIs will happily take your audio. The problem shows up in the text that comes back: a Gulf speaker says one thing, and the transcript reads like a newsreader said something else. This guide walks through transcribing Arabic with the nutq Arabic speech to text API — the request, the response, the cost — and the places where Arabic transcription usually fails, so you can test for them before you ship.

What does an Arabic speech to text API actually do?

It takes a recording and returns the words in it as text. For Arabic that simple job has three hard parts:

  1. Dialect. People in Dubai, Riyadh and Amman do not speak Modern Standard Arabic (MSA) on the phone. Many models are trained mostly on MSA and "correct" dialect into it, which changes what was said. We cover this in detail in Gulf Arabic vs MSA transcription.
  2. Audio quality. Much of the Arabic audio worth transcribing — support calls, sales calls, voice notes — is narrow-band and noisy. Accuracy measured on clean broadcast audio tells you little about a call centre.
  3. Code-switching. Gulf speech mixes in English words mid-sentence. A model that forces one language per file mangles them.

How do you transcribe Arabic audio with one API call?

Create a key in the dashboard, then post the file:

curl -X POST https://nutq.dev/v1/transcribe \
  -H "Authorization: Bearer $NUTQ_API_KEY" \
  -F "file=@call.mp3"

There is no SDK to install. The same request from JavaScript:

const body = new FormData()
body.append('file', file)

const res = await fetch('https://nutq.dev/v1/transcribe', {
  method: 'POST',
  headers: { Authorization: `Bearer ${process.env.NUTQ_API_KEY}` },
  body,
})
const { text, usage } = await res.json()

And from Python:

import requests

with open("call.mp3", "rb") as f:
    res = requests.post(
        "https://nutq.dev/v1/transcribe",
        headers={"Authorization": f"Bearer {NUTQ_API_KEY}"},
        files={"file": f},
    )
print(res.json()["text"])

What comes back?

{
  "text": "هلا وغلا فيك، شحالك اليوم؟",
  "language": "ar",
  "language_source": "detected",
  "language_confidence": 1,
  "usage": { "audio_seconds": 13.23, "cost_usd": 0.00056 },
  "model": "nutq-listen-1"
}

Three fields are worth knowing about:

  • language_source says how the language was chosen: requested (you named it), detected (you sent auto), corrected (you named one, but the recording was clearly another), or fallback (too short to judge, so Arabic).
  • usage is what this request actually cost — not an estimate. Log it next to the transcript and your billing reconciles itself.
  • model names the engine that served the request, so any transcript can be traced to a version.

Which Arabic dialects and languages are supported?

SpeechSupport
Gulf Arabic (Emirati, Saudi, Kuwaiti, Qatari…)Tuned for it; dialect kept as spoken
Levantine ArabicTuned for it; dialect kept as spoken
Modern Standard ArabicSupported
Maghrebi, Yemeni, Sudanese ArabicRecognised, with lower accuracy
13 other languagesen, fr, es, de, it, pt, nl, pl, el, ja, ko, zh, vi

If you know the language, pass it as language=ar (or any code above). If the recording is clearly in another language, the API transcribes what was actually spoken and tells you so, rather than forcing the wrong language onto it.

How accurate is it on real-world audio?

We report one test, because it is the one that matches how Arabic audio is actually recorded. On a genuine narrow-band phone call-in, nutq made 40% fewer character errors than the leading Arabic vendor. The wider evaluation scored 986 clips — broadcast, spontaneous interviews and telephony — against human transcripts. The details, and why phone audio breaks most models, are in Transcribing Arabic phone calls.

Speed matters for batch work too: transcription runs about 22× faster than realtime, so an hour-long recording comes back in under three minutes.

What does Arabic transcription cost?

Price
Speech to text$0.15 per hour of audio
Free credit on sign-up$5, about 33 hours, no card
Failed requestsNot billed

Files can be up to 100 MB and one hour long; longer recordings are split and stitched for you, so you send one request either way.

How do you test an Arabic speech API before committing?

Do not test on clean samples. Take twenty recordings that look like your production traffic — the real phone line, the real accents — and have a native speaker transcribe them by hand. Then compare each API's output against those transcripts, counting character errors. Look specifically at:

  • dialect words turned into their MSA equivalents,
  • English words inside Arabic sentences,
  • names, numbers and product terms,
  • the last seconds of each file, where some pipelines drop audio.

Your $5 of free credit covers far more audio than that test needs.

Can it run on our own servers?

Yes. For banks, government and anyone whose recordings cannot leave the building, the same models run inside your network on a single 24 GB GPU, with no audio sent to us and no per-minute fee. Talk to us about deployment.

Where to go next

Questions people ask

Is there a speech to text API that supports Arabic dialects?

Yes. nutq's speech-to-text model, nutq-listen-1, is tuned for Gulf and Levantine Arabic and writes what was said in dialect rather than converting it to Modern Standard Arabic. Maghrebi, Yemeni and Sudanese Arabic are recognised too, with lower accuracy.

How much does Arabic transcription cost?

$0.15 per hour of audio, billed per second. New accounts get $5 of credit with no card, which covers about 33 hours of transcription. Requests that fail are not billed.

What audio formats and lengths can I send?

Any common audio or video format, up to 100 MB and one hour per file. Long recordings are split and stitched back together for you.

Do I need to tell the API which language the audio is in?

No. The default is auto, which detects the language from the first couple of seconds of speech. You can name one of 14 languages instead, and the response tells you which language was used and how it was decided.

Can I transcribe live audio instead of files?

Yes. A WebSocket endpoint accepts raw PCM and returns a transcript each time a speaker pauses. It uses Deepgram's live protocol, so an existing Deepgram client works by changing the URL.

ShareXLinkedInWhatsApp

Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.

اقرأ بالعربية

Hear it on your own audio.

$5 of credit to start — about 33 hours of transcription. No card.

All articles