Gulf Arabic vs MSA: why most speech models rewrite what your customers said
A caller says شحالك and the transcript says كيف حالك. It looks like a correction. It is a change to the record — and for search, analytics and compliance it matters more than it seems.
nutq teamPublished Updated 4 min read

Ask most Arabic speech recognition systems to transcribe a Gulf customer and you get back clean, formal, grammatical Arabic. That is the problem. The customer did not speak formal Arabic. This article explains Gulf Arabic vs MSA from the point of view of a transcript: what "flattening" looks like, why models do it, and how to tell whether the one you are evaluating does.
What is dialect flattening in Arabic transcription?
Flattening is when a model hears a dialect word and writes its Modern Standard Arabic equivalent instead. The transcript is not garbled — it is tidied. That is what makes it easy to miss: a reviewer who does not have the audio sees a perfectly good sentence.
Take a greeting you would hear on any Emirati support line:
Said: هلا وغلا فيك، شحالك اليوم؟
Flattened: أهلًا وسهلًا بك، كيف حالك اليوم؟
The highlighted words are the dialect a flattening model throws away. Both lines mean roughly the same thing. Only one of them is what the caller said.
Which Gulf words get rewritten most often?
An illustration of common Gulf words and the MSA a flattening model tends to put in their place:
| Gulf Arabic (said) | Flattened to MSA | Meaning |
|---|---|---|
| شحالك | كيف حالك | How are you |
| الحين | الآن | Now |
| وايد | كثيرًا | A lot |
| أبا / أبغى | أريد | I want |
| ليش | لماذا | Why |
| شو / ايش | ماذا | What |
| باكر | غدًا | Tomorrow |
| زين | جيد | Good, fine |
Each substitution is small. Across a thousand calls, they add up to a record of conversations that never happened in that form.
Why does flattening happen?
Because of what models learn from. Written Arabic on the web, in books and in news is overwhelmingly MSA, and so is most of the transcribed Arabic audio available for training — broadcast news and read speech. A model trained on that data learns that "Arabic text" means MSA. When it hears a Gulf word it is less sure of, it reaches for the formal word it has seen thousands of times.
The effect gets worse as audio quality drops. On a clean studio recording, a model may still catch the dialect word. On a narrow-band phone call, where much of the acoustic detail is gone, it leans harder on what it expects — and it expects MSA.
Why does it matter if the meaning is roughly the same?
Because "roughly" is not what a transcript is for.
- Search stops working. If an analyst searches call transcripts for وايد or a local product nickname, a flattened transcript never contains it.
- Analytics describe the wrong customers. Sentiment, topic and intent models trained on how your customers really speak get fed a formalised version of them.
- Quotes become misquotes. In compliance, legal and journalism, the exact words are the point.
- Voice products sound wrong. If you build a Gulf Arabic voice or an agent that replies in dialect, it needs examples of dialect — not a translation of them.
Should a transcript be in dialect or in MSA?
In dialect — the language that was spoken. If a report or a translation needs MSA, convert the accurate transcript as a separate, visible step. A conversion you run on purpose can be reviewed and corrected. One the speech model performs silently cannot be, because nobody knows it happened.
How do you test a model for dialect flattening?
You need a native speaker and about an hour:
- Pick 10–20 recordings of real Gulf speech from your own use case: calls, voice notes, interviews.
- Have a native speaker transcribe them exactly as spoken, dialect included.
- Run the same files through each model you are considering.
- Mark every place where the model's word is a formal synonym of the spoken one. Those are flattening errors, and a word-error-rate figure computed against an MSA reference will not show them.
nutq's speech model, nutq-listen-1, is tuned for Gulf and Levantine Arabic and keeps both as spoken. You can run this test with the $5 of free credit from signing up — that is about 33 hours of audio, far more than the test needs. The request itself is one line; see the Arabic speech to text API guide or the API reference.
The short version
A transcript that reads well is not the same as a transcript that is right. For Gulf Arabic, the most common error is not a garbled word — it is a correct-looking one that the speaker never said.
Questions people ask
What is the difference between Gulf Arabic and Modern Standard Arabic?
Modern Standard Arabic (MSA) is the formal written and broadcast language shared across the Arab world. Gulf Arabic is the spoken dialect of the UAE, Saudi Arabia, Kuwait, Qatar, Bahrain and Oman, with its own vocabulary, pronunciation and grammar — شحالك rather than كيف حالك, الحين rather than الآن.
Why do speech recognition models turn dialect into MSA?
Because most of the Arabic text and transcribed audio they learn from is MSA. When a model is unsure, it falls back to the spelling and words it has seen most, which are the formal ones.
Should a transcript be in dialect or in MSA?
A transcript should record what was said. If you need MSA for a report or a translation, convert the accurate dialect transcript afterwards — that step can be checked. A model that converts silently cannot be.
Which dialects does nutq keep?
nutq-listen-1 is tuned for Gulf and Levantine Arabic and keeps both as spoken. It also recognises Maghrebi, Yemeni and Sudanese Arabic, with lower accuracy.
Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.
اقرأ بالعربيةHear it on your own audio.
$5 of credit to start — about 33 hours of transcription. No card.
Keep reading
All articles
Speech to text3 min read
How to transcribe Arabic voice notes (WhatsApp, Telegram and phone recordings)
In the Gulf, a lot of business happens in voice notes — orders, complaints, instructions. Here is how to turn them into text you can search, forward and act on, in the dialect they were spoken in.

On-premise3 min read
On-premise Arabic speech recognition: keeping every recording inside your network
For a bank or a ministry, the first question about Arabic speech recognition is not accuracy. It is where the recordings go. Here is what running the models inside your own network involves.

Speech to text3 min read
Real-time Arabic transcription over WebSocket (a drop-in for Deepgram clients)
If you have already written a Deepgram client, you have already written most of a real-time Arabic transcriber. Change the URL, and here is exactly what comes back — and what does not.