Arabic text to speech with voice cloning: a Gulf Arabic TTS guide

Six to eight seconds of someone speaking, the words they said, and the text you want read — that is all a Gulf Arabic voice clone needs. How to get a natural result, and the one trick that fixes most mispronunciations.

nutq teamPublished Updated 3 min read

A studio microphone with a pop filter in a small recording booth, with acoustic foam panels and headphones softly out of focus.

Most Arabic text to speech sounds like the evening news: correct, formal, and nothing like the way anyone in the Gulf talks. That is fine for a train announcement. It is wrong for a customer-service agent, a brand voice, or a video voiceover aimed at people in Dubai or Riyadh. This guide covers Arabic text to speech with voice cloning — generating Gulf Arabic speech in a specific voice from a few seconds of audio.

What is Arabic voice cloning?

Voice cloning makes a text-to-speech model speak in a particular person's voice using a short sample of them talking, instead of a fixed stock voice. You give it three things:

  1. reference — a recording of the voice, ideally six to eight seconds.
  2. ref_text — exactly what is said in that recording. The model aligns the audio against it.
  3. text — what you want the voice to say.

The result is a WAV file of your text, in that voice, with a Gulf accent. The model, nutq-speak-1, is trained on Emirati and Saudi speech, so it sounds Gulf rather than like broadcast MSA.

How do you generate Arabic speech with the API?

curl -X POST https://nutq.dev/v1/speech \
  -H "Authorization: Bearer $NUTQ_API_KEY" \
  -F "text=هلا وغلا فيك، شحالك اليوم؟" \
  -F "reference=@voice.wav" \
  -F "ref_text=نص التسجيل المرجعي" \
  --output speech.wav

The body of the response is the audio itself, so usage travels in the headers:

HTTP/1.1 200 OK
Content-Type: audio/wav
X-Nutq-Characters: 57
X-Nutq-Audio-Seconds: 5.672
X-Nutq-Cost-USD: 0.000285
X-Nutq-Model: nutq-speak-1

The same call in Python:

import requests

res = requests.post(
    "https://nutq.dev/v1/speech",
    headers={"Authorization": f"Bearer {NUTQ_API_KEY}"},
    data={"text": "هلا وغلا فيك", "ref_text": "نص التسجيل المرجعي"},
    files={"reference": open("voice.wav", "rb")},
)
open("speech.wav", "wb").write(res.content)

What makes a good reference recording?

The reference clip decides most of the quality. A few rules that matter:

DoAvoid
6–8 seconds of continuous speechClips over 20 seconds, which are rejected
One speaker, close to the microphoneBackground music, a second voice, echo
The tone you want back — calm, warm, energeticA flat read if you want a lively result
A ref_text that matches the audio word for wordA summary or a cleaned-up version of what was said

Longer clips are accepted, but they are trimmed to about eight seconds and re-transcribed, which costs a round trip and can pick a less representative stretch. Trim it yourself.

Clone only voices you have permission to use — your own, or a speaker who has agreed to it.

How do you fix a word that is pronounced wrong?

Arabic is normally written without short vowels and read from context, so a word with two possible readings is a guess. تعد can be تَعُدّ or تُعَدّ; درس can be دَرَسَ or دَرَّسَ. The fix is to add diacritics to that one word:

تُعَدُّ دولة الإمارات من أبرز الدول

Mark only what is ambiguous — the way Arabic is normally written. Marking every word changes the delivery, not just the vowels: speech runs about half again as long and sounds like a recitation. We also tested automatic diacritizers and do not recommend them; they introduce errors of their own, especially in dialect.

How do you trade quality against speed?

ParameterDefaultWhat it does
nfe_step32Quality against speed, 8–96. 64 is noticeably cleaner and about twice as slow
speed1Playback rate of the generated speech
text—Up to 5,000 characters per request

For a live voice agent, lower steps are worth it: in a phone conversation the wait is more audible than the difference in quality. For a voiceover you will publish, use 64.

What does Arabic TTS cost?

$0.05 per 1,000 characters. A 1,000-character script — roughly a minute and a half of speech — costs five cents. New accounts get $5 of free credit with no card at sign-up, and failed requests are never billed.

Can it speak other languages?

Yes, with trained voices instead of cloning. Pass language as en, en-gb, es, fr, hi, it, ja, pt or zh, and a voice name such as af_heart or bm_george. reference and ref_text are ignored for these. GET /v1/voices lists what is available.

Where to use it

Questions people ask

How much audio do I need to clone an Arabic voice?

Six to eight seconds of clear speech is the sweet spot. Longer clips are accepted but trimmed to about eight seconds; anything over twenty seconds is rejected. You also send ref_text, the exact words spoken in the clip.

Does the Arabic voice sound Gulf or like MSA news?

Gulf. nutq-speak-1 is trained on Emirati and Saudi speech, so the accent is Gulf rather than broadcast Modern Standard Arabic, even when the text you give it is written in MSA.

How much does Arabic text to speech cost?

$0.05 per 1,000 characters. A request can carry up to 5,000 characters. Usage is returned in the response headers, and failed requests are not billed.

How do I fix a word that is pronounced wrong?

Arabic is usually written without short vowels, so some words have two possible readings. Add the diacritics to that one word and it will be read the way you wrote it. Avoid marking every word — it makes the delivery slow and formal.

Can it speak languages other than Arabic?

Yes, with trained voices rather than cloning: English (US and UK), Spanish, French, Hindi, Italian, Japanese, Portuguese and Chinese. Pick a voice by name; no reference recording is needed.

ShareXLinkedInWhatsApp

Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.

اقرأ بالعربية

Hear it on your own audio.

$5 of credit to start — about 33 hours of transcription. No card.

All articles