# How to build a Gulf Arabic voice agent with Vapi

> Most voice-agent stacks fall over the moment a caller speaks Gulf Arabic. Two blocks of Vapi configuration fix both ends of the call — hearing the caller and answering them.

- Source: https://nutq.dev/blog/arabic-voice-agent-vapi
- Published: 2026-09-27
- Updated: 2026-09-27
- Author: nutq team (https://nutq.dev)

## Short answer

To build an Arabic voice agent with Vapi, point its custom transcriber at nutq's streaming endpoint so Gulf Arabic callers are understood, and point its custom voice at nutq so the agent answers in a Gulf voice. Each is one block of assistant configuration, with no server to host. Keep replies to one or two sentences, because each reply is synthesised in full before the caller hears it.

An **Arabic voice agent** has to do two things most voice stacks cannot: understand a caller from Dubai or Riyadh speaking dialect over a phone line, and answer in a voice that sounds like it belongs in the Gulf. This guide shows how to wire both halves of a Vapi assistant to nutq, and what to watch for once it is taking real calls.

## Why do Arabic voice agents usually fail?

The weak link is usually hearing, not the language model. Deepgram, the default transcriber on most voice platforms, has no dedicated Arabic model: `nova-2` and `nova-3` reject `ar` outright, and only its multilingual mode accepts it — which mistakes Gulf Arabic for other languages. An agent that cannot hear the caller cannot help them, however good its prompt is.

The other end fails more quietly: a broadcast-MSA voice answering a Gulf caller sounds like a government recording, not a conversation.

## How do you connect the listening half?

Vapi opens a WebSocket and streams the call through it; nutq returns each turn as text when the caller pauses. Add this to the assistant:

```json
"transcriber": {
  "provider": "custom-transcriber",
  "server": {
    "url": "wss://worker.nutq.dev/v1/transcribe/stream?key=$NUTQ_STREAM_KEY&language=auto",
    "timeoutSeconds": 30
  }
}
```

A WebSocket handshake carries no headers you control, so the key goes on the URL. Use a **separate key** here, so it can be revoked on its own if a log ever captures it. Check that your Vapi account accepts `custom-transcriber` before building on it; some accounts reject it at publish time.

With `language=auto`, the first turn long enough to judge sets the call's language, and short replies inherit it — a caller who just says «نعم» is not re-guessed on one word.

## How do you connect the speaking half?

Vapi posts each line its model wants said and plays back the audio you return:

```json
"voice": {
  "provider": "custom-voice",
  "server": {
    "url": "https://nutq.dev/v1/vapi/voice?voice=rashid",
    "headers": { "Authorization": "Bearer $NUTQ_API_KEY" },
    "timeoutSeconds": 45
  }
}
```

| URL setting | What it does |
|---|---|
| `voice` | A Gulf voice from `GET /v1/voices`, or one saved on your account |
| `language` | Use the other-language engine instead of Arabic |
| `nfe_step` | Quality against speed; defaults to 16 here, lower than the studio's 32 |

Settings ride on the URL because Vapi owns the request body. Whatever sample rate Vapi asks for — 8000, 16000, 22050 or 24000 Hz — comes back exactly, as mono 16-bit PCM.

## How do you keep the conversation natural?

Three rules make the difference between a conversation and a hold queue:

1. **Short replies.** Each reply is synthesised in full before the caller hears any of it. Tell the model in its system prompt to answer in one or two sentences.
2. **Let turns end naturally.** A turn ends after about 700 ms of quiet. Pauses mid-sentence are normal and do not split a turn, and anything past 15 seconds is transcribed anyway rather than waiting for a pause that may not come.
3. **Write the prompt in the register you want back.** If the agent should sound Gulf, give it Gulf examples. The voice reads the text it is given; see [Arabic text to speech with voice cloning](/blog/arabic-text-to-speech-voice-cloning) for how to steer pronunciation.

## What does a call cost?

| Half | Billed as | Price |
|---|---|---|
| Listening | Transcription | $0.15 per hour of audio |
| Speaking | Text to speech | $0.05 per 1,000 characters |

Both are billed to the key in the configuration, and failed requests are not billed. New accounts start with $5 of free credit at [sign-up](/sign-up).

## What should you test before going live?

- Real callers on real phone lines — narrow-band audio behaves differently; see [transcribing Arabic phone calls](/blog/transcribe-arabic-phone-calls).
- Callers who switch between Arabic and English mid-call.
- Very short answers («إي», «لا», «تمام») and long complaints.
- Numbers, dates and product names read back to the caller.

You can use either half on its own: nutq for hearing with another voice, or the other way round. The full configuration, including every note above, is in the [API reference](/docs).

## FAQ

### Does Vapi support Arabic?

Vapi lets you plug in a custom transcriber and a custom voice. With nutq on both, a Vapi assistant can hear Gulf Arabic callers and answer them in a Gulf voice.

### Why not use the default transcriber for Arabic?

Deepgram, the default on most voice platforms, has no dedicated Arabic model: nova-2 and nova-3 reject ar, and only its multilingual mode accepts it, which confuses Gulf Arabic with other languages.

### How fast does the agent respond?

A caller's turn ends after about 700 ms of silence. The reply is then synthesised in full before playback, so short replies are what keep the conversation natural.

### Which voice should I use?

Any Gulf voice from GET /v1/voices, passed as voice on the URL, such as rashid. A voice saved on your own account works the same way.

### What does a call cost?

Listening is billed like transcription, $0.15 per hour of audio, and speaking like text to speech, $0.05 per 1,000 characters, both on the key in your configuration.
