# Real-time Arabic transcription over WebSocket (a drop-in for Deepgram clients)

> If you have already written a Deepgram client, you have already written most of a real-time Arabic transcriber. Change the URL, and here is exactly what comes back — and what does not.

- Source: https://nutq.dev/blog/real-time-arabic-transcription
- Published: 2026-09-27
- Updated: 2026-09-27
- Author: nutq team (https://nutq.dev)

## Short answer

For real-time Arabic transcription, open a WebSocket to wss://worker.nutq.dev/v1/listen, stream 16-bit PCM audio, and read a final transcript back each time the speaker pauses, about 700 ms after they stop. It speaks Deepgram's live protocol, so an existing Deepgram client works by changing the URL. It returns final results only — no word-by-word interim text — and is billed per second of audio at $0.15 per hour.

Meeting recorders, note takers, call loggers and live captioning all need the same thing: audio goes in continuously, text comes out as people talk. For Arabic that has been hard to find. This guide covers **real-time Arabic transcription** over nutq's WebSocket — the connection, the messages, and the honest limits worth knowing before you build on it.

## Why is real-time Arabic transcription hard to find?

Most streaming speech APIs were built English-first. Deepgram, the default on many platforms, has no dedicated Arabic model — `nova-2` and `nova-3` reject `ar` — so teams who already built on it have had no Arabic path. nutq's socket speaks **Deepgram's live protocol on purpose**: if you have that client, you change a URL, not an architecture.

## How do you connect?

```text
wss://worker.nutq.dev/v1/listen?encoding=linear16&sample_rate=16000&channels=1&language=auto
```

| Query parameter | Meaning |
|---|---|
| `encoding` | `linear16` — 16-bit little-endian PCM, the only encoding accepted |
| `sample_rate` | 8000 to 48000; defaults to 16000 |
| `channels` | How many interleaved channels you send; defaults to 1 |
| `channel` | Which channel to transcribe, e.g. the far end of a call |
| `language` | `auto` by default; the first long stretch of speech sets it for the socket |
| `model` | Accepted and ignored — there is one |

Authenticate with `Authorization: Token <key>` (what a Deepgram client sends), a Bearer header, or `?key=` on the URL for platforms that cannot set handshake headers. The URL form ends up in access logs, so give it its own key.

## What do you send?

- **Binary frames** of raw PCM at the rate and channel count you declared. Any frame size.
- `{"type":"KeepAlive"}` — accepted; the socket is not closed for being quiet.
- `{"type":"Finalize"}` — return whatever is buffered now, when you stop mid-sentence.
- `{"type":"CloseStream"}` — finish and close; a Metadata frame reports the billed duration.

## What comes back?

```json
{
  "type": "Results",
  "start": 14.30,
  "duration": 14.70,
  "is_final": true,
  "speech_final": true,
  "channel": {
    "alternatives": [{
      "transcript": "تعلمت إن مو لازم كل شي يصير بسرعة",
      "confidence": null,
      "words": [],
      "language": "ar"
    }]
  }
}
```

Note the transcript: Gulf dialect, as spoken, not rewritten into Modern Standard Arabic. Why that matters is covered in [Gulf Arabic vs MSA](/blog/gulf-arabic-vs-msa-transcription).

## What are the limits?

Better to know these before you design an interface around them:

- **No interim results.** Every result is final, about 700 ms after someone stops talking. Words do not appear one by one while they speak.
- **`confidence` is `null`.** The model produces no calibrated score, so none is invented.
- **`words` is empty.** There are no word-level timings, but `start` and `duration` on each result are real, so you can still jump to a moment.
- **Long unbroken speech is cut at about 15 seconds**, at the quietest point near that mark, so a non-stop talker produces a result roughly every fifteen seconds.

For live captions that must track each word as it is spoken, this is the wrong tool. For notes, logs, call records and captions that appear a sentence at a time, it fits.

## Which sample rate and channels should you send?

Send audio at its real rate. Telephone audio is 8000 Hz, so send it as that with `sample_rate=8000`; upsampling adds nothing the model can use. Declare the channel count you actually send — announcing one channel and sending two transcribes every other sample. For call recordings with the customer on one channel and the agent on the other, pick the side you want with `channel`, or open two sockets to get a transcript per speaker.

`linear16` is the only encoding accepted. Anything else is refused at connect time rather than played back as noise, so a misconfigured client fails at once instead of after an hour of empty transcripts.

## A minimal Python client

```python
import asyncio, json, websockets

URL = ("wss://worker.nutq.dev/v1/listen"
       "?encoding=linear16&sample_rate=16000&channels=1&language=auto")

async def main(pcm_frames):
    async with websockets.connect(
        URL, additional_headers={"Authorization": f"Token {NUTQ_API_KEY}"}
    ) as ws:
        async def send():
            for frame in pcm_frames:
                await ws.send(frame)
            await ws.send(json.dumps({"type": "CloseStream"}))

        async def read():
            async for raw in ws:
                m = json.loads(raw)
                if m["type"] == "Results":
                    print(m["channel"]["alternatives"][0]["transcript"])
                elif m["type"] == "Metadata":
                    return

        await asyncio.gather(send(), read())
```

## Live or batch?

| | WebSocket | File upload |
|---|---|---|
| Use for | Meetings, calls in progress, live notes | Recordings you already have |
| Latency | A result per utterance | The whole file at once |
| Price | $0.15 per hour | $0.15 per hour |

For recordings, one POST is simpler — see the [Arabic speech to text API guide](/blog/arabic-speech-to-text-api). For phone agents, the same streaming engine is wired into Vapi in [building an Arabic voice agent](/blog/arabic-voice-agent-vapi). Every message is in the [API reference](/docs), and your first $5 is free at [sign-up](/sign-up).

## FAQ

### Is there a real-time speech to text API for Arabic?

Yes. nutq's WebSocket endpoint accepts live audio and returns Arabic transcripts as each utterance ends, with Gulf and Levantine dialect kept as spoken.

### Can I use my existing Deepgram client?

Yes. The socket speaks Deepgram's live protocol, and it accepts the Authorization: Token header a Deepgram client already sends. Point the client at the new URL.

### Does it return interim results as people speak?

No. Every result is final and arrives once an utterance has ended, about 700 ms after the speaker stops. If your interface shows words appearing mid-sentence, plan for that.

### Which audio format does it accept?

linear16: raw 16-bit little-endian PCM, at 8000 to 48000 Hz, with any number of interleaved channels. You choose which channel to transcribe.

### How is live transcription billed?

Per second of audio, at the same $0.15 per hour as file transcription.
