Real-time Arabic transcription over WebSocket (a drop-in for Deepgram clients)
If you have already written a Deepgram client, you have already written most of a real-time Arabic transcriber. Change the URL, and here is exactly what comes back — and what does not.
nutq teamPublished 3 min read

Meeting recorders, note takers, call loggers and live captioning all need the same thing: audio goes in continuously, text comes out as people talk. For Arabic that has been hard to find. This guide covers real-time Arabic transcription over nutq's WebSocket — the connection, the messages, and the honest limits worth knowing before you build on it.
Why is real-time Arabic transcription hard to find?
Most streaming speech APIs were built English-first. Deepgram, the default on many platforms, has no dedicated Arabic model — nova-2 and nova-3 reject ar — so teams who already built on it have had no Arabic path. nutq's socket speaks Deepgram's live protocol on purpose: if you have that client, you change a URL, not an architecture.
How do you connect?
wss://worker.nutq.dev/v1/listen?encoding=linear16&sample_rate=16000&channels=1&language=auto
| Query parameter | Meaning |
|---|---|
encoding | linear16 — 16-bit little-endian PCM, the only encoding accepted |
sample_rate | 8000 to 48000; defaults to 16000 |
channels | How many interleaved channels you send; defaults to 1 |
channel | Which channel to transcribe, e.g. the far end of a call |
language | auto by default; the first long stretch of speech sets it for the socket |
model | Accepted and ignored — there is one |
Authenticate with Authorization: Token <key> (what a Deepgram client sends), a Bearer header, or ?key= on the URL for platforms that cannot set handshake headers. The URL form ends up in access logs, so give it its own key.
What do you send?
- Binary frames of raw PCM at the rate and channel count you declared. Any frame size.
{"type":"KeepAlive"}— accepted; the socket is not closed for being quiet.{"type":"Finalize"}— return whatever is buffered now, when you stop mid-sentence.{"type":"CloseStream"}— finish and close; a Metadata frame reports the billed duration.
What comes back?
{
"type": "Results",
"start": 14.30,
"duration": 14.70,
"is_final": true,
"speech_final": true,
"channel": {
"alternatives": [{
"transcript": "تعلمت إن مو لازم كل شي يصير بسرعة",
"confidence": null,
"words": [],
"language": "ar"
}]
}
}
Note the transcript: Gulf dialect, as spoken, not rewritten into Modern Standard Arabic. Why that matters is covered in Gulf Arabic vs MSA.
What are the limits?
Better to know these before you design an interface around them:
- No interim results. Every result is final, about 700 ms after someone stops talking. Words do not appear one by one while they speak.
confidenceisnull. The model produces no calibrated score, so none is invented.wordsis empty. There are no word-level timings, butstartanddurationon each result are real, so you can still jump to a moment.- Long unbroken speech is cut at about 15 seconds, at the quietest point near that mark, so a non-stop talker produces a result roughly every fifteen seconds.
For live captions that must track each word as it is spoken, this is the wrong tool. For notes, logs, call records and captions that appear a sentence at a time, it fits.
Which sample rate and channels should you send?
Send audio at its real rate. Telephone audio is 8000 Hz, so send it as that with sample_rate=8000; upsampling adds nothing the model can use. Declare the channel count you actually send — announcing one channel and sending two transcribes every other sample. For call recordings with the customer on one channel and the agent on the other, pick the side you want with channel, or open two sockets to get a transcript per speaker.
linear16 is the only encoding accepted. Anything else is refused at connect time rather than played back as noise, so a misconfigured client fails at once instead of after an hour of empty transcripts.
A minimal Python client
import asyncio, json, websockets
URL = ("wss://worker.nutq.dev/v1/listen"
"?encoding=linear16&sample_rate=16000&channels=1&language=auto")
async def main(pcm_frames):
async with websockets.connect(
URL, additional_headers={"Authorization": f"Token {NUTQ_API_KEY}"}
) as ws:
async def send():
for frame in pcm_frames:
await ws.send(frame)
await ws.send(json.dumps({"type": "CloseStream"}))
async def read():
async for raw in ws:
m = json.loads(raw)
if m["type"] == "Results":
print(m["channel"]["alternatives"][0]["transcript"])
elif m["type"] == "Metadata":
return
await asyncio.gather(send(), read())
Live or batch?
| WebSocket | File upload | |
|---|---|---|
| Use for | Meetings, calls in progress, live notes | Recordings you already have |
| Latency | A result per utterance | The whole file at once |
| Price | $0.15 per hour | $0.15 per hour |
For recordings, one POST is simpler — see the Arabic speech to text API guide. For phone agents, the same streaming engine is wired into Vapi in building an Arabic voice agent. Every message is in the API reference, and your first $5 is free at sign-up.
Questions people ask
Is there a real-time speech to text API for Arabic?
Yes. nutq's WebSocket endpoint accepts live audio and returns Arabic transcripts as each utterance ends, with Gulf and Levantine dialect kept as spoken.
Can I use my existing Deepgram client?
Yes. The socket speaks Deepgram's live protocol, and it accepts the Authorization: Token header a Deepgram client already sends. Point the client at the new URL.
Does it return interim results as people speak?
No. Every result is final and arrives once an utterance has ended, about 700 ms after the speaker stops. If your interface shows words appearing mid-sentence, plan for that.
Which audio format does it accept?
linear16: raw 16-bit little-endian PCM, at 8000 to 48000 Hz, with any number of interleaved channels. You choose which channel to transcribe.
How is live transcription billed?
Per second of audio, at the same $0.15 per hour as file transcription.
Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.
اقرأ بالعربيةHear it on your own audio.
$5 of credit to start — about 33 hours of transcription. No card.
Keep reading
All articles
Speech to text3 min read
How to transcribe Arabic voice notes (WhatsApp, Telegram and phone recordings)
In the Gulf, a lot of business happens in voice notes — orders, complaints, instructions. Here is how to turn them into text you can search, forward and act on, in the dialect they were spoken in.

Speech to text4 min read
Arabic speech to text API: a developer's guide to transcribing Gulf Arabic
One POST request, a file, and Arabic text back — with the dialect intact. What to send, what comes back, what it costs, and where Arabic speech recognition usually goes wrong.

Speech to text4 min read
How to transcribe Arabic phone calls accurately (and why most models fail on them)
A phone line throws away most of the sound a speech model listens to. Here is what that does to Arabic transcription, what we measured, and how to set up call transcription that holds up.