Real-time Arabic transcription over WebSocket (a drop-in for Deepgram clients)

If you have already written a Deepgram client, you have already written most of a real-time Arabic transcriber. Change the URL, and here is exactly what comes back — and what does not.

nutq teamPublished 3 min read

A smartphone face up on a stone café table showing an incoming-call screen, beside a small glass of karak tea.

Meeting recorders, note takers, call loggers and live captioning all need the same thing: audio goes in continuously, text comes out as people talk. For Arabic that has been hard to find. This guide covers real-time Arabic transcription over nutq's WebSocket — the connection, the messages, and the honest limits worth knowing before you build on it.

Why is real-time Arabic transcription hard to find?

Most streaming speech APIs were built English-first. Deepgram, the default on many platforms, has no dedicated Arabic model — nova-2 and nova-3 reject ar — so teams who already built on it have had no Arabic path. nutq's socket speaks Deepgram's live protocol on purpose: if you have that client, you change a URL, not an architecture.

How do you connect?

wss://worker.nutq.dev/v1/listen?encoding=linear16&sample_rate=16000&channels=1&language=auto
Query parameterMeaning
encodinglinear16 — 16-bit little-endian PCM, the only encoding accepted
sample_rate8000 to 48000; defaults to 16000
channelsHow many interleaved channels you send; defaults to 1
channelWhich channel to transcribe, e.g. the far end of a call
languageauto by default; the first long stretch of speech sets it for the socket
modelAccepted and ignored — there is one

Authenticate with Authorization: Token <key> (what a Deepgram client sends), a Bearer header, or ?key= on the URL for platforms that cannot set handshake headers. The URL form ends up in access logs, so give it its own key.

What do you send?

  • Binary frames of raw PCM at the rate and channel count you declared. Any frame size.
  • {"type":"KeepAlive"} — accepted; the socket is not closed for being quiet.
  • {"type":"Finalize"} — return whatever is buffered now, when you stop mid-sentence.
  • {"type":"CloseStream"} — finish and close; a Metadata frame reports the billed duration.

What comes back?

{
  "type": "Results",
  "start": 14.30,
  "duration": 14.70,
  "is_final": true,
  "speech_final": true,
  "channel": {
    "alternatives": [{
      "transcript": "تعلمت إن مو لازم كل شي يصير بسرعة",
      "confidence": null,
      "words": [],
      "language": "ar"
    }]
  }
}

Note the transcript: Gulf dialect, as spoken, not rewritten into Modern Standard Arabic. Why that matters is covered in Gulf Arabic vs MSA.

What are the limits?

Better to know these before you design an interface around them:

  • No interim results. Every result is final, about 700 ms after someone stops talking. Words do not appear one by one while they speak.
  • confidence is null. The model produces no calibrated score, so none is invented.
  • words is empty. There are no word-level timings, but start and duration on each result are real, so you can still jump to a moment.
  • Long unbroken speech is cut at about 15 seconds, at the quietest point near that mark, so a non-stop talker produces a result roughly every fifteen seconds.

For live captions that must track each word as it is spoken, this is the wrong tool. For notes, logs, call records and captions that appear a sentence at a time, it fits.

Which sample rate and channels should you send?

Send audio at its real rate. Telephone audio is 8000 Hz, so send it as that with sample_rate=8000; upsampling adds nothing the model can use. Declare the channel count you actually send — announcing one channel and sending two transcribes every other sample. For call recordings with the customer on one channel and the agent on the other, pick the side you want with channel, or open two sockets to get a transcript per speaker.

linear16 is the only encoding accepted. Anything else is refused at connect time rather than played back as noise, so a misconfigured client fails at once instead of after an hour of empty transcripts.

A minimal Python client

import asyncio, json, websockets

URL = ("wss://worker.nutq.dev/v1/listen"
       "?encoding=linear16&sample_rate=16000&channels=1&language=auto")

async def main(pcm_frames):
    async with websockets.connect(
        URL, additional_headers={"Authorization": f"Token {NUTQ_API_KEY}"}
    ) as ws:
        async def send():
            for frame in pcm_frames:
                await ws.send(frame)
            await ws.send(json.dumps({"type": "CloseStream"}))

        async def read():
            async for raw in ws:
                m = json.loads(raw)
                if m["type"] == "Results":
                    print(m["channel"]["alternatives"][0]["transcript"])
                elif m["type"] == "Metadata":
                    return

        await asyncio.gather(send(), read())

Live or batch?

WebSocketFile upload
Use forMeetings, calls in progress, live notesRecordings you already have
LatencyA result per utteranceThe whole file at once
Price$0.15 per hour$0.15 per hour

For recordings, one POST is simpler — see the Arabic speech to text API guide. For phone agents, the same streaming engine is wired into Vapi in building an Arabic voice agent. Every message is in the API reference, and your first $5 is free at sign-up.

Questions people ask

Is there a real-time speech to text API for Arabic?

Yes. nutq's WebSocket endpoint accepts live audio and returns Arabic transcripts as each utterance ends, with Gulf and Levantine dialect kept as spoken.

Can I use my existing Deepgram client?

Yes. The socket speaks Deepgram's live protocol, and it accepts the Authorization: Token header a Deepgram client already sends. Point the client at the new URL.

Does it return interim results as people speak?

No. Every result is final and arrives once an utterance has ended, about 700 ms after the speaker stops. If your interface shows words appearing mid-sentence, plan for that.

Which audio format does it accept?

linear16: raw 16-bit little-endian PCM, at 8000 to 48000 Hz, with any number of interleaved channels. You choose which channel to transcribe.

How is live transcription billed?

Per second of audio, at the same $0.15 per hour as file transcription.

ShareXLinkedInWhatsApp

Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.

اقرأ بالعربية

Hear it on your own audio.

$5 of credit to start — about 33 hours of transcription. No card.

All articles