On-premise Arabic speech recognition: keeping every recording inside your network
For a bank or a ministry, the first question about Arabic speech recognition is not accuracy. It is where the recordings go. Here is what running the models inside your own network involves.
nutq teamPublished 3 min read

Ask a bank, an insurer or a government department what they need from on-premise Arabic speech recognition and accuracy comes second. The first question is always the same: where do the recordings go? Customer calls, complaints, medical consultations and internal meetings are among the most sensitive data an organisation holds. For many teams in the Gulf, sending them to a third-party cloud is not an option at any price.
What does on-premise speech recognition mean?
It means the speech models run on hardware you control, inside your network. Audio is transcribed where it is already stored, and the text goes wherever your systems send it. No recording is uploaded to a vendor, and there is no vendor-side copy to audit, breach or subpoena.
That is different from a "private cloud" region or a dedicated tenant, where the audio still leaves your network and sits on someone else's servers — just in a separate account.
Who needs Arabic speech recognition on-premise?
| Organisation | Typical audio | Why it cannot leave |
|---|---|---|
| Banks and insurers | Call-centre recordings, complaints | Customer data and internal policy |
| Government and public services | Hotlines, citizen services, meetings | Data must stay in controlled networks |
| Healthcare | Consultations, dictation | Patient confidentiality |
| Legal and compliance teams | Interviews, hearings, investigations | Privilege and chain of custody |
| Telecoms | Support and sales calls at volume | Volume makes per-minute pricing expensive, and data is sensitive |
If your security team has already said no to cloud transcription, you are the audience for this.
What runs inside your network?
nutq deploys the same models it runs in the cloud — not a cut-down edition:
- Speech to text (
nutq-listen-1): Gulf and Levantine Arabic kept as spoken rather than rewritten into Modern Standard Arabic, plus 13 other languages. On a genuine narrow-band phone call it made 40% fewer character errors than the leading Arabic vendor. - Text to speech (
nutq-speak-1): Gulf Arabic voices, for IVR prompts and voice agents that reply in dialect.
The accuracy argument is the one from our article on transcribing Arabic phone calls: call-centre audio is narrow-band and in dialect, and that is where models trained on clean MSA fall apart.
What hardware does it need?
| On-premise | |
|---|---|
| GPU | One, with 24 GB of memory |
| Where it runs | Your hardware, your network |
| Audio egress | None |
| Per-minute fee | None |
A single 24 GB GPU is a modest server, not a cluster. At about 22× faster than realtime, one hour of recordings transcribes in under three minutes, which is what makes batch processing of a day's calls practical on one machine.
How is it different from the hosted API?
The hosted API is the fastest way to start: one request to /v1/transcribe, $0.15 per hour of audio, $5 of free credit, no card. It is the right choice for most teams, and the Arabic speech to text API guide covers it end to end.
On-premise trades that convenience for control. You run the server; we deploy the models on it. Instead of paying per minute, you have a fixed deployment, which also changes the economics once call volume is high.
How do you evaluate on-premise Arabic ASR?
Evaluate it the way you would the cloud version, on your own audio:
- Take 20–50 real recordings: your phone lines, your customers' dialects, your product names.
- Have a native speaker transcribe them exactly as spoken, dialect included.
- Run them through the hosted API first — the model is the same, and your free credit covers the test. See Gulf Arabic vs MSA for what to look for.
- Count character errors and read for dialect words rewritten into MSA.
- If the results hold up, scope the deployment.
Testing on the hosted API with non-sensitive or consented recordings first means your security review only has to approve the deployment, not the evaluation.
Getting started
Tell us about your audio, your volumes and your environment, and we will scope the deployment with you. Talk to us about deployment, or start with the hosted API by creating an account.
Questions people ask
What hardware does on-premise Arabic speech recognition need?
nutq's models run on a single GPU with 24 GB of memory, on your own server, in your own network.
Does any audio leave our network?
No. In an on-premise deployment the recordings are processed where they already live. Nothing is sent to nutq.
How is on-premise priced?
There is no per-minute or per-hour fee, unlike the hosted API. Talk to us about the deployment and we will scope it with you.
Are the on-premise models the same as the cloud API?
Yes. The same models that serve the hosted API — Gulf and Levantine speech to text, and Gulf Arabic text to speech — are deployed inside your network.
Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.
اقرأ بالعربيةHear it on your own audio.
$5 of credit to start — about 33 hours of transcription. No card.
Keep reading
All articles
Speech to text3 min read
How to transcribe Arabic voice notes (WhatsApp, Telegram and phone recordings)
In the Gulf, a lot of business happens in voice notes — orders, complaints, instructions. Here is how to turn them into text you can search, forward and act on, in the dialect they were spoken in.

Speech to text3 min read
Real-time Arabic transcription over WebSocket (a drop-in for Deepgram clients)
If you have already written a Deepgram client, you have already written most of a real-time Arabic transcriber. Change the URL, and here is exactly what comes back — and what does not.

Guides3 min read
How to build a Gulf Arabic voice agent with Vapi
Most voice-agent stacks fall over the moment a caller speaks Gulf Arabic. Two blocks of Vapi configuration fix both ends of the call — hearing the caller and answering them.