On-premise Arabic speech recognition: keeping every recording inside your network

For a bank or a ministry, the first question about Arabic speech recognition is not accuracy. It is where the recordings go. Here is what running the models inside your own network involves.

nutq teamPublished 3 min read

A single server rack with softly glowing status lights in a quiet, glass-walled server room.

Ask a bank, an insurer or a government department what they need from on-premise Arabic speech recognition and accuracy comes second. The first question is always the same: where do the recordings go? Customer calls, complaints, medical consultations and internal meetings are among the most sensitive data an organisation holds. For many teams in the Gulf, sending them to a third-party cloud is not an option at any price.

What does on-premise speech recognition mean?

It means the speech models run on hardware you control, inside your network. Audio is transcribed where it is already stored, and the text goes wherever your systems send it. No recording is uploaded to a vendor, and there is no vendor-side copy to audit, breach or subpoena.

That is different from a "private cloud" region or a dedicated tenant, where the audio still leaves your network and sits on someone else's servers — just in a separate account.

Who needs Arabic speech recognition on-premise?

OrganisationTypical audioWhy it cannot leave
Banks and insurersCall-centre recordings, complaintsCustomer data and internal policy
Government and public servicesHotlines, citizen services, meetingsData must stay in controlled networks
HealthcareConsultations, dictationPatient confidentiality
Legal and compliance teamsInterviews, hearings, investigationsPrivilege and chain of custody
TelecomsSupport and sales calls at volumeVolume makes per-minute pricing expensive, and data is sensitive

If your security team has already said no to cloud transcription, you are the audience for this.

What runs inside your network?

nutq deploys the same models it runs in the cloud — not a cut-down edition:

  • Speech to text (nutq-listen-1): Gulf and Levantine Arabic kept as spoken rather than rewritten into Modern Standard Arabic, plus 13 other languages. On a genuine narrow-band phone call it made 40% fewer character errors than the leading Arabic vendor.
  • Text to speech (nutq-speak-1): Gulf Arabic voices, for IVR prompts and voice agents that reply in dialect.

The accuracy argument is the one from our article on transcribing Arabic phone calls: call-centre audio is narrow-band and in dialect, and that is where models trained on clean MSA fall apart.

What hardware does it need?

On-premise
GPUOne, with 24 GB of memory
Where it runsYour hardware, your network
Audio egressNone
Per-minute feeNone

A single 24 GB GPU is a modest server, not a cluster. At about 22× faster than realtime, one hour of recordings transcribes in under three minutes, which is what makes batch processing of a day's calls practical on one machine.

How is it different from the hosted API?

The hosted API is the fastest way to start: one request to /v1/transcribe, $0.15 per hour of audio, $5 of free credit, no card. It is the right choice for most teams, and the Arabic speech to text API guide covers it end to end.

On-premise trades that convenience for control. You run the server; we deploy the models on it. Instead of paying per minute, you have a fixed deployment, which also changes the economics once call volume is high.

How do you evaluate on-premise Arabic ASR?

Evaluate it the way you would the cloud version, on your own audio:

  1. Take 20–50 real recordings: your phone lines, your customers' dialects, your product names.
  2. Have a native speaker transcribe them exactly as spoken, dialect included.
  3. Run them through the hosted API first — the model is the same, and your free credit covers the test. See Gulf Arabic vs MSA for what to look for.
  4. Count character errors and read for dialect words rewritten into MSA.
  5. If the results hold up, scope the deployment.

Testing on the hosted API with non-sensitive or consented recordings first means your security review only has to approve the deployment, not the evaluation.

Getting started

Tell us about your audio, your volumes and your environment, and we will scope the deployment with you. Talk to us about deployment, or start with the hosted API by creating an account.

Questions people ask

What hardware does on-premise Arabic speech recognition need?

nutq's models run on a single GPU with 24 GB of memory, on your own server, in your own network.

Does any audio leave our network?

No. In an on-premise deployment the recordings are processed where they already live. Nothing is sent to nutq.

How is on-premise priced?

There is no per-minute or per-hour fee, unlike the hosted API. Talk to us about the deployment and we will scope it with you.

Are the on-premise models the same as the cloud API?

Yes. The same models that serve the hosted API — Gulf and Levantine speech to text, and Gulf Arabic text to speech — are deployed inside your network.

ShareXLinkedInWhatsApp

Written by nutq team. nutq is an Arabic speech API built in the UAE: speech to text that keeps Gulf dialect, and Arabic text to speech with voice cloning.

اقرأ بالعربية

Hear it on your own audio.

$5 of credit to start — about 33 hours of transcription. No card.

All articles