Glossary

Speech recognition

Converting speech to text, with accuracy that depends heavily on accent, audio quality and domain.

Also called
ASR, speech to text

Word error rate is meaningful only relative to the audio it was measured on. A figure from clean American English says nothing about a contact centre in Coimbatore.

Domain vocabulary tuning, drug names, product codes, legal terms, is the highest-leverage improvement available.

Audio quality is usually the cheapest lever available. Upgrading handsets or fixing a noisy line often improves recognition more than any model change, and costs less. Measure the audio before blaming the model.

Commonly misunderstood: Code-mixed Indian speech, where speakers switch language mid-sentence, is handled poorly by general models and needs specific testing.

Related terms, in context

The concepts you almost always meet alongside speech recognition.

Speaker diarisation
Identifying who spoke when, so a transcript records attribution rather than just words.
Voice AI agent
A phone agent that holds a real conversation, listening, reasoning and speaking within a conversational turn.
Text to speech
Generating spoken audio from text, with prosody natural enough not to announce itself.

Where this shows up in our work

Speech recognition is not an abstraction for us. It is a decision we make on live projects. It shows up most directly in speech recognition & transcription, voice ai agents, where getting it wrong has a cost someone can measure.

If you are evaluating a vendor on this, the useful question is not whether they can define the term. It is what they measure, what they would refuse to do, and what happens in their system when the assumption behind speech recognition stops holding.

Questions

What is Speech recognition?

Converting speech to text, with accuracy that depends heavily on accent, audio quality and domain.

What do people get wrong about speech recognition?

Code-mixed Indian speech, where speakers switch language mid-sentence, is handled poorly by general models and needs specific testing.

Does Orqent Labs build this?

Yes, Speech Recognition & Transcription and Voice AI Agents. We work across India, covering all 19,238 PIN codes remotely.

Building something that involves speech recognition?

We will tell you honestly whether it is the right approach for your problem.

Or email bd@dtrasglobal.com · call +91 74118 77878