Voice products are easier to compare when the technical language is clear. This Speechfinds voice-tech glossary explains the terms that appear most often in transcription, assistants, creator tools and privacy settings.
Speech recognition and speech-to-text
Automatic speech recognition (ASR) is the technology that converts spoken audio into machine-readable text. Speech-to-text is the common user-facing term for the same outcome. Accuracy can change with accent, language, background noise, microphone quality and subject vocabulary.
Word error rate is a technical measure of transcription mistakes. A lower score is better, but a laboratory score may not predict performance in a Hinglish meeting or noisy classroom.
Transcription, captions and dictation
Transcription turns recorded or live audio into a written document. Captions synchronise text with video or live speech and may include sound information. Dictation lets a person compose text by speaking directly into an app or operating system.
Speaker diarisation separates speakers and labels who spoke when. It is useful for interviews and meetings, although names often need manual correction.
Voice assistants and wake words
A voice assistant interprets commands and completes actions such as search, reminders, calls or smart-home control. A wake word is the phrase that activates listening for a command. Devices may process the wake-word detection locally while sending the following request to cloud services.
Intent recognition identifies what the speaker wants, while natural-language understanding interprets meaning beyond exact command wording.
Text-to-speech and synthetic voices
Text-to-speech (TTS) converts written text into spoken audio. It supports accessibility, navigation, education and content production. A synthetic voice is generated by software rather than recorded one sentence at a time.
Voice cloning creates a synthetic voice resembling a particular person. Responsible use requires clear permission, disclosure and protection against impersonation or fraud. A realistic result does not prove that the named person actually spoke the words.
On-device, cloud and latency
On-device processing happens mainly on the phone, computer or speaker. It can improve speed and privacy for supported tasks. Cloud processing sends data to remote servers and may enable more capable models. Latency is the delay between speaking and receiving a response.
Hybrid products use both approaches. Readers should check what happens offline, what is uploaded and which controls apply to stored recordings.
Indian language and code-switching terms
Code-switching means moving between languages in one conversation, such as Hindi and English. Transliteration represents words from one script in another; Hinglish typed in Latin letters is a common example. Language support on a feature list does not guarantee strong code-switching accuracy.
Privacy and account controls
Retention describes how long a service keeps audio or transcripts. Training-data controls indicate whether user content may improve models. Voice history is the account record of past interactions. Look for deletion, auto-delete, permission and export options.
Speechfinds uses these terms consistently so readers can compare tools without mistaking marketing language for practical capability.