Skip to content
This repository was archived by the owner on Sep 22, 2026. It is now read-only.
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Binary file added authors/assets/notdev-ng-company-dark.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added authors/assets/notdev-ng-company-white.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added authors/assets/notdev-ng.jpg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
15 changes: 15 additions & 0 deletions authors/notdev_ng.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
Author: Notdev Ng
Title: Software Engineer
Description: Notdev Ng is a software engineer focused on developer tooling,
reproducible development environments, and applied AI workflows. He writes
practical guides that help teams turn messy local setups into containers anyone
can run, and he contributes to open-source projects around developer
experience.
Author Image: ![notdev-ng](./assets/notdev-ng.jpg)
Author LinkedIn: [LinkedIn](https://www.linkedin.com/in/notdevng)
Author Twitter: [Twitter](https://twitter.com/notdevng)
Company Name: Codiev
Company Description: Codiev builds AI-assisted development tools that help
engineers turn ideas into working software faster.
Company Logo Dark: ![company-logo dark](./assets/notdev-ng-company-dark.png)
Company Logo White: ![company-logo white](./assets/notdev-ng-company-white.png)
36 changes: 36 additions & 0 deletions definitions/20260920_definition_automatic_speech_recognition.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
---
title: 'Automatic Speech Recognition (ASR)'
description:
'Automatic Speech Recognition is the technology that converts spoken audio
into written text using machine-learning models trained on speech data.'
date: 2026-09-20
author: 'Notdev Ng'
---

# Automatic Speech Recognition (ASR)

## Definition

Automatic Speech Recognition (ASR) is the technology that converts spoken
language in an audio signal into written text. Modern ASR systems are built on
neural networks — most commonly encoder-decoder Transformers — that map an
audio representation, such as a log-Mel spectrogram, to a sequence of text
tokens. Well-known ASR models and services include OpenAI's Whisper family,
ElevenLabs Scribe, Speechmatics, and Google Gemini's audio transcription.

## Context and Usage

ASR sits at the front of almost every voice-driven workflow: meeting and
interview transcription, video captions and subtitles, voice assistants,
call-center analytics, and making audio archives searchable. A typical pipeline
first normalizes the audio (resampling to a fixed sample rate, usually mono),
then runs the recognition model, and optionally post-processes the output for
punctuation, casing, or speaker labels.

For developers, the practical concerns are accuracy (often measured as word
error rate), language and accent coverage, latency, file-size limits on hosted
APIs, and whether audio can leave the machine at all. Local engines such as
Vosk or whisper.cpp trade some accuracy for complete privacy and zero marginal
cost, while hosted APIs offer state-of-the-art accuracy without local model
management. Tools like Sapat wrap many ASR providers behind one CLI so the
choice becomes configuration rather than integration work.
Loading