diff --git a/authors/assets/notdev-ng-company-dark.png b/authors/assets/notdev-ng-company-dark.png new file mode 100644 index 00000000..13677bff Binary files /dev/null and b/authors/assets/notdev-ng-company-dark.png differ diff --git a/authors/assets/notdev-ng-company-white.png b/authors/assets/notdev-ng-company-white.png new file mode 100644 index 00000000..12b794d6 Binary files /dev/null and b/authors/assets/notdev-ng-company-white.png differ diff --git a/authors/assets/notdev-ng.jpg b/authors/assets/notdev-ng.jpg new file mode 100644 index 00000000..d1fe13a2 Binary files /dev/null and b/authors/assets/notdev-ng.jpg differ diff --git a/authors/notdev_ng.md b/authors/notdev_ng.md new file mode 100644 index 00000000..4a3b3394 --- /dev/null +++ b/authors/notdev_ng.md @@ -0,0 +1,15 @@ +Author: Notdev Ng +Title: Software Engineer +Description: Notdev Ng is a software engineer focused on developer tooling, +reproducible development environments, and applied AI workflows. He writes +practical guides that help teams turn messy local setups into containers anyone +can run, and he contributes to open-source projects around developer +experience. +Author Image: ![notdev-ng](./assets/notdev-ng.jpg) +Author LinkedIn: [LinkedIn](https://www.linkedin.com/in/notdevng) +Author Twitter: [Twitter](https://twitter.com/notdevng) +Company Name: Codiev +Company Description: Codiev builds AI-assisted development tools that help +engineers turn ideas into working software faster. +Company Logo Dark: ![company-logo dark](./assets/notdev-ng-company-dark.png) +Company Logo White: ![company-logo white](./assets/notdev-ng-company-white.png) diff --git a/definitions/20260920_definition_automatic_speech_recognition.md b/definitions/20260920_definition_automatic_speech_recognition.md new file mode 100644 index 00000000..124b420d --- /dev/null +++ b/definitions/20260920_definition_automatic_speech_recognition.md @@ -0,0 +1,36 @@ +--- +title: 'Automatic Speech Recognition (ASR)' +description: + 'Automatic Speech Recognition is the technology that converts spoken audio + into written text using machine-learning models trained on speech data.' +date: 2026-09-20 +author: 'Notdev Ng' +--- + +# Automatic Speech Recognition (ASR) + +## Definition + +Automatic Speech Recognition (ASR) is the technology that converts spoken +language in an audio signal into written text. Modern ASR systems are built on +neural networks — most commonly encoder-decoder Transformers — that map an +audio representation, such as a log-Mel spectrogram, to a sequence of text +tokens. Well-known ASR models and services include OpenAI's Whisper family, +ElevenLabs Scribe, Speechmatics, and Google Gemini's audio transcription. + +## Context and Usage + +ASR sits at the front of almost every voice-driven workflow: meeting and +interview transcription, video captions and subtitles, voice assistants, +call-center analytics, and making audio archives searchable. A typical pipeline +first normalizes the audio (resampling to a fixed sample rate, usually mono), +then runs the recognition model, and optionally post-processes the output for +punctuation, casing, or speaker labels. + +For developers, the practical concerns are accuracy (often measured as word +error rate), language and accent coverage, latency, file-size limits on hosted +APIs, and whether audio can leave the machine at all. Local engines such as +Vosk or whisper.cpp trade some accuracy for complete privacy and zero marginal +cost, while hosted APIs offer state-of-the-art accuracy without local model +management. Tools like Sapat wrap many ASR providers behind one CLI so the +choice becomes configuration rather than integration work. diff --git a/guides/20260920_ai_transcription_with_sapat.md b/guides/20260920_ai_transcription_with_sapat.md new file mode 100644 index 00000000..6b3bd111 --- /dev/null +++ b/guides/20260920_ai_transcription_with_sapat.md @@ -0,0 +1,436 @@ +--- +title: 'AI Transcription with Sapat Inside a Daytona Workspace' +description: + 'Build a reproducible speech-to-text pipeline with Sapat inside a Daytona + workspace: ffmpeg conversion, 29 providers, Whisper explained, and batch + correction.' +date: 2026-09-20 +author: 'Notdev Ng' +tags: ['transcription', 'whisper', 'sapat', 'ai', 'devcontainers'] +--- + +# AI Transcription with Sapat Inside a Daytona Workspace + +# Introduction + +Meeting recordings, demo walkthroughs, user interviews, and voice notes are +some of the highest-signal artifacts a team produces — and the hardest to +reuse. Audio cannot be grepped, quoted in a pull request, or fed into a +retrieval pipeline until it is converted to text. That conversion step is where +most transcription setups fall apart: one machine has the right ffmpeg build, +another has the API key, a third was configured by a colleague who left last +quarter. The transcription itself is a solved problem; the environment around +it is not. + +This guide builds that environment once, correctly, inside a +[Daytona](https://www.daytona.io/docs/about/what-is-daytona/) workspace using +[Sapat](https://github.com/nibzard/sapat), an open-source transcription CLI +with a pluggable provider architecture. By the end you will have a +reproducible [development environment](/definitions/20240819_definition_development%20environment.md) +that converts video or audio files to text through any of 29 providers — +hosted Whisper endpoints, multimodal LLMs such as Gemini and Mistral, dedicated +speech vendors like Speechmatics and Gladia, and fully offline engines such as +Vosk and whisper.cpp. Because everything lives in a dev container, the same +setup runs identically on your laptop, a colleague's machine, or a remote +workspace you tear down at the end of the day. + +![Diagram of the Sapat transcription pipeline inside a Daytona workspace: media files are converted and split with ffmpeg, transcribed through an auto-discovered provider, optionally corrected by an LLM pass, and written as sidecar transcripts](assets/20260920_ai_transcription_with_sapat_img1.png) + +You do not need any machine-learning background to follow along. You need +Daytona installed, Docker running, and at least one speech-to-text API key — +or none at all, if you choose a local provider. + +## TL;DR + +- Sapat is a Python CLI that turns media files into `.txt` transcripts. It + converts audio with ffmpeg, routes it to a provider you choose, and can + optionally clean the transcript with an LLM. +- Daytona removes setup drift: the repository ships a dev container that + installs Python, ffmpeg, and dependencies for you. +- Providers are auto-discovered from environment variables — if a key is set, + the provider appears; there is no plugins folder to edit. +- The Whisper step is not magic: understanding language hints, prompt + conditioning, temperature, and file-size limits is what separates a usable + transcript from a hallucinated one. +- Run `sapat recording.mp4 --provider groq --language en` to transcribe, add + `--correct` on providers that support it, and point the CLI at a directory to + batch-process a whole folder of `.mp4` files. + +## Prerequisites + +Before you start, make sure you have: + +- [Docker](https://www.docker.com/) installed and running. +- [Daytona](https://www.daytona.io/docs/installation/installation/) installed + and a default target configured (`daytona target set` if you have never used + it). +- One API key for a hosted provider — `GROQ_API_KEY` or `ELEVENLABS_API_KEY` + is the fastest to obtain (a single variable each) — or a willingness to use + a local engine, which needs no key at all. +- A media file to transcribe. Any `.mp4`, `.mov`, `.mp3`, or `.wav` file works; + a one-minute screen recording is ideal for a first test. + +Familiarity with [environment variables](/definitions/20241126_definition_environment_variables.md) +and basic CLI usage is assumed, but every command below is copy-pasteable. + +## Step 1: Create a Reproducible Workspace with Daytona + +Sapat's repository ships a `.devcontainer/devcontainer.json`, so Daytona can +build the entire toolchain — Python 3.12, ffmpeg, and the Python dependencies — +without you touching a single system package. + +Create the workspace: + +```bash +daytona create https://github.com/nibzard/sapat --code +``` + +The `--code` flag attaches your IDE to the workspace when the build finishes. +While it builds, it is worth opening `.devcontainer/devcontainer.json` in a +browser tab, because this one file is the whole reason the guide is +reproducible: + +```json +{ + "name": "Video Transcription Tool", + "image": "mcr.microsoft.com/devcontainers/python:3.12", + "customizations": { + "vscode": { + "settings": { + "python.defaultInterpreterPath": "/usr/local/bin/python", + "python.linting.enabled": true, + "python.linting.pylintEnabled": true + }, + "extensions": [ + "ms-python.python", + "ms-python.vscode-pylance", + "njpwerner.autodocstring" + ] + } + }, + "mounts": [ + { + "source": "${localEnv:HOME}", + "target": "${containerWorkspaceFolder}/con-home", + "type": "bind" + } + ], + "onCreateCommand": { + "update": "sudo apt update && sudo apt upgrade -y", + "ownership": "sudo chown -R $USER:$USER ${containerWorkspaceFolder}" + }, + "postCreateCommand": { + "ffmpeg": "sudo apt install ffmpeg -y", + "requirements": "pip install -r requirements.txt" + } +} +``` + +Two fields do the heavy lifting. `image` pins the base +[development container](/definitions/20240910_definition_dev_container_feature.md) +— a Python 3.12 image maintained by Microsoft — so every workspace starts from +the same bits. `postCreateCommand` runs after the container is created: it +installs ffmpeg system-wide and then installs Sapat's Python dependencies +(`click`, `requests`, `python-dotenv`, `openai`, `groq`, `build`). Anything +you would otherwise document in a README under "now install these apt +packages" belongs here, where it executes automatically. The remaining fields +are conveniences: `customizations` configures VS Code, `mounts` exposes your +home directory inside the container, and `onCreateCommand` updates apt +packages and fixes workspace ownership. + +If you are integrating Sapat into an existing project instead of cloning the +repository, add the same two commands to your own project's dev container — +`apt install ffmpeg` plus `pip install +"git+https://github.com/nibzard/sapat.git"` — and you get the same result. + +One thing `requirements.txt` does not do is install Sapat itself — the +package is not published to PyPI — so once the workspace is up, install it +editable from the repository root: + +```bash +pip install -e . +``` + +The editable install puts the `sapat` command on `PATH` and also pulls in +`halo`, the one runtime dependency `requirements.txt` leaves out. + +**Verify the workspace.** Inside the container, run: + +```bash +ffmpeg -version | head -1 +python -m pip show sapat +``` + +Both should succeed before you continue. If ffmpeg is missing, rebuild the +workspace rather than installing it by hand — manual fixes are exactly the +drift Daytona exists to eliminate. + +## Step 2: Configure a Speech-to-Text Provider + +Sapat does not hard-code a provider list in its CLI. Each provider lives in +`sapat/providers/.py` and declares what it needs — required environment +variables and, for local engines, required Python packages. At startup the +registry walks that directory, checks each provider's requirements, and +registers only the ones that can actually run. In practice: **the provider you +have credentials for is the provider that shows up.** + +Copy the example file and edit it inside the workspace: + +```bash +cp .env.example .env +``` + +`.env` is listed in `.gitignore`, so credentials stay out of version control. +Set only the variables for providers you plan to use. A minimal configuration: + +```bash +# Groq — fast hosted Whisper +GROQ_API_KEY=gsk_... + +# ElevenLabs Scribe — strong multilingual accuracy +ELEVENLABS_API_KEY=sk_... +``` + +Azure OpenAI needs a small block because it also powers the correction pass: + +```bash +AZURE_OPENAI_API_KEY=your_azure_api_key +AZURE_OPENAI_ENDPOINT=https://.openai.azure.com +AZURE_OPENAI_STT_MODEL_NAME=whisper +AZURE_OPENAI_STT_API_VERSION=2024-06-01 +AZURE_OPENAI_DEPLOYMENT_NAME_CHAT=gpt-4o +AZURE_OPENAI_API_VERSION_CHAT=2023-03-15-preview +``` + +A few providers deserve a second look before you pick one at random: + +| Provider | Kind | Needs | Correction | Notes | +| --- | --- | --- | --- | --- | +| `azure` | Hosted Whisper | 3+ env vars | Yes | Also corrects via a GPT deployment | +| `groq` | Hosted Whisper | `GROQ_API_KEY` | No | Very fast, generous free tier | +| `elevenlabs` | Vendor ASR | `ELEVENLABS_API_KEY` | No | Scribe v2, 90+ languages | +| `gemini` | Multimodal LLM | `GOOGLE_API_KEY` | No | Transcribes via a general model | +| `mistral` | Multimodal LLM | `MISTRAL_API_KEY` | Yes | Voxtral speech model | +| `vosk` | Local engine | `vosk` package | No | Fully offline, small models | +| `whisper_cpp` | Local engine | `whisper-cli` binary | No | Whisper on CPU, no API key | + +The table lists only the highlights; `.env.example` documents all 29 providers, +including Speechmatics, Gladia, Soniox, NVIDIA NIM, Baidu, Yandex, and several +OpenAI-compatible aggregators (DeepInfra, Together, Venice, xAI). Local +providers need their engine installed — either a Python package via extras +(`pip install -e ".[vosk]"` from the repository root) or a binary on `PATH` +(`whisper-cli` for `whisper_cpp`, `whisperx` for `whisperx`). + +## Step 3: Transcribe Your First File + +With credentials in place, transcribe a file: + +```bash +sapat recording.mp4 --provider groq --language en +``` + +Sapat prints the provider, model, and language, then runs a fixed pipeline: +convert, optionally split, transcribe, optionally correct, write the result. +When it finishes, `recording.txt` sits next to `recording.mp4` — a plain-text +transcript ready for diffing, committing, or feeding to another tool. + +Every flag maps to one stage of that pipeline: + +- `--provider`, `-p` — which service transcribes. Defaults to `azure` when + configured, otherwise the first available provider. +- `--model`, `-m` — the provider's model name, e.g. + `--model whisper-large-v3-turbo` on Groq. Each provider defines its own + default, so you only set this to override it. +- `--language`, `-l` — the spoken language as an ISO-639-1 code (`en`, `es`, + `de`). Default `en`. Whisper-family APIs treat this as a hard hint that skips + language detection. +- `--transcription-prompt`, `-t` — free text passed to the model to bias + vocabulary and style (more on this in the next section). +- `--temperature`, `-temp` — sampling temperature, `0`–`1`. Lower is more + deterministic; `0` is the safe default. +- `--quality`, `-q` — the ffmpeg conversion preset applied before upload: + `L` (22.05 kHz mono, 96 kbps), `M` (44.1 kHz mono, 96 kbps), or `H` + (44.1 kHz stereo, 160 kbps). Default `L`, which is plenty for speech. +- `--correct` — a second pass where a chat model fixes names, punctuation, and + casing. Only enabled on providers that implement it (currently `azure` and + `mistral`); elsewhere Sapat prints a warning and skips it. + +## Step 4: What Happens During Transcription + +The issue that inspired this guide asked for one thing above all: **explain +the Whisper part properly.** Most of Sapat's providers call a Whisper-family +model, so understanding it pays off regardless of which provider you pick. + +Whisper is an encoder-decoder Transformer trained by OpenAI on 680,000 hours +of weakly supervised audio. Sapat does not run Whisper itself — it prepares +audio and calls an API — but the API parameters exist because of how the model +works, and knowing the "why" tells you when to change them. + +**Audio becomes a spectrogram.** Whisper never sees waveforms. ffmpeg first +decodes your file, then the provider (or the local engine) converts it to a +log-Mel spectrogram — a grid of frequency-over-time that the encoder consumes +in 30-second windows. This is why `--quality` exists at all: the model was +trained on 16 kHz mono speech, so Sapat downsamples aggressively. `H` stereo +buys you almost nothing for single-speaker audio; it mostly buys upload time. +When in doubt, stay on `L`. + +**The decoder conditions on three things.** Whisper's decoder is trained to +predict text tokens *given* a language token, a task token (transcribe or +translate), and optional context text. The hosted APIs expose the first and +third as `--language` and `--transcription-prompt`. Setting `--language es` +does not translate Spanish — it tells the decoder to emit Spanish text instead +of spending capacity detecting the language. If you omit it, most APIs +auto-detect, which works well on clean audio and badly on the first seconds of +a noisy recording. + +**`--transcription-prompt` is the most underused flag.** The prompt is fed to +the decoder as if it were preceding transcript text, so the model inherits its +vocabulary and style. If your recording says "SapAT", "Daytona", and a +coworker's name that Whisper keeps mangling, write the prompt as one sentence +that uses them correctly: + +```bash +sapat standup.mp4 --provider groq \ + --language en \ + --transcription-prompt "Daily standup for the Sapat project with Nibzard." +``` + +The same prompt improves punctuation style on long dictation and keeps product +names spelled the way your documentation spells them. It does not add words — +it biases decoding toward vocabulary that already appears in the prompt. + +**Temperature controls how brave the decoder is.** At `0`, Whisper picks the +highest-probability token at every step; the open-source decoder also retries +with higher temperatures when a window's average log-probability is too low or +its compression ratio too high — the model's built-in escape hatch from +degenerate output. Hosted APIs accept a single temperature value. Keep `0` +unless a transcript comes out weirdly flat on expressive speech; raising it +above ~0.5 invites creative transcription you did not ask for. + +**Hallucinations have recognizable causes.** Whisper loops ("thank you thank +you thank you"), invents sentences during silence, and hallucinates subtitle +text ("Subtitles by Amara.org") on quiet audio because its training data +contained subtitled video. The practical mitigations: trim long silences or +split chapters before transcribing, keep temperature low, and sanity-check +files that come back suspiciously short or repetitive. Sapat's `--correct` +pass can patch names and casing, but it cannot fix a transcript that was never +grounded in the audio. + +**Size limits are why Sapat splits.** Hosted Whisper endpoints cap uploads at +roughly 25 MB. When a converted file exceeds the provider's limit, Sapat +slices it into time-based chunks with ffmpeg, transcribes each chunk, and +concatenates the result. If a long file produces a transcript with a visible +seam every N minutes, that is the chunk boundary — give each chunk context by +keeping `--transcription-prompt` generic rather than referencing specific +moments. + +## Step 5: Batch Processing and Transcript Correction + +Point Sapat at a directory instead of a file and it processes every `.mp4` +inside, with a progress bar: + +```bash +sapat ./recordings --provider elevenlabs --language en +``` + +Two operational notes save real time on batches. First, Sapat skips the +ffmpeg step when `recording.mp3` already exists next to `recording.mp4`. That +mainly helps retries after a crashed run, or files you converted yourself +beforehand — a successful run deletes the `.mp3` it used, so a clean re-run +reconverts from the source file. Second, transcripts land as sidecar `.txt` +files — and both `*.mp3` and `*.txt` are in the repo's `.gitignore`, so a +batch run never pollutes a commit. + +On providers with `supports_correction`, `--correct` runs a second pass: the +raw transcript goes to a chat model with a prompt that asks only for spelling, +punctuation, and casing fixes. Use it for content headed to documentation or +captions; skip it for search indexing, where the extra latency buys nothing. + +## Step 6: Choose the Right Provider + +With everything working, the provider choice is an engineering trade-off, not +a coin flip: + +- **Speed and cost**: Groq is typically the fastest hosted Whisper endpoint + and has a usable free tier — the right default for batch jobs. It ships with + `whisper-large-v3`; pass `--model whisper-large-v3-turbo` for the faster + turbo variant. +- **No data leaves the box**: `vosk`, `whisper_cpp`, `whisperx`, and + `moonshine` run entirely offline. Expect lower accuracy than the hosted + models, and budget for model files (hundreds of MB). +- **Already-paying clouds**: Azure OpenAI and Oracle Cloud AI Speech reuse + credentials and quotas teams often already have; `azure` is also the + reference provider for `--correct`. +- **Accuracy on hard audio**: ElevenLabs Scribe and Speechmatics are dedicated + ASR vendors and tend to beat general models on accents, crosstalk, and domain + jargon. + +A reliable way to choose: transcribe the same difficult two-minute clip with +two providers and `diff` the sidecar files. Ten minutes of testing beats an +hour of reading marketing pages. + +## Common Issues and Troubleshooting + +**Problem:** `sapat: command not found`, or `ModuleNotFoundError: No module +named 'halo'`. + +**Solution:** The dev container installs Sapat's dependencies but not Sapat +itself. Run `pip install -e .` from the repository root inside the workspace. + +**Problem:** `Provider 'x' is not available.` + +**Solution:** The registry only lists providers whose requirements are met. +Check that the env vars are actually in `.env` (not just exported in another +shell), and for local engines that the package or binary is installed in the +container. + +**Problem:** `ffmpeg: command not found` or conversion fails. + +**Solution:** Rebuild the workspace (`daytona delete`, then `daytona create` +again) so `postCreateCommand` reruns. Outside dev containers, install ffmpeg +with your system package manager. + +**Problem:** Transcription fails on a large file with a 413 or size error. + +**Solution:** The provider's `max_file_size_mb` gate should have split it — +confirm the split happened (look for "splitting into chunks"). If it did not, +convert down with `--quality L` or split the media manually with `ffmpeg -f +segment`. + +**Problem:** `--correct` prints a warning and does nothing. + +**Solution:** Your provider does not implement correction. Use `azure` or +`mistral`, or post-process the `.txt` with your own LLM call. + +**Problem:** The transcript loops or invents text during quiet sections. + +**Solution:** Classic Whisper hallucination on silence. Trim silence with +`ffmpeg -af silenceremove`, keep `--temperature 0`, and add a +`--transcription-prompt` describing the recording. + +## Conclusion + +You now have a transcription environment that any teammate can reproduce in +one command: `daytona create` builds Python and ffmpeg from the dev container, +a `.env` file selects among 29 auto-discovered providers, and `sapat` turns +media into sidecar transcripts — with `--correct` available when the output is +headed somewhere humans will read it. + +The natural next step is wiring transcripts into something bigger: commit them +to a docs repo, index them for search, or feed them to an LLM for summaries and +action items. Whatever you build on top, the transcript is now the easy part. + +## References + +- [Sapat on GitHub](https://github.com/nibzard/sapat) — source, provider list, + and `.env.example` +- [Daytona documentation](https://www.daytona.io/docs) — workspaces, targets, + and dev container support +- [OpenAI Whisper](https://github.com/openai/whisper) — the model family behind + most providers used here +- [Whisper API reference](https://platform.openai.com/docs/api-reference/audio) + — `language`, `prompt`, `temperature`, and `response_format` parameters +- [How to Set Up a Project with Cal.com and Supabase](https://www.daytona.io/dotfiles/how-to-set-up-a-project-with-cal-com-and-supabase) + — a shorter Daytona walkthrough in the same style +- [Dev Containers specification](https://containers.dev/) — the + `devcontainer.json` fields used in Step 1 diff --git a/guides/assets/20260920_ai_transcription_with_sapat_img1.png b/guides/assets/20260920_ai_transcription_with_sapat_img1.png new file mode 100644 index 00000000..31030008 Binary files /dev/null and b/guides/assets/20260920_ai_transcription_with_sapat_img1.png differ