A Python implementation of the Visual Microphone algorithm, which recovers sound from high-speed video by analyzing sub-pixel surface vibrations. When sound hits an object, it makes the surface vibrate by tiny amounts. They are far too small to see with the naked eye, but they show up in the phase of complex wavelet coefficients. This tool measures those vibrations and turns them back into sound, so a bag of chips or a plant leaf becomes a microphone.
The original work by Davis et al. (MIT CSAIL, SIGGRAPH 2014) used Complex Steerable Pyramids for the video decomposition. This project uses 2D Dual-Tree Complex Wavelet Transform (DTCWT) instead, which is ~5x more computationally efficient while still providing reliable phase information for motion estimation. We test against the same high-speed videos provided by MIT CSAIL.
The sample videos can be downloaded from here. For example, Chips1-2200Hz-Mary_Had-input.avi is a high-speed video of a bag of chips vibrating to "Mary Had A Little Lamb" (704x704, 22,859 frames captured at 2200 fps, ~14 GB). Note: the AVI container reports ~30 fps, but the actual capture rate is 2200 Hz. Use --fps 2200 when running visualmic.py to get the correct audio sample rate.
- Setup
- Usage
- GPU Acceleration
- Part 1: The Original Work (Davis et al., SIGGRAPH 2014)
- Part 2: Our Implementation (2D DTCWT)
- Part 3: Literature Survey
- Future Work
- Development
- References
git clone https://github.com/joeljose/Visual-Mic.git
cd Visual-Mic
pip install -r requirements.txt
python visualmic.py -i testvid.avi -o recovered_audio.wavRequirements: Python 3.11 (the version the Docker images and CI use; older versions are untested)
# Build
./docker-build.sh
# Run
docker run --rm --user "$(id -u):$(id -g)" -v /path/to/videos:/data \
visual-mic:latest \
-i /data/testvid.avi -o /data/sound.wavRequires nvidia-container-toolkit.
# Build
./docker-build-gpu.sh
# Run
docker run --rm --gpus all --user "$(id -u):$(id -g)" -v /path/to/videos:/data \
visual-mic-gpu:latest \
--gpu -i /data/Chips1-2200Hz-Mary_Had-input.avi \
-o /data/sound.wav --fps 2200 --batch-size 32The --batch-size flag controls how many frames are processed per GPU batch (default: 16). Larger batches are faster but use more GPU memory. At 704x704, each frame uses ~10 MB of GPU memory, so --batch-size 32 needs ~660 MB including overhead.
Note: GPU mode produces very similar but not bit-identical output compared to CPU mode, due to float32 vs float64 precision differences and different DTCWT implementations (pytorch_wavelets vs dtcwt).
# Basic usage
python visualmic.py -i testvid.avi -o recovered_audio.wav
# With temporal bandpass filter
python visualmic.py -i testvid.avi -fl 80 -fh 1000
# With ROI (focus on vibrating object)
python visualmic.py -i testvid.avi --roi 100,50,200,150
# Override frame rate for high-speed video
python visualmic.py -i Chips1-2200Hz-Mary_Had-input.avi --fps 2200
# GPU acceleration
python visualmic.py -i testvid.avi --gpu --batch-size 32
# Custom wavelet filters
python visualmic.py -i testvid.avi --biort near_sym_a --qshift qshift_a| Flag | Default | Description |
|---|---|---|
-i / --input |
(required) | Input video path |
-o / --output |
sound.wav |
Output audio path |
-fl / --freq-low |
fps/40, kept within 20 to 100 Hz | Lower cutoff frequency (Hz) for temporal bandpass filter |
-fh / --freq-high |
none | Upper cutoff frequency (Hz) for temporal bandpass filter |
--no-filter |
off | Disable the default high-pass (raw phase signals) |
--denoise |
off | Spectral subtraction of stationary noise (hum, light flicker, sensor noise) |
--fps |
from the video | Override video frame rate (Hz) for audio sample rate |
--roi |
whole frame | Region of interest as x,y,w,h |
--gpu |
off | Use GPU-accelerated DTCWT (requires CUDA + pytorch_wavelets) |
--batch-size |
16 | Frames per GPU batch (GPU mode only) |
--device |
cuda |
PyTorch device for --gpu: cuda, cuda:N, or cpu to run the PyTorch path without a GPU |
--nlevels |
3 | Number of DTCWT decomposition levels |
--biort |
near_sym_b |
Biorthogonal wavelet filter for DTCWT level 1 |
--qshift |
qshift_b |
Quarter-shift wavelet filter for DTCWT levels 2+ |
--version |
Show program version and exit |
Available wavelet filters:
--biort:antonini,legall,near_sym_a,near_sym_b--qshift:qshift_06,qshift_a,qshift_b,qshift_c,qshift_d
A Butterworth high-pass is applied to the phase signals by default, at 1/20 of the Nyquist frequency (fps/40), kept within 20 to 100 Hz. That is the rule Davis et al. used, and it gives 55 Hz at 2200 fps. It rejects slow drift, which would otherwise swamp the audio. -fl and -fh set the band explicitly; --no-filter turns it off.
When --denoise is specified, stationary noise is removed by spectral subtraction. The noise level of each frequency is estimated as its median over the whole clip, so steady hum and light flicker (multiples of 50/60 Hz) are suppressed while the recovered sound, which comes and goes, is kept.
When --fps is specified, the given value is used as the audio sample rate instead of the frame rate reported by the video container. High-speed footage often needs it, because the container frame rate can be wrong. Below 500 fps the tool prints a warning, since that is usually the sign of a wrong rate.
The filter settings are checked against the frame rate before any frame is processed. A cutoff of zero or less, or a -fl at or above the Nyquist frequency, stops the run with an error. Frames are read until the video ends, so a wrong frame count in the file doesn't drop or invent audio.
Errors and warnings go to stderr. The output path is checked before any frame is processed, so a typo in -o fails at once instead of after the whole run.
When --roi is specified, each frame is cropped to the given rectangle before the DTCWT decomposition. This reduces computation and can improve SNR by focusing on the vibrating object.
- Always pass the true capture rate with
--fps; the high-pass cutoff and output sample rate depend on it. - Add
--denoisefor listening; leave it off when you need the raw motion signal. - Use
--roito crop to the vibrating object. It cuts computation, and it can raise SNR when much of the frame is background. - For MIT CSAIL videos, always use
--fps 2200(the container reports ~30 fps incorrectly). - Use
--gpufor large videos. The DTCWT forward pass takes most of the time. - If GPU runs out of memory, reduce
--batch-size. - The audio can only hold frequencies up to half the frame rate. At 2200 fps that is 1100 Hz: enough for the pitch of a voice or a tune, not for clear consonants.
The --gpu flag enables GPU-accelerated processing via PyTorch and pytorch_wavelets. The GPU path replaces the CPU DTCWT forward transform with a CUDA-accelerated batched equivalent while keeping the same algorithmic pipeline. Temporal postprocessing (filtering, combining the sub-bands, denoising) stays on the CPU in NumPy, because it works on the small phase signal array, not on the full wavelet coefficients.
The GPU path processes frames in configurable batches rather than one at a time:
- Batch accumulation: Grayscale frames are collected into batches of
--batch-sizeframes (default: 16). The last batch can be smaller, and it is processed too - GPU transfer: The batch is stacked into a
(B, 1, H, W)float32 tensor and sent to GPU - Batched forward DTCWT:
pytorch_wavelets.DTCWTForwardprocesses all frames in the batch simultaneously, producingYh[level]with shape(B, 1, 6, H_l, W_l, 2)where the last dimension is real/imaginary - Phase extraction: Performed on-GPU for the entire batch (see below)
- Transfer back: Only the small phase signal array
(B, nlevels, 6)is copied back to the CPU. The full wavelet coefficients are discarded - Memory cleanup: GPU tensors are explicitly deleted (
del batch_tensor, Yl, Yh) after each batch
This streaming architecture means GPU memory usage is proportional to batch_size, not frame_count, so a video of any length fits.
The CPU path uses NumPy's complex number support (np.angle(coeffs * prev_conj)). The GPU path must handle complex arithmetic manually because pytorch_wavelets represents coefficients as real/imaginary pairs in the last dimension:
Yh[level] shape: (B, 1, 6, H, W, 2)
└── [0]=real, [1]=imag
Conjugate multiplication (phase change since the previous frame):
# (c + id)(a - ib) = (ca + db) + i(da - cb)
prod_real = c_real * r_real + c_imag * r_imag
prod_imag = c_imag * r_real - c_real * r_imagPhase and amplitude-squared weighting (vectorized over entire batch):
phase_diff = torch.atan2(prod_imag, prod_real)
amp_sq = c_real * c_real + c_imag * c_imag
weighted = (amp_sq * phase_diff).sum(dim=(-2, -1)) # sum over H, WThis produces the same
Previous frame across batches: each frame is compared with the frame before it. Inside a batch that is the previous entry; for the first frame of a batch it is the last frame of the batch before, kept as a small GPU tensor. The very first frame is compared with itself, so it contributes zero.
Pre-flight VRAM estimation: Before processing, estimate_vram() calculates peak VRAM usage based on batch size and frame dimensions. The DTCWT forward transform requires ~15x the input frame size in working memory (filter banks, intermediate convolutions). If estimated usage exceeds 70% of available VRAM, a warning is printed with suggestions to reduce --batch-size, use --roi, or switch to CPU mode.
OOM handling: If a CUDA out-of-memory error occurs during processing, the tool catches it and exits with an actionable error message rather than a raw PyTorch traceback.
CPU/GPU output differences: The GPU path uses float32 (PyTorch/CUDA standard) while the CPU path uses float64. Combined with different DTCWT implementations (pytorch_wavelets vs dtcwt), outputs are not bit-identical. On a synthetic clip with 4 px of drift, the two outputs correlate at 1.0.
Benchmarked on Chips2-2200Hz-Mary_MIDI-input.avi (704x400, 38,083 frames, 2200 fps) on a laptop with an RTX 4050 (6 GB VRAM):
| Configuration | Time | Notes |
|---|---|---|
| GPU, default settings | 3m 50s | --batch-size 32, near_sym_b/qshift_b |
GPU, --nlevels 2 |
3m 49s | Fewer decomposition levels |
| GPU, old filters | 3m 30s | --biort near_sym_a --qshift qshift_a |
| GPU, with ROI | 2m 46s | --roi 100,50,400,300 (smaller region) |
| CPU, default settings | 24m 23s | one process, v3.1.0, Docker image; the machine was also running other jobs |
Hardware requirements (GPU path):
- NVIDIA GPU with CUDA 12.1+ support
- Minimum ~1 GB VRAM for typical videos (scales with
--batch-sizeand resolution) nvidia-container-toolkitfor Docker GPU support
Paper: "The Visual Microphone: Passive Recovery of Sound from Video" Authors: Abe Davis, Michael Rubinstein, Neal Wadhwa, Gautham J. Mysore, Frédo Durand, William T. Freeman Venue: ACM Transactions on Graphics (SIGGRAPH 2014), Vol 33, No 4 Institutions: MIT CSAIL, Microsoft Research, Adobe Research
When sound travels through air, it creates pressure waves. When these waves hit an object's surface, they make it vibrate. The displacements are about a micrometer or less (the paper measured this with a laser vibrometer). These vibrations are far too small to see with the naked eye, but a high-speed camera recording thousands of frames per second can capture them as subtle pixel-level changes.
So if we can measure those sub-pixel displacements over time, we have a recording of the sound. The object has become a microphone.
Example: Playing "Mary Had A Little Lamb" near a bag of chips causes the bag's surface to vibrate at the frequencies of the music. A high-speed camera (the paper used 2 kHz to 20 kHz) pointed at the bag captures these vibrations as tiny frame-to-frame changes.
You might think: "Just compute optical flow between frames and track the motion." The problem:
-
The motions are sub-pixel, typically
$\frac{1}{100}$ to$\frac{1}{1000}$ of a pixel. Standard optical flow fails at this scale. - Noise dominates. Sensor noise, quantization and lighting flicker are all larger than the vibration in any one pixel.
- Every frame counts. Each frame is one audio sample, so the motion has to be measured in every frame, not smoothed over several.
Solution: Instead of tracking pixels in the spatial domain, work in the frequency domain using the phase of complex wavelet/pyramid coefficients. Phase is far more sensitive to small motions than amplitude.
Consider a 1D signal shifted by a small displacement
In the Fourier domain, a spatial shift becomes a phase shift:
So if you decompose an image into frequency bands and track how the phase of each band changes over time, you're directly measuring local displacement at that spatial frequency.
For a band-pass filtered signal at spatial frequency
where
| Property | Amplitude |
Phase |
|---|---|---|
| Physical meaning | "How much texture is here" | "Where exactly is this texture positioned" |
| Response to small motion | Relatively stable | Shifts linearly with displacement |
| Sub-pixel sensitivity | Poor | Excellent, detects fractions of a pixel |
The original paper uses a Complex Steerable Pyramid to decompose each video frame.
A multi-scale, multi-orientation filter bank that decomposes an image into:
- Multiple scales (frequency bands): coarse
$\rightarrow$ fine detail - Multiple orientations at each scale: e.g.,
$0°, 30°, 60°, 90°, 120°, 150°$ (for 6 orientations) - A lowpass residual (the blurry base image)
- A highpass residual (the finest details)
Each sub-band produces complex-valued coefficients. At each spatial location
where:
-
$A$ = amplitude (how strong the texture is at this location/scale/orientation) -
$\phi$ = phase (the precise position of the texture pattern)
The filters can be rotated to any orientation analytically, without recomputing. That gives fine control over direction and avoids aliasing artifacts.
- Translation equivariant: shifting the input shifts the coefficients predictably (phase changes linearly)
-
Overcomplete (~21x for 8 orientations): more coefficients than pixels
$\rightarrow$ redundancy helps with noise - Shift-invariant: no downsampling artifacts that would corrupt phase measurements
- High-speed video
$V(x, y, t)$ with$N$ frames at$F$ fps - Grayscale frames (color not needed for vibration)
For each frame
This gives complex coefficients at
For each coefficient:
Choose a reference frame
This phase difference is proportional to how much the texture at location
Why subtract the reference? The absolute phase values encode the texture pattern itself (which we don't care about). By subtracting the reference, we keep only the change, and the change is the vibration.
For each scale
Why weight by
- Regions with strong texture (high amplitude) give reliable phase measurements
- Regions with weak/no texture (low amplitude) have noisy/random phase, and we want to suppress them
-
$A^2$ weighting is a reliability-weighted average: trustworthy measurements count more
This produces one 1D time signal per
Different scales and orientations may have phase offsets relative to each other. Align them using cross-correlation:
- Pick a reference sub-band (e.g., scale 0, orientation 0):
$\text{ref} = \Phi(0, 0, t)$ - For each other
$(s, \theta)$ , find the time shift that maximizes correlation:
- Shift each sub-band signal by its optimal lag
$\tau$ .
Averaging also removes noise. The vibration is the same in every sub-band, so it adds up, while the noise differs from band to band and partly cancels.
The paper high-passes the result at 20 to 100 Hz (for most examples 1/20 of the Nyquist frequency) to remove low-frequency noise that isn't sound. For very noisy videos it filters each sub-band before the alignment instead. Then it denoises: spectral subtraction when the goal is accuracy, or a speech enhancement method (Loizou 2005) when the goal is intelligibility. Every result in the paper is denoised.
Maps the signal to
Write as WAV file with sampling rate = video FPS.
If the video is 2200 fps, the audio is sampled at 2200 Hz. By the Nyquist theorem it holds frequencies up to 1100 Hz. That covers the fundamental frequency of most voices and of low musical notes.
High-speed cameras are expensive. Most ordinary cameras have a rolling shutter: the sensor reads its rows one after another, not all at once, so each row is a picture of a slightly different moment.
The paper uses this to get audio from a normal 60 fps camera. If the object moves horizontally, the horizontal shift of each row is one audio sample, so the sample rate becomes the row rate instead of the frame rate. Their Pentax K-01 recorded 1280x720 at 60 fps with a row delay of 16 microseconds. That gives 61,920 samples per second. About 30% of them are missing, because the sensor pauses for 5 ms between frames, and the paper fills the gaps with an audio interpolation method.
The exposure time sets the real limit. A row exposed for 1/2000 s blurs any vibration faster than about 2000 Hz. The recovered sound (a recording of "The Raven" played near a bag of candy) is shown as a spectrogram in the paper. This project does not implement the rolling shutter method.
- Requires high-speed video for good quality (2000+ fps ideal; 60 fps with rolling shutter is limited)
- Object must have visible texture. Smooth, featureless surfaces give poor results
- Signal-to-noise ratio depends on the object's material, the distance, and the sound volume
- Computationally expensive. The paper reports 2 to 3 hours per video in MATLAB
- Global averaging loses spatial information. All vibrations are mixed together
The DTCWT was developed by Nick Kingsbury (Cambridge, late 1990s) as an improvement over the standard Discrete Wavelet Transform (DWT).
- Not shift-invariant: shifting input by 1 pixel completely changes the coefficients
- Poor directional selectivity: only separates horizontal, vertical and diagonal, with no finer orientations
- Oscillating coefficients: makes phase extraction unreliable
Run two parallel DWT filter banks (two "trees"):
-
Tree
$a$ : uses one set of filters$\rightarrow$ produces real part -
Tree
$b$ : uses a slightly different (quarter-sample shifted) set of filters$\rightarrow$ produces imaginary part
The filters are designed so that Tree
This complex wavelet is approximately analytic (has energy only on one side of the frequency spectrum), which provides:
- Approximate shift invariance (2x oversampling eliminates most aliasing)
- Clean phase information (no oscillation artifacts)
For 2D images, the DTCWT produces 6 complex sub-bands per scale, oriented at approximately:
This is fewer orientations than a typical steerable pyramid (which might use 8+), but:
- Only ~4x overcomplete (vs ~21x for steerable pyramid)
- Much faster to compute
- Still provides good directional selectivity
- Phase information is reliable for motion estimation
| Property | Complex Steerable Pyramid | 2D DTCWT |
|---|---|---|
| Shift invariant | Yes (exactly) | Approximately |
| Orientations per scale | Configurable (typically 8) | 6 (fixed) |
| Overcompleteness | ~21x (8 orientations) | ~4x |
| Computation speed | Slow (frequency domain) | Fast (filter banks) |
| Phase quality | Excellent | Good |
| Python library | pyrtools |
dtcwt |
| Reconstruction | Perfect | Near-perfect |
Trade-off: DTCWT is ~5x more computationally efficient at the cost of slightly fewer orientations and approximate (rather than exact) shift invariance. For the visual microphone that is a good trade: the phase is still good enough to detect sub-pixel vibrations.
This is how visualmic.py implements the pipeline. The snippets are simplified from the source.
Frames are streamed from the video file. Each frame is read, transformed and discarded, so only one raw frame is in memory at a time, and a video of any length fits in memory. If an ROI is specified, each frame is cropped before the DTCWT decomposition, reducing computation and focusing on the vibrating object.
def extract_audio(cap, frame_count, nlevels, n_orient, fps, ..., roi=None, denoise=False):
transform = dtcwt.Transform2d()
prev_conj = None
phase_signals = []
while True: # read until the video ends
ret, raw_frame = cap.read()
if not ret or raw_frame is None:
break
gray = cv2.cvtColor(raw_frame, cv2.COLOR_BGR2GRAY)
if roi is not None:
rx, ry, rw, rh = roi
gray = gray[ry:ry+rh, rx:rx+rw]
highpasses = transform.forward(gray, nlevels=nlevels).highpasses
if prev_conj is None:
prev_conj = [np.conj(h) for h in highpasses]
frame_phases = np.zeros((nlevels, n_orient))
for level in range(nlevels):
coeffs = highpasses[level]
amp = np.abs(coeffs)
phase_diff = np.angle(coeffs * prev_conj[level]) # change since the previous frame
frame_phases[level, :] = np.sum(amp * amp * phase_diff, axis=(0, 1))
phase_signals.append(frame_phases)
prev_conj = [np.conj(h) for h in highpasses]
phase_signals = accumulate_phase(phase_signals) # running sum, shape (frames, nlevels, n_orient)Vectorized NumPy operations on entire 2D spatial slices:
| Operation | Code | Corresponds to |
|---|---|---|
| Extract amplitude |
np.abs(coeffs) |
Step 2 of original |
| Phase change since the previous frame | np.angle(coeffs * prev_conj[level]) |
Step 3 of original (see below) |
|
|
amp * amp * phase_diff |
Step 4 of original |
| Spatial sum |
np.sum(..., axis=(0, 1)) |
Step 4 of original |
| Running sum over frames | np.cumsum(..., axis=0) |
replaces the fixed reference frame |
The conjugate multiplication coeffs * conj(prev) gives the phase difference directly: angle(z * conj(w)) = angle(z) - angle(w), wrapped to
Why frame to frame, not against frame 0? Phase is only known up to a multiple of
Result: phase_signals[fc, level, angle]
A 4th-order Butterworth filter is applied to all 18 phase signals at once. By default it is a high-pass at default_freq_low(fps) (fps/40, kept within 20 to 100 Hz); -fl/-fh override it and --no-filter disables it:
phase_signals = signal.sosfiltfilt(sos, phase_signals, axis=0)sosfiltfiltapplies the filter forward and backward (zero-phase), so no time delay is introduced- The filter rejects low-frequency drift (camera shake, thermal effects) and high-frequency noise
- Upper cutoff is automatically clamped to 99% of Nyquist to avoid instability
- Clips too short for
sosfiltfilt's usual edge padding (under 28 frames for a band-pass) get a shorter pad - If only
-flis given, acts as highpass; if only-fh, acts as lowpass
Each orientation band measures the image motion projected onto its own direction. The six DTCWT orientations point at 15°, 45° and 75°, and at 105°, 135° and 165°. The last three respond like −75°, −45° and −15°, so they see vertical motion upside down. The bands are combined into horizontal and vertical motion, and the audio is the motion along the dominant vibration direction:
ORIENT_ANGLES = np.deg2rad([15, 45, 75, -75, -45, -15])
dx = (phase_signals * np.cos(ORIENT_ANGLES)).sum(axis=(1, 2))
dy = (phase_signals * np.sin(ORIENT_ANGLES)).sum(axis=(1, 2))
u = vibration_direction(dx, dy, fps) # principal axis after removing stationary noise
sound_raw = u[0] * dx + u[1] * dyWhy not align sub-bands like the original? Davis et al. shift each band in time to best match a reference band (Step 5 of the original). On the MIT Chips2 video that picked arbitrary lags, up to the full clip length, and cost 7 dB of SNR: a deforming bag's sub-bands aren't time-shifted copies of each other. Plain summing works on the bag but cancels rigid vertical motion. Projecting the 2D motion handles both. The direction is estimated after spectral subtraction, so hum and light flicker don't decide it.
Spectral subtraction (Boll 1979), as in the original's Section 3.3, with the noise of each frequency estimated as its median power over the whole clip:
noise = np.median(power, axis=1, keepdims=True)
gain = np.sqrt(np.maximum(1 - alpha * noise / power, beta ** 2)) # alpha=2, beta=0.05p_min = np.min(sound_raw)
p_max = np.max(sound_raw)
if p_max == p_min:
sound_data = np.zeros_like(sound_raw) # silent output if no motion
else:
sound_data = ((2 * sound_raw) - (p_min + p_max)) / (p_max - p_min)def save_wav(samples, output_name, sample_rate):
waveform_integers = np.int16(samples * 32767)
write(output_name, sample_rate, waveform_integers)The sample rate is the frame rate (from the video, or from --fps), rounded to a whole number. One frame becomes one audio sample.
| Parameter | Default | Meaning |
|---|---|---|
nlevels |
3 | Number of wavelet decomposition scales |
n_orient |
6 | Number of orientations per scale (fixed by DTCWT) |
freq_low |
fps/40, within 20 to 100 Hz | High-pass cutoff |
alpha, beta |
2.0, 0.05 | Over-subtraction and gain floor of --denoise |
biort |
near_sym_b |
Biorthogonal filter for level 1 |
qshift |
qshift_b |
Quarter-shift filter for levels 2+ |
With 3 levels of DTCWT decomposition:
| Level | Spatial Frequency | What It Captures | Spatial Resolution |
|---|---|---|---|
| 0 (finest) | High | Fine textures, edges, sharp details | |
| 1 (middle) | Medium | Medium-scale patterns | |
| 2 (coarsest) | Low | Broad structures, large features |
The vibration signal is present across all scales (the whole surface moves), but the signal-to-noise ratio varies:
-
Fine scales: more spatial locations to average
$\rightarrow$ better denoising - Coarse scales: fewer locations but stronger phase response to motion
| Year | Paper | Key Contribution |
|---|---|---|
| 2012 | Eulerian Video Magnification (SIGGRAPH) | Predecessor: amplifies color/intensity changes to visualize motion |
| 2013 | Phase-Based Video Motion Processing (SIGGRAPH) | Established that phase of complex steerable pyramid coefficients = local motion |
| 2014 | Riesz Pyramids (ICCP) | Compact 4x overcomplete pyramid for real-time phase-based processing |
| 2014 | The Visual Microphone (SIGGRAPH) | Recovered sound from video using phase-based motion analysis |
| 2016 | Visual Vibration Analysis (PhD Thesis, Abe Davis) | Extended to modal analysis, material properties, damping estimation |
| Year | Work | Advance |
|---|---|---|
| 2018 | Local Visual Microphones | Local vibration aggregation (not global averaging), 100 to 1000x speedup, sound direction estimation |
| 2022 | Effect of Video Resolution | Studies resolution impact on recovery quality; frame-wise denoising preprocessing |
| 2023 | Event-Based Visual Microphone (CVPR Workshop) | Neuromorphic event cameras for cheap, efficient vibration capture |
| 2024 | PSO-CNN Hybrid | Particle Swarm Optimization + CNN for enhanced sound restoration |
| 2025 | Single-Pixel Visual Microphone (Optica) | Single-pixel imaging with a spatial light modulator, so no expensive high-speed camera is needed |
| Implementation | Technique | Language |
|---|---|---|
| MIT Original | Complex Steerable Pyramid | MATLAB |
| dsforza96/visual-mic | Complex Steerable Pyramid (pyrtools) |
Python |
| This repo (Visual-Mic) | 2D DTCWT (dtcwt) |
Python |
-
Multiprocessing across frames: each frame only needs itself and the frame before it. Reading frames remains sequential (VideoCapture limitation), but the DTCWT + phase extraction can be parallelized across CPU cores using batch processing with
multiprocessing.Pool, giving ~Nx speedup on an N-core machine. -
Closing the gap to the original: on Chips2, coherence with the played sound is 0.49 against 0.60 for MIT's result. Candidates: more scales, a different motion estimate, or equalising the object's frequency response (original Section 4.3).
The MIT dataset includes the sound that was played for each video (*-input.wav) and MIT's own result (*-recovered.wav). scripts/eval_audio.py aligns a recovered WAV to the played sound (lag and sign) and reports SNR, segmental SNR and coherence between 100 and 1000 Hz. Coherence ignores the object's frequency response, so it is the fairest single number.
python scripts/eval_audio.py sound.wav --ref Chips2-2200Hz-Mary_MIDI-input.wavScored on two MIT videos at 2200 fps, run on the GPU. The table has each video's SNR / segSNR (dB) / coherence:
| Output | Chips2 (bag of chips) | Plant (leaves) |
|---|---|---|
| v2.0.0, default | −11.1 / −9.2 / 0.49 | −26.5 / −9.9 / 0.40 |
| v3.0.1, default | −3.8 / −3.2 / 0.49 | |
| current, default | −2.9 / −1.5 / 0.49 | −12.5 / −9.9 / 0.45 |
current, --denoise |
−2.9 / −1.1 / 0.49 | −5.3 / −4.0 / 0.46 |
MIT recovered.wav (denoised) |
−4.0 / −1.1 / 0.60 | −4.7 / −4.0 / 0.46 |
On Plant, the current version with --denoise matches MIT's published result on segmental SNR and coherence. On Chips2 its coherence is still lower (0.49 against 0.60).
All tests run inside Docker, so you need no local Python dependencies:
# CPU: lint + unit tests (builds image automatically if not found)
./test.sh
# GPU: lint + unit tests (requires nvidia-container-toolkit)
./test.sh gpu
# Force rebuild before testing
./test.sh --build
./test.sh gpu --buildCPU tests (tests/test_visualmic.py) cover:
- Utility functions (
format_duration,save_wav,default_freq_low,denoise_spectral) - Phase signal postprocessing (normalization, Butterworth filter)
- VRAM estimation arithmetic
- Full
extract_audiopipeline on synthetic 256x256 video (shape, finiteness) - All CLI validation error paths
Recovery tests (tests/test_recovery.py) move a texture by a known sub-pixel amount, following a 100 to 1000 Hz chirp, and check that the recovered signal matches it. They cover horizontal, vertical and diagonal motion at the motion sizes the paper measured (0.005 to 0.01 px), drift of up to 10 px, wrong frame counts in the file, and the paper's rule that SNR rises 6 dB when the motion doubles.
tests/test_eval_audio.py checks that scripts/eval_audio.py finds a known delay, sign and SNR.
PyTorch path tests (tests/test_visualmic_gpu.py) cover:
- DTCWTForward shapes, finiteness, and custom filter selection
- Full
extract_audio_gpupipeline on synthetic video, including wrong frame counts - Agreement with the NumPy path on a clip with 4 px of drift (correlation above 0.999)
- The
--gpu --devicecommand line, end to end
They run on CUDA when a GPU is present and on the CPU otherwise, and skip only when PyTorch or pytorch_wavelets is missing. CI runs them on a CPU-only PyTorch build on every push.
Version is tracked in a VERSION file at the project root. visualmic.py has __version__ baked into the source (updated at release time).
To cut a release:
- Update
VERSIONwith the new version number - Update
__version__invisualmic.py - Update
CHANGELOG.md: move items from[Unreleased]to[X.Y.Z] - YYYY-MM-DD - Commit:
Release vX.Y.Z - Tag:
git tag -a vX.Y.Z -m "Release vX.Y.Z" - Push:
git push && git push origin vX.Y.Z - Rebuild Docker images:
./docker-build.sh && ./docker-build-gpu.sh
visualmic.py # CLI tool (CPU + GPU paths)
Dockerfile # CPU Docker image (python:3.11-slim)
Dockerfile.gpu # GPU Docker image (python:3.11-slim + torch 2.1.2 CUDA wheels)
docker-build.sh # Build + tag CPU image
docker-build-gpu.sh # Build + tag GPU image
test.sh # Run lint + tests (Docker, supports cpu/gpu mode)
requirements.txt # CPU runtime dependencies
requirements-gpu.txt # GPU runtime dependencies
requirements-dev.txt # Dev dependencies (pytest, ruff)
scripts/
eval_audio.py # Score recovered audio against the played sound
tests/
test_visualmic.py # CPU unit tests
test_visualmic_gpu.py # GPU unit tests (CUDA-only, skip on CPU)
test_recovery.py # Recovery of known synthetic motion
test_eval_audio.py # Check of the evaluation script
docs/design/ # Architecture decision records
visualmic-hardening.md # Hardening design doc
VERSION # Single source of truth for version
CHANGELOG.md # Release history
CONTRIBUTING.md # Contribution guidelines
-
Davis, A., Rubinstein, M., Wadhwa, N., Mysore, G.J., Durand, F., & Freeman, W.T. (2014). The Visual Microphone: Passive Recovery of Sound from Video. ACM Transactions on Graphics (SIGGRAPH), 33(4). Paper PDF | Project Page
-
Wadhwa, N., Rubinstein, M., Durand, F., & Freeman, W.T. (2013). Phase-Based Video Motion Processing. ACM Transactions on Graphics (SIGGRAPH). Project Page
-
Wadhwa, N., Rubinstein, M., Durand, F., & Freeman, W.T. (2014). Riesz Pyramids for Fast Phase-Based Video Magnification. IEEE ICCP. Project Page
-
Selesnick, I.W., Baraniuk, R.G., & Kingsbury, N.G. (2005). The Dual-Tree Complex Wavelet Transform. IEEE Signal Processing Magazine, 22(6), 123-151. Tutorial PDF
-
Davis, A. (2016). Visual Vibration Analysis. PhD Thesis, MIT. Thesis PDF
-
Shen, M. & Bhatt, S. (2018). Local Visual Microphones: Improved Sound Extraction from Silent Video. arXiv:1801.09436
-
Niwa, T., Fushimi, T., Yamamoto, S., & Ochiai, Y. (2023). Live Demonstration: Event-based Visual Microphone. CVPR Workshop on Event-based Vision. Paper PDF
-
dtcwt Python library. Documentation | GitHub
-
MIT CSAIL Visual Microphone Dataset. Download
