WSL2 Microphone Access Troubleshooting Guide
Overview
The voice-mode MCP server uses the Python sounddevice library for audio recording, which relies on PortAudio for cross-platform audio I/O. WSL2 has known limitations with audio device access, particularly for microphone input.
The Problem
WSL2 does not natively support audio devices. When running voice-mode in WSL2, you may encounter:
No audio devices detected:
sounddevice.PortAudioError: Error querying device -1Empty device list:
sd.query_devices() # Returns nothingRecording failures:
Error: Could not record audio
Root Causes
- WSL2 Architecture: WSL2 runs in a lightweight VM and doesn't have direct access to Windows hardware devices
- Missing Audio Subsystem: WSL2 doesn't include ALSA or PulseAudio by default
- No Native Audio Drivers: WSL2 kernel doesn't include sound card drivers
Solutions
Solution 1: Windows Permissions + WSL2 Packages (Recommended for WSL 2.3.26.0+)
This solution has been confirmed working on recent WSL2 versions with WSLg support.
Prerequisites
- WSL version: 2.3.26.0 or higher
- WSLg enabled (comes with recent WSL2)
- Windows 10/11 with latest updates
Steps
Enable Windows Microphone Permissions:
- Go to Windows Settings → Privacy & security → Microphone
- Turn ON "Let desktop apps access your microphone"
- Ensure your terminal app (Windows Terminal, etc.) has permission
Install Required Packages in WSL2:
sudo apt update sudo apt install -y libasound2-plugins pulseaudioStart PulseAudio (if not auto-started):
pulseaudio --startTest Audio Devices:
# Test with pactl pactl info pactl list sources short # Test with Python python3 -c "import sounddevice as sd; print(sd.query_devices())"Set Default Audio Device (if needed):
# List devices python3 -m sounddevice # Set default input device (replace X with device number) export VOICEMODE_INPUT_DEVICE=X
Solution 2: USB Microphone with USB/IP
For USB microphones, you can use USB/IP to share the device from Windows to WSL2.
Steps
Install USB/IP on Windows:
- Download from usbipd-win releases
- Install the MSI package
Install USB/IP in WSL2:
sudo apt install linux-tools-generic hwdata sudo update-alternatives --install /usr/local/bin/usbip usbip /usr/lib/linux-tools/*-generic/usbip 20Share USB Device:
# In Windows PowerShell (as Administrator) usbipd list usbipd bind --busid <BUSID> usbipd attach --wsl --busid <BUSID>Verify in WSL2:
lsusb
Solution 3: Network Audio Streaming
Use a network audio solution to stream audio from Windows to WSL2.
Option A: PulseAudio Server on Windows
Install PulseAudio for Windows:
- Download from PulseAudio Windows builds
- Extract to
C:\pulseaudio
Configure PulseAudio (
C:\pulseaudio\etc\pulse\default.pa):load-module module-native-protocol-tcp auth-ip-acl=127.0.0.1;172.16.0.0/12 load-module module-waveout sink_name=output source_name=inputStart PulseAudio on Windows:
C:\pulseaudio\bin\pulseaudio.exe -DConfigure WSL2 to use Windows PulseAudio:
# In WSL2 export PULSE_SERVER=tcp:$(cat /etc/resolv.conf | grep nameserver | awk '{print $2}')
Solution 4: LiveKit Transport (Recommended Alternative)
Instead of using local microphone access, use LiveKit for room-based audio:
Set up LiveKit (see LiveKit setup guide)
Configure voice-mode:
export LIVEKIT_URL="wss://your-app.livekit.cloud" export LIVEKIT_API_KEY="your-api-key" export LIVEKIT_API_SECRET="your-api-secret"Use LiveKit transport:
# Force LiveKit transport converse("Hello", transport="livekit")
Debugging Steps
Check WSL Version:
wsl --versionVerify Audio Subsystem:
# Check if PulseAudio is running ps aux | grep pulse # Check ALSA aplay -l # Check sounddevice python3 -m sounddeviceEnable Debug Mode:
export VOICEMODE_DEBUG=true python3 -m voice_modeTest Basic Audio:
import sounddevice as sd import numpy as np # List devices print(sd.query_devices()) # Try recording duration = 1 # seconds fs = 44100 recording = sd.rec(int(duration * fs), samplerate=fs, channels=1) sd.wait() print(f"Recorded {len(recording)} samples")
Known Limitations
- Audio Latency: Network-based solutions (PulseAudio over TCP) introduce latency
- CPU Usage: Audio streaming can increase CPU usage
- Reliability: Network audio can be less reliable than native access
- WSL1 vs WSL2: WSL1 has better device access but worse overall performance
Recommendations
- For Development: Use LiveKit transport to bypass local audio requirements
- For Testing: Run voice-mode on native Linux or macOS for best results
- For Production: Deploy to a proper Linux environment or use cloud-based STT/TTS
Alternative Approaches
- Docker with Audio: Run voice-mode in Docker with proper audio device mapping
- Remote Development: Develop on WSL2 but run voice-mode on a remote Linux server
- Dual Boot: Use native Linux for audio-intensive development
Environment Variables for Troubleshooting
# Force specific audio device
export VOICEMODE_INPUT_DEVICE=0
# Increase audio buffer size
export VOICEMODE_AUDIO_BUFFER_SIZE=4096
# Use different sample rate
export VOICEMODE_SAMPLE_RATE=16000
# Enable verbose audio debugging
export VOICEMODE_DEBUG=true
export VOICEMODE_AUDIO_DEBUG=true