Forensic Audio Research Audio Voice Recovery Best Practices
Comprehensive audio forensics and voice recovery guide providing CSI-level capabilities for recovering voice from low-quality, low-volume, or damaged audio recordings. Contains 45 rules across 8 categories, prioritized by impact to guide audio enhancement, forensic analysis, and transcription workflows.
When to Apply
Reference these guidelines when:
- Recovering voice from noisy or low-quality recordings
- Enhancing audio for transcription or legal evidence
- Performing forensic audio authentication
- Analyzing recordings for tampering or splices
- Building automated audio processing pipelines
- Transcribing difficult or degraded speech
Rule Categories by Priority
| Priority |
Category |
Impact |
Prefix |
Rules |
| 1 |
Signal Preservation & Analysis |
CRITICAL |
signal- |
5 |
| 2 |
Noise Profiling & Estimation |
CRITICAL |
noise- |
5 |
| 3 |
Spectral Processing |
HIGH |
spectral- |
6 |
| 4 |
Voice Isolation & Enhancement |
HIGH |
voice- |
7 |
| 5 |
Temporal Processing |
MEDIUM-HIGH |
temporal- |
5 |
| 6 |
Transcription & Recognition |
MEDIUM |
transcribe- |
5 |
| 7 |
Forensic Authentication |
MEDIUM |
forensic- |
5 |
| 8 |
Tool Integration & Automation |
LOW-MEDIUM |
tool- |
7 |
Quick Reference
1. Signal Preservation & Analysis (CRITICAL)
signal-preserve-original - Never modify original recording
signal-lossless-format - Use lossless formats for processing
signal-sample-rate - Preserve native sample rate
signal-bit-depth - Use maximum bit depth for processing
signal-analyze-first - Analyze before processing
2. Noise Profiling & Estimation (CRITICAL)
noise-profile-silence - Extract noise profile from silent segments
noise-identify-type - Identify noise type before reduction
noise-adaptive-estimation - Use adaptive estimation for non-stationary noise
noise-snr-assessment - Measure SNR before and after
noise-avoid-overprocessing - Avoid over-processing and musical artifacts
3. Spectral Processing (HIGH)
spectral-subtraction - Apply spectral subtraction for stationary noise
spectral-wiener-filter - Use Wiener filter for optimal noise estimation
spectral-notch-filter - Apply notch filters for tonal interference
spectral-band-limiting - Apply frequency band limiting for speech
spectral-equalization - Use forensic equalization to restore intelligibility
spectral-declip - Repair clipped audio before other processing
4. Voice Isolation & Enhancement (HIGH)
voice-rnnoise - Use RNNoise for real-time ML denoising
voice-dialogue-isolate - Use source separation for complex backgrounds
voice-formant-preserve - Preserve formants during pitch manipulation
voice-dereverb - Apply dereverberation for room echo
voice-enhance-speech - Use AI speech enhancement services for quick results
voice-vad-segment - Use VAD for targeted processing
voice-frequency-boost - Boost frequency regions for specific phonemes
5. Temporal Processing (MEDIUM-HIGH)
temporal-dynamic-range - Use dynamic range compression for level consistency
temporal-noise-gate - Apply noise gate to silence non-speech segments
temporal-time-stretch - Use time stretching for intelligibility
temporal-transient-repair - Repair transient damage (clicks, pops, dropouts)
temporal-silence-trim - Trim silence and normalize before export
6. Transcription & Recognition (MEDIUM)
transcribe-whisper - Use Whisper for noise-robust transcription
transcribe-multipass - Use multi-pass transcription for difficult audio
transcribe-segment - Segment audio for targeted transcription
transcribe-confidence - Track confidence scores for uncertain words
transcribe-hallucination - Detect and filter ASR hallucinations
7. Forensic Authentication (MEDIUM)
forensic-enf-analysis - Use ENF analysis for timestamp verification
forensic-metadata - Extract and verify audio metadata
forensic-tampering - Detect audio tampering and splices
forensic-chain-custody - Document chain of custody for evidence
forensic-speaker-id - Extract speaker characteristics for identification
8. Tool Integration & Automation (LOW-MEDIUM)
tool-ffmpeg-essentials - Master essential FFmpeg audio commands
tool-sox-commands - Use SoX for advanced audio manipulation
tool-python-pipeline - Build Python audio processing pipelines
tool-audacity-workflow - Use Audacity for visual analysis and manual editing
tool-install-guide - Install audio forensic toolchain
tool-batch-automation - Automate batch processing workflows
tool-quality-assessment - Measure audio quality metrics
Essential Tools
| Tool |
Purpose |
Install |
| FFmpeg |
Format conversion, filtering |
brew install ffmpeg |
| SoX |
Noise profiling, effects |
brew install sox |
| Whisper |
Speech transcription |
pip install openai-whisper |
| librosa |
Python audio analysis |
pip install librosa |
| noisereduce |
ML noise reduction |
pip install noisereduce |
| Audacity |
Visual editing |
brew install audacity |
Workflow Scripts (Recommended)
Use the bundled scripts to generate objective baselines, create a workflow plan, and verify results.
scripts/preflight_audio.py - Generate a forensic preflight report (JSON or Markdown).
scripts/plan_from_preflight.py - Create a workflow plan template from the preflight report.
scripts/compare_audio.py - Compare objective metrics between baseline and processed audio.
Example usage:
# 1) Analyze and capture baseline metrics
python3 skills/.experimental/audio-voice-recovery/scripts/preflight_audio.py evidence.wav --out preflight.json
# 2) Generate a workflow plan template
python3 skills/.experimental/audio-voice-recovery/scripts/plan_from_preflight.py --preflight preflight.json --out plan.md
# 3) Compare baseline vs processed metrics
python3 skills/.experimental/audio-voice-recovery/scripts/compare_audio.py \
--before evidence.wav \
--after enhanced.wav \
--format md \
--out comparison.md
Forensic Preflight Workflow (Do This Before Any Changes)
Align preflight with SWGDE Best Practices for the Enhancement of Digital Audio (20-a-001) and SWGDE Best Practices for Forensic Audio (08-a-001).
Establish an objective baseline state and plan the workflow so processing does not introduce clipping, artifacts, or false "done" confidence.
Use scripts/preflight_audio.py to capture baseline metrics and preserve the report with the case file.
Capture and record before processing:
- Record evidence identity and integrity: path, filename, file size, SHA-256 checksum, source, format/container, codec
- Record signal integrity: sample rate, bit depth, channels, duration
- Measure baseline loudness and levels: LUFS/LKFS, true peak, peak, RMS, dynamic range, DC offset
- Detect clipping and document clipped-sample percentage, peak headroom, exact time ranges
- Identify noise profile: stationary vs non-stationary, dominant noise bands, SNR estimate
- Locate the region of interest (ROI) and document time ranges and changes over time
- Inspect spectral content and estimate speech-band energy and intelligibility risk
- Scan for temporal defects: dropouts, discontinuities, splices, drift
- Evaluate channel correlation and phase anomalies (if stereo)
- Extract and preserve metadata: timestamps, device/model tags, embedded notes
Procedure:
- Prepare a forensic working copy, verify hashes, and preserve the original untouched.
- Locate ROI and target signal; document exact time ranges and changes across the recording.
- Assess challenges to intelligibility and signal quality; map challenges to mitigation strategies.
- Identify required processing and plan a workflow order that avoids unwanted artifacts.
Generate a plan draft with
scripts/plan_from_preflight.py and complete it with case-specific decisions.
- Measure baseline loudness and true peak per ITU-R BS.1770 / EBU R 128 and record peak/RMS/DC offset.
- Detect clipping and dropouts; if clipping is present, declip first or pause and document limitations.
- Inspect spectral content and noise type; collect representative noise profile segments and estimate SNR.
- If stereo, evaluate channel correlation and phase; document anomalies.
- Create a baseline listening log (multiple devices) and define success criteria for intelligibility and listenability.
Failure-pattern guardrails:
- Do not process until every preflight field is captured.
- Document every process, setting, software version, and time segment to enable repeatability.
- Compare each processed output to the unprocessed input and assess progress toward intelligibility and listenability.
- Avoid over-processing; review removed signal (filter residue) to avoid removing target signal components.
- Keep intermediate files uncompressed and preserve sample rate/bit depth when moving between tools.
- Perform a final review against the original; if unsatisfactory, revise or stop and report limitations.
- If the request is not achievable, communicate limitations and do not declare completion.
- Require objective metrics and A/B listening before declaring completion.
- Do not rely solely on objective metrics; corroborate with critical listening.
- Take listening breaks to avoid ear fatigue during extended reviews.
Quick Enhancement Pipeline
# 1. Analyze original (run preflight and capture baseline metrics)
python3 skills/.experimental/audio-voice-recovery/scripts/preflight_audio.py evidence.wav --out preflight.json
# 2. Create working copy with checksum
cp evidence.wav working.wav
sha256sum evidence.wav > evidence.sha256
# 3. Apply enhancement
ffmpeg -i working.wav -af "\
highpass=f=80,\
adeclick=w=55:o=75,\
afftdn=nr=12:nf=-30:nt=w,\
equalizer=f=2500:t=q:w=1:g=3,\
loudnorm=I=-16:TP=-1.5:LRA=11\
" enhanced.wav
# 4. Transcribe
whisper enhanced.wav --model large-v3 --language en
# 5. Verify original unchanged
sha256sum -c evidence.sha256
# 6. Verify improvement (objective comparison + A/B listening)
python3 skills/.experimental/audio-voice-recovery/scripts/compare_audio.py \
--before evidence.wav \
--after enhanced.wav \
--format md \
--out comparison.md
How to Use
Read individual reference files for detailed explanations and code examples:
- Section definitions - Category structure and impact levels
- Rule template - Template for adding new rules
Reference Files
| File |
Description |
| AGENTS.md |
Complete compiled guide with all rules |
| references/_sections.md |
Category definitions and ordering |
| assets/templates/_template.md |
Template for new rules |
| metadata.json |
Version and reference information |
1---2name: audio-voice-recovery3description: Forensic Audio Research Audio Voice Recovery Best Practices4---5# Forensic Audio Research Audio Voice Recovery Best Practices67Comprehensive audio forensics and voice recovery guide providing CSI-level capabilities for recovering voice from low-quality, low-volume, or damaged audio recordings. Contains 45 rules across 8 categories, prioritized by impact to guide audio enhancement, forensic analysis, and transcription workflows.89## When to Apply1011Reference these guidelines when:12- Recovering voice from noisy or low-quality recordings13- Enhancing audio for transcription or legal evidence14- Performing forensic audio authentication15- Analyzing recordings for tampering or splices16- Building automated audio processing pipelines17- Transcribing difficult or degraded speech1819## Rule Categories by Priority2021| Priority | Category | Impact | Prefix | Rules |22|----------|----------|--------|--------|-------|23| 1 | Signal Preservation & Analysis | CRITICAL | `signal-` | 5 |24| 2 | Noise Profiling & Estimation | CRITICAL | `noise-` | 5 |25| 3 | Spectral Processing | HIGH | `spectral-` | 6 |26| 4 | Voice Isolation & Enhancement | HIGH | `voice-` | 7 |27| 5 | Temporal Processing | MEDIUM-HIGH | `temporal-` | 5 |28| 6 | Transcription & Recognition | MEDIUM | `transcribe-` | 5 |29| 7 | Forensic Authentication | MEDIUM | `forensic-` | 5 |30| 8 | Tool Integration & Automation | LOW-MEDIUM | `tool-` | 7 |3132## Quick Reference3334### 1. Signal Preservation & Analysis (CRITICAL)3536- [`signal-preserve-original`](references/signal-preserve-original.md) - Never modify original recording37- [`signal-lossless-format`](references/signal-lossless-format.md) - Use lossless formats for processing38- [`signal-sample-rate`](references/signal-sample-rate.md) - Preserve native sample rate39- [`signal-bit-depth`](references/signal-bit-depth.md) - Use maximum bit depth for processing40- [`signal-analyze-first`](references/signal-analyze-first.md) - Analyze before processing4142### 2. Noise Profiling & Estimation (CRITICAL)4344- [`noise-profile-silence`](references/noise-profile-silence.md) - Extract noise profile from silent segments45- [`noise-identify-type`](references/noise-identify-type.md) - Identify noise type before reduction46- [`noise-adaptive-estimation`](references/noise-adaptive-estimation.md) - Use adaptive estimation for non-stationary noise47- [`noise-snr-assessment`](references/noise-snr-assessment.md) - Measure SNR before and after48- [`noise-avoid-overprocessing`](references/noise-avoid-overprocessing.md) - Avoid over-processing and musical artifacts4950### 3. Spectral Processing (HIGH)5152- [`spectral-subtraction`](references/spectral-subtraction.md) - Apply spectral subtraction for stationary noise53- [`spectral-wiener-filter`](references/spectral-wiener-filter.md) - Use Wiener filter for optimal noise estimation54- [`spectral-notch-filter`](references/spectral-notch-filter.md) - Apply notch filters for tonal interference55- [`spectral-band-limiting`](references/spectral-band-limiting.md) - Apply frequency band limiting for speech56- [`spectral-equalization`](references/spectral-equalization.md) - Use forensic equalization to restore intelligibility57- [`spectral-declip`](references/spectral-declip.md) - Repair clipped audio before other processing5859### 4. Voice Isolation & Enhancement (HIGH)6061- [`voice-rnnoise`](references/voice-rnnoise.md) - Use RNNoise for real-time ML denoising62- [`voice-dialogue-isolate`](references/voice-dialogue-isolate.md) - Use source separation for complex backgrounds63- [`voice-formant-preserve`](references/voice-formant-preserve.md) - Preserve formants during pitch manipulation64- [`voice-dereverb`](references/voice-dereverb.md) - Apply dereverberation for room echo65- [`voice-enhance-speech`](references/voice-enhance-speech.md) - Use AI speech enhancement services for quick results66- [`voice-vad-segment`](references/voice-vad-segment.md) - Use VAD for targeted processing67- [`voice-frequency-boost`](references/voice-frequency-boost.md) - Boost frequency regions for specific phonemes6869### 5. Temporal Processing (MEDIUM-HIGH)7071- [`temporal-dynamic-range`](references/temporal-dynamic-range.md) - Use dynamic range compression for level consistency72- [`temporal-noise-gate`](references/temporal-noise-gate.md) - Apply noise gate to silence non-speech segments73- [`temporal-time-stretch`](references/temporal-time-stretch.md) - Use time stretching for intelligibility74- [`temporal-transient-repair`](references/temporal-transient-repair.md) - Repair transient damage (clicks, pops, dropouts)75- [`temporal-silence-trim`](references/temporal-silence-trim.md) - Trim silence and normalize before export7677### 6. Transcription & Recognition (MEDIUM)7879- [`transcribe-whisper`](references/transcribe-whisper.md) - Use Whisper for noise-robust transcription80- [`transcribe-multipass`](references/transcribe-multipass.md) - Use multi-pass transcription for difficult audio81- [`transcribe-segment`](references/transcribe-segment.md) - Segment audio for targeted transcription82- [`transcribe-confidence`](references/transcribe-confidence.md) - Track confidence scores for uncertain words83- [`transcribe-hallucination`](references/transcribe-hallucination.md) - Detect and filter ASR hallucinations8485### 7. Forensic Authentication (MEDIUM)8687- [`forensic-enf-analysis`](references/forensic-enf-analysis.md) - Use ENF analysis for timestamp verification88- [`forensic-metadata`](references/forensic-metadata.md) - Extract and verify audio metadata89- [`forensic-tampering`](references/forensic-tampering.md) - Detect audio tampering and splices90- [`forensic-chain-custody`](references/forensic-chain-custody.md) - Document chain of custody for evidence91- [`forensic-speaker-id`](references/forensic-speaker-id.md) - Extract speaker characteristics for identification9293### 8. Tool Integration & Automation (LOW-MEDIUM)9495- [`tool-ffmpeg-essentials`](references/tool-ffmpeg-essentials.md) - Master essential FFmpeg audio commands96- [`tool-sox-commands`](references/tool-sox-commands.md) - Use SoX for advanced audio manipulation97- [`tool-python-pipeline`](references/tool-python-pipeline.md) - Build Python audio processing pipelines98- [`tool-audacity-workflow`](references/tool-audacity-workflow.md) - Use Audacity for visual analysis and manual editing99- [`tool-install-guide`](references/tool-install-guide.md) - Install audio forensic toolchain100- [`tool-batch-automation`](references/tool-batch-automation.md) - Automate batch processing workflows101- [`tool-quality-assessment`](references/tool-quality-assessment.md) - Measure audio quality metrics102103## Essential Tools104105| Tool | Purpose | Install |106|------|---------|---------|107| FFmpeg | Format conversion, filtering | `brew install ffmpeg` |108| SoX | Noise profiling, effects | `brew install sox` |109| Whisper | Speech transcription | `pip install openai-whisper` |110| librosa | Python audio analysis | `pip install librosa` |111| noisereduce | ML noise reduction | `pip install noisereduce` |112| Audacity | Visual editing | `brew install audacity` |113114## Workflow Scripts (Recommended)115116Use the bundled scripts to generate objective baselines, create a workflow plan, and verify results.117118- `scripts/preflight_audio.py` - Generate a forensic preflight report (JSON or Markdown).119- `scripts/plan_from_preflight.py` - Create a workflow plan template from the preflight report.120- `scripts/compare_audio.py` - Compare objective metrics between baseline and processed audio.121122Example usage:123124```bash125# 1) Analyze and capture baseline metrics126python3 skills/.experimental/audio-voice-recovery/scripts/preflight_audio.py evidence.wav --out preflight.json127128# 2) Generate a workflow plan template129python3 skills/.experimental/audio-voice-recovery/scripts/plan_from_preflight.py --preflight preflight.json --out plan.md130131# 3) Compare baseline vs processed metrics132python3 skills/.experimental/audio-voice-recovery/scripts/compare_audio.py \133 --before evidence.wav \134 --after enhanced.wav \135 --format md \136 --out comparison.md137```138139## Forensic Preflight Workflow (Do This Before Any Changes)140141Align preflight with SWGDE Best Practices for the Enhancement of Digital Audio (20-a-001) and SWGDE Best Practices for Forensic Audio (08-a-001).142Establish an objective baseline state and plan the workflow so processing does not introduce clipping, artifacts, or false "done" confidence.143Use `scripts/preflight_audio.py` to capture baseline metrics and preserve the report with the case file.144145Capture and record before processing:146- Record evidence identity and integrity: path, filename, file size, SHA-256 checksum, source, format/container, codec147- Record signal integrity: sample rate, bit depth, channels, duration148- Measure baseline loudness and levels: LUFS/LKFS, true peak, peak, RMS, dynamic range, DC offset149- Detect clipping and document clipped-sample percentage, peak headroom, exact time ranges150- Identify noise profile: stationary vs non-stationary, dominant noise bands, SNR estimate151- Locate the region of interest (ROI) and document time ranges and changes over time152- Inspect spectral content and estimate speech-band energy and intelligibility risk153- Scan for temporal defects: dropouts, discontinuities, splices, drift154- Evaluate channel correlation and phase anomalies (if stereo)155- Extract and preserve metadata: timestamps, device/model tags, embedded notes156157Procedure:1581. Prepare a forensic working copy, verify hashes, and preserve the original untouched.1592. Locate ROI and target signal; document exact time ranges and changes across the recording.1603. Assess challenges to intelligibility and signal quality; map challenges to mitigation strategies.1614. Identify required processing and plan a workflow order that avoids unwanted artifacts.162 Generate a plan draft with `scripts/plan_from_preflight.py` and complete it with case-specific decisions.1635. Measure baseline loudness and true peak per ITU-R BS.1770 / EBU R 128 and record peak/RMS/DC offset.1646. Detect clipping and dropouts; if clipping is present, declip first or pause and document limitations.1657. Inspect spectral content and noise type; collect representative noise profile segments and estimate SNR.1668. If stereo, evaluate channel correlation and phase; document anomalies.1679. Create a baseline listening log (multiple devices) and define success criteria for intelligibility and listenability.168169Failure-pattern guardrails:170- Do not process until every preflight field is captured.171- Document every process, setting, software version, and time segment to enable repeatability.172- Compare each processed output to the unprocessed input and assess progress toward intelligibility and listenability.173- Avoid over-processing; review removed signal (filter residue) to avoid removing target signal components.174- Keep intermediate files uncompressed and preserve sample rate/bit depth when moving between tools.175- Perform a final review against the original; if unsatisfactory, revise or stop and report limitations.176- If the request is not achievable, communicate limitations and do not declare completion.177- Require objective metrics and A/B listening before declaring completion.178- Do not rely solely on objective metrics; corroborate with critical listening.179- Take listening breaks to avoid ear fatigue during extended reviews.180181## Quick Enhancement Pipeline182183```bash184# 1. Analyze original (run preflight and capture baseline metrics)185python3 skills/.experimental/audio-voice-recovery/scripts/preflight_audio.py evidence.wav --out preflight.json186187# 2. Create working copy with checksum188cp evidence.wav working.wav189sha256sum evidence.wav > evidence.sha256190191# 3. Apply enhancement192ffmpeg -i working.wav -af "\193 highpass=f=80,\194 adeclick=w=55:o=75,\195 afftdn=nr=12:nf=-30:nt=w,\196 equalizer=f=2500:t=q:w=1:g=3,\197 loudnorm=I=-16:TP=-1.5:LRA=11\198" enhanced.wav199200# 4. Transcribe201whisper enhanced.wav --model large-v3 --language en202203# 5. Verify original unchanged204sha256sum -c evidence.sha256205206# 6. Verify improvement (objective comparison + A/B listening)207python3 skills/.experimental/audio-voice-recovery/scripts/compare_audio.py \208 --before evidence.wav \209 --after enhanced.wav \210 --format md \211 --out comparison.md212```213214## How to Use215216Read individual reference files for detailed explanations and code examples:217218- [Section definitions](references/_sections.md) - Category structure and impact levels219- [Rule template](assets/templates/_template.md) - Template for adding new rules220221## Reference Files222223| File | Description |224|------|-------------|225| [AGENTS.md](AGENTS.md) | Complete compiled guide with all rules |226| [references/_sections.md](references/_sections.md) | Category definitions and ordering |227| [assets/templates/_template.md](assets/templates/_template.md) | Template for new rules |228| [metadata.json](metadata.json) | Version and reference information |