AudioJuly 24, 2026· 8 min read

How Voice Actors Export Audio: Professional Delivery Formats Guide

The exact formats, sample rates, and export settings professional voice actors use to deliver broadcast-ready files to clients.

How Voice Actors Export Audio: Professional Delivery Formats Guide

Here's the thing about voice acting: the performance is only half the job. The other half? Delivering files that don't make audio engineers want to cry.

I've worked with voice actors who nail the read but send over 8kHz MP3s recorded through laptop mics. And I've seen beginners deliver pristine 96kHz/32-bit float files that are overkill for a 15-second radio spot. Both waste everyone's time.

So let's talk about what the pros actually do. Not the gear-obsessed forum answers, but the real-world formats and specs that get you hired again.

The Default: 48kHz/24-bit WAV, Mono

If a client doesn't specify, this is your safest bet. Here's why:

  • 48kHz matches video production standards. Most film, TV, and commercial work runs at 48kHz because it syncs perfectly with 24fps and 30fps video. Audio editors don't want to resample your files.
  • 24-bit gives headroom. It captures more dynamic range than 16-bit, so if you recorded a bit quiet, there's room to boost without introducing noise. (Not that you should record quiet, but it's insurance.)
  • WAV is universal. Every DAW, every video editor, every platform opens WAV files. No codec issues, no compatibility headaches.
  • Mono is correct for single-voice work. Unless you're doing ASMR or intentional stereo effects, a voice is a mono source. Stereo just doubles the file size for no reason.

This is the format I see in 80% of professional voice-over briefs. It's the "just works" choice.

Audiobooks: ACX Has Specific Requirements

If you're narrating audiobooks for Audible (via ACX), you can't just wing it. They have strict technical specs:

  • Format: MP3, constant bitrate (CBR) at 192kbps
  • Sample rate: 44.1kHz
  • Bit depth: 16-bit (before exporting to MP3)
  • Channels: Mono
  • RMS level: Between -23dB and -18dB
  • Peak levels: No louder than -3dB
  • Noise floor: -60dB or lower

ACX will auto-reject files that don't meet these specs. So before you upload, run your files through their ACX Check tool or use RX or similar mastering plugins to hit those targets.

The RMS requirement is the tricky one. If you're too quiet, listeners crank their volume and hear your mouth clicks. Too loud, and you're fatiguing. Use a limiter, normalize to around -20dB RMS, and keep peaks below -3dB. (And yes, this means light compression and careful editing—audiobook narration is mastering work, not just raw reads.)

Video Game Dialogue: Dry and Organized

Game audio is a different beast. You're often delivering hundreds or thousands of lines, and the game engine handles all the mixing, reverb, and dynamics. What the audio team wants from you:

  • Format: WAV, 48kHz/24-bit, mono
  • Processing: None. Deliver dry, uncompressed, no reverb, no EQ, no heavy compression. They'll process it in-engine.
  • Naming: Follow their exact naming convention. Usually something like CharacterName_LineID_TakeNumber.wav
  • Length: Keep files tight. No 5-second lead-in silence. Start within 0.5 seconds of audio, end within 0.5 seconds. They'll add fades programmatically.

Game studios often send a spreadsheet with line IDs. Your job is to record each line, label it correctly, and deliver it clean. If they ask for alternates (angry version, whisper version, hurt version), label those clearly: Zara_042_Angry_T1.wav.

And here's a pro tip: batch your exports. If you're delivering 300 lines, don't manually export each one. Use your DAW's batch export or region export function. In Reaper, for example, you can set regions, name them, and render all at once. Saves hours.

Commercials and Broadcast: It Depends

TV and radio commercials vary by client and country. In the U.S., most broadcast specs look like this:

  • Format: WAV, 48kHz/24-bit, mono
  • Peak level: -3dB to -6dB (leaves room for the mix engineer to add music and SFX)
  • Processing: Light. A bit of EQ for clarity, light compression to even out dynamics, maybe a de-esser if you're sibilant. Don't over-process—engineers prefer raw and will ask if they want it polished.

Some agencies want the final mixed spot, some want just the VO stem. Always ask. And if they're doing the final mix, send them a few takes so they have options.

For radio, older stations might still ask for 44.1kHz/16-bit (CD quality), but 48kHz is becoming universal. When in doubt, go higher—it's easy to downsample, impossible to upsample quality you didn't record.

Animation and Dubbing: Sync Matters

If you're dubbing or doing ADR (automated dialogue replacement), timing is everything. You're often replacing existing dialogue, so your delivery needs to match the mouth movements or original pacing.

Standard delivery:

  • Format: WAV, 48kHz/24-bit, mono
  • Sync reference: If they send you a video reference, export your audio to match the exact frame. Use timecode if provided.
  • Processing: Minimal. The mix engineer will match the sonic profile to the existing production.

Some dubbing studios want you to include a "guide track" (the original audio for reference) in a separate channel or file. Check the brief.

Podcasts and YouTube: More Flexible

Podcast and YouTube voice-over work is less strict because the final product is digital-only (no broadcast regulations). But professionalism still matters.

Common specs:

  • Format: WAV or MP3
  • Sample rate: 44.1kHz or 48kHz
  • Bit depth: 16-bit or 24-bit (for WAV)
  • Bitrate: 192kbps or 320kbps (for MP3)
  • Channels: Mono, unless the podcast has stereo music and they want you in a specific channel

For podcast ads, some networks want you to deliver a fully mixed :30 or :60 spot. Others just want the raw VO and they'll drop it into the feed. Ask.

YouTube creators are often less technical, so if they don't specify, send 48kHz/24-bit WAV and you're covered. If file size is an issue (long narration, slow internet), compress to 320kbps MP3—it's transparent enough for YouTube's compression pipeline.

File Naming: The Thing Everyone Forgets

You can deliver perfect audio in the perfect format and still annoy clients with bad file names. Here's the structure pros use:

ProjectName_ScriptID_TakeNumber_YourName_Date.wav

Example:

CosmoBrand_Commercial30_T3_JaneSmith_072426.wav

Why this works:

  • Project name first so files group together alphabetically
  • Script or scene ID so the editor knows which line this is
  • Take number so they can compare multiple reads
  • Your name (crucial if multiple actors are on the project)
  • Date in MMDDYY format (no slashes, keeps it filesystem-safe)

Avoid vague names like FinalVO_v2.wav or Take1.wav. When the client has 50 files from 10 different actors, your file should be self-explanatory.

Delivery Method: How to Send the Files

Email attachments are fine for small files (under 10MB). Beyond that, you'll hit size limits and annoy clients. Here's what pros use:

  • Google Drive or Dropbox: Upload, share a link with download access. Clean, reliable, clients can grab files anytime.
  • WeTransfer: Free for files up to 2GB, expires after 7 days. Good for one-time deliveries.
  • Frame.io or other review platforms: Some clients use these for script review and file delivery. You upload, they comment and approve.
  • FTP/SFTP: Some studios have their own servers. They'll give you login details.

Whatever you use, include a README.txt file with:

  • Your name and contact info
  • Project name
  • File specs (sample rate, bit depth, format)
  • Any notes (e.g., "Take 3 has the client-requested pacing change")

It's a small thing, but it shows you're organized and professional.

When the Client Asks for "The Highest Quality Possible"

Sometimes clients don't know what they actually need and just say "send the best quality." Here's what that means in practice:

Don't send them 96kHz/32-bit float files. That's overkill unless they're doing forensic audio restoration or sound design for a Christopher Nolan film. Your voice doesn't need 96kHz. The Nyquist theorem says 48kHz captures all frequencies humans can hear (up to 24kHz). Going higher just bloats the file.

Do send 48kHz/24-bit WAV, mono, with clean edits and minimal processing. That's professional broadcast quality and works for 99% of projects.

If they push back and want higher, fine—send 96kHz/24-bit. But it's rarely necessary and often slows down their workflow (bigger files = longer upload/download times, more storage, harder to edit on older systems).

Quick Format Cheat Sheet

Project TypeFormatSample RateBit DepthChannels
Commercial / BroadcastWAV48kHz24-bitMono
Audiobook (ACX)MP3 (192kbps CBR)44.1kHz16-bit (source)Mono
Video Game DialogueWAV48kHz24-bitMono
Animation / DubbingWAV48kHz24-bitMono
Podcast / YouTubeWAV or MP3 (320kbps)44.1 or 48kHz16 or 24-bitMono

And look, if a client sends you specific specs, follow those. This guide is for when they don't—or when you're setting your default workflow.

The goal is simple: deliver files that sound great, import cleanly, and don't create extra work for the person on the other end. Do that consistently, and you'll get hired again. That's the real measure of professional delivery.

Need to convert your audio files to the right format? Or maybe trim silence from the ends? KokoConvert handles all the common export tasks in your browser—no software install needed.

Frequently Asked Questions

What audio format should I use for voice-over deliverables?
For most professional work, deliver WAV files at 48kHz/24-bit. This is the industry standard for broadcast, film, and video production. For audiobooks, ACX requires MP3 at 192kbps CBR. Game audio often needs both WAV (for implementation) and compressed formats for runtime playback.
Should I deliver mono or stereo voice files?
Mono is standard for most voice-over work. A single voice source doesn't benefit from stereo, and mono files are half the size. The exception: if you're delivering atmospheric narration, audiobook chapters with intentional spatial effects, or client specifically requests stereo.
Do I need to normalize or compress my audio before delivery?
It depends on the client and project type. Audiobooks need strict mastering specs (ACX requires -23 to -18 dB RMS, with peaks below -3 dB). Commercial spots often want light compression and normalization. For game dialogue, deliver dry, uncompressed audio—the game engine handles dynamics. Always ask the client for their spec sheet.
What sample rate should voice actors record at?
Record at 48kHz/24-bit for video, broadcast, and games. This matches professional video frame rates and provides headroom for editing. For audiobooks, 44.1kHz/16-bit is acceptable (CD quality), but 48kHz gives you more flexibility. Never record below 44.1kHz—you can't upsample quality you didn't capture.
How should I name my voice-over files?
Use a consistent naming convention: ProjectName_ScriptID_Take_YourName_Date. For example: CosmoBrand_Commercial30_T3_JaneSmith_072426.wav. For games with hundreds of lines, include character name and line ID: ZombieSlayer_Zara_Line042_T1.wav. Clear naming prevents confusion when clients manage multiple takes and actors.