Why Capturing System Audio on Mac Feels Complicated
Ask anyone who has tried to record a browser tab, a game, or a video call on a Mac and you will hear the same story: the screen recorder captures everything except the sound. That is not a bug in one particular app. It is a deliberate consequence of how macOS isolates audio between applications.
On many other operating systems, a program can ask the system for whatever is playing on the speakers and usually gets an answer without much friction. macOS takes a different stance. Applications are sandboxed from each other's audio streams, which is excellent for privacy and stability. A misbehaving app cannot quietly listen to your voice call, and one crashing audio engine does not take down another. The trade-off is that recording what you hear is not a first-class feature.
The practical solution is to insert something between the apps and your speakers: a virtual audio device. Instead of fighting the operating system, you give it a new sound card that exists only in software. You tell macOS to send audio to that device, and you point your recorder at it. Once that mental model clicks, the rest of the setup is mostly configuration.
This guide walks through the whole chain: choosing an approach, installing a virtual device, building a multi-output setup, recording cleanly, monitoring without feedback, and fixing the handful of problems that account for most failed sessions. It also covers how this workflow fits into modern video production, including screen recordings and AI-assisted video projects where you need reliable internal audio.
The Mental Model: Output, Input, and the Device in Between
Almost every mistake in system audio recording comes from confusing three roles. Separate them before you touch a single setting.
Output devices decide where sound goes
Your Mac sends playback to exactly one output device at a time through the standard settings. Normally that is your speakers or headphones. The moment you switch the default output to something else, every app that follows the system default follows along.
Input devices decide what recorders can hear
Recording applications do not capture the speakers. They capture an input device. If you want a recorder to hear what an app is playing, the audio has to arrive on an input that the recorder can select. This is the gap that a virtual device fills.
Virtual devices play both roles
A virtual audio device appears in the system as both an output and an input. Apps send audio into it as if it were a speaker, and recorders read from it as if it were a microphone. Because the device is software, the signal travels entirely inside your machine with no cable, no adapter, and no loss from analog conversion.
The catch: if you set the virtual device as your only output, you stop hearing anything yourself. Playback goes into the void of the recorder. That is why the standard setup pairs the virtual device with your real speakers inside a multi-output device, so the same signal reaches both destinations.
One more property is worth understanding: clocking. A virtual device has no hardware clock of its own; it borrows timing from the system. When two devices run at different nominal sample rates, the operating system resamples one of them continuously. Over ten minutes that is invisible. Over a two-hour live session, it can drift by an audible amount. Keeping every device at one sample rate removes that entire category of problem before it starts.
Choosing Your Approach Before You Install Anything
Not every project needs the same level of infrastructure. Match the tool to the job.
Quick and occasional recordings
If you need to capture a short clip once a month, the built-in screen recording tool plus a virtual device is enough. The setup takes a few minutes, and you can leave the configuration in place for later sessions.
Repeated, high-fidelity recording
If you record regularly, invest in a routing application that lets you build a matrix of sources, apply gain, and save presets. This pays for itself the first time you need to blend a microphone with internal audio and keep the levels repeatable across sessions.
Multi-source sessions
For interviews, tutorials with live narration, or production work, plan for an aggregate device that combines a microphone, the virtual device, and any hardware interface into a single input. Recorders then see one device with multiple channels, which keeps everything aligned in one file.
Decision criteria, in order of importance: how often you record, whether you need to monitor while recording, how many sources you must blend, and whether you need per-source level control after the fact. If you answer often, yes, more than two, and yes, build the full routing setup. Otherwise keep it simple and add complexity only when a specific project demands it. Beginning with an elaborate matrix you do not understand is the fastest way to spend an evening debugging silence.
Setting Up a Virtual Audio Device: The Standard Walkthrough
This is the core procedure most people end up using. The example uses a free virtual device, but the steps are nearly identical for commercial alternatives.
Step 1: Install the virtual device
Download the installer, run it, and restart if prompted. During setup, macOS will ask you to allow a system extension. Approve it in Privacy and Security, because without that approval the device will appear in the list but never pass audio. After installation, open Audio MIDI Setup from Applications, then Utilities, and confirm the device shows up in the sidebar.
Step 2: Create a Multi-Output Device
In Audio MIDI Setup, click the plus button at the bottom left and choose Create Multi-Output Device. In the right panel, check two boxes: your real output (speakers or headphones) and the virtual device. Order matters for latency, not for correctness, but putting your physical output first usually keeps monitoring tight.
Then right-click the new multi-output device and choose to use it for sound output, or select it in System Settings, then Sound, then Output. Now anything playing on your Mac reaches both your ears and the recorder.
Step 3: Set sample rate and channels
Open the virtual device properties and match its sample rate to the rest of your chain. Common choices are 44.1 kHz for music-adjacent work and 48 kHz for video. Mixing rates forces the system to resample, which can introduce artifacts and drift over long sessions. Set the channel count to two unless you specifically need multichannel capture. If your recorder offers mono, do not downsample at capture time; keep stereo and fold down later if you need to.
Step 4: Select the device in your recorder
In QuickTime Player, choose File, then New Audio Recording, then click the arrow next to the record button and pick the virtual device as the input. In other recorders, the setting might be called input source, audio device, or capture device. Some apps default to the system microphone, so verify the selection every single time.
Step 5: Test before you commit
Play ten seconds of audio, record it, and listen back with headphones. Confirm three things: you heard the audio live, the recording contains it, and there is no echo. If you hear an echo, you are likely capturing both the microphone and the internal audio at once, or your monitoring is routing back into the recording path.
Before any real session, run this short checklist:
- Virtual device installed and approved
- Multi-output device set as the system output
- Sample rate consistent across devices
- Recorder input pointing at the virtual device
- Headphones connected to avoid feedback
- Ten-second test clip recorded and reviewed
- Destination folder and file naming decided in advance
Recording the Audio: Recorder Options Compared
The routing is done; now choose the tool that captures it.
QuickTime Player
Best for straightforward audio-only or screen-with-audio capture. It is already installed, it handles long recordings, and it exports clean files. Its weak points are limited level metering and no built-in noise processing. Record to a lossless or high-bitrate format and handle processing afterward.
The built-in screen recording tool
The screenshot toolbar (Command-Shift-5) can record a selected area, a window, or the full screen, and it can include system audio in recent macOS versions. It is convenient for quick demos but offers almost no control over gain, channel layout, or codec. For anything you plan to edit, prefer a dedicated recorder.
Dedicated audio editors
Applications built for audio give you metering, punch-in recording, takes, and non-destructive editing. If you are capturing narration plus internal audio, an editor that supports multiple tracks is worth the learning curve.
Streaming and capture suites
If your workflow already includes scene composition, overlays, and multiple sources, a capture suite can handle system audio as one of several inputs. Set the virtual device as an audio input source, add a separate microphone input, and monitor through headphones only. Avoid routing the monitor output back into the virtual device, which creates a feedback loop.
Whichever tool you pick, record at a healthy level. Aim for peaks around -12 dBFS and an average near -18 dBFS. Digital clipping is unrecoverable; a slightly quiet file is easy to fix later.
Mixing and Monitoring Without Feedback
Monitoring is where new users get into trouble. The ideal setup: your headphones receive the multi-output signal, the recorder receives the virtual device, and the microphone is never monitored through speakers.
If you must monitor through speakers, mute the microphone input in the recorder while it is armed, or use a routing app to build separate monitor and record paths. A routing app lets you send source A to your headphones, source B to the recorder, and both to a mix bus without any chance of the loop closing.
Level balance matters too. Internal audio from apps is often mastered loud and compressed, while a live microphone is dynamic and quiet by comparison. If you mix them at the same gain, the internal audio will dominate and your voice will sit under it. Bring the internal audio down by 6 to 10 dB and raise it only when the voice is not present. Ducking, where music and game audio drop automatically under narration, is a fast way to keep a mix intelligible without riding faders manually.
Post-Production: Levels, Cleanup, and Sync
Loudness targeting
For web video, a common target is around -14 LUFS integrated with true peaks below -1 dBTP. Podcasts often sit closer to -16 LUFS. Measure the finished mix rather than guessing from meters, and apply a limiter only after you have controlled dynamics.
Removing hum and clicks
A high-pass filter around 80 Hz clears rumble and handling noise without thinning a voice. Use a short de-click pass for keyboard and mouse noise, and apply gentle noise reduction only where the noise floor is audible. Over-processing produces watery artifacts that are harder to fix than the original noise.
Fixing drift
If audio and video slowly separate over a long recording, the cause is usually a sample rate mismatch between capture and project settings. Conforming the audio to the project rate and re-syncing at a known marker fixes it. Prevent it next time by keeping every device at the same rate.
Troubleshooting: The Five Problems That Account for Most Failures
The recording is silent. The recorder is almost certainly listening to the wrong input. Check the input selection first, then confirm the virtual device is approved in Privacy and Security.
You hear nothing while recording. Your output is the virtual device alone. Create a multi-output device so sound reaches your headphones too.
Audio is doubled or echoed. Two capture paths are active, for example a microphone input plus internal audio, or the monitor signal is feeding back. Disable the extra input.
The virtual device disappears after sleep. Some drivers need to be re-initialized. Unplug and replug headphones, toggle the output device, or switch devices once to restart the audio service.
Levels are too hot or too quiet. Adjust at the source, not with a limiter. Lower the app volume, then raise recorder gain if needed.
Applying This to Voice-Over and AI-Assisted Video Projects
Modern video work increasingly combines generated footage, screen captures, and human narration. Internal audio capture shows up in all three.
For screen-recorded tutorials, you need the app's own sound plus your commentary, cleanly separated so you can re-record narration without redoing the screen capture. For generated video clips, you often need to lay a reference track, a scratch voice-over, or sound effects against visuals that have no native audio. For review sessions with clients or collaborators, capturing the call audio is sometimes the fastest way to produce an accurate transcript or a highlight reel.
A routing matrix earns its keep here because you can save a preset for each scenario: one for tutorial capture, one for voice-over with a music bed, one for call recording. Consistency between sessions is what makes editing predictable, and predictable audio is what keeps a project on schedule.
If your pipeline involves multiple generated clips, keep audio routing separate from the generation step. Record reference audio, align it once, and then treat the audio timeline as the spine of the edit. Visuals can be swapped; a misaligned audio bed forces you to rebuild timing decisions from scratch.
FAQ
Can I record system audio without installing anything? In recent macOS versions the built-in screen recording tool can capture internal audio in some configurations, but control is limited and reliability varies by app. A virtual device remains the dependable route.
Do I need a paid routing application? No. A free virtual device plus a multi-output device covers most single-source needs. Paid routing software becomes worthwhile when you need gain control, presets, and multi-source matrices.
Can I record a microphone and system audio at the same time? Yes. Combine both into an aggregate device so the recorder sees multiple channels, or add each as a separate input inside a capture suite. Keep them on separate tracks so you can fix one without touching the other.
Will this affect call quality in meetings? Only if you set the virtual device as your default input for a call app. Keep your microphone as the input for meetings unless you specifically want to send internal audio to participants.
Why is my recording mono? Either the source is mono or your recorder collapsed the channels. Check the device channel configuration and the recorder's track settings.
How long can I record? Hours, assuming disk space and stable devices. Long sessions are where sample rate mismatches reveal themselves, so monitor sync during the first few minutes.
Does this work with headphones that have their own device? Yes. Add both the headphone output and the virtual device to the multi-output device. If the headphones use Bluetooth, expect extra latency and possible quality reduction.
Can I capture only one app's audio? With a basic virtual device, no; you capture everything going to the output. Routing applications that support per-app sources can isolate a single application if that precision matters.
What format should I record in? Record lossless or high-bitrate WAV or AIFF for editing, and export compressed audio only at the end. Storage is cheap; re-recording a ruined take is not.
Why does the recording sound worse than what I heard? Usually a sample rate mismatch, aggressive noise reduction, or a codec with a low bitrate. Compare the raw capture to the exported file to isolate the stage that introduced the problem.
Wrapping Up
Recording Mac system audio is less about finding a hidden switch and more about building a short chain: a virtual device to receive the signal, a multi-output device so you can still hear it, and a recorder pointed at the right input. Once those three pieces are in place, the workflow is repeatable and boring in the best way.
Start with the simplest setup that solves your problem. Add routing software only when you need per-app control or repeatable presets. Test with a ten-second clip before every important session, keep your sample rates aligned, and monitor with headphones. Do that, and internal audio capture stops being a source of frustration and becomes just another step in your production routine.



