YouTube's Single Gain Coefficient: Evidence, Targets, Limits

```html

TakeawayDetail
YouTube's normalization is a ceiling, not a targetGain is applied downward only, pulling louder uploads toward roughly -14 LUFS integrated while never boosting quiet audio, so a music-heavy mix can certify at -14 LUFS while dialogue plays near -19 LUFS — a deficit that matters when conversational intelligibility shifts about 17.2% per dB of signal-to-noise.
Spontaneous speech sits close to a measurable comprehension cliffAcross 384 conversational sentences and 32 normally-hearing listeners tested at signal-to-noise ratios from -8 to 0 dB, the average 50% speech reception threshold was -5.2 dB — the point where half of spontaneous sentences stop being understood.
AI repair passes can restore buried consonants before uploadOne enhancement pipeline reported background noise reduced by -24 dB, clarity boosted +18%, and SNR improved +12 dB in 3.2 seconds of processing, with guidance to start at the Medium strength preset and reserve Light for already-clean recordings.
Speech-first masters belong below the ceiling, not on itDelivering dialogue-led content at -16 LUFS leaves headroom under the -14 normalization point, keeping voices away from the -5.2 dB SNR region where average listeners decode just 50% of conversational sentences.

A video can integrate at -14 LUFS and still put dialogue in viewers' ears at roughly -19 LUFS. YouTube's loudness normalization works in one direction only: louder files are pulled down toward approximately -14 LUFS integrated, and quieter ones are left untouched. Because the meter weighs the entire mix — music beds, effects, and voice together — it can certify an upload whose actual speech runs several decibels quieter than its headline number suggests.

The listening science explains why that gap hurts. In a benchmark of 384 spontaneously produced conversational sentences played to 32 young, normally-hearing listeners over stationary noise, the average 50% speech reception threshold landed at -5.2 dB signal-to-noise ratio, with intelligibility shifting about 17.2% per dB. Close to that threshold, a handful of decibels separates effortless comprehension from strained guesswork — exactly the territory where buried dialogue lives.

That asymmetry reframes delivery. Chasing -14 LUFS buys nothing, because the platform will never restore loudness it did not remove; the durable move is to master speech-first content to -16 LUFS, measure dialogue separately from integrated loudness, and prove the dialogue-to-integrated gap before upload. Where recordings need rescue, denoise and clarity processing — one reported pipeline lifted clarity +18% — should begin at a Medium preset rather than an aggressive one.

YouTube's Single Gain Coefficient

Inside YouTube's Normalizer

YouTube stores exactly one number about your audio: a single gain coefficient, computed once from the uploaded file and frozen for the life of the video. The measurement follows ITU-R BS.1770-4 to the letter — a K-weighted filter shapes the signal to approximate hearing, an absolute gate at -70 LUFS discards silence, and a relative gate excludes passages sitting well below the running level. Momentary loudness updates in short blocks, short-term loudness in 3 s windows, and the integrated figure that survives both gates sets that one stored gain.

Playback applies that gain in one direction only. Content above the ceiling — independently measured at roughly -13 to -14 LUFS — is attenuated down to meet it; content below passes through untouched. A -16 LUFS master therefore streams at exactly file loudness with 0.0 dB applied. That asymmetry is what kills the persistent advice to "master to -14 because it's YouTube's target": -14 is a ceiling the platform enforces by turning audio down, never a level it boosts toward, and the integrated figure behind it counts every music bed in the mix.

Dialogue loss makes the second half of that advice measurable. Dialogue-gated LUFS is the same BS.1770-4 measurement restricted to speech spans, and dialogue loss equals integrated LUFS minus that gated figure. Energy summation supplies the mechanism: an un-ducked bed 3 LU hotter than the voice carries twice the speech's power, lifting combined loudness by roughly 4.8 LU above dialogue alone. Followed through, a file whose integrated meter reads -14 can hold dialogue gated near -18.8 LUFS. The meter isn't lying — it answers a total-energy question nobody actually listens to.

Exported masterNormalizer verdictGain appliedWhat ships to viewers
-12 LUFS integratedAbove ceilingAttenuated ≈2 dB-14 LUFS; loudness matched, mix balance untouched
-14 LUFS integrated, bed 3 LU hotAt ceiling≈0 dB-14 LUFS with dialogue gated near -18.8
-16 LUFS integrated, dialogue within 2 LUBelow ceiling0.0 dBExactly -16 LUFS, file-faithful
-20 LUFS integratedBelow ceiling0.0 dBExactly -20 LUFS, untouched

Peaks fail separately from loudness. YouTube applies no true-peak limiter at playback, and its Opus and AAC transcodes can overshoot a hot master by up to about 1 dB — headroom you have to spend yourself at export. Pair the -16 target with a true-peak ceiling of -1 dBTP or lower. Verification tooling confirms the split: according to Soundandgo's Loudness Analyzer & LUFS Normalizer, the workflow exposes separate "Normalize Target LUFS" and "Ceiling (dBTP)" fields and returns a post-run True Peak flag, because level compliance and peak compliance are independent failure modes.

The 2026 delivery surface raises the stakes. The identical file plays through phone speakers, laptop DACs, TV soundbars, and Bluetooth chains, and because nothing below the ceiling is ever altered, no device receives a corrected signal — every device-to-device difference traces back to the mix, not the platform. A mix carrying 4.8 LU of dialogue loss plays badly everywhere; a clean -16 master degrades gracefully everywhere.

That leaves exactly one lever under your control: the exported file. For anything at or under the ceiling, playback loudness equals file loudness, which makes the export measurement — not any platform-side number — the real specification. Before upload, verify -16 LUFS integrated, dialogue-gated loudness within 2 LU of it, and true peak at or below -1 dBTP; if the dialogue gap runs wider, rebalance the bed instead of adding limiting to chase -14.

Inside YouTube's Normalizer — YouTube's Single Gain Coefficient

The Receipts

Ian Shepherd's published tests on Production Advice are the origin point for everything else on this list. He uploaded deliberately loud masters and measured what came back: YouTube pulled them down to land near -13 to -14 LUFS integrated, while quieter files passed through untouched. That asymmetry — attenuation without boost — is the entire attenuation-only ceiling model, demonstrated publicly before most creators had even heard the term "loudness normalization."

The replication wave ran from 2021 through 2024. According to measurement write-ups from Mastering The Mix and the team behind the Youlean Loudness Meter, independent uploads kept reproducing the same roughly -14 LUFS ceiling, with one detail that matters more than the headline number: no playback peak limiting. Set that against the platforms those write-ups measured alongside it. Spotify normalizes to -14 LUFS and applies -1 dBTP limiting; Apple Music sits at -16 LUFS with the same -1 dBTP limiter. A limiter is active processing — it reshapes transients. YouTube's normalizer turns audio down and does nothing else, so a file under the ceiling plays back at its file loudness, full stop.

Then the standards bodies. AES TD1004.1.15-10, the AES Streaming Loudness Recommendation, specifies -16 LUFS plus or minus 1 dB for stereo podcasts and speech-first streaming. Read that as a deliberate rejection of -14 for voice-led programs: the committee had access to the same meters everyone else does and chose lower anyway. Against this record, the persistent advice to "master to -14 because that's YouTube's target" fails twice — -14 is a ceiling enforced only downward, and the integrated figure behind it counts every music bed in the mix, which is precisely how the dialogue shortfall quantified earlier hides inside a compliant-looking number.

Netflix goes further than any platform. Its Sound Mix Specifications require dialogue-gated loudness to measure -27 LUFS plus or minus 2 LU, with true peak at or below -2 dBTP. The largest single buyer of audio mixes does not police integrated loudness alone — it gates the dialogue and measures that, because dialogue is what listeners actually reach for the volume control over. The absolute value is irrelevant to your upload; the gated-versus-integrated mechanism is the transferable part.

Underneath all of it sits EBU R128 at -23 LUFS plus or minus 0.5 LU, with its companion technical documents defining the momentary, short-term, and integrated measurement modes and loudness range. Every streaming target descends from this broadcast baseline, which is quietly practical news: any meter implementing the EBU stack reads every number in the table below natively.

The convergence is the argument. Platform ceilings (-14 YouTube, -16 Apple), standards targets (-16 AES, -23 EBU), and dialogue mandates (Netflix) all point one direction — measure the speech, mind the gap between gated and integrated, and stop optimizing for the loudest permitted number. Of these receipts, the AES recommendation is the one that directly answers the upload question, and Netflix supplies the verification method for it.

SourceSpecificationFiguresWhat it settles
Ian Shepherd, Production AdviceUpload-and-measure playback testsLoud masters land near -13 to -14 LUFS; quiet files untouchedCeiling acts downward only
Mastering The Mix & Youlean write-ups, 2021–2024Independent replicationsRoughly -14 LUFS ceiling; no playback peak limitingModel holds across many uploads
SpotifyNormalize plus limiter-14 LUFS with -1 dBTP limitingSome platforms reshape peaks; YouTube does not
Apple MusicNormalize plus limiter-16 LUFS with -1 dBTP limitingA major streamer already targets -16
AES TD1004.1.15-10Streaming loudness recommendation-16 LUFS ±1 dB, speech-first stereoStandards chose -16 over -14 for voice
Netflix Sound Mix SpecificationsDelivery mandateDialogue-gated -27 LUFS ±2 LU; true peak ≤ -2 dBTPThe gated-to-integrated gap gets policed
EBU R128 + companion tech docsBroadcast baseline-23 LUFS ±0.5 LU; momentary, short-term, integrated, LRA definedOne meter stack reads every target above
The Receipts — YouTube's Single Gain Coefficient

Target Selection

Target selection is not picking a number — it is auditing whether the mix has earned one. Three strategies compete for a 2026 speech-first upload, and only one survives the attenuation-only behavior covered earlier in this guide: the platform turns audio down above its ceiling and never boosts anything below it. The persistent advice to master to -14 "because that is YouTube's target" fails twice here — the ceiling only shaves, and the integrated figure happily counts a music bed while the dialogue beneath it sits far quieter than the number implies. Score the field:

StrategyIntegrated targetExpected dialogue-gated resultYouTube playback gainLRA retainedTrue-peak safety
A: Chase -14 via limiting-14 LUFS, forcedFar below integrated — the bed owns the number (the gap quantified earlier)Attenuated; coefficient goes negativeTypically 4-5 LUMarginal; dense limiting crowds the -1 dBTP line
B: Hybrid — -14 integrated, -16 dialogue floor-14 LUFS via ducking and automationPinned near -16 under speechNear zero; parked at the ceilingMiddle; macro moves survive, beds pumpModerate; limiting still required to force -14
C: -16 integrated, dialogue loss ≤2 LU — winner-16 LUFS, naturalWithin 2 LU of integratedZero; file loudness passes through untouched8 LU or moreComfortable; roughly 2-3 dB of spare gain-reduction budget keeps peaks under -1 dBTP

Strategy C wins outright for speech-first content, and the two decision variables below tell you whether your mix qualifies.

Decision variable one: measure dialogue loss before any mastering move. Take the integrated reading, then run a gated, speech-weighted pass over the same file, and subtract. A gap of 2 LU or less means ship at -16 as-is. More than 2 LU means rebalance — duck the beds, trim music-only intros that inflate the integrated read — because no limiter can restore speech prominence the mix removed. A limiter applies one scalar gain curve to the entire program: it drags bed and voice down together and cannot re-rank them. If your meter reads -15.9 integrated against -18.6 gated, that 2.7 LU gap is a fader problem, not a ceiling-clipper problem. The perceptual stakes are not abstract: published listening work (PMC7060086) sat 32 young normal-hearing listeners against stationary noise at five signal-to-noise ratios from -8 to 0 dB — precisely the regime where a hot bed buries speech long before any integrated meter objects.

Decision variable two: headroom accounting. Abandoning the -14 chase for a -16 target hands back roughly 2-3 dB of limiter gain-reduction budget, and gain reduction is where loudness range goes to die. When a limiter works near-constantly, macro-dynamics flatten — which is why crushed -14 masters typically retain only 4-5 LU of LRA while the same program at -16 holds 8 LU or more. The same relief buys true-peak margin: thinner density makes the -1 dBTP ceiling cheap to honor instead of a permanent fight.

Boundary condition: music-led uploads. A performance capture or music video has no dialogue to lose — the integrated figure is the product. Those legitimately sit at -14.5 to -15 LUFS, just under the ceiling: close enough that the stored coefficient stays near zero, far enough to avoid being shaved. The -16 rule is scoped to speech-first content and nothing else.

The mono case. Summing correlated stereo program raises energy by roughly 3 LU, so a compliant -16 stereo master can fold down near -13 — over the line — with the dialogue gate shifting by a similar amount. Phone speakers and single-driver soundbars live in this regime. Re-run the integrated-plus-dialogue-gate check on the summed signal instead of assuming the stereo numbers transfer unchanged.

Close every upload with a written delivery note recording target integrated, measured dialogue-gated, true peak, and LRA, so the next episode is compared against numbers rather than memory:

FieldConventionExample entry
Target integratedLocked at ship-16.0 LUFS
Measured dialogue-gatedGated, speech-weighted pass-17.2 LUFS
Dialogue lossIntegrated (-16.1) minus gated1.1 LU — pass; budget is 2 LU
True peakOversampled reading-1.3 dBTP
LRAFull program9.4 LU
VerdictShip or rebalanceShip

Fill the right-hand column before export, not after publish — the log becomes the show's institutional memory, and a drifting dialogue gap shows up in the numbers episodes before it shows up in listener complaints.

Target Selection — YouTube's Single Gain Coefficient

What the Data Doesn't Tell You

The evidence behind the rule above is narrower than the confidence it projects, and knowing precisely where it runs out is what separates a defensible master from a lucky one. Three gaps matter: the measurements stop at the file, individual programs scatter widely around the pattern, and there are genuine cases where the discipline stops paying off.

Start with the deepest limitation: every platform-side measurement terminates at the file, while the thing creators actually care about terminates at a listener. The published attenuation tests verify what a frozen gain coefficient does to an uploaded waveform — nothing more. They cannot tell you whether anyone understood the dialogue. Intelligibility is a property of audiences, and audiences are not uniform. According to the intelligibility-by-age model study, the analysis determined quantile-specific ages of steepest growth, and growth rates were also quantified per quantile — in plain terms, the fastest changes in speech-intelligibility trajectories arrive at different points for different segments of the listening population. A single integrated figure is an average stretched over listeners whose needs diverge sharply. The dialogue-gated verification step is the best available proxy for protecting them; it is not a guarantee that every ear hears every word.

Variance across cases is the second gap. Two uploads mastered identically can behave differently after the platform's transcoder, because lossy encoding overshoot depends on spectral content — dense, bright material stresses the ceiling more readily than dry speech. The verification step exists precisely because the gated-versus-integrated relationship is stable for some programs and unstable for others:

CaseSource of varianceVerification consequenceVerdict
Dry solo narrationGate and integrator track the same signalOne pre-upload check sufficesRule holds cleanly
Interview plus sparse bedBed fader rides shift the floor between gatesRe-verify after every mix revisionHolds with re-checking
Score-heavy documentaryContinuous music inflates the integrated readGated gap often breaches the 2 LU budgetRebalance the bed; never limit toward -14
Multi-mic live panelMic distance and gain drift per speakerCheck the quietest speaker, not the averageConditional until fixed
Close-mic whispered speechUnusual crest behavior stresses gating assumptionsTreat the check as advisory; audition on a phoneEdge case — proceed cautiously

Then there are the cases where the rule itself breaks. It is scoped to speech-first uploads; a music performance video sits outside its jurisdiction entirely, and forcing dialogue-gated logic onto one is a category error. It also assumes the playback chain ends at the platform's coefficient — it does not. Television apps and Bluetooth receivers apply their own downstream volume processing after the frozen gain, so the rule predicts file-relative loudness, not the level reaching a specific living room. And for extremely dynamic programs where closing the gated gap would demand audible pumping, the honest reading is that the mix has not yet earned its target: the remedy is rebalancing, and if that is truly impossible, shipping with the shortfall disclosed beats disguising it beneath a louder integrated number — the exact failure mode the receipts above document.

None of these caveats rehabilitate the persistent "-14 because that is the platform's target" advice; every edge case here either leaves the rule untouched or strengthens the case for verifying dialogue before trusting any integrated figure. Treat the framework as time-stamped: this guide is a current-year reference document, and platform pipelines change without announcement. Before each major upload cycle, run one controlled probe — a short reference file mastered to -16 with true peak at or below -1 dBTP — measure what returns, and confirm the coefficient still behaves as described before betting a season of content on it.

What the Data Doesn't Tell You — YouTube's Single Gain Coefficient

What the Numbers Can't Promise

No page on YouTube's own servers states a normalization target. A review of the retrieved source set turned up no YouTube-official specification confirming the platform's loudness ceiling — every figure in this guide, the attenuation threshold included, is reverse-engineered from uploaded files and measured playbacks. Community measurements have landed anywhere from -13 to -15 LUFS across years and content types, so the ceiling behaves like an estimate with a built-in error bar, not a published spec. Anyone quoting it to a single decimal place is selling false precision.

Attenuation-only behavior is an observed pattern, not a published contract. Isolated uploader reports of upward normalization on very quiet uploads exist, and the platform has never committed in writing to either direction. The engineering response is symmetry: build a master that survives both. A file with true peaks pressed against the -1 dBTP guard clips instantly if the normalizer ever pushes instead of pulling; keeping genuine headroom beneath that guard means a hypothetical upward ride degrades gracefully. The robust master is the one whose dialogue survives a cut and a boost.

The dialogue-gated measurement itself wobbles. Gated readings shift by 1-2 LU depending on the speech detector — Whisper word timestamps, an energy-based VAD, and hand-marked spans each draw the speech boundary differently — so the 2 LU dialogue-loss budget effectively carries plus or minus 1 LU of instrument uncertainty. That wobble is structural, not sloppiness: according to the Everyday Conversational Sentences in Noise study, even a laboratory pipeline with an explicit intelligibility-normalization step left an average RMS error of 4.7% between measured scores and the fitted psychometric function. The working protocol is dual-detector verification — rebalance only when two independent detectors agree the gap breaches budget.

There is also a format where this whole argument mostly dissolves. Documentary-style mixes with continuously ducked scores show dialogue gaps of only 1-2 LU; there, -14 and -16 masters converge at playback and chasing the ceiling costs almost nothing. The rule earns its keep on bed-heavy, intro-heavy, ad-read-heavy formats:

Upload formatDialogue-to-integrated relationshipVerdict at -16
Bed-free talking headDialogue is the entire program; gated and integrated readings sit togetherTarget choice barely matters; verification is cheap insurance
Documentary, continuously ducked scoreGap holds at 1-2 LU-14 and -16 converge; ceiling chase costs almost nothing
Produced intro/outro stingsSting bursts inflate integrated above dialogueRebalance stings first, target second
Ad-read-heavy commentaryReads run hotter than surrounding talkRule binds hardest; fix segment levels, not limiting
Music-supervised vlogContinuous beds compete with voiceHighest breach risk against the 2 LU budget

Even a perfect number promises less than it appears to. Phone-speaker versus soundbar playback can swing perceived speech intelligibility by more than the 2 LU at stake, and no integrated or gated figure captures spectral losses on 2 cm drivers — loudness normalization fixes level, not timbre. The perceptual literature frames both halves: according to the speech-reception study archived in PubMed Central, the average psychometric slope near the 50%-correct reception threshold runs 17.2% per dB, so level errors are expensive; the same body of work shows everyday speech shifting between normal, raised, and loud vocal effort with the acoustic environment, variability that fixed-level playback tests fail to reproduce. Equal loudness is a level guarantee, never an intelligibility guarantee.

The feedback loop is blind, too. YouTube Studio exposes retention and average view duration but not volume-ride events, so a viewer who rode the volume up under your music bed and back down for your ad read is statistically invisible. Improvements from fixing dialogue loss are therefore inferable only from proxies, and retention sampled at music-to-speech transitions is the cleanest one. Mark those timestamps before you rebalance, publish, and compare the retention curve at identical positions — if the dip shrinks, the budget fix worked. Anything you cannot tie to such a proxy is faith, not measurement.

What the Numbers Can't Promise — YouTube's Single Gain Coefficient

Worked Case

A 22-minute interview integrating at -12.6 LUFS reads as "1.4 dB too hot" on any dashboard, and the reflexive response is a limiter. Gate the same file for speech and the story inv

```

Frequently Asked Questions

If my upload is quieter than -14 LUFS, will YouTube boost it up to match?

No — gain is applied downward only, so content below the ceiling passes through untouched and a -16 LUFS master streams at exactly its file loudness with 0.0 dB applied.

Does YouTube ever recalculate the loudness adjustment on my video after it's published?

No — YouTube stores exactly one gain coefficient, computed once from the uploaded file and frozen for the life of the video.

Do I still need to worry about true peaks if my file sits under the loudness ceiling?

Yes — YouTube applies no true-peak limiter at playback and its Opus and AAC transcodes can overshoot a hot master by up to about 1 dB, so pair the -16 target with a true-peak ceiling of -1 dBTP or lower.

My mix integrates at -14 LUFS, so why does the dialogue sound buried?

Because an un-ducked music bed 3 LU hotter than the voice carries twice the speech's power and lifts combined loudness by roughly 4.8 LU above dialogue alone, a file reading -14 LUFS integrated can hold dialogue gated near -18.8 LUFS.

At what signal-to-noise ratio do average listeners actually stop understanding conversational speech?

Across 384 spontaneously produced conversational sentences played to 32 normally-hearing listeners over stationary noise, the average 50% speech reception threshold was -5.2 dB SNR, with intelligibility shifting about 17.2% per dB.

How aggressive should I set denoise and clarity processing on a noisy recording before uploading?

One reported pipeline reduced background noise by -24 dB, boosted clarity +18%, and improved SNR +12 dB in 3.2 seconds of processing, and its guidance is to start at the Medium strength preset and reserve Light for already-clean recordings.

Quick answers

How does YouTube's single stored gain coefficient treat content above versus below its loudness ceiling?Content above the ceiling, independently measured at roughly -13 to -14 LUFS, is attenuated down to meet it, while content below passes through untouched — so a -16 LUFS master streams at exactly file loudness with 0.0 dB applied.
What did the conversational speech benchmark find about where comprehension breaks down?Across 384 spontaneously produced conversational sentences played to 32 young, normally-hearing listeners over stationary noise, the average 50% speech reception threshold landed at -5.2 dB signal-to-noise ratio, with intelligibility shifting about 17.2% per dB.
What results did the reported AI repair pipeline achieve before upload?One enhancement pipeline reported background noise reduced by -24 dB, clarity boosted +18%, and SNR improved +12 dB in 3.2 seconds of processing, with guidance to start at the Medium strength preset and reserve Light for already-clean recordings.
How large can the gap between integrated loudness and actual dialogue level become?An un-ducked bed 3 LU hotter than the voice carries twice the speech's power, lifting combined loudness by roughly 4.8 LU above dialogue alone, so a file whose integrated meter reads -14 LUFS can hold dialogue gated near -18.8 LUFS.
What export measurements should creators verify before upload?Verify -16 LUFS integrated, dialogue-gated loudness within 2 LU of it, and true peak at or below -1 dBTP, because YouTube applies no true-peak limiter at playback and its Opus and AAC transcodes can overshoot a hot master by up to about 1 dB.

Also worth reading: 2026 A/B Test: -14 LUFS Boosts YouTube Watch Time by 12%: 2026 A/B Test: -14 LUFS · Reels Loudness: Why -14 LUFS Is a Gate, Not a Creative Choice: Reels Loudness: Why -14 LUFS · Gated LUFS Showdown: 4 Podcast Masters, 2 Normalizers, 1 File: Gated LUFS Showdown: 4 Podcast

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Audobox editorial desk (About, Contact, Privacy).

Related answers