Commit a86dc8bd2a for ffmpeg

commit a86dc8bd2a25ad00e2a13fbcbca9dbbae72e2c29
Author: Lynne <dev@lynne.ee>
Date:   Wed Sep 23 20:32:13 2026 +0900

    avcodec/aaccoder_nmr, aacenc, aacpsy: nd-target VBR

    -q:a with the NMR coder becomes a quality-domain VBR: bisect the pooled
    lambda until the mean log2 dist/real-mask over coded bands (floored at
    -4/band, so a flat sf-chain over-coding quiet bands cannot let a coarser
    lambda read as a finer statistic) meets a target. The target is anchored
    so -q:a 1 lands near 128 kbps stereo; higher q is finer, as with the
    other coders, at -1.2 per q doubling (about a third more rate), clamped
    at a reachable floor that q above about 6 reaches: the stat clips at
    -4/band, targets below the clamp are asymptotically unreachable, the
    bisect rails at the finest lambda and the frame exceeds the decoder
    buffer. The anchor is calibrated with the broadband quality-mode mask
    floor from earlier in the series; that floor measured -8.8% / -5.0% /
    -0.9% Zimtohrli (ViSQOL +0.013 / +0.018 / +0.004) at matched rate on
    three corpora, worst sample -34%, per-sample quality spread 2.41 ->
    1.74.

    The psy model runs its PE reduction at a fixed reference quality for
    this coder (psy.unbounded_pe): thresholds keep their perceptual shaping
    but stay independent of -q:a, which otherwise moves thr_real under the
    coder's own statistic and inverts the quality ladder.

    Frame classes the statistic mishandles:
    - Short frames: the grouped-band statistic is inflated, so they bisect
      against an offset target and never write solver state. A fixed +3
      offset was a content-class policy in disguise - broadband beats
      (velvet) need it, LF/mid tonal decays audibly do not and paid 1.7-2.8x
      the CBR error through every post-hit decay. The offset is measured
      per frame instead: evaluate the grouped stat at the long-anchor lambda
      (continuity says true quality is on target there, the excess is this
      content's inflation), weighted by the frame's HF energy share - the
      stat reads inflated on both classes, audibility does not. Post-hit
      decays 1.67x -> 1.13x CBR error (worst 2.82x -> 1.44x).
    - Continuous shorts: the offset-bisect assumes isolated transients
      between long anchors; on dense onsets the per-frame stat is bisect
      noise and lambda careened across decades (0.16 -> 82 -> 2 -> 43), an
      audible whoosh. Shorts hold to the CBR run slew, clamped after the
      last solve (PASS 2 re-bisects on a fresh +-32x bracket and otherwise
      rails at exactly lam*32); dives toward finer stay free.
    - Operating points: only bisected long frames whose stat actually
      landed within 0.5 of the target write slew/warm-start state. Cheap
      frames (room tone, gaps) rail fine at no cost; the first frames of a
      quiet channel rail coarse from no warm state - writing either builds
      a slew prison the channel climbs out of one rung per frame (a muffled
      stream head on fade-ins).
    - Decoupled solo solves keep per-channel slew state: a shared lam_slew
      clamped each channel against the OTHER channel's operating lambda and
      landed the damage asymmetrically.

    The window history the slew rules read (previous-frame short, frames
    since the last short run) used to advance only inside the CBR reservoir
    bookkeeping; it now advances once per frame in every rate-control mode,
    so the run slew and the pooled short-TNS hysteresis see real history
    here too. Measured neutral (IgorC-1 VBR q=1: -0.3% Zimtohrli).

    Legality is absolute: a frame may never exceed the decoder buffer.
    - The cap is the format limit regardless of -q:a. It used to scale with
      the outer lambda, i.e. with -q:a itself, so the finest settings had
      the smallest cap (531 bits/channel at q=0.1 under the old direction)
      and lost bands where they wanted the most bits.
    - Dense short frames at high rates can exceed 6144 bits/channel even at
      the coarsest candidate grid (the 1e4 rail still emitted 11-13k
      spectral bits); the legality path sheds the highest bands across all
      window groups (act[] runs group by group) until the frame fits,
      keeping the scalefactor-chain anchor.
    - The bits a solve reports for the next frame's side-information
      estimate count only the surviving bands, and PNS filters the active
      list instead of rebuilding it from all bands. Counting shed bands
      drove the estimate negative (one frame emitting 1256 bits counted
      13825), and a negative estimate enlarged the next frame's cap: 13056
      payload bits against a 12288 limit. A negative estimate can no longer
      enlarge the cap.
    - The outer loop no longer skips the size check under a constant
      quality: an oversized frame is retried at a lower lambda, which
      shrinks the coders' caps. This applies to every coder, and only to
      frames that would otherwise be illegal.
    A fine-quality encode test (fate-aac-nmr-vbr-encode) pins the per-packet
    sizes; its largest frame is 11560 payload bits.

    Encoder side: the bandwidth ladder takes the expected per-channel rate
    from the quality ladder (q=1 ~ 64.5 kbps/ch, x1.27 per doubling), and
    Qavg reports the coder's quality-mode slew state: the lambda of the last
    long operating-point solve. The encoder documentation describes the
    NMR meaning of -q:a.

    CBR output is unchanged.

diff --git a/doc/encoders.texi b/doc/encoders.texi
index b36d7dee21..98c5f22da5 100644
--- a/doc/encoders.texi
+++ b/doc/encoders.texi
@@ -44,6 +44,13 @@ Set quality for variable bit rate (VBR) mode. This option is valid only using
 the @command{ffmpeg} command-line tool. For library interface users, use
 @option{global_quality}.

+With the default @samp{nmr} coder, VBR holds a constant noise-to-mask target
+and lets the bitrate follow the content. Higher values are finer: @code{1}
+lands near 128 kbps for typical stereo music, and each doubling of the value
+raises the rate by roughly a third; values above about 6 make no further
+difference. A quality setting takes precedence over @option{b} and
+@option{aac_rc}.
+
 @item cutoff
 Set cutoff frequency. If unspecified will allow the encoder to dynamically
 adjust the cutoff to improve clarity on low bitrates.
diff --git a/libavcodec/aaccoder_nmr.h b/libavcodec/aaccoder_nmr.h
index 255675a4f3..398cc91059 100644
--- a/libavcodec/aaccoder_nmr.h
+++ b/libavcodec/aaccoder_nmr.h
@@ -89,6 +89,20 @@
  * escalation step; the residual overage rides as reservoir debt. */
 #define NMR_RC_CAPK   3.0f

+/* Quality-target calibration anchor: the nd set-point of -q:a 1, which lands
+ * near 128-136 kbps stereo on the tuning corpora with the q ladder's own
+ * bandwidth. It moved -2.1 -> -0.75 with the quality-target mask floor
+ * (PSY_THRFL_QUALITY), which lowers thr_real and so lifts the whole statistic.
+ * The reachable-range clamp is set by the stat's -4/band sub-mask clip, which
+ * the mask floor does not touch: targets below it are asymptotically
+ * unreachable - the bisect rails at the finest lambda, the frame exceeds the
+ * decoder buffer and the outer re-encode loop never converges. */
+#define NMR_VBR_ANCHOR (-0.75f)
+#define NMR_VBR_TMIN   (-3.8f)
+
+/* VBR shorts: the measured stat inflation counts fully once this share of
+ * the frame's energy sits above 6 kHz */
+#define NMR_INFL_HF 0.10f
 /* Reservoir half-window (bits/ch); swept 512/1536/3072, 1536 optimal. */
 #define NMR_CBR_BUF   1536
 /* Slew limit on the FINAL operating lambda per frame; bits deviate instead,
@@ -556,6 +570,55 @@ static float nmr_solve_slots(AACEncContext *s, NMRSlot *const *sl, int nsl, int
     return lam;
 }

+/* Mean log2 achieved dist/real-mask over the active bands of all slots at
+ * their current chosen[] - average dB-distance from the mask (Brandenburg
+ * mean-NMR). Log compression keeps unavoidably-loud bands from swamping the
+ * statistic. NAN when nothing is coded. */
+static float nmr_nd_stat(AACEncContext *s, NMRSlot *const *sl, int nsl)
+{
+    float ndsum = 0.0f;
+    int n = 0;
+    for (int k = 0; k < nsl; k++) {
+        NMRSlot *t = sl[k];
+        const float (*ndk)[NMR_NCAND] = (const float (*)[NMR_NCAND])s->nmr->nd[t->si];
+        for (int b_ = 0; b_ < t->nact; b_++) {
+            int b = t->act[b_], bi = t->bidx[b];
+            if (t->thr_real[bi] > 0.0f && t->thr[bi] > 0.0f) {
+                /* floor the sub-mask credit: a flat sf-chain over-codes
+                 * quiet bands and their unbounded negative log2 would let a
+                 * COARSER lambda read as a finer statistic */
+                ndsum += FFMAX(log2f(FFMAX(ndk[b][t->chosen[b]] * t->thr[bi] /
+                                           t->thr_real[bi], 1e-6f)), -4.0f);
+                n++;
+            }
+        }
+    }
+    return n ? ndsum / n : NAN;
+}
+
+/* Bisect ONE shared lambda across the slots so the achieved noise-to-mask
+ * statistic meets the target: quality-domain VBR. Monotone: coarser lambda ->
+ * more distortion. */
+static float nmr_solve_slots_nd(AACEncContext *s, NMRSlot *const *sl, int nsl,
+                                int step, float target, float lo_l, float hi_l,
+                                int iters)
+{
+    float lam = 1.0f;
+    for (int it = 0; it < iters; it++) {
+        float st;
+        lam = sqrtf(lo_l * hi_l);
+        nmr_eval_slots(s, sl, nsl, step, lam);
+        st = nmr_nd_stat(s, sl, nsl);
+        if (isnan(st) || it == iters - 1)
+            break;
+        if (st > target)
+            hi_l = lam;
+        else
+            lo_l = lam;
+    }
+    return lam;
+}
+
 /* Write a solved slot back into its channel: band types, scalefactors, and the
  * SCALE_MAX_DIFF legality fixups. Verbatim from the pre-pool single-channel tail. */
 static void nmr_commit_channel(AACEncContext *s, NMRSlot *t)
@@ -575,19 +638,10 @@ static void nmr_commit_channel(AACEncContext *s, NMRSlot *t)
     }


-    {   /* record the bits this solve accounted for; the encoder compares them
-         * against the channel's real output to keep the budget honest */
-        int tot = 0, prevb = -1;
-        for (int b = 0; b < t->nbnd; b++) {
-            if (t->is_pns[b])
-                continue;
-            tot += nb[b][t->chosen[b]];
-            if (prevb >= 0)
-                tot += NMR_SFBITS((t->blo[b]+t->chosen[b]*NMR_STEP) - (t->blo[prevb]+t->chosen[prevb]*NMR_STEP));
-            prevb = b;
-        }
-        s->nmr->counted[t->cur_ch] = tot;
-    }
+    /* record the bits this solve accounted for; the encoder compares them
+     * against the channel's real output to keep the budget honest. Only the
+     * surviving active bands are coded: bands shed for legality are not. */
+    s->nmr->counted[t->cur_ch] = nmr_slot_bits(t, nb, NMR_STEP);

     /* SCALE_MAX_DIFF condition:
      * re-clamp, codebook fixup, drop uncodeable, set global gain
@@ -657,10 +711,23 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,
     int is8_any = 0;
     float lam;
     float rc_off = 1.0f, lam_dem = 0.0f;
+    /* the outer loop's starting lambda: legality caps scale with lambda
+     * relative to it, so they only shrink on a hard-overflow retry */
+    const float lam_ref = (avctx->flags & AV_CODEC_FLAG_QSCALE) && avctx->global_quality > 0 ?
+                          avctx->global_quality : 120.0f;

     for (int k = 0; k < nsl; k++)
         is8_any |= sl[k]->is8;

+    /* decoupled solo solves keep per-channel slew state: sharing one
+     * lam_slew clamps each channel against the OTHER's operating lambda */
+    float *vslew_st = &s->nmr->lam_slew;
+    if (nsl == 1 && s->channels > 1) {
+        vslew_st = &s->nmr->lam_slew_ch[sl[0]->cur_ch & 15];
+        if (*vslew_st <= 0.0f)
+            *vslew_st = s->nmr->lam_slew;
+    }
+
     if (s->psy.bitres.alloc >= 0)
         destbits = s->psy.bitres.alloc *
                    (lambda / (avctx->global_quality ? avctx->global_quality : 120)) * chans;
@@ -686,7 +753,87 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,
         destbits = av_clip(destbits + FFMIN(extra, avail), 64, 6800 * chans);
     }

-    if (rc_global) {
+    int vbr = (avctx->flags & AV_CODEC_FLAG_QSCALE) && s->nmr;
+    int vbr_subst = 1;
+    float vbr_t = 0.0f;
+    if (vbr) {
+        /* nd-target VBR: constant achieved noise-to-mask, bits float. */
+        /* target in log2(dist/real-mask), anchored so -q:a 1 lands near
+         * 128 kbps stereo on the tuning corpus; higher q is finer, each q
+         * doubling ~ -1.2 */
+        vbr_t = FFMAX(NMR_VBR_ANCHOR - 1.2f * log2f(avctx->global_quality > 0 ?
+                                            avctx->global_quality / (float)FF_QP2LAMBDA : 1.0f),
+                      NMR_VBR_TMIN);
+    }
+    if (vbr && is8_any) {
+        /* short frames: the grouped-band statistic is inflated (same reason
+         * nd_ema excludes shorts), so they bisect against an offset target
+         * rather than the long-frame one - and never write solver state
+         * (transient-dense content would otherwise starve the long-frame
+         * warm start of updates). Warm off the surrounding operating point. */
+        float lam0 = *vslew_st > 0.0f ? FFMIN(*vslew_st, 1e4f) : 0.0f;
+        float *infl = &s->nmr->vbr_infl[sl[0]->cur_ch & 15];
+        if (lam0 > 0.0f) {
+            /* measured stat offset, decided at shorts-run entry: evaluate
+             * the grouped stat at the long-anchor lambda - continuity says
+             * true quality is on target there, so any excess IS this
+             * content's inflation. Broadband beats (velvet) measure +2..+4,
+             * tonal musical decays measure none; the old fixed +3 picked
+             * one class and audibly damaged the other. Re-measured every
+             * short frame: run-entry-only attribution tars a whole tonal
+             * decay run with its mini-attack entry frame. */
+            float st0, ehf = 0.0f, etot = 0.0f, hfw;
+            nmr_eval_slots(s, sl, nsl, cstep, lam0);
+            st0 = nmr_nd_stat(s, sl, nsl);
+            /* the stat reads inflated on BOTH classes; audibility does not.
+             * Coarse shorts hide under HF-dominant transients (hats, claps)
+             * and stick out on LF/mid tonal material - weight the offset by
+             * the frame's HF energy share. */
+            for (int k = 0; k < nsl; k++)
+                for (int b_ = 0; b_ < sl[k]->nact; b_++) {
+                    int b = sl[k]->act[b_], bi = sl[k]->bidx[b];
+                    float en = sl[k]->pener[bi];
+                    int bin = sl[k]->bst[b] - sl[k]->bw[b]*128;
+                    etot += en;
+                    if (bin * (avctx->sample_rate * 4.0f) / 1024.0f > 6000.0f)
+                        ehf += en;
+                }
+            hfw = etot > 0.0f ? av_clipf(ehf / (etot * NMR_INFL_HF), 0.0f, 1.0f) : 1.0f;
+            *infl = isnan(st0) ? 0.0f : av_clipf(st0 - vbr_t, 0.0f, 4.0f) * hfw;
+        }
+        if (lam0 > 0.0f)
+            lam = nmr_solve_slots_nd(s, sl, nsl, cstep, vbr_t + *infl,
+                                     lam0/32.0f, FFMIN(lam0*32.0f, 1e4f), NMR_CWARM + 2);
+        else
+            lam = nmr_solve_slots_nd(s, sl, nsl, cstep, vbr_t, 1e-9f, 1e4f, NMR_ITERS);
+        if (lam0 > 0.0f) {
+            /* the offset-bisect assumes isolated transients between long
+             * anchors; on continuous-shorts content the per-frame grouped
+             * stat is bisect noise and lambda careens across decades -
+             * audible as whooshing. Hold shorts to the same near-constant
+             * run slew CBR uses (dives toward finer stay free: transient
+             * bits float by design). */
+            float kup = s->nmr->prev_was_short ? NMR_SLEW_RUN : NMR_SLEW;
+            if (lam > lam0 * kup || lam < lam0 / 4.0f) {
+                lam = av_clipf(lam, lam0 / 4.0f, lam0 * kup);
+                nmr_eval_slots(s, sl, nsl, cstep, lam);
+            }
+        }
+        vbr_subst = 0;
+    } else if (vbr) {
+        /* warm-started off the previous frame's lambda */
+        float lam0 = s->nmr->lam[sl[0]->cur_ch];
+        lam = 1.0f;
+        if (lam0 > 0.0f) {
+            lam0 = FFMIN(lam0, 1e4f);
+            lam = nmr_solve_slots_nd(s, sl, nsl, cstep, vbr_t,
+                                     lam0/32.0f, FFMIN(lam0*32.0f, 1e4f), NMR_CWARM + 2);
+            if (lam < lam0/16.0f || lam > lam0*16.0f)
+                lam0 = 0.0f;
+        }
+        if (lam0 <= 0.0f)
+            lam = nmr_solve_slots_nd(s, sl, nsl, cstep, vbr_t, 1e-9f, 1e4f, NMR_ITERS);
+    } else if (rc_global) {
         /* corridor bisect around the servoed centre; pressure = stateless
          * rc_off multiplier (folding it into lam_rc winds up) */
         float R = avctx->bit_rate * 1024.0 / avctx->sample_rate;
@@ -787,7 +934,16 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,
             }
         }
         /* fine pass: narrow corridor around the coarse solve */
-        if (rc_global)
+        if (vbr && !vbr_subst) {
+            float infl2 = is8_any ? s->nmr->vbr_infl[sl[0]->cur_ch & 15] : 0.0f;
+            lam = nmr_solve_slots_nd(s, sl, nsl, NMR_STEP, vbr_t + infl2,
+                                     lam/32.0f, FFMIN(lam*32.0f, 1e4f), NMR_IFINE);
+        } else if (vbr)
+            /* wide re-bisect: the fine grid sits systematically finer than
+             * the coarse grid at equal lambda, well past a 1-octave bracket */
+            lam = nmr_solve_slots_nd(s, sl, nsl, NMR_STEP, vbr_t,
+                                     lam/32.0f, FFMIN(lam*32.0f, 1e4f), NMR_IFINE);
+        else if (rc_global)
             lam = nmr_solve_slots(s, sl, nsl, NMR_STEP, destbits, lam/2.0f, lam*2.0f, NMR_RC_FITERS);
         else
             lam = nmr_solve_slots(s, sl, nsl, NMR_STEP, destbits, lam/16.0f, lam*16.0f, NMR_IFINE);
@@ -795,6 +951,103 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,

     lam_dem = lam;   /* demand-solved lambda, pre bucket clamp: what content wants */

+    if (vbr) {
+        /* quality slew: lambda moves smoothly, bits follow content (the same
+         * anti-flutter rule the CBR path uses; unconstrained per-frame lambda
+         * reads as framerate-rate quality modulation). Low-content frames
+         * (silence, stop-gaps) rail lambda high legitimately - they must
+         * neither be clamped nor write back solver state, or every gap
+         * poisons the slew/warm-start and the solve re-traps at the rail. */
+        int subst = 0;
+        float st_fin = nmr_nd_stat(s, sl, nsl);
+        for (int k = 0; k < nsl; k++)
+            subst += sl[k]->nact;
+        /* a frame counts as an operating point only if it bisected the
+         * statistic (not a short frame holding lambda) and the solve
+         * actually REACHED the target: cheap frames (room tone, gaps) sit
+         * far below it at any lambda and rail high while costing nothing */
+        /* two-sided: a frame railed COARSE (st far above target, e.g. the
+         * first content frames of a quiet channel bisecting from nothing)
+         * is no more an operating point than a railed-fine cheap frame -
+         * writing its rail lambda builds a slew prison the channel then
+         * climbs out of one rung per frame, muffling the stream head */
+        subst = vbr_subst && subst >= 8 && !isnan(st_fin) &&
+                st_fin > vbr_t - 0.5f && st_fin < vbr_t + 0.5f;
+        if (subst) {
+            if (*vslew_st > 0.0f) {
+                /* only bisected LONG frames land here (shorts run free); on
+                 * steady-tonal content per-frame lambda jitter reads as
+                 * framerate-rate noise modulation (see NMR_RC_CORR) */
+                if (lam > *vslew_st * NMR_SLEW || lam < *vslew_st / NMR_SLEW) {
+                    lam = av_clipf(lam, *vslew_st / NMR_SLEW, *vslew_st * NMR_SLEW);
+                    nmr_eval_slots(s, sl, nsl, NMR_STEP, lam);
+                }
+            }
+            *vslew_st = s->nmr->lam_slew = lam;
+        } else if (is8_any && *vslew_st > 0.0f) {
+            /* short frames: PASS 2 re-bisects the inflated grouped stat on
+             * a fresh +-32x bracket, undoing any earlier bound - the clamp
+             * must sit here, after the last solve (love.flac whoosh: lambda
+             * railed at exactly lam*32 through an all-shorts passage) */
+            float kup  = s->nmr->prev_was_short ? NMR_SLEW_RUN : NMR_SLEW;
+            float lam0 = FFMIN(*vslew_st, 1e4f);
+            if (lam > lam0 * kup || lam < lam0 / 4.0f) {
+                lam = av_clipf(lam, lam0 / 4.0f, lam0 * kup);
+                nmr_eval_slots(s, sl, nsl, NMR_STEP, lam);
+            }
+        }
+        vbr_subst = subst;
+    }
+    if (vbr) {
+        /* legality only: a frame may never exceed the bit reservoir bound.
+         * Side bits ride on top of the trellis count; a negative side
+         * estimate must never enlarge the cap. The cap is the format limit
+         * at the reference lambda (the -q:a quality itself for VBR) and
+         * only shrinks on an outer hard-overflow retry, so the retry
+         * converges. */
+        int hardcap = av_clip((int)(5800.f * FFMIN(1.f, lambda / lam_ref)), 256, 5800) * chans;
+        int tot = 0;
+        if (s->nmr->side_inited)
+            hardcap = FFMAX(hardcap - (int)(FFMAX(s->nmr->side_ema, 0.0f) * chans / s->channels), 256);
+        for (int k = 0; k < nsl; k++)
+            tot += nmr_slot_bits(sl[k], s->nmr->nb[sl[k]->si], NMR_STEP);
+        if (tot > hardcap) {
+            lam = nmr_solve_slots(s, sl, nsl, NMR_STEP, hardcap, lam, 1e4f, NMR_RC_ITERS);
+            tot = 0;
+            for (int k = 0; k < nsl; k++)
+                tot += nmr_slot_bits(sl[k], s->nmr->nb[sl[k]->si], NMR_STEP);
+            /* dense short frames at high rates can exceed the decoder
+             * buffer even at the coarsest candidate grid; an illegal frame
+             * traps the outer re-encode loop forever. Shed the highest
+             * bands until the frame fits. act[] runs window group by window
+             * group, so pick the highest band across all groups; the first
+             * band anchors the scalefactor chain and stays. */
+            while (tot > hardcap) {
+                int dropped = 0;
+                for (int k = 0; k < nsl; k++) {
+                    NMRSlot *t = sl[k];
+                    SingleChannelElement *sce = t->sce;
+                    int hi = 1, b, w0, g;
+                    if (t->nact <= 1)
+                        continue;
+                    for (int b_ = 2; b_ < t->nact; b_++)
+                        if (t->bg[t->act[b_]] >= t->bg[t->act[hi]])
+                            hi = b_;
+                    b  = t->act[hi];
+                    w0 = t->bw[b]; g = t->bg[b];
+                    for (int w2 = 0; w2 < sce->ics.group_len[w0]; w2++)
+                        sce->zeroes[(w0+w2)*16+g] = 1;
+                    memmove(&t->act[hi], &t->act[hi + 1], (t->nact - hi - 1) * sizeof(t->act[0]));
+                    t->nact--;
+                    dropped = 1;
+                }
+                if (!dropped)
+                    break;
+                tot = nmr_eval_slots(s, sl, nsl, NMR_STEP, lam);
+            }
+        }
+    }
+
     if (rc_global) {
         /* reservoir cap on the fine grid (see the coarse-pass bound above),
          * then the quality slew limiter */
@@ -839,7 +1092,7 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,
             }
         }
         s->nmr->lam_slew = lam;
-    } else if (rc_eligible) {
+    } else if (rc_eligible && !vbr) {
         /* corridor not yet bootstrapped: the same bucket-full floor, so a
          * fade-in bootstraps lam_rc at a spending operating point instead
          * of memorizing the intro's starvation lambda */
@@ -852,8 +1105,9 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,
             lam = nmr_solve_slots(s, sl, nsl, NMR_STEP, fbits, lam / 64.0f, lam, NMR_RC_ITERS);
     }

-    for (int k = 0; k < nsl; k++)
-        s->nmr->lam[sl[k]->cur_ch] = lam;   /* warm start for the next frame */
+    if (!vbr || vbr_subst)
+        for (int k = 0; k < nsl; k++)
+            s->nmr->lam[sl[k]->cur_ch] = lam;   /* warm start for the next frame */
     {   /* nd: mean achieved dist/real-mask (dimensionless starvation +
          * noise-class signal) */
         float ndsum = 0.0f; int ndn = 0;
@@ -926,7 +1180,9 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,
                 int bi = t->bidx[b];
                 float spread = t->pspread[bi];
                 float nmr_pns, cost_keep, cost_pns, frac;
-                if (!t->sce->can_pns[bi])
+                /* zeroes[] is only set on a coded band when it was shed for
+                 * legality; such a band must not come back as noise */
+                if (!t->sce->can_pns[bi] || t->sce->zeroes[bi])
                     continue;

                 int was  = s->nmr->pns_prev[t->cur_ch & 15][bi];
@@ -979,10 +1235,13 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,
                 }
             }
             if (pns_count) {
-                t->nact = 0;
-                for (int b = 0; b < t->nbnd; b++)
-                    if (!t->is_pns[b])
-                        t->act[t->nact++] = b;
+                /* filter the active list rather than rebuilding it from
+                 * nbnd: bands shed for legality must stay out */
+                int n = 0;
+                for (int b_ = 0; b_ < t->nact; b_++)
+                    if (!t->is_pns[t->act[b_]])
+                        t->act[n++] = t->act[b_];
+                t->nact = n;
             }
             pns_total += pns_count;
         }
@@ -990,7 +1249,7 @@ static void nmr_solve_group(AVCodecContext *avctx, AACEncContext *s,
             /* re-solve over the survivors: at fixed lambda the allocation is
              * the same except for the repaired sf-delta chain; in bisection
              * mode re-spend the freed budget */
-            if (rc_global)
+            if (rc_global || vbr)
                 nmr_eval_slots(s, sl, nsl, NMR_STEP, lam);
             else
                 nmr_solve_slots(s, sl, nsl, NMR_STEP, destbits - pns_total * NMR_PNS_BITS,
@@ -1050,10 +1309,14 @@ static void search_for_quantizers_nmr(AVCodecContext *avctx,
         /* latch the RC mode per frame: a mid-frame bootstrap must not flip
          * the CPE defer logic between channels */
         n->rc_gl = rc_eligible && n->lam_rc > 0.0f;
-
-        /* Transient burst run state: set at run start and held across the run so
+    }
+    if (avctx->frame_num != n->win_frame_num) {
+        /* Window history, once per frame in every rate-control mode (the
+         * quality-target slew and the ABR servo read it too).
+         * Transient burst run state: set at run start and held across the run so
          * coding stays uniform; repaid from the reservoir's steady stretches. */
         int is_short = sce->ics.window_sequence[0] == EIGHT_SHORT_SEQUENCE;
+        n->win_frame_num = avctx->frame_num;
         if (is_short) {
             if (!n->prev_was_short) {           /* run start */
                 if (n->frames_since_short >= NMR_BURST_GAP) {
diff --git a/libavcodec/aacenc.c b/libavcodec/aacenc.c
index 9398467bc7..e415f54247 100644
--- a/libavcodec/aacenc.c
+++ b/libavcodec/aacenc.c
@@ -1537,8 +1537,21 @@ static int aac_encode_frame(AVCodecContext *avctx, AVPacket *avpkt,
         }

         if (avctx->flags & AV_CODEC_FLAG_QSCALE) {
-            /* When using a constant Q-scale, don't mess with lambda */
-            break;
+            /* When using a constant Q-scale, don't mess with lambda, unless
+             * the frame does not fit the decoder buffer: retry coarser (the
+             * coders' legality caps shrink with lambda) */
+            frame_bits = put_bits_count(&s->pb);
+            if (frame_bits < 6144 * s->channels - 3 || its >= 16)
+                break;
+            s->lambda *= FFMIN(0.9f, (6144.0f * s->channels - 3) / frame_bits);
+            for (i = 0; i < s->chan_map[0]; i++) {
+                chans = s->chan_map[i + 1] == TYPE_CPE ? 2 : 1;
+                for (ch = 0; ch < chans; ch++)
+                    memcpy(s->cpe[i].ch[ch].coeffs, s->cpe[i].ch[ch].pcoeffs,
+                           sizeof(s->cpe[i].ch[ch].coeffs));
+            }
+            its++;
+            continue;
         }

         frame_bits = put_bits_count(&s->pb);
@@ -1616,6 +1629,8 @@ static int aac_encode_frame(AVCodecContext *avctx, AVPacket *avpkt,
             break;
         }
     } while (1);
+    if (avctx->flags & AV_CODEC_FLAG_QSCALE)
+        s->lambda = avctx->global_quality > 0 ? avctx->global_quality : 120;

     /* tool-usage stats over the final per-band decisions of this frame */
     for (i = 0; i < s->chan_map[0]; i++) {
@@ -1676,7 +1691,15 @@ static int aac_encode_frame(AVCodecContext *avctx, AVPacket *avpkt,
     }
     avpkt->size            = put_bytes_output(&s->pb);

-    s->lambda_sum += (s->nmr && s->nmr->lam_rc > 0.0f) ? s->nmr->lam_rc : s->lambda;
+    /* NMR reports its real operating lambda: the corridor centre in CBR,
+     * the quality-mode slew state in VBR - the lambda of the last long
+     * operating-point solve, which short frames and legality re-solves do
+     * not update (the outer-loop s->lambda is never touched for this coder
+     * and would pin Qavg at its 120 init) */
+    s->lambda_sum += (s->nmr && s->nmr->lam_slew > 0.0f &&
+                      (avctx->flags & AV_CODEC_FLAG_QSCALE)) ?
+                     s->nmr->lam_slew :
+                     (s->nmr && s->nmr->lam_rc > 0.0f) ? s->nmr->lam_rc : s->lambda;
     s->lambda_count++;

     ret = ff_af_queue_remove(&s->afq, avctx->frame_size, avpkt);
@@ -1877,9 +1900,20 @@ static av_cold int aac_encode_init(AVCodecContext *avctx)
     if (avctx->cutoff > 0) {
         s->bandwidth = avctx->cutoff;
     } else {
-        int frame_br = (avctx->flags & AV_CODEC_FLAG_QSCALE) ?
-                       (avctx->bit_rate / 2.0f * (s->lambda / 120.f) * 1.5f) :
-                       (avctx->bit_rate / avctx->ch_layout.nb_channels);
+        int frame_br;
+        if (avctx->flags & AV_CODEC_FLAG_QSCALE) {
+            if (s->options.coder == AAC_CODER_NMR) {
+                /* nd-target VBR: expected per-channel rate from the quality
+                 * ladder (measured: q=1 ~ 64.5 kbps/ch, x1.27 per doubling) */
+                float q = avctx->global_quality > 0 ?
+                          avctx->global_quality / (float)FF_QP2LAMBDA : 1.0f;
+                frame_br = 66000 * powf(q, 0.29f);
+            } else {
+                frame_br = avctx->bit_rate / 2.0f * (s->lambda / 120.f) * 1.5f;
+            }
+        } else {
+            frame_br = avctx->bit_rate / avctx->ch_layout.nb_channels;
+        }

         if (s->options.coder == AAC_CODER_NMR && frame_br >= 24000) {
             /* Ear-tuned, not metric-tuned: Zim rewards HF presence and cannot
@@ -1930,9 +1964,12 @@ static av_cold int aac_encode_init(AVCodecContext *avctx)
     lengths[1] = ff_aac_num_swb_128[s->samplerate_index];
     for (i = 0; i < s->chan_map[0]; i++)
         grouping[i] = s->chan_map[i + 1] == TYPE_CPE;
+    s->psy.unbounded_pe = (avctx->flags & AV_CODEC_FLAG_QSCALE) &&
+                          s->options.coder == AAC_CODER_NMR;
     if ((ret = ff_psy_init(&s->psy, avctx, 2, sizes, lengths,
                            s->chan_map[0], grouping, s->bandwidth)) < 0)
         return ret;
+
     ff_lpc_init(&s->lpc, 2*avctx->frame_size, TNS_MAX_ORDER, FF_LPC_TYPE_LEVINSON);
     s->random_state = 0x1f2e3d4c;

diff --git a/libavcodec/aacenc.h b/libavcodec/aacenc.h
index 08bc0fb961..9d2ba89056 100644
--- a/libavcodec/aacenc.h
+++ b/libavcodec/aacenc.h
@@ -238,8 +238,11 @@ typedef struct AACNMRCurves {
     int     rc_sat_frame;                        ///< the current frame hit a saturated overage
     int     frames_since_short;                  ///< long-block frames since the last short run (the "gap"): large = isolated transient
     int     prev_was_short;                      ///< previous frame was a short block (for run-start detection)
+    int64_t win_frame_num;                       ///< frame the window history was last advanced for
     float   run_burst;                           ///< transient bit-burst factor, set at run start and held across the short run
     float   lam_slew;                            ///< final operating lambda of the previous RC frame (slew-limiter state)
+    float   lam_slew_ch[16];                     ///< per-channel slew state for decoupled solo solves (a shared slew cross-contaminates the pair)
+    float   vbr_infl[16];                        ///< per-channel grouped-stat inflation measured at shorts-run entry
     float   nd_ema;                              ///< smoothed achieved distortion/real-mask over long-frame coded bands (1 = at threshold; >>1 flags psy-unreliable noise-class content)
     float   press;                               ///< rate-pressure ramp [0,1]: lambda EMA against anchors that scale up when nd_ema flags noise-class content (psy masks unreliable there, lambda reads inflated)
     float   lam_short_ema;                       ///< smoothed operating lambda of short frames
diff --git a/libavcodec/aacpsy.c b/libavcodec/aacpsy.c
index ceb978d0a6..0fd0c87b85 100644
--- a/libavcodec/aacpsy.c
+++ b/libavcodec/aacpsy.c
@@ -807,7 +807,17 @@ static void psy_3gpp_analyze_channel(FFPsyContext *ctx, int channel,

     /* 5.6.1.3.2 "Calculation of the desired perceptual entropy" */
     ctx->ch[channel].entropy = pe;
-    if (ctx->avctx->flags & AV_CODEC_FLAG_QSCALE) {
+    if (ctx->unbounded_pe) {
+        /* quality-target coder: run the PE reduction at a FIXED reference
+         * quality so thresholds keep their perceptual shaping but stay
+         * independent of the user's -q:a (the coder's own noise-to-mask
+         * target is the sole quality authority) */
+        desired_pe = pe * 120.0f / (2 * 2.5f * 120.0f);
+        desired_bits = FFMIN(2560, PSY_3GPP_PE_TO_BITS(desired_pe));
+        desired_pe = PSY_3GPP_BITS_TO_PE(desired_bits);
+        pctx->pe.max = FFMAX(pe, pctx->pe.max);
+        pctx->pe.min = FFMIN(pe, pctx->pe.min);
+    } else if (ctx->avctx->flags & AV_CODEC_FLAG_QSCALE) {
         /* (2.5 * 120) achieves almost transparent rate, and we want to give
          * ample room downwards, so we make that equivalent to QSCALE=2.4
          */
diff --git a/libavcodec/psymodel.h b/libavcodec/psymodel.h
index 302d25e1a2..464b206008 100644
--- a/libavcodec/psymodel.h
+++ b/libavcodec/psymodel.h
@@ -94,6 +94,7 @@ typedef struct FFPsyContext {
     FFPsyChannelGroup *group;         ///< channel group information
     int num_groups;                   ///< number of channel groups
     int cutoff;                       ///< lowpass frequency cutoff for analysis
+    int unbounded_pe;                 ///< PE reduction at a fixed reference quality, not the rate (quality-target coder owns the rate)

     uint8_t **bands;                  ///< scalefactor band sizes for possible frame sizes
     int     *num_bands;               ///< number of scalefactor bands for possible frame sizes
diff --git a/tests/fate/aac.mak b/tests/fate/aac.mak
index ec49b83b7d..b1f49aff2a 100644
--- a/tests/fate/aac.mak
+++ b/tests/fate/aac.mak
@@ -263,6 +263,11 @@ fate-aac-yoraw-encode: REF = $(SAMPLES)/audio-reference/yo.raw-short.wav
 fate-aac-yoraw-encode: CMP_TARGET = 226
 fate-aac-yoraw-encode: FUZZ = 17

+# NMR quality-target modes, per-packet sizes pinned: VBR at the finest quality
+# drives frames against the decoder-buffer limit
+FATE_AAC_ENCODE_TRANSCODE += fate-aac-nmr-vbr-encode
+fate-aac-nmr-vbr-encode: CMD = transcode wav $(TARGET_SAMPLES)/audio-reference/luckynight_2ch_44kHz_s16.wav adts "-af aresample -c:a aac -q:a 8 -t 2" "-c copy"
+
 tests/data/fate/aac-5_1_2.adts: TAG = GEN
 tests/data/fate/aac-5_1_2.adts: tests/data/asynth-44100-8.wav
 tests/data/fate/aac-5_1_2.adts: ffmpeg$(PROGSSUF)$(EXESUF) | tests/data/fate
@@ -354,11 +359,12 @@ FATE_AAC_FRAMECRC-$(call FRAMECRC, MOV, AAC, ARESAMPLE_FILTER PAN_FILTER PCM_F32

 FATE_AAC_ENCODE-$(call TRANSCODE, AAC, MP4 MOV, WAV_MUXER WAV_DEMUXER PCM_S16LE_DECODER ARESAMPLE_FILTER) += $(FATE_AAC_ENCODE)
 FATE_AAC_ENCODE_FFPROBE-$(call TRANSCODE, AAC, ADTS AAC, WAV_DEMUXER PCM_S16LE_DECODER ARESAMPLE_FILTER) += $(FATE_AAC_ENCODE_FFPROBE)
+FATE_AAC_ENCODE_TRANSCODE-$(call TRANSCODE, AAC, ADTS AAC, WAV_DEMUXER PCM_S16LE_DECODER ARESAMPLE_FILTER) += $(FATE_AAC_ENCODE_TRANSCODE)

 FATE_AAC_BSF-$(call FRAMECRC, AAC MATROSKA, AAC, AAC_PARSER AAC_ADTSTOASC_BSF MATROSKA_MUXER) += fate-aac-autobsf-adtstoasc

-FATE_SAMPLES_FFMPEG += $(FATE_AAC_ALL) $(FATE_AAC_FRAMECRC-yes) $(FATE_AAC_ENCODE-yes) $(FATE_AAC_BSF-yes)
+FATE_SAMPLES_FFMPEG += $(FATE_AAC_ALL) $(FATE_AAC_FRAMECRC-yes) $(FATE_AAC_ENCODE-yes) $(FATE_AAC_ENCODE_TRANSCODE-yes) $(FATE_AAC_BSF-yes)
 FATE_SAMPLES_FFMPEG_FFPROBE += $(FATE_AAC_ENCODE_FFPROBE-yes)

-fate-aac: $(FATE_AAC_ALL) $(FATE_AAC_FRAMECRC-yes) $(FATE_AAC_ENCODE-yes) $(FATE_AAC_ENCODE_FFPROBE-yes) $(FATE_AAC_BSF-yes)
+fate-aac: $(FATE_AAC_ALL) $(FATE_AAC_FRAMECRC-yes) $(FATE_AAC_ENCODE-yes) $(FATE_AAC_ENCODE_FFPROBE-yes) $(FATE_AAC_ENCODE_TRANSCODE-yes) $(FATE_AAC_BSF-yes)
 fate-aac-latm: $(FATE_AAC_LATM-yes)
diff --git a/tests/ref/fate/aac-nmr-vbr-encode b/tests/ref/fate/aac-nmr-vbr-encode
new file mode 100644
index 0000000000..b3403bed90
--- /dev/null
+++ b/tests/ref/fate/aac-nmr-vbr-encode
@@ -0,0 +1,95 @@
+d97fad39ae2d7044cd12da83580d8909 *tests/data/fate/aac-nmr-vbr-encode.adts
+113631 tests/data/fate/aac-nmr-vbr-encode.adts
+#tb 0: 1/28224000
+#media_type 0: audio
+#codec_id 0: aac
+#sample_rate 0: 44100
+#channel_layout_name 0: stereo
+0,          0,          0,   655360,     1432, 0xa47d9aba
+0,     655360,     655360,   655360,     1448, 0xccd8bcb2
+0,    1310720,    1310720,   655360,     1432, 0xd087ef11
+0,    1966080,    1966080,   655360,     1387, 0x52b1d993
+0,    2621440,    2621440,   655360,     1452, 0x9fb1a7ba
+0,    3276800,    3276800,   655360,     1412, 0xe912e636
+0,    3932160,    3932160,   655360,     1445, 0x312bdfd6
+0,    4587520,    4587520,   655360,     1408, 0xf391e39b
+0,    5242880,    5242880,   655360,     1400, 0x7b44fad2
+0,    5898240,    5898240,   655360,     1433, 0xcb84f594
+0,    6553600,    6553600,   655360,     1423, 0x131be20d
+0,    7208960,    7208960,   655360,     1438, 0x8510fa98
+0,    7864320,    7864320,   655360,     1406, 0x7c65bd4e
+0,    8519680,    8519680,   655360,     1423, 0x11ddf5c4
+0,    9175040,    9175040,   655360,     1448, 0xf566feef
+0,    9830400,    9830400,   655360,     1358, 0xc04ec806
+0,   10485760,   10485760,   655360,     1451, 0x1660fe6e
+0,   11141120,   11141120,   655360,     1451, 0xf81b1064
+0,   11796480,   11796480,   655360,     1396, 0x3f9cfb0f
+0,   12451840,   12451840,   655360,     1382, 0x0bb2ec24
+0,   13107200,   13107200,   655360,     1395, 0xcb71dcb6
+0,   13762560,   13762560,   655360,     1345, 0x5defc602
+0,   14417920,   14417920,   655360,     1396, 0x5fc7e2ba
+0,   15073280,   15073280,   655360,     1370, 0x3d24eb78
+0,   15728640,   15728640,   655360,     1356, 0x2773e5aa
+0,   16384000,   16384000,   655360,     1312, 0x109dc0f9
+0,   17039360,   17039360,   655360,     1341, 0x21d3d81c
+0,   17694720,   17694720,   655360,     1319, 0x0d90c7b2
+0,   18350080,   18350080,   655360,     1355, 0x0c53e537
+0,   19005440,   19005440,   655360,     1352, 0xb69edf91
+0,   19660800,   19660800,   655360,     1352, 0x788ed5eb
+0,   20316160,   20316160,   655360,     1344, 0x3af9d252
+0,   20971520,   20971520,   655360,     1296, 0x5eafc05d
+0,   21626880,   21626880,   655360,     1352, 0xbe22d67a
+0,   22282240,   22282240,   655360,     1358, 0x6569dcaf
+0,   22937600,   22937600,   655360,     1371, 0x21f2d28c
+0,   23592960,   23592960,   655360,     1354, 0xf068d8dc
+0,   24248320,   24248320,   655360,     1348, 0x8990d8c6
+0,   24903680,   24903680,   655360,     1350, 0x6c12e7dc
+0,   25559040,   25559040,   655360,     1338, 0x2225c0e1
+0,   26214400,   26214400,   655360,     1368, 0x6b11db0e
+0,   26869760,   26869760,   655360,     1352, 0xeda9c2b9
+0,   27525120,   27525120,   655360,     1376, 0x1b92da50
+0,   28180480,   28180480,   655360,     1263, 0x8af89e03
+0,   28835840,   28835840,   655360,     1374, 0x466ae490
+0,   29491200,   29491200,   655360,     1335, 0xcc46dd61
+0,   30146560,   30146560,   655360,     1335, 0x7417c793
+0,   30801920,   30801920,   655360,     1337, 0xd4a4c744
+0,   31457280,   31457280,   655360,     1356, 0xd2f1e61e
+0,   32112640,   32112640,   655360,     1360, 0x2cfacf15
+0,   32768000,   32768000,   655360,     1374, 0x853bf9c4
+0,   33423360,   33423360,   655360,     1339, 0x1ed6df4e
+0,   34078720,   34078720,   655360,     1333, 0x3443cb14
+0,   34734080,   34734080,   655360,     1391, 0xaf6ddeb5
+0,   35389440,   35389440,   655360,     1349, 0x966dfa91
+0,   36044800,   36044800,   655360,     1322, 0x60bccbb9
+0,   36700160,   36700160,   655360,     1356, 0xf363e54c
+0,   37355520,   37355520,   655360,     1330, 0x6439ce4a
+0,   38010880,   38010880,   655360,     1350, 0xeadfd452
+0,   38666240,   38666240,   655360,     1328, 0x2e4dcf0b
+0,   39321600,   39321600,   655360,     1300, 0x943cc8f7
+0,   39976960,   39976960,   655360,     1316, 0x9071cda4
+0,   40632320,   40632320,   655360,     1325, 0x43eda4ce
+0,   41287680,   41287680,   655360,     1397, 0x820ac663
+0,   41943040,   41943040,   655360,     1127, 0x03e73eef
+0,   42598400,   42598400,   655360,     1125, 0xd8fa42ed
+0,   43253760,   43253760,   655360,     1071, 0x864622e0
+0,   43909120,   43909120,   655360,     1159, 0x45d160c4
+0,   44564480,   44564480,   655360,     1037, 0xd8e624d0
+0,   45219840,   45219840,   655360,     1056, 0x70b422bd
+0,   45875200,   45875200,   655360,     1027, 0xf33a1216
+0,   46530560,   46530560,   655360,     1281, 0x4e0fb0a1
+0,   47185920,   47185920,   655360,     1293, 0x3b98adc6
+0,   47841280,   47841280,   655360,     1027, 0x9bbe0e89
+0,   48496640,   48496640,   655360,     1060, 0x6bda21fb
+0,   49152000,   49152000,   655360,     1060, 0xf3312f1c
+0,   49807360,   49807360,   655360,     1088, 0x01193f8b
+0,   50462720,   50462720,   655360,     1103, 0x023f43bd
+0,   51118080,   51118080,   655360,     1109, 0x991f49fc
+0,   51773440,   51773440,   655360,     1072, 0xb2cb21e3
+0,   52428800,   52428800,   655360,     1040, 0xeb492506
+0,   53084160,   53084160,   655360,     1355, 0x53669cf5
+0,   53739520,   53739520,   655360,     1014, 0x10750d2b
+0,   54394880,   54394880,   655360,     1029, 0xdd471f62
+0,   55050240,   55050240,   655360,     1221, 0x61598299
+0,   55705600,   55705600,   655360,     1445, 0x65e0e77d
+0,   56360960,   56360960,   655360,     1123, 0x975437bb
+0,   57016320,   57016320,   655360,       14, 0x3f2f0844