Commit d477de3c for whisper.cpp

commit d477de3ce5aa696a991382df2a2958200ad620a8
Author: Georgi Gerganov <ggerganov@gmail.com>
Date:   Mon Oct 5 11:36:25 2026 +0300

    llama : fix unexpected graph reallocation in the k-pool models (llama/29958)

    * llama : fix unexpected graph reallocation in the k-pool models

    Both k-pool models built a graph shape that depends on state the
    full-context reserve cannot know:

    - qwen4exp branched on inp->cache_safe, which turns false as soon as
      llama_memory_seq_cp shares cells (e.g. batched-bench -pps): the QSA
      layers swapped scatter+gather for fill+concat and dropped the
      new_pool_rep leaf, so the decode graph had 12 fewer nodes than the
      reserved one
    - glm5-next branched on gather = n_tokens <= 16 && n_kv > n_sel, so the
      TG decode built the gather shape (7564 nodes) while the last reserve,
      the PP one, had the dense shape (7762 nodes)

    Either mismatch forces a decode-time re-reserve that drops the
    worst-case sizing and bakes in the current state, so the next state
    growth (n_pool, n_kv, n_new) needs more room at an unchanged graph size
    and aborts under GGML_SCHED_DEBUG_REALLOC=1. Reproduce with, e.g.:

      GGML_SCHED_DEBUG_REALLOC=1 ./bin/llama-batched-bench \
        -hf ggml-org/GLM-5.3-Flash-GGUF:Q2_K -npp 2500 -ntg 32 -npl 1,2 \
        -c 32768 -pps -kvu

    Always scatter+gather the pooled keys, and pick gather from context
    constants only: n_ubatch bounds every ubatch, top_k + kpool - 1 bounds
    n_sel. Every graph of a context then shares one shape, which the
    reserve covers, and the dense path measured faster than the gather path
    at 2.5k and 16k context.

    Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

    * llama : drop the unused k-pool cache_safe graph API

    The k-pool graphs no longer branch on cache_safe, so nothing reads
    get_kpool_cache_safe() or the conditional new_pool_rep any more: both
    models always pass the scatter target, which set_input_kpool now
    requires instead of merely preferring.

    Also drop the cache_safe copy in kpool_build_sizes(), a sizes-only
    helper. The layout and state flag itself stays, it still decides which
    pools a layout with shared cells must re-pool.

    Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

    * tests : add a shared-seq graph reserve regression test

    Decode a prompt into seq 0, share its cells with seq 1 via
    llama_memory_seq_cp (what llama-batched-bench does for -pps), then keep
    decoding both sequences. For the k-pool models sharing clears
    cache_safe, which changes the graph topology while the pools keep
    growing, so a scheduler that re-reserves with the current state
    instead of the worst-case one aborts under GGML_SCHED_DEBUG_REALLOC=1.
    The test registration sets that flag, and the test aborts on both
    k-pool models before 2220411ec1.

    kimi-linear and minimax-01 are skipped: they reserve the final pp graph
    with n_seqs = 1 (see [TAG_RESERVE_DIAG_DECAY] in llama-context.cpp), so
    every multi-seq graph has a different layout and re-reserves by design.

    Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

    * cont : add TODOs

    * cont : fix comment

    * cuda: match the moe weighted reduction on empty ubatches

    ggml_cuda_match_moe_weighted_reduction rejected tensors with zero
    rows. A ubatch without outputs shrinks the last layer to zero rows
    through inp_out_ids, so graph_optimize dropped its alloc dep there and
    the scheduler graph lost one node compared to the reserved one. The
    scheduler then re-reserved at the size of that ubatch, and the next
    ubatch with the same node count but larger tensors aborted under
    GGML_SCHED_DEBUG_REALLOC=1.

    The compute loop already skips empty nodes before trying any fusion,
    so the guard only made the alloc deps depend on the row count.

    * tests: build the rollback test only where internal symbols link

    The shared-seq case calls llm_arch_from_string, which libllama does
    not export through LLAMA_API, so linking test-recurrent-state-rollback
    fails on Windows with shared libraries. Its build now sits in the
    NOT WIN32 OR NOT BUILD_SHARED_LIBS block, next to test-llama-archs and
    the test registration it already lives under.

    * tests: skip archs by name in the shared-seq reserve test

    The skip of kimi-linear and minimax-01 went through llm_arch_from_string,
    which libllama does not export through LLAMA_API, so the test could not
    link on Windows with shared libraries. It now compares the
    general.architecture string directly, and the test builds on every
    platform again.

    ---------

    Co-authored-by: Pascal <admin@serveurperso.com>

diff --git a/ggml/src/ggml-cuda/ggml-cuda.cu b/ggml/src/ggml-cuda/ggml-cuda.cu
index 303d67e4..c7241961 100644
--- a/ggml/src/ggml-cuda/ggml-cuda.cu
+++ b/ggml/src/ggml-cuda/ggml-cuda.cu
@@ -3164,7 +3164,7 @@ static bool ggml_cuda_match_moe_weighted_reduction(

     const int     n_expert_used = (int) weighted->ne[1];
     const int64_t n_tokens      = weighted->ne[2] * weighted->ne[3];
-    if (n_expert_used < 2 || n_expert_used > MOE_WEIGHTED_REDUCTION_MAX_EXPERTS || n_tokens <= 0) {
+    if (n_expert_used < 2 || n_expert_used > MOE_WEIGHTED_REDUCTION_MAX_EXPERTS) {
         return false;
     }