Commit dd4c286f3 for llama.cpp

commit dd4c286f3809e73be087f374d9d865ca97c205e4
Author: Yash Raj Pandey <55940078+devYRPauli@users.noreply.github.com>
Date:   Fri Oct 2 10:37:11 2026 -0400

    ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (#27096)

    * ggml-cpu : fix soft_max_back wrong output when dst aliases src1

    GGML_OP_SOFT_MAX_BACK is listed in ggml_op_can_inplace, so the graph
    allocator may assign dst to alias either src0 (dy) or src1 (y).

    The result was built in several steps:

        ggml_vec_cpy_f32  (nc, dx, dy);
        ggml_vec_acc1_f32 (nc, dx, -dot_y_dy);
        ggml_vec_mul_f32  (nc, dx, dx, y);
        ggml_vec_scale_f32(nc, dx, scale);

    When dst aliases src1, the first step overwrites y and the third step
    then reads the overwritten values, so the output is silently wrong.
    Aliasing dst with src0 is unaffected. The CUDA kernel completes its
    reduction before writing and is already safe.

    Replace the sequence with a single fused loop that reads both sources
    before writing, which is correct under either aliasing.

    Add a regression test that marks dy as a graph output so the allocator
    is forced to alias dst with y, asserts that the alias actually
    happened, and compares against values computed on the host.

    * cont : remove comment

    ---------

    Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

diff --git a/ggml/src/ggml-cpu/ops.cpp b/ggml/src/ggml-cpu/ops.cpp
index 3e377298f..16a916d20 100644
--- a/ggml/src/ggml-cpu/ops.cpp
+++ b/ggml/src/ggml-cpu/ops.cpp
@@ -6012,11 +6012,10 @@ static void ggml_compute_forward_soft_max_ext_back_f32(

         // linear runtime, no additional memory
         float dot_y_dy = 0;
-        ggml_vec_dot_f32  (nc, &dot_y_dy, 0, y, 0, dy, 0, 1);
-        ggml_vec_cpy_f32  (nc, dx, dy);
-        ggml_vec_acc1_f32 (nc, dx, -dot_y_dy);
-        ggml_vec_mul_f32  (nc, dx, dx, y);
-        ggml_vec_scale_f32(nc, dx, scale);
+        ggml_vec_dot_f32(nc, &dot_y_dy, 0, y, 0, dy, 0, 1);
+        for (int i = 0; i < nc; i++) {
+            dx[i] = scale * (dy[i] - dot_y_dy) * y[i];
+        }

 #ifndef NDEBUG
         for (int i = 0; i < nc; ++i) {