Commit 1167d3f42 for llama.cpp

commit 1167d3f42c2b4dc5469bf4e696cb8c678ca79110
Author: hey-gm <heitorgm@outlook.com>
Date:   Thu Oct 8 14:29:34 2026 +0100

    CUDA: fix CCCL version guard breaking on major version rollover (#29453)

    * CUDA: fix CCCL version guard breaking on major version rollover

    The guard compared the major and minor components independently:

        CCCL_MAJOR_VERSION >= 3 && CCCL_MINOR_VERSION >= 1

    Minor resets to 0 whenever a new major series is cut, so on CCCL 4.x
    this evaluates as 4 >= 3 && 0 >= 1, i.e. false. STRIDED_ITERATOR_AVAILABLE
    stops being defined and argsort silently falls back to the
    init_offsets path. Nothing warns and the build still succeeds, so the
    regression is a quiet performance loss rather than a compile error.

    CCCL already exposes the version as a single packed integer in
    MMMmmmpp form, which is what its own version header uses:

        CCCL_VERSION = MAJOR * 1000000 + MINOR * 1000 + PATCH

    so 3.4.3 is 3004003 and ">= 3.1" is a plain ">= 3001000". One
    comparison, with no component arithmetic left to get wrong.

    Checked against a hand-written "version >= 3.1" reference over 2.9.9,
    3.0.0, 3.1.0, 3.1.99, 3.2.0, 3.4.3, 3.9.9, 3.99.99, 4.0.0, 4.2.7 and
    5.0.0: no divergences. The old guard disagreed at 4.0.0 and 5.0.0.

    Verified on RTX 4070 (sm_89), CUDA 13.4, CCCL 3.4.3:

      - cmake --build build --config Release: exit 0
      - test-backend-ops test -o ARGSORT -b CUDA0: 98/98 passed, CUDA0 OK

    Note that a passing regression test does not on its own prove the guard
    is still taken, since the fallback path passes too. Preprocessing the
    real translation unit confirms the strided-iterator branch is the one
    compiled in: counting_iterator is present, init_offsets is not.

    Signed-off-by: Heitor <heitorgm@outlook.com>

    * Update ggml/src/ggml-cuda/argsort.cu

    * Apply suggestion from @ORippler

    ---------

    Signed-off-by: Heitor <heitorgm@outlook.com>
    Co-authored-by: Oliver Simons <osimons@nvidia.com>

diff --git a/ggml/src/ggml-cuda/argsort.cu b/ggml/src/ggml-cuda/argsort.cu
index f6a850dda..589b99ff0 100644
--- a/ggml/src/ggml-cuda/argsort.cu
+++ b/ggml/src/ggml-cuda/argsort.cu
@@ -2,7 +2,8 @@

 #ifdef GGML_CUDA_USE_CUB
 #    include <cub/cub.cuh>
-#    if (CCCL_MAJOR_VERSION >= 3 && CCCL_MINOR_VERSION >= 1)
+    // strided_iterator was added in CCCL 3.1
+#    if (CCCL_MAJOR_VERSION > 3 || (CCCL_MAJOR_VERSION == 3 && CCCL_MINOR_VERSION >= 1))
 #        define STRIDED_ITERATOR_AVAILABLE
 #        include <cuda/iterator>
 #    endif