Commit 1167d3f42 for llama.cpp
commit 1167d3f42c2b4dc5469bf4e696cb8c678ca79110
Author: hey-gm <heitorgm@outlook.com>
Date: Thu Oct 8 14:29:34 2026 +0100
CUDA: fix CCCL version guard breaking on major version rollover (#29453)
* CUDA: fix CCCL version guard breaking on major version rollover
The guard compared the major and minor components independently:
CCCL_MAJOR_VERSION >= 3 && CCCL_MINOR_VERSION >= 1
Minor resets to 0 whenever a new major series is cut, so on CCCL 4.x
this evaluates as 4 >= 3 && 0 >= 1, i.e. false. STRIDED_ITERATOR_AVAILABLE
stops being defined and argsort silently falls back to the
init_offsets path. Nothing warns and the build still succeeds, so the
regression is a quiet performance loss rather than a compile error.
CCCL already exposes the version as a single packed integer in
MMMmmmpp form, which is what its own version header uses:
CCCL_VERSION = MAJOR * 1000000 + MINOR * 1000 + PATCH
so 3.4.3 is 3004003 and ">= 3.1" is a plain ">= 3001000". One
comparison, with no component arithmetic left to get wrong.
Checked against a hand-written "version >= 3.1" reference over 2.9.9,
3.0.0, 3.1.0, 3.1.99, 3.2.0, 3.4.3, 3.9.9, 3.99.99, 4.0.0, 4.2.7 and
5.0.0: no divergences. The old guard disagreed at 4.0.0 and 5.0.0.
Verified on RTX 4070 (sm_89), CUDA 13.4, CCCL 3.4.3:
- cmake --build build --config Release: exit 0
- test-backend-ops test -o ARGSORT -b CUDA0: 98/98 passed, CUDA0 OK
Note that a passing regression test does not on its own prove the guard
is still taken, since the fallback path passes too. Preprocessing the
real translation unit confirms the strided-iterator branch is the one
compiled in: counting_iterator is present, init_offsets is not.
Signed-off-by: Heitor <heitorgm@outlook.com>
* Update ggml/src/ggml-cuda/argsort.cu
* Apply suggestion from @ORippler
---------
Signed-off-by: Heitor <heitorgm@outlook.com>
Co-authored-by: Oliver Simons <osimons@nvidia.com>
diff --git a/ggml/src/ggml-cuda/argsort.cu b/ggml/src/ggml-cuda/argsort.cu
index f6a850dda..589b99ff0 100644
--- a/ggml/src/ggml-cuda/argsort.cu
+++ b/ggml/src/ggml-cuda/argsort.cu
@@ -2,7 +2,8 @@
#ifdef GGML_CUDA_USE_CUB
# include <cub/cub.cuh>
-# if (CCCL_MAJOR_VERSION >= 3 && CCCL_MINOR_VERSION >= 1)
+ // strided_iterator was added in CCCL 3.1
+# if (CCCL_MAJOR_VERSION > 3 || (CCCL_MAJOR_VERSION == 3 && CCCL_MINOR_VERSION >= 1))
# define STRIDED_ITERATOR_AVAILABLE
# include <cuda/iterator>
# endif