Commit 1a3011cc0 for llama.cpp
commit 1a3011cc0c183f184849b5fa4fa15ab0399f17d8
Author: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>
Date: Tue Oct 6 10:41:23 2026 +0200
llama : re-reserve the sched when the nextn extraction flags change (#30020)
The speculative MTP init enables NextN extraction on the target and draft
contexts after both were created and their schedulers reserved. With
unmasked extraction the trunk graph keeps every token through the last
layer instead of cropping to the output rows, so the first decode
reallocates to that batch's shape and the next, wider batch trips
GGML_SCHED_DEBUG_REALLOC. Invalidate the reserve when the flags change so
the next compute re-reserves with the new graph shape.
Assisted-by: Claude
diff --git a/src/llama-context.cpp b/src/llama-context.cpp
index 70b7c21af..ff2ea461c 100644
--- a/src/llama-context.cpp
+++ b/src/llama-context.cpp
@@ -1234,6 +1234,11 @@ void llama_context::set_embeddings(bool value) {
void llama_context::set_embeddings_nextn(bool value, bool masked) {
LLAMA_LOG_DEBUG("%s: value = %d, masked = %d\n", __func__, value, masked);
+ if (cparams.embeddings_nextn != value || cparams.embeddings_nextn_masked != masked) {
+ // these flags change the graph shape
+ sched_need_reserve = true;
+ }
+
cparams.embeddings_nextn = value;
cparams.embeddings_nextn_masked = masked;
}