Commit 60199339b for llama.cpp

commit 60199339bcff9092dd7273d371b52308c85a92de
Author: y198 <90976397+y198nt@users.noreply.github.com>
Date:   Wed Sep 16 18:03:11 2026 +0700

    rpc : invalidate cached compute graph when a referenced buffer is freed (#24292)

    The server caches the most recent compute graph per device so that
    GRAPH_RECOMPUTE can re-execute it without resending tensor data. The
    cached graph nodes hold direct pointers to backend buffers that were
    live at graph_compute() time. If any of those buffers is later
    released via FREE_BUFFER, the next GRAPH_RECOMPUTE re-executes the
    cached graph through the dangling pointers (use-after-free).

    The bug is reachable by an unauthenticated remote client. The
    dangling pointers point into chunks an attacker can reshape via
    subsequent ALLOC_BUFFER/SET_TENSOR commands, and the resulting
    read/write through the cached graph is sufficient to leak libc
    addresses and hijack the buffer iface vtable used by BUFFER_CLEAR,
    yielding remote code execution.

    Discard all cached graphs in free_buffer(). The existing null-check
    in graph_recompute() then rejects the request and the client falls
    back to GRAPH_COMPUTE on the next call.

    No protocol or API change.

diff --git a/ggml/src/ggml-rpc/ggml-rpc.cpp b/ggml/src/ggml-rpc/ggml-rpc.cpp
index adb88a245..c24caad77 100644
--- a/ggml/src/ggml-rpc/ggml-rpc.cpp
+++ b/ggml/src/ggml-rpc/ggml-rpc.cpp
@@ -1301,6 +1301,11 @@ bool rpc_server::free_buffer(const rpc_msg_free_buffer_req & request) {
         GGML_LOG_ERROR("[%s] buffer not found\n", __func__);
         return false;
     }
+    // Discard all cached graphs to avoid use-after-free in graph_recompute,
+    // since their nodes may hold pointers to the buffer being freed.
+    for (auto & sg : stored_graphs) {
+        sg.graph = nullptr;
+    }
     ggml_backend_buffer_free(buffer);
     buffers.erase(buffer);
     return true;
@@ -1752,7 +1757,6 @@ bool rpc_server::graph_compute(const std::vector<uint8_t> & input) {
         int64_t id;
         memcpy(&id, &nodes[i], sizeof(id));
         graph->nodes[i] = create_node(id, ctx, tensor_ptrs, tensor_map);
-
         // Check if create_node failed for a *non-zero* ID.
         // If id was 0, create_node returning nullptr is expected.
         // If id was non-zero and create_node returned nullptr, it indicates a deserialization error.