Commit b9a5a00b8 for llama.cpp

commit b9a5a00b86fd285a445916086a0b1dc35bee6d66
Author: Ravi Panchumarthy <ravi.panchumarthy@intel.com>
Date:   Mon Oct 5 23:57:04 2026 -0700

    ggml-openvino: fix CI tests; fix GPU regressions. (#30037)

    * ggml-openvino: skip unselected graph branches and support DUP

    Upstream #29622 adds a mixed token/embd branch to every input
    embedding graph through ggml_build_forward_select(). Its nodes are
    not flagged for compute, but the backend translated them anyway,
    and the DUP in that branch was unsupported, so the scheduler split
    the graph and passed the embeddings across the split with a fixed
    token count. The first single-token decode then failed
    (test-thread-safety on CPU and GPU).

    Build the OV model from the compute nodes only, and translate a
    same-type contiguous DUP like CONT so the graph stays on one backend.

    * ggml-openvino: make inp_scale_rows token dim dynamic

    #29622 also moves the per-token embedding scale (gemma3, gemma3n,
    gemma4) into a new [1, n_tokens] input. Give it a dynamic token dim
    and pad it per chunk on the static (NPU) path.

    * ggml-openvino: skip GPU MUL_MAT op tests with unbound Q4_1/Q4_K weights

    Op tests build Q4_1/Q4_K weights as u4 with an f16 zero point. The GPU
    plugin fails to compile that form for some row counts with "clFinish,
    error code: -5 CL_OUT_OF_RESOURCES", which aborts test-backend-ops on
    the MUL_MAT cases added in #29869 (e.g. m=1000, n=2, k=1024). Model
    weights use a u4 zero point and are not affected.

    Report these cases as unsupported on GPU until the plugin is fixed.
    Op tests check support before allocating, so the check matches unbound
    weights only; model loading probes with a dummy buffer and keeps its
    weights on the GPU.

    * ggml-openvino: create FILL in the output type

    translate_fill always built an f32 constant, so an f16 FILL produced
    f32 data and the copy back overran the f16 output buffer. Use the
    output type for the constant.

    * ggml-openvino: reject CONCAT with a quantized type

    Quantized inputs are dequantized when translated, so the backend cannot
    write a quantized CONCAT output. Report it as unsupported, as for CPY
    to a quantized type.

    * ggml-openvino: handle the single recurrent state gather of build_rs

    #29856 changed build_rs to gather all recurrent states with one GET_ROWS
    on the s_copy leaf and take the ubatch and extra states as views of it.
    The stateful path matched only the previous form, a GET_ROWS per view of
    s_copy, so Qwen3.5 failed with stateful execution on CPU and GPU
    ("is_axis_valid(axis, r)" in a Concat).

    For a single-slot cache, treat the GET_ROWS on the s_copy leaf as the
    active-state gather, keep the rank-4 layout of reshapes that read a view
    of it, and map the copy of the empty extra-state view to the single-slot
    remainder writeback. Do not warn about the dynamic dim of empty views.

    * openvino: align eltwise operand ranks to work around a GPU-plugin defect

    * openvino: match the MoE fusion on the rank-3 stateful graph

    * ggml-openvino: do not unsqueeze an RMS norm output in AlignEltwiseOperandRanks

    The pass unsqueezes the lower-rank operand of an Add/Multiply/Subtract
    whose operand ranks differ. In gemma-3 the lower-rank operand of the
    post-attention residual add is the norm output, and unsqueezing it makes
    the GPU plugin compute the layer wrongly: gemma-3 returns empty answers
    on GPU with stateful execution.

    Skip the rewrite when the lower-rank operand is an RMS norm output.

    * docs : update OpenVINO validated models

    ---------

    Co-authored-by: Mustafa Cavus <mustafa.cavus@intel.com>

diff --git a/docs/backend/OPENVINO.md b/docs/backend/OPENVINO.md
index c9bbc9a9e..decc588a9 100644
--- a/docs/backend/OPENVINO.md
+++ b/docs/backend/OPENVINO.md
@@ -113,13 +113,13 @@ Although, the validated models below were tested with `llama-cli` using the `Q4_
 | [bartowski/Qwen_Qwen3-1.7B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3-1.7B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
 | [Qwen/Qwen3-4B-Q4_K_M](https://huggingface.co/Qwen/Qwen3-4B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
 | [lm-kit/Qwen3-8B-Q4_K_M](https://huggingface.co/lm-kit/qwen-3-8b-instruct-gguf) | ✓ / ✓ | ✓ / ✓ | ✓ |
-| [bartowski/Qwen_Qwen3.5-0.8B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-0.8B-GGUF) | ✓ / ✓ | ✓ / ~ | ✗ |
-| [bartowski/Qwen_Qwen3.5-2B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-2B-GGUF) | ✓ / ✓ | ✓ / ~ | ✗ |
-| [bartowski/Qwen_Qwen3.5-4B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-4B-GGUF) | ✓ / ✓ | ✓ / ~ | ✗ |
-| [lmstudio-community/Qwen3.5-9B-Q4_K_M](https://huggingface.co/lmstudio-community/Qwen3.5-9B-GGUF) | ✓ / ✓ | ✓ / ~ | ✗ |
+| [bartowski/Qwen_Qwen3.5-0.8B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-0.8B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✗ |
+| [bartowski/Qwen_Qwen3.5-2B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-2B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✗ |
+| [bartowski/Qwen_Qwen3.5-4B-Q4_K_M](https://huggingface.co/bartowski/Qwen_Qwen3.5-4B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✗ |
+| [lmstudio-community/Qwen3.5-9B-Q4_K_M](https://huggingface.co/lmstudio-community/Qwen3.5-9B-GGUF) | ✓ / ✓ | ✓ / ✓ | ✗ |
 |  |  |  |  |
 | [unsloth/gemma-3-4b-it-Q4_K_M](https://huggingface.co/unsloth/gemma-3-4b-it-GGUF) | ✓ / ✓ | ✓ / ✓ | ✓ |
-| [bartowski/google_gemma-4-E2B-it-Q4_K_M](https://huggingface.co/bartowski/google_gemma-4-E2B-it-GGUF) | ✓ / ✓ | ✓ / ~ | ~ |
+| [bartowski/google_gemma-4-E2B-it-Q4_K_M](https://huggingface.co/bartowski/google_gemma-4-E2B-it-GGUF) | ✓ / ✓ | ✓ / ✓ | ~ |
 | [bartowski/google_gemma-4-E4B-it-Q4_K_M](https://huggingface.co/bartowski/google_gemma-4-E4B-it-GGUF) | ✓ / ✓ | ✗ / ✗ | ✓ |
 | [bartowski/gemma-4-12B-it-Q4_K_M](https://huggingface.co/bartowski/gemma-4-12B-it-GGUF) | ✓ / ✓ | ✓ / ✓ | ✗ |
 |  |  |  |  |
diff --git a/ggml/src/ggml-openvino/ggml-decoder.cpp b/ggml/src/ggml-openvino/ggml-decoder.cpp
index ae36a0959..10e82f3e6 100644
--- a/ggml/src/ggml-openvino/ggml-decoder.cpp
+++ b/ggml/src/ggml-openvino/ggml-decoder.cpp
@@ -336,7 +336,9 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
         }
         if (op_case == 1 && m_is_stateful) {
             // Recurrent convolution and GDN gates retain their rank-4 layout.
-            bool recurrent = src->op == GGML_OP_GET_ROWS && is_recurrent_cache(src->src[0]);
+            // The gathered states may come through a view of the one gather of all states (see build_rs).
+            const auto * gather = src->op == GGML_OP_VIEW ? src->src[0] : src;
+            bool recurrent = gather->op == GGML_OP_GET_ROWS && is_recurrent_cache(gather->src[0]);
             for (int i = 0; i < m_cgraph->n_nodes && !recurrent; ++i) {
                 const auto * consumer = m_cgraph->nodes[i];
                 if (consumer->op == GGML_OP_GATED_DELTA_NET) {
@@ -411,15 +413,19 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
         break;
     }
     case GGML_OP_GET_ROWS: {
-        if (node->src[1]->op == GGML_OP_VIEW) {
-            // GET_ROWS gathering recurrent state cache rows via the inp->s_copy index list:
-            // src[0] is a reshape of cache_r/cache_s, src[1] is a view of the s_copy leaf.
-            // op_case 1/2: active/extra rows of a multi-slot cache
-            // op_case 3/4: active/extra rows of a single-slot cache
-            if (node->src[0]->op == GGML_OP_RESHAPE && node->src[0]->src[0] != nullptr &&
-                is_recurrent_cache(node->src[0]->src[0])) {
-                const bool single_slot = node->src[0]->src[0]->ne[1] == 1;
+        // GET_ROWS gathering recurrent state cache rows via the inp->s_copy index list:
+        // src[0] is a reshape of cache_r/cache_s, src[1] is a view of the s_copy leaf, or the leaf itself
+        // when one gather covers all states (see build_rs).
+        // op_case 1/2: active/extra rows of a multi-slot cache
+        // op_case 3/4: active/extra rows of a single-slot cache
+        if (node->src[0]->op == GGML_OP_RESHAPE && node->src[0]->src[0] != nullptr &&
+            is_recurrent_cache(node->src[0]->src[0])) {
+            const bool single_slot = node->src[0]->src[0]->ne[1] == 1;
+            if (node->src[1]->op == GGML_OP_VIEW) {
                 op_case = (node->src[1]->view_offs == 0 ? 1 : 2) + (single_slot ? 2 : 0);
+            } else if (single_slot) {
+                // a single-slot cache holds exactly the one state the gather selects
+                op_case = 3;
             }
         }
         break;
@@ -579,6 +585,12 @@ int GgmlOvDecoder::compute_op_case(const ggml_tensor * node) const {
                        node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src == node->view_src) {
                 op_case = 4;
                 break;
+            } else if (node->src[0]->src[0]->op == GGML_OP_GET_ROWS && node->src[1] != nullptr &&
+                       node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src != nullptr &&
+                       is_recurrent_cache(node->src[1]->view_src) && node->src[1]->view_src->ne[1] == 1) {
+                // defrag remainder writeback of a single-slot cache, taken from a view of the one gather of
+                // all states (see build_rs)
+                op_case = 9;
             }
         } else if (node->src[0]->op == GGML_OP_GET_ROWS && node->src[1] != nullptr &&
                    node->src[1]->op == GGML_OP_VIEW && node->src[1]->view_src != nullptr &&
@@ -1097,6 +1109,9 @@ ov::PartialShape GgmlOvDecoder::get_graph_input_shape(const ggml_tensor * op,
         input_shape = m_is_static ? ov::PartialShape{1, 1, input->ne[1], m_prefill_chunk_size} :
                                     ov::PartialShape{1, 1, -1, -1};

+    } else if (is_inp_scale_rows(input, op)) {
+        input_shape = ov::PartialShape{1, 1, m_is_static ? (m_is_prefill ? m_prefill_chunk_size : 1) : -1, 1};
+
     } else if (is_inp_mask(input, op)) {
         // mask
         if (m_is_static) {
@@ -2101,6 +2116,10 @@ void GgmlOvDecoder::compute_node_dynamic_dims() {
                     m_node_dynamic_dims[src] = 0;
                     continue;
                 }
+                if (is_inp_scale_rows(src, node)) {
+                    m_node_dynamic_dims[src] = 1;
+                    continue;
+                }
                 if (node->op == GGML_OP_VIEW && src->op == GGML_OP_NONE && !is_stateful() && !m_model_is_splitted) {
                     m_node_dynamic_dims[src] = 1;
                     continue;
@@ -2181,8 +2200,11 @@ void GgmlOvDecoder::compute_node_dynamic_dims() {
                 }
                 if (m_node_dynamic_dims[node] != -1 && dynamic_dim_value != node->ne[m_node_dynamic_dims[node]]) {
                     m_node_dynamic_dims[node] = -1;
-                    GGML_LOG_WARN("ggml-openvino: dynamic dim value mismatch for VIEW node '%s', src[0]: '%s'\n",
-                                  node->name, node->src[0]->name);
+                    // an empty view, e.g. the extra states of a single-slot recurrent cache, always mismatches
+                    if (ggml_nelements(node) > 0) {
+                        GGML_LOG_WARN("ggml-openvino: dynamic dim value mismatch for VIEW node '%s', src[0]: '%s'\n",
+                                      node->name, node->src[0]->name);
+                    }
                 }
             }
             break;
@@ -2307,6 +2329,7 @@ void GgmlOvDecoder::compute_node_dynamic_dims() {
         case GGML_OP_DIAG:
         case GGML_OP_TRI:
         case GGML_OP_REPEAT:
+        case GGML_OP_DUP:
         // Shape-preserving elementwise ops: the dynamic dim is unchanged from src[0].
         // DIV/CLAMP are used in the MoE routing-weight normalization
         // (sum_rows -> clamp -> div). If they are left untracked here the dynamic
diff --git a/ggml/src/ggml-openvino/ggml-decoder.h b/ggml/src/ggml-openvino/ggml-decoder.h
index 33b95340a..a614d291a 100644
--- a/ggml/src/ggml-openvino/ggml-decoder.h
+++ b/ggml/src/ggml-openvino/ggml-decoder.h
@@ -393,6 +393,11 @@ public:
                op->src[0] != nullptr && op->src[0]->op != GGML_OP_NONE;
     }

+    // per-token embedding scale [1, n_tokens] (inp->scale_rows in llama-graph.cpp)
+    static bool is_inp_scale_rows(const ggml_tensor * tensor, const ggml_tensor * op) {
+        return op->op == GGML_OP_MUL && tensor == op->src[1] && strcmp(tensor->name, "inp_scale_rows") == 0;
+    }
+
     static bool is_rope_freqs_weight(const ggml_tensor * tensor, const ggml_tensor * op) {
         return op->op == GGML_OP_ROPE && tensor == op->src[2];
     }
diff --git a/ggml/src/ggml-openvino/ggml-openvino-extra.cpp b/ggml/src/ggml-openvino/ggml-openvino-extra.cpp
index 0257e23db..14a0b0382 100644
--- a/ggml/src/ggml-openvino/ggml-openvino-extra.cpp
+++ b/ggml/src/ggml-openvino/ggml-openvino-extra.cpp
@@ -131,6 +131,7 @@ void ggml_openvino_device_config::init() {
         "GGML_OPENVINO_DISABLE_KV_SLICE",
         "GGML_OPENVINO_ENABLE_FALLBACK",
         "GGML_OPENVINO_MANUAL_GQA_ATTN",
+        "GGML_OPENVINO_DISABLE_ELTWISE_RANK_ALIGN",
         "GGML_OPENVINO_MOE_OP",
         "GGML_OPENVINO_MEMORY_OPTIMIZE",
         "GGML_OPENVINO_RELEASE_WEIGHTS",
diff --git a/ggml/src/ggml-openvino/ggml-openvino.cpp b/ggml/src/ggml-openvino/ggml-openvino.cpp
index 87ee096fa..22b2f45e2 100644
--- a/ggml/src/ggml-openvino/ggml-openvino.cpp
+++ b/ggml/src/ggml-openvino/ggml-openvino.cpp
@@ -1369,6 +1369,10 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
         if (op->type == GGML_TYPE_I64) {
             return {false, "CONCAT with I64 type is not supported"};
         }
+        // quantized inputs are dequantized, so the output cannot be written in the quantized type
+        if (ggml_is_quantized(op->type)) {
+            return {false, "CONCAT with quantized type is not supported"};
+        }
         if (ggml_openvino_is_gpu() && op->type == GGML_TYPE_BF16 && has_view_op_input(op)) {
             return {false, "CONCAT with BF16 type and VIEW input is not supported on GPU"};
         }
@@ -1529,6 +1533,13 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
         }
         break;
     }
+    case GGML_OP_DUP: {
+        // translated as CONT, so only a plain copy
+        if (op->type != op->src[0]->type || !ggml_are_same_shape(op, op->src[0]) || !ggml_is_contiguous(op->src[0])) {
+            return {false, "DUP with type conversion or non-contiguous src is not supported"};
+        }
+        break;
+    }
     case GGML_OP_CPY: {
         if (op->src[0]->type != GGML_TYPE_BF16 && op->src[1]->type == GGML_TYPE_BF16) {
             return {false, "CPY with BF16 src[1] type is not supported"};
@@ -1568,6 +1579,14 @@ static ggml_openvino_op_support is_op_supported_case(const ggml_tensor * op) {
             (op->src[0]->buffer == nullptr || op->src[0]->buffer->usage != GGML_BACKEND_BUFFER_USAGE_WEIGHTS)) {
             return {false, "MUL_MAT scalar dot product with non-weight src[0] on GPU is not supported"};
         }
+        // The GPU plugin fails to compile u4 weights with an f16 zero point for some row counts
+        // (clFinish CL_OUT_OF_RESOURCES). Op tests build Q4_1/Q4_K weights in that form; model weights use a
+        // u4 zero point. Op tests check support before allocating, while model loading checks with a dummy
+        // buffer, so only unbound weights are excluded. Remove once the GPU plugin is fixed.
+        if (ggml_openvino_is_gpu() && (op->src[0]->type == GGML_TYPE_Q4_1 || op->src[0]->type == GGML_TYPE_Q4_K) &&
+            op->src[0]->buffer == nullptr) {
+            return {false, "MUL_MAT with unbound Q4_1/Q4_K src[0] on GPU is not supported"};
+        }
         if (op->src[0]->ne[3] != op->src[1]->ne[3] && op->src[0]->ne[3] != 1 && op->src[1]->ne[3] != 1) {
             return {false, "MUL_MAT with incompatible broadcast on ne[3]: src0->ne[3]=" + std::to_string(op->src[0]->ne[3]) +
                            ", src1->ne[3]=" + std::to_string(op->src[1]->ne[3])};
diff --git a/ggml/src/ggml-openvino/openvino/op/fill.cpp b/ggml/src/ggml-openvino/openvino/op/fill.cpp
index db2fecb53..87358e12e 100644
--- a/ggml/src/ggml-openvino/openvino/op/fill.cpp
+++ b/ggml/src/ggml-openvino/openvino/op/fill.cpp
@@ -20,7 +20,7 @@ OutputVector translate_fill(const NodeContext & context) {

     auto shape = context.get_input_shape(0).to_shape();

-    auto val = ov::op::v0::Constant::create(ov::element::f32, {}, {c});
+    auto val = ov::op::v0::Constant::create(context.get_output_type(), {}, {c});
     auto target_shape = ov::op::v0::Constant::create(ov::element::i64, {shape.size()},
         std::vector<int64_t>(shape.begin(), shape.end()));
     auto res = std::make_shared<ov::op::v3::Broadcast>(val, target_shape);
diff --git a/ggml/src/ggml-openvino/openvino/op_table.cpp b/ggml/src/ggml-openvino/openvino/op_table.cpp
index 12a0953e3..23b3a2cff 100644
--- a/ggml/src/ggml-openvino/openvino/op_table.cpp
+++ b/ggml/src/ggml-openvino/openvino/op_table.cpp
@@ -37,6 +37,7 @@ std::unordered_map<std::string, CreatorFunction> get_supported_ops() {
         {"GGML_OP_ADD_ID",          op::translate_add_id                           },
         {"GGML_OP_CONCAT",          op::translate_concat                           },
         {"GGML_OP_CONT",            op::translate_cont                             },
+        {"GGML_OP_DUP",             op::translate_cont                             },
         {"GGML_OP_DIV",             op::translate_div                              },
         {"GGML_OP_FILL",            op::translate_fill                             },
         {"GGML_OP_GET_ROWS",        op::translate_get_rows                         },
diff --git a/ggml/src/ggml-openvino/openvino/pass/align_eltwise_ranks.cpp b/ggml/src/ggml-openvino/openvino/pass/align_eltwise_ranks.cpp
new file mode 100644
index 000000000..eae5b2c4e
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/pass/align_eltwise_ranks.cpp
@@ -0,0 +1,96 @@
+#include "align_eltwise_ranks.h"
+
+#include <numeric>
+#include <openvino/op/add.hpp>
+#include <openvino/op/constant.hpp>
+#include <openvino/op/divide.hpp>
+#include <openvino/op/multiply.hpp>
+#include <openvino/op/sqrt.hpp>
+#include <openvino/op/subtract.hpp>
+#include <openvino/op/unsqueeze.hpp>
+#include <openvino/pass/pattern/op/wrap_type.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace pass {
+
+namespace {
+
+// True for an RMS-norm output, x * (1 / sqrt(mean(x^2) + eps)) as translate_rms_norm builds it, optionally
+// scaled by the norm weight.
+bool is_rms_norm_output(const ov::Output<ov::Node> & value, int depth = 1) {
+    const auto * node = value.get_node();
+    if (!ov::is_type<ov::op::v1::Multiply>(node)) {
+        return false;
+    }
+    for (const auto & input : node->input_values()) {
+        const auto * src = input.get_node();
+        if (ov::is_type<ov::op::v1::Divide>(src) && ov::is_type<ov::op::v0::Sqrt>(src->get_input_node_ptr(1))) {
+            return true;
+        }
+        if (depth > 0 && is_rms_norm_output(input, depth - 1)) {
+            return true;
+        }
+    }
+    return false;
+}
+
+}  // namespace
+
+AlignEltwiseOperandRanks::AlignEltwiseOperandRanks() {
+    auto eltwise_m = ov::pass::pattern::wrap_type<ov::op::v1::Add, ov::op::v1::Multiply, ov::op::v1::Subtract>();
+
+    const auto callback = [this](ov::pass::pattern::Matcher & m) {
+        auto node = m.get_match_root();
+        if (node->get_input_size() != 2) {
+            return false;
+        }
+
+        auto lhs = node->input_value(0);
+        auto rhs = node->input_value(1);
+
+        // A tensor-vs-Constant mismatch is the norm's own eps / 1-over-sqrt arithmetic. That
+        // folds into the `rms` primitive itself instead of becoming a fused eltwise post-op,
+        // so it is not affected and is left alone.
+        if (ov::is_type<ov::op::v0::Constant>(lhs.get_node()) ||
+            ov::is_type<ov::op::v0::Constant>(rhs.get_node())) {
+            return false;
+        }
+
+        const auto lhs_rank = lhs.get_partial_shape().rank();
+        const auto rhs_rank = rhs.get_partial_shape().rank();
+        if (lhs_rank.is_dynamic() || rhs_rank.is_dynamic() || lhs_rank == rhs_rank) {
+            return false;
+        }
+
+        const size_t shorter_idx = lhs_rank.get_length() < rhs_rank.get_length() ? 0 : 1;
+        const auto & shorter = shorter_idx == 0 ? lhs : rhs;
+
+        // The defect is with the norm output as the higher-rank operand. When the norm output is the
+        // lower-rank one (gemma-3 adds it to a rank-4 residual), unsqueezing it puts the Unsqueeze between
+        // `rms` and its post-op, and the GPU plugin then computes the layer wrongly.
+        if (is_rms_norm_output(shorter)) {
+            return false;
+        }
+        const int64_t diff = std::abs(lhs_rank.get_length() - rhs_rank.get_length());
+
+        std::vector<int64_t> axes(static_cast<size_t>(diff));
+        std::iota(axes.begin(), axes.end(), 0);
+        auto unsqueeze = std::make_shared<ov::op::v0::Unsqueeze>(
+            shorter, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{ axes.size() }, axes));
+
+        node->input(shorter_idx).replace_source_output(unsqueeze->output(0));
+        register_new_node(unsqueeze);
+        return true;
+    };
+
+    register_matcher(
+        std::make_shared<ov::pass::pattern::Matcher>(eltwise_m, "ov::frontend::ggml::pass::AlignEltwiseOperandRanks"),
+        callback);
+}
+
+}  // namespace pass
+}  // namespace ggml
+}  // namespace frontend
+}  // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/pass/align_eltwise_ranks.h b/ggml/src/ggml-openvino/openvino/pass/align_eltwise_ranks.h
new file mode 100644
index 000000000..04f59cce6
--- /dev/null
+++ b/ggml/src/ggml-openvino/openvino/pass/align_eltwise_ranks.h
@@ -0,0 +1,39 @@
+#pragma once
+
+#include <openvino/pass/matcher_pass.hpp>
+
+namespace ov {
+namespace frontend {
+namespace ggml {
+namespace pass {
+
+// Give a binary eltwise op's two operands the same rank, by unsqueezing leading axes onto the
+// shorter one.
+//
+// This is semantically a no-op -- NUMPY broadcasting already left-pads the lower-rank operand
+// with 1s, and the result rank is max(rank_a, rank_b) either way. It exists purely to work
+// around an OpenVINO GPU-plugin defect: an eltwise op whose operands differ in rank is computed
+// incorrectly once the plugin fuses it as a post-op into an `rms` primitive. gemma-4 dense hits
+// this under stateful execution, where the layer tail adds a rank-3 residual to the rank-4
+// RMS-norm output; the model then degenerates into repeated tokens on GPU while CPU is correct.
+// Equalising the ranks keeps the fusion and makes it compute the right answer.
+//
+// Deliberately a graph pass rather than something the op translators do: rank is load-bearing
+// during translation (several translators and later passes read operand ranks), and rewriting
+// operands mid-translate breaks the attention path. Running after the graph is complete avoids
+// that entirely.
+//
+// Only applies when both operands are real tensors -- a tensor-vs-Constant mismatch is the
+// RMS norm's own eps/rsqrt arithmetic, which folds into the `rms` primitive rather than
+// becoming a fused post-op, and is not affected by the defect. The norm output itself is never
+// unsqueezed: when it is the lower-rank operand (gemma-3), that breaks the fused path instead.
+class AlignEltwiseOperandRanks : public ov::pass::MatcherPass {
+public:
+    OPENVINO_MATCHER_PASS_RTTI("ov::frontend::ggml::pass::AlignEltwiseOperandRanks")
+    AlignEltwiseOperandRanks();
+};
+
+}  // namespace pass
+}  // namespace ggml
+}  // namespace frontend
+}  // namespace ov
diff --git a/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.cpp b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.cpp
index db8fc6133..3041a6ca7 100644
--- a/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.cpp
+++ b/ggml/src/ggml-openvino/openvino/pass/fuse_moe_compressed.cpp
@@ -321,22 +321,39 @@ FuseMoeCompressedFusedGateUp::FuseMoeCompressedFusedGateUp() {
     auto gate_up_w_m = any_input();
     auto ids_gate_up_m = any_input();
     auto bgm_fused_m = wrap_type<ov::op::internal::GatherMatmul>({ a_m, gate_up_w_m, ids_gate_up_m, any_input() });
-    auto gu_u_m = optional<ov::op::v0::Convert>({ wrap_type<ov::op::v0::Unsqueeze>(
-        { wrap_type<ov::op::v1::Transpose>({ bgm_fused_m, any_input() }), any_input() }) });
+    // Two graph shapes reach here. The rank-4 (stateless) graph restores the batch dim with an
+    // Unsqueeze after the Transpose and may Convert afterwards:
+    //     GatherMatmul -> Transpose -> Unsqueeze -> [Convert] -> Slice
+    // The rank-3 (stateful) graph never drops to a batch dim at all, so there is no Unsqueeze
+    // here - it appears later, between the GEGLU and the down Reshape - and the Convert sits on
+    // the other side of the Transpose:
+    //     GatherMatmul -> Convert -> Transpose -> Slice
+    // optional<T>({a, b}) matches T(a, b) or bare a, so one pattern covers both.
+    auto gu_t_m =
+        wrap_type<ov::op::v1::Transpose>({ optional<ov::op::v0::Convert>({ bgm_fused_m }), any_input() });
+    auto gu_u_m =
+        optional<ov::op::v0::Convert>({ optional<ov::op::v0::Unsqueeze>({ gu_t_m, any_input() }) });

     auto gate_slice_m = wrap_type<ov::op::v8::Slice>({ gu_u_m, any_input(), any_input(), any_input(), any_input() });
     auto up_slice_m = wrap_type<ov::op::v8::Slice>({ gu_u_m, any_input(), any_input(), any_input(), any_input() });
     auto gelu_m = wrap_type<ov::op::v7::Gelu>({ gate_slice_m });
     auto geglu_m = wrap_type<ov::op::v1::Multiply>({ gelu_m, up_slice_m });

+    // The rank-3 graph inserts the Unsqueeze the gate/up branch lacked right here, before the
+    // Reshape that feeds the down projection; the rank-4 graph goes straight from the GEGLU
+    // into the Reshape.
     auto d_t_m = wrap_type<ov::op::v1::Transpose>(
-        { optional<ov::op::v0::Convert>({ wrap_type<ov::op::v1::Reshape>({ geglu_m, any_input() }) }),
+        { optional<ov::op::v0::Convert>({ wrap_type<ov::op::v1::Reshape>(
+              { optional<ov::op::v0::Unsqueeze>({ geglu_m, any_input() }), any_input() }) }),
           any_input() });
     auto down_w_m = any_input();
     auto ids_down_m = any_input();
     auto bgm_down_m = wrap_type<ov::op::internal::GatherMatmul>({ d_t_m, down_w_m, ids_down_m, any_input() });
-    auto down_u_m = optional<ov::op::v0::Convert>({ wrap_type<ov::op::v0::Unsqueeze>(
-        { wrap_type<ov::op::v1::Transpose>({ bgm_down_m, any_input() }), any_input() }) });
+    // Same two shapes as the gate/up branch above.
+    auto down_t_m =
+        wrap_type<ov::op::v1::Transpose>({ optional<ov::op::v0::Convert>({ bgm_down_m }), any_input() });
+    auto down_u_m =
+        optional<ov::op::v0::Convert>({ optional<ov::op::v0::Unsqueeze>({ down_t_m, any_input() }) });

     // gemma-4 applies an extra per-expert output scale to the down projection before the
     // router-weight multiply (llama-graph.cpp's ffn_down_exps.scale); FuseMoeCompressed's
@@ -421,13 +438,20 @@ FuseMoeCompressedFusedGateUp::FuseMoeCompressedFusedGateUp() {
         }
         const size_t top_k = ids_pshape[ids_pshape.rank().get_length() - 1].get_length();

-        // routing weights arrive as [1, n_tokens, top_k, 1]; the op wants [..., top_k]
-        auto routing = pm.at(routing_m);
-        const auto routing_pshape = routing.get_partial_shape();
-        if (routing_pshape.rank().is_dynamic() || routing_pshape.rank().get_length() != 4 ||
-            routing_pshape[3] != 1) {
+        // Routing weights arrive as [1, n_tokens, top_k, 1] on the rank-4 (stateless) graph and
+        // as [n_tokens, top_k, 1] on the rank-3 (stateful) one, which is the same thing without
+        // the leading batch dim. Normalise the rank-3 form up to the rank-4 one so everything
+        // below - and the op's own config - stays in the shape that is already validated on the
+        // stateless path; the batch dim is taken back off the result at the end.
+        const auto routing_in = pm.at(routing_m);
+        const auto routing_pshape = routing_in.get_partial_shape();
+        const auto routing_rank = routing_pshape.rank();
+        if (routing_rank.is_dynamic() || (routing_rank.get_length() != 4 && routing_rank.get_length() != 3) ||
+            routing_pshape[routing_rank.get_length() - 1] != 1) {
             return false;
         }
+        const bool batchless = routing_rank.get_length() == 3;
+
         // Fold gemma-4's per-expert output scale into the routing weights: the reduction is
         // sum_e(routing[e] * scale[e] * down_out[e]), and MOECompressed only takes one
         // per-expert weight, so pre-multiply it into routing here (same [.., top_k, 1] shape).
@@ -435,7 +459,11 @@ FuseMoeCompressedFusedGateUp::FuseMoeCompressedFusedGateUp() {
         if (down_scale.get_partial_shape() != routing_pshape) {
             return false;
         }
-        routing = std::make_shared<ov::op::v1::Multiply>(routing, down_scale);
+        ov::Output<ov::Node> routing = std::make_shared<ov::op::v1::Multiply>(routing_in, down_scale);
+        if (batchless) {
+            routing = std::make_shared<ov::op::v0::Unsqueeze>(
+                routing, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{ 1 }, { 0 }));
+        }
         routing = std::make_shared<ov::op::v0::Squeeze>(
             routing, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{ 1 }, { 3 }));
         if (ids_pshape.rank().get_length() == 2) {
@@ -494,10 +522,19 @@ FuseMoeCompressedFusedGateUp::FuseMoeCompressedFusedGateUp() {
         auto moe = std::make_shared<ov::op::internal::MOECompressed>(args, config);

         ov::Output<ov::Node> result = moe->output(0);
+        // The op was fed the batched form, so it produces [1, n_tokens, hidden]. On the rank-3
+        // graph the ReduceSum being replaced is [n_tokens, hidden], so drop the batch dim again.
+        if (batchless) {
+            result = std::make_shared<ov::op::v0::Squeeze>(
+                result, ov::op::v0::Constant::create(ov::element::i64, ov::Shape{ 1 }, { 0 }));
+        }
         const auto root_type = m.get_match_root()->get_output_element_type(0);
         if (result.get_element_type() != root_type) {
             result = std::make_shared<ov::op::v0::Convert>(result, root_type);
         }
+        if (result.get_partial_shape() != m.get_match_root()->get_output_partial_shape(0)) {
+            return false;
+        }

         result.get_node_shared_ptr()->set_friendly_name(m.get_match_root()->get_friendly_name());
         ov::copy_runtime_info(m.get_matched_nodes(), result.get_node_shared_ptr());
diff --git a/ggml/src/ggml-openvino/openvino/translate_session.cpp b/ggml/src/ggml-openvino/openvino/translate_session.cpp
index 5a3d11f2d..ec3596de2 100644
--- a/ggml/src/ggml-openvino/openvino/translate_session.cpp
+++ b/ggml/src/ggml-openvino/openvino/translate_session.cpp
@@ -7,6 +7,7 @@
 #include "input_model.h"
 #include "pass/fuse_argsort_topk.h"
 #include "pass/fuse_moe_router.h"
+#include "pass/align_eltwise_ranks.h"
 #include "pass/fuse_moe_compressed.h"
 #include "pass/fuse_to_conv.h"
 #include "pass/kv_state_seq_axis.h"
@@ -498,6 +499,17 @@ std::shared_ptr<Model> TranslateSession::apply_transformations(std::shared_ptr<M
             manager.register_pass<pass::FuseMoeCompressedFusedGateUp>();
         }

+        // Workaround for an OpenVINO GPU-plugin defect: an eltwise op whose operands differ in
+        // rank is computed wrongly once the plugin fuses it as a post-op into `rms`. gemma-4
+        // dense under stateful execution adds a rank-3 residual to the rank-4 norm output and
+        // decodes garbage on GPU while CPU is correct. Equalising the ranks is a no-op for
+        // NUMPY broadcasting and makes the fused path correct.
+        // Remove once the plugin guards that fusion. Opt out with
+        // GGML_OPENVINO_DISABLE_ELTWISE_RANK_ALIGN=1.
+        if (ggml_openvino_is_gpu() && !ggml_openvino_getenv_int("GGML_OPENVINO_DISABLE_ELTWISE_RANK_ALIGN")) {
+            manager.register_pass<pass::AlignEltwiseOperandRanks>();
+        }
+
         if (ggml_model_decoder->is_stateful()) {
             const auto kv_param_res_names = ggml_model_decoder->get_kv_param_res_names();
             const auto kv_param_res_pairs = get_kv_param_res_pairs(model, kv_param_res_names);
diff --git a/ggml/src/ggml-openvino/utils.cpp b/ggml/src/ggml-openvino/utils.cpp
index 67a4123cb..f07e180a4 100644
--- a/ggml/src/ggml-openvino/utils.cpp
+++ b/ggml/src/ggml-openvino/utils.cpp
@@ -676,6 +676,15 @@ ov::Tensor get_ov_input_tensor_static_prefill(const std::shared_ptr<GgmlOvDecode
         return input_tensor;
     }

+    if (GgmlOvDecoder::is_inp_scale_rows(ggml_tensor, op)) {
+        ov::Tensor input_tensor(ov::element::f32, ov::Shape{1, 1, chunk_size, 1});
+        auto * dst = input_tensor.data<float>();
+        const auto * src = static_cast<const float *>(ggml_tensor->data) + chunk_index * chunk_size;
+        std::copy(src, src + chunk_valid_size, dst);
+        std::fill(dst + chunk_valid_size, dst + chunk_size, 1.0f);
+        return input_tensor;
+    }
+
     if (GgmlOvDecoder::is_inp_mean(ggml_tensor, op)) {
         const size_t n_seqs = ggml_tensor->ne[1];
         const size_t src_stride = ggml_tensor->ne[0];
@@ -1726,6 +1735,25 @@ enum ggml_status ov_graph_compute_static(ggml_cgraph * cgraph, const std::shared
 }
 }  // namespace

+// Nodes on the unselected branches of ggml_build_forward_select() stay in the graph but must not be
+// computed. Keep them out of the OV model, or their inputs become parameters with fixed shapes.
+static ggml_cgraph * get_compute_graph(ggml_cgraph * cgraph, ov_runtime_context & r_ctx) {
+    auto is_skipped = [](const ggml_tensor * node) {
+        return node->op != GGML_OP_NONE && !(node->flags & GGML_TENSOR_FLAG_COMPUTE);
+    };
+    if (std::none_of(cgraph->nodes, cgraph->nodes + cgraph->n_nodes, is_skipped)) {
+        return cgraph;
+    }
+    auto & compute = r_ctx.compute_graphs[cgraph];
+    compute.nodes.clear();
+    std::copy_if(cgraph->nodes, cgraph->nodes + cgraph->n_nodes, std::back_inserter(compute.nodes),
+                 [&](const ggml_tensor * node) { return !is_skipped(node); });
+    compute.graph = *cgraph;
+    compute.graph.nodes = compute.nodes.data();
+    compute.graph.n_nodes = (int) compute.nodes.size();
+    return &compute.graph;
+}
+
 // Both execution paths use two cache levels:
 // 1. Reuse this backend's decoder/request via graph_key and compatibility checks.
 // 2. On a local miss, look up compiled_graph_key in the shared compilation cache,
@@ -1744,6 +1772,7 @@ enum ggml_status ov_graph_compute(ggml_cgraph * cgraph, ggml_backend_t backend)
         GGML_ASSERT(ctx->runtime_context != nullptr);
         std::shared_ptr<ov_runtime_context> r_ctx = std::static_pointer_cast<ov_runtime_context>(ctx->runtime_context);
         std::lock_guard<std::mutex> execution_lock(r_ctx->execution_mutex);
+        cgraph = get_compute_graph(cgraph, *r_ctx);

         return is_static ? ov_graph_compute_static(cgraph, r_ctx) : ov_graph_compute_dynamic(cgraph, r_ctx);
     } catch (const ov::Exception & e) {
diff --git a/ggml/src/ggml-openvino/utils.h b/ggml/src/ggml-openvino/utils.h
index 491dbadc2..c34441f5e 100644
--- a/ggml/src/ggml-openvino/utils.h
+++ b/ggml/src/ggml-openvino/utils.h
@@ -119,6 +119,12 @@ struct ov_runtime_context {
     std::unordered_map<graph_key, std::vector<std::string>, graph_key_hash> ov_output_names_cache;
     size_t stateful_kv_size;
     std::map<std::string, std::string> kv_state_input_name_map;
+    // compute-only copies of graphs that carry unselected ggml_build_forward_select() branches
+    struct compute_graph {
+        ggml_cgraph graph;
+        std::vector<ggml_tensor *> nodes;
+    };
+    std::unordered_map<const ggml_cgraph *, compute_graph> compute_graphs;

     ov_runtime_context() : device("CPU"), stateful(false), stateful_kv_size(0) {}