Commit b016f461b for llama.cpp

commit b016f461be57949f4d0749ec04caaee1c16c5b60
Author: Swigler <124839156+Swigler@users.noreply.github.com>
Date:   Wed Sep 30 20:30:21 2026 +0300

    convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324)

    * convert: fix LoRA conversion crash for Qwen3.5 V-head reorder

    _reorder_v_heads does reshape+permute+reshape to reorder V heads from
    grouped to tiled order.  LoraTorchTensor.reshape() cannot split its
    row dimension (A matrix), so converting Qwen3.5 LoRA adapters that
    target out_proj crashes with NotImplementedError.

    Fix: detect LoRA tensors and apply the equivalent index permutation
    directly — column reorder (dim=last) permutes A's columns, row
    reorder (dim=0) permutes B's rows.  This is mathematically identical:
      (B @ A)[:, perm] == B @ A[:, perm]
      (B @ A)[perm, :] == B[perm, :] @ A

    Verified: both paths produce exactly zero diff against the full-tensor
    reorder on random (rank=32, 4096×4096) matrices.

    Fixes #21125

    Signed-off-by: Radu Swigler <radu@swigler.com>

    * convert: add ty: ignore for hasattr-guarded LoRA call

    Assisted-By: Claude Opus 4.6 <noreply@anthropic.com>

    * fix comment

    * nowrap

    ---------

    Signed-off-by: Radu Swigler <radu@swigler.com>
    Co-authored-by: Radu Swigler <radu@swigler.com>
    Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
    Assisted-by: Claude Opus 4.6 <noreply@anthropic.com>

diff --git a/conversion/qwen.py b/conversion/qwen.py
index 6b87ff25e..95a41fb3a 100644
--- a/conversion/qwen.py
+++ b/conversion/qwen.py
@@ -469,6 +469,21 @@ class _LinearAttentionVReorderBase(Qwen3NextModel):
         shape = list(tensor.shape)
         if dim < 0:
             dim += len(shape)
+
+        # LoRA tensors (W ≈ B @ A) cannot reshape their row dimension.
+        # Instead, build a permutation index and apply it to A (column reorder) or B (row reorder) directly.
+        if hasattr(tensor, 'get_lora_A_B'):
+            n = shape[dim]
+            idx = torch.arange(n).reshape(num_k_heads, num_v_per_k, head_dim)
+            idx = idx.permute(1, 0, 2).contiguous().reshape(n)
+            lora_A, lora_B = tensor.get_lora_A_B()  # ty: ignore[call-non-callable]
+            if dim == len(shape) - 1:
+                return type(tensor)(lora_A[:, idx], lora_B)
+            elif dim == 0:
+                return type(tensor)(lora_A, lora_B[idx])
+            else:
+                raise NotImplementedError(f"_reorder_v_heads on dim={dim} not supported for LoRA tensors")
+
         new_shape = shape[:dim] + [num_k_heads, num_v_per_k, head_dim] + shape[dim + 1:]
         tensor = tensor.reshape(*new_shape)
         perm = list(range(len(new_shape)))