Commit b016f461b for llama.cpp
commit b016f461be57949f4d0749ec04caaee1c16c5b60
Author: Swigler <124839156+Swigler@users.noreply.github.com>
Date: Wed Sep 30 20:30:21 2026 +0300
convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324)
* convert: fix LoRA conversion crash for Qwen3.5 V-head reorder
_reorder_v_heads does reshape+permute+reshape to reorder V heads from
grouped to tiled order. LoraTorchTensor.reshape() cannot split its
row dimension (A matrix), so converting Qwen3.5 LoRA adapters that
target out_proj crashes with NotImplementedError.
Fix: detect LoRA tensors and apply the equivalent index permutation
directly — column reorder (dim=last) permutes A's columns, row
reorder (dim=0) permutes B's rows. This is mathematically identical:
(B @ A)[:, perm] == B @ A[:, perm]
(B @ A)[perm, :] == B[perm, :] @ A
Verified: both paths produce exactly zero diff against the full-tensor
reorder on random (rank=32, 4096×4096) matrices.
Fixes #21125
Signed-off-by: Radu Swigler <radu@swigler.com>
* convert: add ty: ignore for hasattr-guarded LoRA call
Assisted-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix comment
* nowrap
---------
Signed-off-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Assisted-by: Claude Opus 4.6 <noreply@anthropic.com>
diff --git a/conversion/qwen.py b/conversion/qwen.py
index 6b87ff25e..95a41fb3a 100644
--- a/conversion/qwen.py
+++ b/conversion/qwen.py
@@ -469,6 +469,21 @@ class _LinearAttentionVReorderBase(Qwen3NextModel):
shape = list(tensor.shape)
if dim < 0:
dim += len(shape)
+
+ # LoRA tensors (W ≈ B @ A) cannot reshape their row dimension.
+ # Instead, build a permutation index and apply it to A (column reorder) or B (row reorder) directly.
+ if hasattr(tensor, 'get_lora_A_B'):
+ n = shape[dim]
+ idx = torch.arange(n).reshape(num_k_heads, num_v_per_k, head_dim)
+ idx = idx.permute(1, 0, 2).contiguous().reshape(n)
+ lora_A, lora_B = tensor.get_lora_A_B() # ty: ignore[call-non-callable]
+ if dim == len(shape) - 1:
+ return type(tensor)(lora_A[:, idx], lora_B)
+ elif dim == 0:
+ return type(tensor)(lora_A, lora_B[idx])
+ else:
+ raise NotImplementedError(f"_reorder_v_heads on dim={dim} not supported for LoRA tensors")
+
new_shape = shape[:dim] + [num_k_heads, num_v_per_k, head_dim] + shape[dim + 1:]
tensor = tensor.reshape(*new_shape)
perm = list(range(len(new_shape)))