Commit c479922ac for llama.cpp

commit c479922ac520a08969b4c1dc154d7bbb3c386d85
Author: Max Krasnyansky <maxk@qti.qualcomm.com>
Date:   Tue Oct 6 15:24:17 2026 -0700

    hexagon: CPY/CONCAT/CONT/DUP overhaul to use DMA/HVX for all cases (#30067)

    * hex-cpy: replace more paths with dma and simplify l2flush

    * hex-cpy: use DMA in all sametype paths

    * hex-cpy: rewrite the rest of the copy paths (diff type) to use dma

    * hex-concat: use dma for multi-dev path which also removes the need for l2-line alignment

    * hex-concat: proper support for mdev splitting

    * hex-concat: cleanup ctx and kern params usage

    * hex-cpy: cleanup contex and remove left-over non-dma checks

    * hex-cpy: clean dma_cpy naming

    * hex-cpy: proper kernel params and kernel selection

    * hex-build: resolve left-over rebase conflicts

    * hex-cpy: update dev guide to clarify 128 byte alignment requirement

    * hex-concat: make sure we go through mdev barrier

    * hex-concat: make sure to flush dma-queue

    * hex-cpy/concat: cleanup kparams and vtcm layout handling

    * hex-cpy: remove dead check for contig (routed to diff kernel) and update comments

    * hex-dev: update developer guide based on latest changes

    * hex-cpy: safe skip of noop copies

    * hex-dup: route DUP to CPY

diff --git a/docs/backend/snapdragon/developer.md b/docs/backend/snapdragon/developer.md
index 378d47653..3608250dc 100644
--- a/docs/backend/snapdragon/developer.md
+++ b/docs/backend/snapdragon/developer.md
@@ -37,9 +37,10 @@ In llama.cpp/GGML, each Hexagon session is mapped to a single GGML backend devic
 `GGML_HEXAGON_DEVICES`, or `HTP0`, `HTP1` in legacy mode).

 To support running models larger than 3.5GB on a single device, the Hexagon backend dynamically maps and unmaps buffers:
-- Buffers are allocated in shared DDR (RPCMEM) via file descriptors (`fastrpc_mmap` using `FASTRPC_MAP_FD_DELAYED`).
+- Buffers are allocated in shared DDR (RPCMEM) and mapped through FastRPC file descriptors. Non-pinned buffers use delayed
+  mappings (`FASTRPC_MAP_FD_DELAYED` or `FASTRPC_MAP_FD_DELAYED_EXTENDED`).
 - Pinned buffers (such as KV cache and active compute buffers) remain mapped throughout execution.
-- Inactive weight buffers are dynamically mapped into the NPU session via `HAP_mmap()` during batch buffer preparation
+- Inactive weight buffers are dynamically mapped into the NPU session during batch buffer preparation
   (`prep_op_bufs()` in `htp/main.c`) and unmapped via `htp_iface_munmap()` when no longer needed by the active batch.
 - This dynamic sliding window allows a single NPU session to execute models that exceed the 3.5GB window.

@@ -55,6 +56,9 @@ Writing high-performance operators for Hexagon requires following specific guide

 - Strongly prefer the `DDR -> DMA -> VTCM -> compute (HVX/HMX) -> VTCM -> DMA -> DDR` data flow.
 - Direct HVX reads/writes from/to DDR are less efficient and should only be used as a fallback.
+- Use `dma_addr_t` only for DMA base and final addresses. Form a final address by adding a 32-bit byte offset to a
+  `dma_addr_t` tensor base address. This permits a 64-bit mapped base address on newer platforms while retaining 32-bit
+  relative addressing.
 - The DMA queue is a strict FIFO where operations must be pushed and popped in strict order.
 - Follow the pipelined multi-buffering sequence properly (typically 2x to 16x buffering) so every push has a corresponding pop:

@@ -66,7 +70,7 @@ Writing high-performance operators for Hexagon requires following specific guide
 - Because every push must be matched by a pop, `dma_queue_flush()` is not required when the pipeline sequence is followed
   properly. Flushing is only used in rare exceptions where a batch of operations is pushed without individual pops.
 - Use the DMA queue interface from [`dma-queue.h`](../../../ggml/src/ggml-hexagon/htp/dma-queue.h)
-  (`dma_queue_push_ddr_to_vtcm`, `dma_queue_pop`, `dma_queue_push_vtcm_to_ddr`).
+  (`dma_queue_push()`, `dma_queue_pop()`, and `dma_queue_flush()`).
   See [`cumsum-ops.c`](../../../ggml/src/ggml-hexagon/htp/cumsum-ops.c) and
   [`act-ops.c`](../../../ggml/src/ggml-hexagon/htp/act-ops.c) for reference implementations.

@@ -125,7 +129,6 @@ Writing high-performance operators for Hexagon requires following specific guide
 - Do not add defensive NULL checks or assertions for internal framework pointers or required graph operands and outputs.
   Internal pointers include `ctx`, `octx`, local context structs like `*ctx`, `kparams`, and worker callback `data`.
 - These pointers are architectural invariants during kernel execution and host-side graph preparation.
-  Graph compute receives allocated nodes with valid required `node->src[N]` and `node->data` pointers.
 - Do not turn an invariant violation into an unsupported operation or missed fusion.
   Checks such as `if (!octx || !octx->ctx)` clutter the code, obscure intent, and hide upstream errors.
 - **Distinction**: `octx->src[N]` pointers *can* be NULL by design and must be checked when optional.
@@ -177,26 +180,28 @@ sessions.

 - Shared tensor buffers reside in DDR (RPCMEM) with a 128-byte cache line granularity
   (`HEX_L2_LINE_SIZE` = 128 bytes, `HTP_TENSOR_MDEV_LINE_SIZE`).
-- **Rule**: Multi-device work partitions must align destination write regions to 128-byte cache line boundaries so distinct
-  devices never share or overwrite the same cache line.
+- **Rule**: Multi-device work partitions that write directly to DDR through HVX/L2 must align destination write regions to
+  128-byte cache line boundaries so distinct devices never share or overwrite the same cache line.
+- DMA writes to DDR are not subject to this cache-line ownership rule. They may use smaller non-overlapping destination
+  ranges when the operator only writes through DMA.

 ### Partitioning Helpers in `htp-tensor.h`

 Common partitioning logic is factored into reusable inline helpers in
 [`htp-tensor.h`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h):

-1. [`htp_tensor_mdev_rows_per_chunk`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L67):
+1. [`htp_tensor_mdev_rows_per_chunk`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L71):
    Determines the minimum number of rows per chunk so that the chunk byte size is a multiple of 128 bytes:

    ```
    rows_per_chunk = 128 / hex_gcd_u32(row_size, 128)
    ```

-   If row stride `nb[1]` is already a multiple of 128 bytes, `rows_per_chunk = 1`.
+   If the active row and outer strides are already multiples of 128 bytes, `rows_per_chunk = 1`.
    Returns `false` if the tensor cannot be safely row-partitioned (such as unaligned base pointer, permuted layout,
    or non-128-byte aligned outer strides).

-2. [`htp_tensor_mdev_partition`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L94):
+2. [`htp_tensor_mdev_partition`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L98):
    Calculates the per-device work range `struct htp_tensor_mdev_range { uint32_t start; uint32_t count; }` given
    `total_units`, `units_per_chunk`, `mdev_idx`, `mdev_count`, and the precomputed `mdev_count_div`.
    Handles chunk distribution across devices, assigns remainder units to the last device, and automatically triggers
@@ -204,11 +209,10 @@ Common partitioning logic is factored into reusable inline helpers in

 ### Row-Partitioned Operators

-For row-wise operators
+For row-wise operators that write directly to DDR
 (such as activations in [`act-ops.c`](../../../ggml/src/ggml-hexagon/htp/act-ops.c),
 binary ops in [`binary-ops.c`](../../../ggml/src/ggml-hexagon/htp/binary-ops.c),
-unary ops in [`unary-ops.c`](../../../ggml/src/ggml-hexagon/htp/unary-ops.c), and
-sameshape copies in [`cpy-ops.c`](../../../ggml/src/ggml-hexagon/htp/cpy-ops.c)):
+and unary ops in [`unary-ops.c`](../../../ggml/src/ggml-hexagon/htp/unary-ops.c)):

 ```c
 const uint32_t total_rows   = ne01 * ne02 * ne03;
@@ -233,20 +237,19 @@ if (nrows == 0) {

 ### Element-Partitioned Operators

-For flat element-wise operations (such as reshape copies in
-[`cpy-ops.c`](../../../ggml/src/ggml-hexagon/htp/cpy-ops.c)):
+For flat element-wise operations that write directly to DDR:
 - Partition total linear elements N = ne0 * ne1 * ne2 * ne3 in 128-byte cache line chunks (`elems_per_line = (elem_size == 4) ? 32 : 64`).
 - Requires strict 1D contiguity:
-  [`htp_tensor_is_contiguous(dst, elem_size)`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L28)
+  [`htp_tensor_is_contiguous(dst, elem_size)`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L32)
   and 128-byte aligned destination pointer
-  [`htp_tensor_mdev_data_aligned(dst)`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L47).
+  [`htp_tensor_mdev_data_aligned(dst)`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L51).
 - If contiguous and aligned, pass `elems_per_line` to
-  [`htp_tensor_mdev_partition`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L94);
+  [`htp_tensor_mdev_partition`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L98);
   otherwise pass 0 to trigger Device 0 fallback.

 ### Single-Device Fallback (Device 0)

-- Fallback to Device 0 (`mdev.idx == 0`) when partitioning would cause cache line tearing or when work cannot be evenly distributed.
+- Fallback to Device 0 (`mdev.idx == 0`) when partitioning would cause cache line tearing or there are too few aligned chunks.
 - Triggers:
   1. Destination tensor cannot be safely partitioned (`rows_per_chunk == 0` or non-contiguous/unaligned buffer).
   2. Total aligned chunks < `mdev_count`.
@@ -303,7 +306,7 @@ Multi-device execution synchronizes worker sessions through atomic fence slots a
   (Input Prep)                                 (Input Prep)
        |                                            |
   Pre-Op Barrier ----------------------------- Pre-Op Barrier
-  (mdev_sync_fence)                            (mdev_sync_fence)
+  (htp_mdev_group_barrier)                     (htp_mdev_group_barrier)
        |                                            |
   Kernel Execution                             Kernel Execution
   (Output Slice 0)                             (Output Slice 1)
@@ -326,10 +329,10 @@ Multi-device execution synchronizes worker sessions through atomic fence slots a
   atomic_uint * my_fence = htp_mdev_fence_slot(fence_base, mdev_idx);
   ```

-- **Writing to fence ([`htp_fence_write`](../../../ggml/src/ggml-hexagon/htp/htp-fence.h#L18))**:
+- **Writing to fence ([`htp_fence_write`](../../../ggml/src/ggml-hexagon/htp/htp-fence.h#L17))**:
   Stores `seq` and `status`, issues a `syncht` thread synchronization barrier, and flushes/invalidates the line
   using `Q6_dccleaninva_A(fence)`.
-- **Reading from peer fence ([`htp_fence_read`](../../../ggml/src/ggml-hexagon/htp/htp-fence.h#L26))**:
+- **Reading from peer fence ([`htp_fence_read`](../../../ggml/src/ggml-hexagon/htp/htp-fence.h#L25))**:
   Executes `Q6_dccleaninva_A(fence)` and `syncht` before reading atomic values to ensure fresh data from DDR.

 ### Deterministic Monotonic Sequence Numbers
@@ -348,7 +351,7 @@ Multi-device execution synchronizes worker sessions through atomic fence slots a

 - In the kernel, ensure all pushed DMA operations have been popped in strict FIFO order to drain the queue.
 - Use [`htp_tensor_flush_all()`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h) to flush specific dirty tensors back to DDR:
-  - [`htp_tensor_flush_all()`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h) flushes only modified tensor address ranges,
-    ensuring peer devices and the host CPU observe consistent data in DDR.
+  - [`htp_tensor_flush_all()`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h) flushes modified tensor address ranges, or the
+    full D-cache when their total size exceeds the flush threshold, ensuring peer devices and the host CPU observe consistent
+    data in DDR.
 - Never signal completion before all DMA transfers are drained and dirty tensor flushes have completed.
-
diff --git a/ggml/src/ggml-hexagon/ggml-hexagon.cpp b/ggml/src/ggml-hexagon/ggml-hexagon.cpp
index 1e15e2eb7..8c650b85c 100644
--- a/ggml/src/ggml-hexagon/ggml-hexagon.cpp
+++ b/ggml/src/ggml-hexagon/ggml-hexagon.cpp
@@ -71,6 +71,8 @@
 #include "htp/ssm-conv.h"
 #include "htp/gated-delta-net-ops.h"
 #include "htp/argsort-ops.h"
+#include "htp/concat-ops.h"
+#include "htp/cpy-ops.h"
 #include "htp_iface.h"
 #include "htp-drv.h"

@@ -400,6 +402,18 @@ static void ggml_hexagon_precompute_pool_2d_params(
     bool is_pool_1d
 );

+static bool ggml_hexagon_precompute_concat_params(
+    const struct ggml_hexagon_session * sess,
+    const struct ggml_tensor * op,
+    struct htp_concat_kernel_params * kparams
+);
+
+static bool ggml_hexagon_precompute_cpy_params(
+    const struct ggml_hexagon_session * sess,
+    const struct ggml_tensor * op,
+    struct htp_copy_kernel_params * kparams
+);
+
 static void ggml_hexagon_precompute_fused_mmnx_params(
     const struct ggml_hexagon_session * sess,
     const struct ggml_tensor * src0,
@@ -4260,6 +4274,11 @@ void ggml_hexagon_session::enqueue_cpy(const ggml_tensor * src, ggml_tensor * ds
     if (with_fence) {
         cpy_node.name = "CPY+FENCE";
     }
+    const bool ok = ggml_hexagon_precompute_cpy_params(this, node, (struct htp_copy_kernel_params *) cpy_node.kernel_params);
+    const auto * kparams = (const struct htp_copy_kernel_params *) cpy_node.kernel_params;
+    if (ok && !with_fence && kparams->total_elems == 0) {
+        return;
+    }
     this->enqueue_op(cpy_node);
 }

@@ -6350,6 +6369,208 @@ static void ggml_hexagon_precompute_pool_2d_params(
     kparams->inv_kernel_area = 1.0f / (float) (kparams->kernel_x * kparams->kernel_y);
 }

+static bool ggml_hexagon_precompute_concat_params(
+    const struct ggml_hexagon_session * sess,
+    const struct ggml_tensor * op,
+    struct htp_concat_kernel_params * kparams
+) {
+    memset(kparams, 0, sizeof(*kparams));
+    kparams->kernel_type = HTP_CONCAT_KERNEL_UNSUPPORTED;
+
+    const struct ggml_tensor * src0 = op->src[0];
+    const struct ggml_tensor * src1 = op->src[1];
+    const struct ggml_tensor * dst  = op;
+
+    if (!src0 || !src1 || !dst) {
+        return false;
+    }
+
+    int dim = ((const int32_t *) op->op_params)[0];
+    if (dim < 0 || dim >= GGML_MAX_DIMS) {
+        return false;
+    }
+    kparams->dim = dim;
+
+    if (dst->type != GGML_TYPE_F32 && dst->type != GGML_TYPE_F16 && dst->type != GGML_TYPE_I32) {
+        return false;
+    }
+    if (src0->type != dst->type || src1->type != dst->type) {
+        return false;
+    }
+
+    const uint32_t type_size = ggml_type_size(dst->type);
+
+    for (int d = 0; d < GGML_MAX_DIMS; d++) {
+        const int64_t ne_d = (d == dim) ? src0->ne[d] + src1->ne[d] : src0->ne[d];
+        if (dst->ne[d] != ne_d || (d != dim && src1->ne[d] != dst->ne[d])) {
+            return false;
+        }
+    }
+
+    const bool dma_strides_ok = (src0->nb[0] == type_size && src1->nb[0] == type_size && dst->nb[0] == type_size);
+
+    if (dma_strides_ok) {
+        kparams->kernel_type = HTP_CONCAT_KERNEL_REGULAR;
+        kparams->n_threads   = 1;
+        return true;
+    }
+
+    const bool is_src1_transposed = (src1->nb[0] > src1->nb[1]);
+    const bool is_src0_transposed = (src0->nb[0] > src0->nb[1]);
+    const bool transposed_rows_ok = (src0->nb[0] == type_size && src1->nb[1] == type_size && dst->nb[0] == type_size);
+
+    if (dim == 0 && is_src1_transposed && !is_src0_transposed && transposed_rows_ok &&
+        (dst->type == GGML_TYPE_F32 || dst->type == GGML_TYPE_F16)) {
+
+        const uint32_t n_threads = sess->n_threads > 0 ? (uint32_t) sess->n_threads : 8;
+        struct htp_concat_transposed_vtcm_layout layout;
+        htp_concat_transposed_vtcm_layout_build(&layout, (uint32_t) src0->ne[0], (uint32_t) src1->ne[0], type_size, n_threads);
+
+        if (sess->vtcm_size > 0 && layout.total_bytes > sess->vtcm_size) {
+            return false;
+        }
+
+        kparams->kernel_type           = HTP_CONCAT_KERNEL_TRANSPOSED;
+        kparams->n_threads             = n_threads;
+        kparams->vtcm_size             = layout.total_bytes;
+        kparams->spad0_size_per_thread = layout.src0_spad_size_per_thread;
+        kparams->spad1_size_per_thread = layout.src1_spad_size_per_thread;
+        return true;
+    }
+
+    return false;
+}
+
+static bool ggml_hexagon_precompute_cpy_params(
+    const struct ggml_hexagon_session * sess,
+    const struct ggml_tensor * op,
+    struct htp_copy_kernel_params * kparams
+) {
+    memset(kparams, 0, sizeof(*kparams));
+    kparams->kernel_type = HTP_COPY_KERNEL_UNSUPPORTED;
+
+    const struct ggml_tensor * src0 = op->src[0];
+    const struct ggml_tensor * dst  = op;
+
+    if (!src0 || !dst) {
+        return false;
+    }
+
+    if (src0->type != GGML_TYPE_F32 && src0->type != GGML_TYPE_F16 && src0->type != GGML_TYPE_I32) {
+        return false;
+    }
+    if (dst->type != GGML_TYPE_F32 && dst->type != GGML_TYPE_F16 && dst->type != GGML_TYPE_I32) {
+        return false;
+    }
+
+    const int64_t nelem_src = ggml_nelements(src0);
+    const int64_t nelem_dst = ggml_nelements(dst);
+    if (nelem_src != nelem_dst || nelem_src < 0) {
+        return false;
+    }
+
+    const uint32_t src_type_size = ggml_type_size(src0->type);
+    const uint32_t dst_type_size = ggml_type_size(dst->type);
+
+    kparams->src0_type_size = (uint8_t) src_type_size;
+    kparams->dst_type_size  = (uint8_t) dst_type_size;
+    kparams->total_elems    = (uint32_t) nelem_src;
+
+    if (nelem_src == 0) {
+        kparams->kernel_type = HTP_COPY_KERNEL_1D_CONTIG;
+        return true;
+    }
+
+    if (nelem_src == 1) {
+        if (src0->type == dst->type) {
+            kparams->kernel_type = HTP_COPY_KERNEL_SCALAR;
+            return true;
+        }
+        if ((src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_I32) ||
+            (src0->type == GGML_TYPE_I32 && dst->type == GGML_TYPE_F32) ||
+            (src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_F16) ||
+            (src0->type == GGML_TYPE_F16 && dst->type == GGML_TYPE_F32)) {
+            kparams->kernel_type = HTP_COPY_KERNEL_SCALAR;
+            return true;
+        }
+        return false;
+    }
+
+    const bool sametype = (src0->type == dst->type);
+
+    bool same_extents = true;
+    for (int d = 0; d < GGML_MAX_DIMS; d++) {
+        if (src0->ne[d] != dst->ne[d]) {
+            same_extents = false;
+            break;
+        }
+    }
+
+    const bool transposed = (src0->nb[0] > src0->nb[1])    || (dst->nb[0] > dst->nb[1]) ||
+                            (src0->nb[0] != src_type_size) || (dst->nb[0] != dst_type_size) ||
+                            (src0->nb[1] < (size_t) src0->ne[0] * src_type_size) || (dst->nb[1] < (size_t) dst->ne[0] * dst_type_size);
+    const bool sameshape  = same_extents && !transposed;
+
+    const bool src_is_contiguous = ggml_is_contiguous(src0);
+    const bool dst_is_contiguous = ggml_is_contiguous(dst);
+
+    if (sametype) {
+        if (src_is_contiguous && dst_is_contiguous) {
+            kparams->kernel_type = HTP_COPY_KERNEL_1D_CONTIG;
+            return true;
+        }
+
+        if (sameshape) {
+            kparams->kernel_type = HTP_COPY_KERNEL_SAMESHAPE_SAMETYPE;
+            kparams->total_rows  = (uint32_t) (src0->ne[1] * src0->ne[2] * src0->ne[3]);
+            return true;
+        }
+
+        kparams->kernel_type = HTP_COPY_KERNEL_RESHAPE;
+        kparams->n_threads   = sess->n_threads > 0 ? (uint8_t) sess->n_threads : 4;
+        kparams->u.reshape.div_ne0            = init_fastdiv_values((uint32_t) dst->ne[0]);
+        kparams->u.reshape.div_ne1_ne0        = init_fastdiv_values((uint32_t) (dst->ne[1] * dst->ne[0]));
+        kparams->u.reshape.div_ne2_ne1_ne0    = init_fastdiv_values((uint32_t) (dst->ne[2] * dst->ne[1] * dst->ne[0]));
+        kparams->u.reshape.div_ne00           = init_fastdiv_values((uint32_t) src0->ne[0]);
+        kparams->u.reshape.div_ne01_ne00      = init_fastdiv_values((uint32_t) (src0->ne[1] * src0->ne[0]));
+        kparams->u.reshape.div_ne02_ne01_ne00 = init_fastdiv_values((uint32_t) (src0->ne[2] * src0->ne[1] * src0->ne[0]));
+        return true;
+    }
+
+    if (!sameshape) {
+        return false;
+    }
+
+    const bool valid_conversion = (src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_F16) ||
+                                  (src0->type == GGML_TYPE_F16 && dst->type == GGML_TYPE_F32) ||
+                                  (src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_I32) ||
+                                  (src0->type == GGML_TYPE_I32 && dst->type == GGML_TYPE_F32);
+    if (!valid_conversion) {
+        return false;
+    }
+
+    const uint32_t n_threads = sess->n_threads > 0 ? (uint32_t) sess->n_threads : 4;
+    struct htp_copy_convert_vtcm_layout layout;
+    htp_copy_convert_vtcm_layout_build(&layout, (uint32_t) src0->ne[0], src_type_size, dst_type_size, n_threads);
+
+    if (sess->vtcm_size > 0 && layout.total_bytes > sess->vtcm_size) {
+        return false;
+    }
+
+    kparams->kernel_type             = HTP_COPY_KERNEL_SAMESHAPE_CONVERT;
+    kparams->total_rows              = (uint32_t) (src0->ne[1] * src0->ne[2] * src0->ne[3]);
+    kparams->n_threads               = (uint8_t) n_threads;
+    kparams->vtcm_size               = layout.total_bytes;
+    kparams->u.convert.src0_buf_size = layout.src0_buf_size;
+    kparams->u.convert.dst_buf_size  = layout.dst_buf_size;
+    kparams->u.convert.spad0_size_per_thread = layout.spad0_size_per_thread;
+    kparams->u.convert.spad1_size_per_thread = layout.spad1_size_per_thread;
+    kparams->u.convert.div_ne01      = init_fastdiv_values((uint32_t) src0->ne[1]);
+    kparams->u.convert.div_ne02_ne01 = init_fastdiv_values((uint32_t) (src0->ne[2] * src0->ne[1]));
+
+    return true;
+}
+
 static void ggml_hexagon_precompute_fused_mmnx_params(
     const struct ggml_hexagon_session * sess,
     const struct ggml_tensor * src0, // W0
@@ -7416,6 +7637,7 @@ static htp_op_code op_remap_to_htp(const ggml_tensor * t) {
         case GGML_OP_ADD_ID:          return HTP_OP_ADD_ID;
         case GGML_OP_SUB:             return HTP_OP_SUB;
         case GGML_OP_DIV:             return HTP_OP_DIV;
+        case GGML_OP_DUP:
         case GGML_OP_CPY:             return HTP_OP_CPY;
         case GGML_OP_CONT:            return HTP_OP_CPY;
         case GGML_OP_GET_ROWS:        return HTP_OP_GET_ROWS;
@@ -7716,8 +7938,21 @@ static ggml_status ggml_backend_hexagon_graph_compute(ggml_backend_t backend, gg
                 ggml_hexagon_precompute_pool_2d_params(
                     sess, node.node->src[0], node.dst(),
                     (struct htp_pool_2d_kernel_params *)node.kernel_params,
-                    node.opcode == HTP_OP_POOL_1D
+                    node.opcode == HTP_OP_POOL_1D);
+            } else if (node.opcode == HTP_OP_CONCAT) {
+                ggml_hexagon_precompute_concat_params(sess,
+                    node.node,
+                    (struct htp_concat_kernel_params *) node.kernel_params
                 );
+            } else if (node.opcode == HTP_OP_CPY || node.opcode == HTP_OP_CPY_FENCE) {
+                const bool ok = ggml_hexagon_precompute_cpy_params(sess,
+                    node.node,
+                    (struct htp_copy_kernel_params *) node.kernel_params
+                );
+                const auto * kparams = (const struct htp_copy_kernel_params *) node.kernel_params;
+                if (ok && node.opcode == HTP_OP_CPY && kparams->total_elems == 0) {
+                    continue;
+                }
             }
             computed_nodes.push_back(std::move(node));
         }
@@ -8346,49 +8581,13 @@ static ggml_backend_buffer_type_t ggml_backend_hexagon_device_get_host_buffer_ty
 }

 static bool ggml_hexagon_supported_cpy(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
-    GGML_UNUSED(sess);
-
-    const struct ggml_tensor * src0 = op->src[0];
-    const struct ggml_tensor * dst  = op;
-
-    if (src0->type != GGML_TYPE_F32 && src0->type != GGML_TYPE_F16 &&
-        src0->type != GGML_TYPE_I32) return false;
-    if (dst->type != GGML_TYPE_F32 && dst->type != GGML_TYPE_F16 &&
-        dst->type != GGML_TYPE_I32) return false;
-
-    const bool is_scalar  = (ggml_nelements(src0) == 1 && ggml_nelements(dst) == 1);
-    const bool sametype   = (src0->type == dst->type);
-    const bool transposed = !is_scalar && (ggml_is_transposed(src0) || ggml_is_transposed(dst));
-    const bool sameshape  = is_scalar || (!transposed && ggml_are_same_shape(src0, dst));
-
-    if (src0->type == GGML_TYPE_I32 || dst->type == GGML_TYPE_I32) {
-        if (!sameshape) return false;
-        if (sametype) return true;
-        if ((src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_I32) ||
-            (src0->type == GGML_TYPE_I32 && dst->type == GGML_TYPE_F32)) {
-            return true;
-        }
-        return false;
-    }
-
-    // can handle any shape and any same-type (pretty slow if reshaping is required)
-    if (sametype) return true;
-
-    // cannot handle re-shaping and type conversion at the same time
-    if (!sameshape) return false;
-
-    return true;
+    struct htp_copy_kernel_params kparams;
+    return ggml_hexagon_precompute_cpy_params(sess, op, &kparams);
 }

 static bool ggml_hexagon_supported_cont(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
-    GGML_UNUSED(sess);
-    const struct ggml_tensor * src0 = op->src[0];
-
-    // CONT is same-type only and supports F32, F16, and I32.
-    if (src0->type != GGML_TYPE_F32 && src0->type != GGML_TYPE_F16 &&
-        src0->type != GGML_TYPE_I32) return false;
-
-    return true;
+    struct htp_copy_kernel_params kparams;
+    return ggml_hexagon_precompute_cpy_params(sess, op, &kparams);
 }

 static bool ggml_hexagon_supported_repeat(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
@@ -8415,23 +8614,8 @@ static bool ggml_hexagon_supported_repeat(const struct ggml_hexagon_session * se
 }

 static bool ggml_hexagon_supported_concat(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
-    int dim = ((const int32_t *) op->op_params)[0];
-    if (dim < 0 || dim >= GGML_MAX_DIMS) {
-        return false;
-    }
-
-    for (int i = 0; i < GGML_MAX_SRC; ++i) {
-        const struct ggml_tensor * src = op->src[i];
-        if (!src) {
-            continue;
-        }
-        if (src->type != GGML_TYPE_F32 && src->type != GGML_TYPE_I32 && src->type != GGML_TYPE_F16) {
-            return false;
-        }
-    }
-
-    return true;
-    GGML_UNUSED(sess);
+    struct htp_concat_kernel_params kparams;
+    return ggml_hexagon_precompute_concat_params(sess, op, &kparams);
 }

 static bool ggml_hexagon_supported_fill(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
@@ -8601,6 +8785,7 @@ static bool ggml_backend_hexagon_device_supports_op(ggml_backend_dev_t dev, cons
             supp = ggml_hexagon_supported_get_rows(sess, op);
             break;

+        case GGML_OP_DUP:
         case GGML_OP_CPY:
             supp = ggml_hexagon_supported_cpy(sess, op);
             break;
diff --git a/ggml/src/ggml-hexagon/htp/concat-ops.c b/ggml/src/ggml-hexagon/htp/concat-ops.c
index 4dc146393..f7dda1408 100644
--- a/ggml/src/ggml-hexagon/htp/concat-ops.c
+++ b/ggml/src/ggml-hexagon/htp/concat-ops.c
@@ -1,11 +1,13 @@
+#include "concat-ops.h"
 #include "dma-queue.h"
 #include "hex-common.h"
-#include "hex-cpy-dma.h"
+#include "dma-copy.h"
 #include "hex-fastdiv.h"
 #include "hex-profile.h"
 #include "hexagon_protos.h"
 #include "hexagon_types.h"
 #include "htp-ctx.h"
+#include "htp-fence.h"
 #include "htp-ops.h"
 #include "htp-tensor.h"
 #include "htp-vtcm.h"
@@ -16,15 +18,14 @@

 struct htp_concat_context {
     struct htp_ops_context * octx;
-    uint32_t dim;
-    uint32_t nrows_per_thread;
-    uint32_t row_start;
-    uint32_t nrows;
-    uint32_t elem_start;
-    uint32_t nelems;
-    uint32_t nplanes;
-    struct fastdiv_values div_ne0;
-    struct fastdiv_values div_ne1;
+    uint8_t * spad0_base;
+    uint8_t * spad1_base;
+    uint32_t  spad0_size_per_thread;
+    uint32_t  spad1_size_per_thread;
+    uint32_t  row_start;
+    uint32_t  nrows;
+    uint32_t  nrows_per_thread;
+    uint32_t  nplanes;
     struct fastdiv_values div_ne2;
 };

@@ -52,8 +53,8 @@ static void concat_2d_f32_transposed(unsigned int nth, unsigned int ith, void *

     dma_queue * dma_q = octx->ctx->dma[ith];

-    uint8_t * spad0_base = octx->src0_spad.data + ith * octx->src0_spad.size_per_thread;
-    uint8_t * spad1_base = octx->src1_spad.data + ith * octx->src1_spad.size_per_thread;
+    uint8_t * spad0_base = cctx->spad0_base + ith * cctx->spad0_size_per_thread;
+    uint8_t * spad1_base = cctx->spad1_base + ith * cctx->spad1_size_per_thread;

     const uint32_t block_i = 32;
     const uint32_t spad1_stride = block_i * sizeof(float);
@@ -127,6 +128,7 @@ static void concat_2d_f32_transposed(unsigned int nth, unsigned int ith, void *
         p = np;
         i = ni;
     }
+    dma_queue_flush(dma_q);
 }

 static void concat_2d_f16_transposed(unsigned int nth, unsigned int ith, void * data) {
@@ -147,8 +149,8 @@ static void concat_2d_f16_transposed(unsigned int nth, unsigned int ith, void *

     dma_queue * dma_q = octx->ctx->dma[ith];

-    uint8_t * spad0_base = octx->src0_spad.data + ith * octx->src0_spad.size_per_thread;
-    uint8_t * spad1_base = octx->src1_spad.data + ith * octx->src1_spad.size_per_thread;
+    uint8_t * spad0_base = cctx->spad0_base + ith * cctx->spad0_size_per_thread;
+    uint8_t * spad1_base = cctx->spad1_base + ith * cctx->spad1_size_per_thread;

     const uint32_t block_i = 64;
     const uint32_t spad1_stride = block_i * sizeof(__fp16);
@@ -222,219 +224,123 @@ static void concat_2d_f16_transposed(unsigned int nth, unsigned int ith, void *
         p = np;
         i = ni;
     }
+    dma_queue_flush(dma_q);
 }

-static void concat_generic(unsigned int nth, unsigned int ith, void * data) {
-    struct htp_concat_context * cctx = (struct htp_concat_context *) data;
-    struct htp_ops_context * octx = cctx->octx;
-
-    const struct htp_tensor * src0 = octx->src[0];
-    const struct htp_tensor * src1 = octx->src[1];
-    const struct htp_tensor * dst  = octx->dst;
-
-    const int dim = cctx->dim;
-    const uint32_t type_size = (dst->type == HTP_TYPE_F32 || dst->type == HTP_TYPE_I32) ? 4 : 2;
-
-    const uint32_t ne[4] = {dst->ne[0], dst->ne[1], dst->ne[2], dst->ne[3]};
-
-    // Per-device element range aligned to prevent false sharing
-    const uint32_t elem_start = cctx->elem_start;
-    const uint32_t nelems     = cctx->nelems;
-    const uint32_t chunk_size = fastdiv(nelems + nth - 1, &octx->n_threads_div);
-
-    const uint32_t start_idx = MIN(elem_start + ith * chunk_size, elem_start + nelems);
-    const uint32_t end_idx   = MIN(start_idx + chunk_size, elem_start + nelems);
-
-    // Naive scalar element-wise copy
-    for (uint32_t idx = start_idx; idx < end_idx; idx++) {
-        uint32_t idx_div_ne0 = fastdiv(idx, &cctx->div_ne0);
-        uint32_t i0 = idx - idx_div_ne0 * ne[0];
-
-        uint32_t idx_div_ne01 = fastdiv(idx_div_ne0, &cctx->div_ne1);
-        uint32_t i1 = idx_div_ne0 - idx_div_ne01 * ne[1];
-
-        uint32_t idx_div_ne012 = fastdiv(idx_div_ne01, &cctx->div_ne2);
-        uint32_t i2 = idx_div_ne01 - idx_div_ne012 * ne[2];
-        uint32_t i3 = idx_div_ne012;
-
-        uint8_t * dst_ptr = (uint8_t *)dst->data + i3 * dst->nb[3] + i2 * dst->nb[2] + i1 * dst->nb[1] + i0 * dst->nb[0];
-
-        uint32_t idx_dim = 0;
-        if (dim == 0) idx_dim = i0;
-        else if (dim == 1) idx_dim = i1;
-        else if (dim == 2) idx_dim = i2;
-        else if (dim == 3) idx_dim = i3;
-
-        const struct htp_tensor * src = (idx_dim < src0->ne[dim]) ? src0 : src1;
-
-        uint32_t s0 = i0;
-        uint32_t s1 = i1;
-        uint32_t s2 = i2;
-        uint32_t s3 = i3;
-
-        if (dim == 0 && src == src1) s0 -= src0->ne[0];
-        if (dim == 1 && src == src1) s1 -= src0->ne[1];
-        if (dim == 2 && src == src1) s2 -= src0->ne[2];
-        if (dim == 3 && src == src1) s3 -= src0->ne[3];
-
-        uint8_t * src_ptr = (uint8_t *)src->data + s3 * src->nb[3] + s2 * src->nb[2] + s1 * src->nb[1] + s0 * src->nb[0];
-
-        if (type_size == 4) {
-            *(float*)dst_ptr = *(float*)src_ptr;
-        } else {
-            *(__fp16*)dst_ptr = *(__fp16*)src_ptr;
-        }
-    }
-}
-
-static bool concat_dma(struct htp_ops_context * octx, int dim, uint32_t type_size) {
-    if (dim < 0 || dim >= HTP_OP_MAX_DIMS) {
-        return false;
-    }
-
+static int concat_regular(struct htp_ops_context * octx, int dim, uint32_t type_size) {
     const struct htp_tensor * src0 = octx->src[0];
     const struct htp_tensor * src1 = octx->src[1];
     const struct htp_tensor * dst  = octx->dst;

-    // Not partitioned across devices: the row/element-split paths handle that.
-    if (octx->ctx->mdev.count > 1 ||
-        (dst->type != HTP_TYPE_F32 && dst->type != HTP_TYPE_F16 && dst->type != HTP_TYPE_I32) ||
-        src0->type != dst->type || src1->type != dst->type || src0->nb[0] != type_size || src1->nb[0] != type_size ||
-        dst->nb[0] != type_size || (size_t) dst->ne[0] * type_size > DMA_MAX_SIZE_24B ||
-        dst->nb[1] > DMA_MAX_STRIDE_24B || src0->nb[1] > DMA_MAX_STRIDE_24B || src1->nb[1] > DMA_MAX_STRIDE_24B) {
-        return false;
-    }
-
-    for (int d = 0; d < HTP_OP_MAX_DIMS; d++) {
-        const uint32_t ne_d = (d == dim) ? src0->ne[d] + src1->ne[d] : src0->ne[d];
-        if (dst->ne[d] != ne_d || (d != dim && src1->ne[d] != dst->ne[d])) {
-            return false;
-        }
-    }
-
-    // The two views of dst, shaped like the sources.
     struct htp_tensor view0 = *dst;
     struct htp_tensor view1 = *dst;
     for (int d = 0; d < HTP_OP_MAX_DIMS; d++) {
         view0.ne[d] = src0->ne[d];
         view1.ne[d] = src1->ne[d];
     }
-    view1.data += (uint64_t) src0->ne[dim] * dst->nb[dim];
+    view1.data += src0->ne[dim] * dst->nb[dim];
+
+    const uint32_t total_rows_0 = src0->ne[1] * src0->ne[2] * src0->ne[3];
+    const uint32_t total_rows_1 = src1->ne[1] * src1->ne[2] * src1->ne[3];
+
+    uint32_t rstart0 = 0, nrows0 = total_rows_0;
+    uint32_t rstart1 = 0, nrows1 = total_rows_1;
+
+    if (octx->ctx->mdev.count > 1) {
+        const struct htp_tensor_mdev_range range0 = htp_tensor_mdev_partition(
+            total_rows_0, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+        rstart0 = range0.start;
+        nrows0  = range0.count;
+
+        const struct htp_tensor_mdev_range range1 = htp_tensor_mdev_partition(
+            total_rows_1, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+        rstart1 = range1.start;
+        nrows1  = range1.count;
+    }

     dma_queue * q = octx->ctx->dma[0];

-    cpy_dma_sametype_sameshape(q, &view0, src0, type_size);
-    cpy_dma_sametype_sameshape(q, &view1, src1, type_size);
+    dma_cpy_sametype_sameshape_range(q, &view0, src0, type_size, rstart0, nrows0);
+    dma_cpy_sametype_sameshape_range(q, &view1, src1, type_size, rstart1, nrows1);
     dma_queue_flush(q);
-    return true;
+
+    return HTP_STATUS_OK;
 }

-int op_concat(struct htp_ops_context * octx) {
-    int dim = octx->op_params[0];
-    if (dim < 0 || dim >= HTP_OP_MAX_DIMS) {
-        return HTP_STATUS_NO_SUPPORT;
+static int concat_transposed(struct htp_ops_context * octx, const struct htp_concat_kernel_params * kparams, uint32_t type_size) {
+    if (!htp_ops_context_set_n_threads(octx, kparams->n_threads)) {
+        return HTP_STATUS_INVAL_PARAMS;
     }

-    const struct htp_tensor * src0 = octx->src[0];
-    const struct htp_tensor * src1 = octx->src[1];
-    const struct htp_tensor * dst  = octx->dst;
+    const struct htp_tensor * dst = octx->dst;

-    const uint32_t type_size = (dst->type == HTP_TYPE_F32 || dst->type == HTP_TYPE_I32) ? 4 : 2;
-    bool is_src1_transposed  = (src1->nb[0] > src1->nb[1]);
-    bool is_src0_transposed  = (src0->nb[0] > src0->nb[1]);
+    const uint32_t total_rows = dst->ne[1];
+    uint32_t row_start = 0;
+    uint32_t nrows     = total_rows;
+    if (octx->ctx->mdev.count > 1) {
+        const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_rows, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+        row_start = range.start;
+        nrows     = range.count;
+    }

-    if (concat_dma(octx, dim, type_size)) {
+    if (nrows == 0 || dst->ne[2] == 0 || dst->ne[3] == 0) {
         return HTP_STATUS_OK;
     }

-    uint32_t n_threads = octx->n_threads;
-    struct htp_concat_context cctx;
-    cctx.octx = octx;
-    cctx.dim = dim;
-    cctx.div_ne0 = init_fastdiv_values(dst->ne[0]);
-    cctx.div_ne1 = init_fastdiv_values(dst->ne[1]);
-    cctx.div_ne2 = init_fastdiv_values(dst->ne[2]);
-
-    void (*worker_func)(unsigned int, unsigned int, void *) = concat_generic;
-
-    const bool rows_ok = src0->nb[0] == type_size && src1->nb[1] == type_size && dst->nb[0] == type_size;
-
-    if (dim == 0 && is_src1_transposed && !is_src0_transposed && rows_ok) {
-        const uint32_t total_rows = dst->ne[1];
-        const size_t dst_data_row_size = dst->ne[0] * type_size;
-        uint32_t row_start = 0;
-        uint32_t nrows     = total_rows;
-        if (octx->ctx->mdev.count > 1) {
-            uint32_t rows_per_chunk = 0;
-            htp_tensor_mdev_rows_per_chunk(dst, type_size, (uint32_t) dst_data_row_size, &rows_per_chunk);
-            const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_rows, rows_per_chunk, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
-            row_start = range.start;
-            nrows     = range.count;
-        }
-
-        if (nrows == 0) {
-            return HTP_STATUS_OK;
-        }
-
-        cctx.row_start = row_start;
-        cctx.nrows     = nrows;
-        cctx.nplanes   = dst->ne[2] * dst->ne[3];
-
-        uint32_t block_i = (type_size == 4) ? 32 : 64;
-
-        cctx.nrows_per_thread = fastdiv(nrows + n_threads - 1, &octx->n_threads_div);
+    if (kparams->vtcm_size > octx->ctx->vtcm_size) {
+        return HTP_STATUS_VTCM_TOO_SMALL;
+    }

-        // Allocate VTCM
-        uint32_t spad1_stride = block_i * type_size;
+    const uint32_t n_threads = octx->n_threads;

-        uint32_t src1_ne0_padded = hex_round_up(src1->ne[0], block_i);
-        // src0 row is right-aligned to VLEN so the gathered src1 part starts aligned
-        uint32_t spad0_row_bytes = hex_round_up(src0->ne[0] * type_size, VLEN) + src1_ne0_padded * type_size;
+    // layout precomputed on host; kept for reference:
+    // struct htp_concat_transposed_vtcm_layout layout;
+    // htp_concat_transposed_vtcm_layout_build(&layout, octx->src[0]->ne[0], octx->src[1]->ne[0], type_size, n_threads);

-        octx->src0_spad.size_per_thread = block_i * spad0_row_bytes;
-        octx->src1_spad.size_per_thread = src1_ne0_padded * spad1_stride;
+    uint8_t * vtcm_base = (uint8_t *) octx->ctx->vtcm_base;

-        octx->src0_spad.size = n_threads * octx->src0_spad.size_per_thread;
-        octx->src1_spad.size = n_threads * octx->src1_spad.size_per_thread;
+    struct htp_concat_context cctx;
+    cctx.octx                  = octx;
+    cctx.spad0_base            = vtcm_base;
+    cctx.spad1_base            = vtcm_base + n_threads * kparams->spad0_size_per_thread;
+    cctx.spad0_size_per_thread = kparams->spad0_size_per_thread;
+    cctx.spad1_size_per_thread = kparams->spad1_size_per_thread;
+    cctx.row_start             = row_start;
+    cctx.nrows                 = nrows;
+    cctx.nplanes               = dst->ne[2] * dst->ne[3];
+    cctx.div_ne2               = init_fastdiv_values(dst->ne[2]);
+    cctx.nrows_per_thread      = fastdiv(nrows + n_threads - 1, &octx->n_threads_div);
+
+    work_queue_func_t worker_func = (type_size == 4) ? concat_2d_f32_transposed : concat_2d_f16_transposed;
+    work_queue_run(octx->ctx->work_queue, worker_func, &cctx, n_threads);
+    return HTP_STATUS_OK;
+}

-        if (octx->src0_spad.size + octx->src1_spad.size > octx->ctx->vtcm_size) {
-            return HTP_STATUS_VTCM_TOO_SMALL;
-        }
+int op_concat(struct htp_ops_context * octx) {
+    const struct htp_concat_kernel_params * kparams = (const struct htp_concat_kernel_params *) octx->kernel_params;
+    const struct htp_tensor * dst = octx->dst;
+    const uint32_t type_size = (dst->type == HTP_TYPE_F32 || dst->type == HTP_TYPE_I32) ? 4 : 2;

-        octx->src0_spad.data = octx->ctx->vtcm_base;
-        octx->src1_spad.data = octx->src0_spad.data + octx->src0_spad.size;
-        octx->src0_spad.src  = NULL;
-        octx->src1_spad.src  = NULL;
+    int status = HTP_STATUS_OK;
+    switch (kparams->kernel_type) {
+        case HTP_CONCAT_KERNEL_REGULAR:
+            status = concat_regular(octx, kparams->dim, type_size);
+            break;

-        if (type_size == 4) {
-            worker_func = concat_2d_f32_transposed;
-        } else {
-            worker_func = concat_2d_f16_transposed;
-        }
-    } else {
-        if (htp_tensor_is_extended(src0) || htp_tensor_is_extended(src1) || htp_tensor_is_extended(dst)) {
-            return HTP_STATUS_NO_SUPPORT;
-        }
+        case HTP_CONCAT_KERNEL_TRANSPOSED:
+            status = concat_transposed(octx, kparams, type_size);
+            break;

-        const uint32_t total_elements = dst->ne[0] * dst->ne[1] * dst->ne[2] * dst->ne[3];
-        uint32_t elem_start = 0;
-        uint32_t nelems     = total_elements;
-        if (octx->ctx->mdev.count > 1) {
-            const uint32_t elems_per_chunk = HEX_L2_LINE_SIZE / type_size;
-            const bool can_split = htp_tensor_mdev_data_aligned(dst) && htp_tensor_is_contiguous(dst, type_size) && !htp_tensor_is_permuted(dst);
-            const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_elements, can_split ? elems_per_chunk : 0, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
-            elem_start = range.start;
-            nelems     = range.count;
-        }
+        default:
+            status = HTP_STATUS_NO_SUPPORT;
+            break;
+    }

-        if (nelems == 0) {
-            return HTP_STATUS_OK;
-        }
+    htp_ops_context_set_status(octx, status);

-        cctx.elem_start = elem_start;
-        cctx.nelems     = nelems;
+    if (octx->ctx->mdev.count > 1) {
+        htp_mdev_group_barrier(octx);
     }

-    work_queue_run(octx->ctx->work_queue, worker_func, &cctx, n_threads);
-    return HTP_STATUS_OK;
+    return octx->status;
 }
diff --git a/ggml/src/ggml-hexagon/htp/concat-ops.h b/ggml/src/ggml-hexagon/htp/concat-ops.h
new file mode 100644
index 000000000..1d7218441
--- /dev/null
+++ b/ggml/src/ggml-hexagon/htp/concat-ops.h
@@ -0,0 +1,53 @@
+#ifndef HTP_CONCAT_OPS_H
+#define HTP_CONCAT_OPS_H
+
+#include "hex-common.h"
+#include <stdint.h>
+
+enum htp_concat_kernel_type {
+    HTP_CONCAT_KERNEL_UNSUPPORTED = 0,
+    HTP_CONCAT_KERNEL_REGULAR     = 1,
+    HTP_CONCAT_KERNEL_TRANSPOSED  = 2,
+};
+
+struct htp_concat_kernel_params {
+    uint8_t  kernel_type;
+    uint8_t  dim;
+    uint8_t  n_threads;
+    uint8_t  pad;
+
+    uint32_t vtcm_size;
+    uint32_t spad0_size_per_thread;
+    uint32_t spad1_size_per_thread;
+};
+
+#if defined(__cplusplus)
+static_assert(sizeof(struct htp_concat_kernel_params) <= 128, "htp_concat_kernel_params is too large for kernel_params blob");
+#else
+_Static_assert(sizeof(struct htp_concat_kernel_params) <= 128, "htp_concat_kernel_params is too large for kernel_params blob");
+#endif
+
+struct htp_concat_transposed_vtcm_layout {
+    uint32_t src0_spad_size_per_thread;
+    uint32_t src1_spad_size_per_thread;
+    uint32_t total_bytes;
+};
+
+static inline void htp_concat_transposed_vtcm_layout_build(
+    struct htp_concat_transposed_vtcm_layout * layout,
+    uint32_t src0_ne0,
+    uint32_t src1_ne0,
+    uint32_t type_size,
+    uint32_t n_threads) {
+
+    uint32_t block_i = (type_size == 4) ? 32 : 64;
+    uint32_t spad1_stride = block_i * type_size;
+    uint32_t src1_ne0_padded = hex_round_up(src1_ne0, block_i);
+    uint32_t spad0_row_bytes = hex_round_up(src0_ne0 * type_size, 128) + src1_ne0_padded * type_size;
+
+    layout->src0_spad_size_per_thread = block_i * spad0_row_bytes;
+    layout->src1_spad_size_per_thread = src1_ne0_padded * spad1_stride;
+    layout->total_bytes = n_threads * (layout->src0_spad_size_per_thread + layout->src1_spad_size_per_thread);
+}
+
+#endif // HTP_CONCAT_OPS_H
diff --git a/ggml/src/ggml-hexagon/htp/cpy-ops.c b/ggml/src/ggml-hexagon/htp/cpy-ops.c
index df38d3eed..d366acb75 100644
--- a/ggml/src/ggml-hexagon/htp/cpy-ops.c
+++ b/ggml/src/ggml-hexagon/htp/cpy-ops.c
@@ -11,7 +11,8 @@

 #define GGML_COMMON_DECL_C
 #include "ggml-common.h"
-#include "hex-cpy-dma.h"
+#include "cpy-ops.h"
+#include "dma-copy.h"
 #include "htp-ctx.h"
 #include "htp-fence.h"
 #include "htp-ops.h"
@@ -19,34 +20,19 @@
 #include "hvx-utils.h"

 struct htp_copy_context {
-    struct htp_ops_context * octx;
+    struct htp_ops_context *              octx;
+    const struct htp_copy_kernel_params * kparams;

-    uint32_t          src0_type_size;
-    uint32_t          src0_block_size;
+    uint32_t row_start;
+    uint32_t nrows;
+    uint32_t src0_nrows_per_thread;

-    uint32_t          dst_type_size;
-    uint32_t          dst_block_size;
+    uint32_t elem_start;
+    uint32_t nelem;
+    uint32_t elem_per_thread;

-    uint32_t          src0_blocks_per_row;
-    uint32_t          dst_blocks_per_row;
-
-    uint32_t          elem_start;
-    uint32_t          nelem;
-    uint32_t          elem_per_thread;
-
-    uint32_t          src0_nrows_per_thread;
-    uint32_t          row_start;
-    uint32_t          nrows;
-
-    struct fastdiv_values div_ne01;
-    struct fastdiv_values div_ne02_ne01;
-
-    struct fastdiv_values div_ne0;
-    struct fastdiv_values div_ne1_ne0;
-    struct fastdiv_values div_ne2_ne1_ne0;
-    struct fastdiv_values div_ne00;
-    struct fastdiv_values div_ne01_ne00;
-    struct fastdiv_values div_ne02_ne01_ne00;
+    uint8_t * vtcm_src0;
+    uint8_t * vtcm_dst;
 };

 #define cpy_preamble                              \
@@ -73,52 +59,6 @@ struct htp_copy_context {
     const uint32_t  nb2 = dst->nb[2];             \
     const uint32_t  nb3 = dst->nb[3];

-#define DEFINE_CPY_SAMESHAPE(NAME, ELEM_TYPE, ELEM_SIZE)                                                           \
-static void cpy_thread_##NAME##_sameshape(unsigned int nth, unsigned int ith, void * data) {                       \
-    struct htp_copy_context * ct = (struct htp_copy_context *) data;                                               \
-    struct htp_ops_context * octx = ct->octx;                                                                      \
-    cpy_preamble;                                                                                                  \
-    const uint32_t dr  = ct->src0_nrows_per_thread;                                                                \
-    const uint32_t ir0 = ct->row_start + dr * ith;                                                                 \
-    const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);                                                 \
-    if (ir0 >= ir1) return;                                                                                        \
-    const bool contiguous = htp_tensor_is_contiguous(src0, ELEM_SIZE) && htp_tensor_is_contiguous(dst, ELEM_SIZE); \
-    if (contiguous) {                                                                                              \
-        dma_queue * dma_q = octx->ctx->dma[ith];                                                                   \
-        dma_addr_t dst_addr  = dst->data  + ir0 * ne00 * ELEM_SIZE;                                                \
-        dma_addr_t src0_addr = src0->data + ir0 * ne00 * ELEM_SIZE;                                                \
-        cpy_dma_sametype_reshape_contig(dma_q, dst_addr, src0_addr, (ir1 - ir0) * ne00 * ELEM_SIZE);               \
-        dma_queue_flush(dma_q);                                                                                    \
-        return;                                                                                                    \
-    }                                                                                                              \
-    const uint32_t ne02_ne01 = ne02 * ne01;                                                                        \
-    uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);                                                               \
-    uint32_t rem = ir0 - i03 * ne02_ne01;                                                                          \
-    uint32_t i02 = fastdiv(rem, &ct->div_ne01);                                                                    \
-    uint32_t i01 = rem - i02 * ne01;                                                                               \
-    uint8_t * dst_ptr  = (uint8_t *) dst->data  + i01*nb1  + i02*nb2  + i03*nb3;                                   \
-    uint8_t * src0_ptr = (uint8_t *) src0->data + i01*nb01 + i02*nb02 + i03*nb03;                                  \
-    for (uint32_t r = ir0; r < ir1; r++) {                                                                         \
-        hex_l2fetch(src0_ptr, ne00 * ELEM_SIZE, nb01, 2);                                                          \
-        hvx_copy_uu(dst_ptr, src0_ptr, ne00, ELEM_SIZE);                                                           \
-        dst_ptr  += nb1;                                                                                           \
-        src0_ptr += nb01;                                                                                          \
-        if (++i01 == ne01) {                                                                                       \
-            i01 = 0;                                                                                               \
-            if (++i02 == ne02) {                                                                                   \
-                i02 = 0;                                                                                           \
-                i03++;                                                                                             \
-            }                                                                                                      \
-            dst_ptr  = (uint8_t *) dst->data  + i02*nb2  + i03*nb3;                                                \
-            src0_ptr = (uint8_t *) src0->data + i02*nb02 + i03*nb03;                                               \
-        }                                                                                                          \
-    }                                                                                                              \
-}
-
-DEFINE_CPY_SAMESHAPE(f32,  float, 4)
-DEFINE_CPY_SAMESHAPE(f16, __fp16, 2)
-DEFINE_CPY_SAMESHAPE(i32, int32_t, 4)
-
 #define DEFINE_CPY_RESHAPE(NAME, ELEM_TYPE, ELEM_SIZE)                                                \
 static void cpy_thread_##NAME##_reshape(unsigned int nth, unsigned int ith, void * data) {            \
     struct htp_copy_context * ct = (struct htp_copy_context *) data;                                  \
@@ -129,52 +69,45 @@ static void cpy_thread_##NAME##_reshape(unsigned int nth, unsigned int ith, void
     const uint32_t th_end   = MIN(th_start + th_nelem, ct->elem_start + ct->nelem);                   \
     if (th_start >= th_end) return;                                                                   \
                                                                                                       \
-    if (htp_tensor_is_contiguous(src0, ELEM_SIZE) && htp_tensor_is_contiguous(dst, ELEM_SIZE)) {      \
-        dma_queue * dma_q = octx->ctx->dma[ith];                                                      \
-        dma_addr_t dst_addr  = dst->data  + th_start * ELEM_SIZE;                                     \
-        dma_addr_t src0_addr = src0->data + th_start * ELEM_SIZE;                                     \
-        cpy_dma_sametype_reshape_contig(dma_q, dst_addr, src0_addr, (th_end - th_start) * ELEM_SIZE); \
-        dma_queue_flush(dma_q);                                                                       \
-        return;                                                                                       \
-    }                                                                                                 \
+    dma_queue * dma_q = octx->ctx->dma[ith];                                                          \
                                                                                                       \
     const uint32_t ne01_ne00      = ne01 * ne00;                                                      \
     const uint32_t ne02_ne01_ne00 = ne02 * ne01_ne00;                                                 \
     const uint32_t ne1_ne0        = ne1 * ne0;                                                        \
     const uint32_t ne2_ne1_ne0    = ne2 * ne1_ne0;                                                    \
                                                                                                       \
+    const struct htp_copy_reshape_params * rsh = &ct->kparams->u.reshape;                             \
     uint32_t e = th_start;                                                                            \
-    uint32_t i13 = fastdiv(e, &ct->div_ne2_ne1_ne0);                                                  \
+    uint32_t i13 = fastdiv(e, &rsh->div_ne2_ne1_ne0);                                                 \
     uint32_t rem = e - i13 * ne2_ne1_ne0;                                                             \
-    uint32_t i12 = fastdiv(rem, &ct->div_ne1_ne0);                                                    \
+    uint32_t i12 = fastdiv(rem, &rsh->div_ne1_ne0);                                                   \
     uint32_t rem2 = rem - i12 * ne1_ne0;                                                              \
-    uint32_t i11 = fastdiv(rem2, &ct->div_ne0);                                                       \
+    uint32_t i11 = fastdiv(rem2, &rsh->div_ne0);                                                      \
     uint32_t i10 = rem2 - i11 * ne0;                                                                  \
                                                                                                       \
-    uint32_t i03 = fastdiv(e, &ct->div_ne02_ne01_ne00);                                               \
+    uint32_t i03 = fastdiv(e, &rsh->div_ne02_ne01_ne00);                                              \
     uint32_t rem_s = e - i03 * ne02_ne01_ne00;                                                        \
-    uint32_t i02 = fastdiv(rem_s, &ct->div_ne01_ne00);                                                \
+    uint32_t i02 = fastdiv(rem_s, &rsh->div_ne01_ne00);                                               \
     uint32_t rem2_s = rem_s - i02 * ne01_ne00;                                                        \
-    uint32_t i01 = fastdiv(rem2_s, &ct->div_ne00);                                                    \
+    uint32_t i01 = fastdiv(rem2_s, &rsh->div_ne00);                                                   \
     uint32_t i00 = rem2_s - i01 * ne00;                                                               \
                                                                                                       \
-    char * dst_ptr        = (char *)       dst->data  + i10*nb0  + i11*nb1  + i12*nb2  + i13*nb3;     \
-    const char * src0_ptr = (const char *) src0->data + i00*nb00 + i01*nb01 + i02*nb02 + i03*nb03;    \
+    dma_addr_t dst_addr  = dst->data  + i10*nb0  + i11*nb1  + i12*nb2  + i13*nb3;                     \
+    dma_addr_t src0_addr = src0->data + i00*nb00 + i01*nb01 + i02*nb02 + i03*nb03;                    \
                                                                                                       \
     const bool rows_contig = (nb00 == ELEM_SIZE) && (nb0 == ELEM_SIZE);                               \
                                                                                                       \
     while (e < th_end) {                                                                              \
-        uint32_t run = 1;                                                                             \
+        const uint32_t run = MIN(MIN(ne00 - i00, ne0 - i10), th_end - e);                             \
         if (rows_contig) {                                                                            \
-            run = MIN(MIN(ne00 - i00, ne0 - i10), th_end - e);                                        \
-            hvx_copy_uu((uint8_t *) dst_ptr, (const uint8_t *) src0_ptr, run, ELEM_SIZE);             \
+            dma_cpy_sametype_reshape_contig(dma_q, dst_addr, src0_addr, run * ELEM_SIZE);             \
         } else {                                                                                      \
-            *((ELEM_TYPE *) dst_ptr) = *((const ELEM_TYPE *) src0_ptr);                               \
+            dma_cpy_push_2d_chunked(dma_q, dst_addr, src0_addr, nb0, nb00, ELEM_SIZE, run);           \
         }                                                                                             \
         e += run;                                                                                     \
                                                                                                       \
-        dst_ptr += run * nb0;                                                                         \
-        i10     += run;                                                                               \
+        dst_addr += run * nb0;                                                                        \
+        i10      += run;                                                                              \
         if (i10 == ne0) {                                                                             \
             i10 = 0;                                                                                  \
             if (++i11 == ne1) {                                                                       \
@@ -184,11 +117,11 @@ static void cpy_thread_##NAME##_reshape(unsigned int nth, unsigned int ith, void
                     i13++;                                                                            \
                 }                                                                                     \
             }                                                                                         \
-            dst_ptr = (char *) dst->data + i11*nb1 + i12*nb2 + i13*nb3;                               \
+            dst_addr = dst->data + i11*nb1 + i12*nb2 + i13*nb3;                                       \
         }                                                                                             \
                                                                                                       \
-        src0_ptr += run * nb00;                                                                       \
-        i00      += run;                                                                              \
+        src0_addr += run * nb00;                                                                      \
+        i00       += run;                                                                             \
         if (i00 == ne00) {                                                                            \
             i00 = 0;                                                                                  \
             if (++i01 == ne01) {                                                                      \
@@ -198,371 +131,350 @@ static void cpy_thread_##NAME##_reshape(unsigned int nth, unsigned int ith, void
                     i03++;                                                                            \
                 }                                                                                     \
             }                                                                                         \
-            src0_ptr = (const char *) src0->data + i01*nb01 + i02*nb02 + i03*nb03;                    \
+            src0_addr = src0->data + i01*nb01 + i02*nb02 + i03*nb03;                                  \
         }                                                                                             \
     }                                                                                                 \
+    dma_queue_flush(dma_q);                                                                           \
 }

 DEFINE_CPY_RESHAPE(f32,  float, 4)
 DEFINE_CPY_RESHAPE(f16, __fp16, 2)
 DEFINE_CPY_RESHAPE(i32, int32_t, 4)

-static void cpy_thread_f16_f32_sameshape(unsigned int nth, unsigned int ith, void * data) {
-    struct htp_copy_context * ct = (struct htp_copy_context *) data;
-    struct htp_ops_context * octx = ct->octx;
-    cpy_preamble;
-
-    const uint32_t dr  = ct->src0_nrows_per_thread;
-    const uint32_t ir0 = ct->row_start + dr * ith;
-    const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);
-    if (ir0 >= ir1) return;
-
-    const uint32_t ne02_ne01 = ne02 * ne01;
-    uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);
-    uint32_t rem = ir0 - i03 * ne02_ne01;
-    uint32_t i02 = fastdiv(rem, &ct->div_ne01);
-    uint32_t i01 = rem - i02 * ne01;
-
-    uint8_t* dst_ptr  = (uint8_t*) dst->data  + i01*nb1  + i02*nb2  + i03*nb3;
-    uint8_t* src0_ptr = (uint8_t*) src0->data + i01*nb01 + i02*nb02 + i03*nb03;
-
-    for (uint32_t r = ir0; r < ir1; r++) {
-        hex_l2fetch(src0_ptr, ne00 * sizeof(float), nb01, 2);
-        hvx_copy_f16_f32_uu(dst_ptr, src0_ptr, ne00);
-        dst_ptr  += nb1;
-        src0_ptr += nb01;
-        if (++i01 == ne01) {
-            i01 = 0;
-            if (++i02 == ne02) {
-                i02 = 0;
-                i03++;
-            }
-            dst_ptr  = (uint8_t*) dst->data  + i02*nb2  + i03*nb3;
-            src0_ptr = (uint8_t*) src0->data + i02*nb02 + i03*nb03;
-        }
-    }
+#define DEFINE_CPY_CONVERT_SAMESHAPE(NAME, CONV_FUNC)                                        \
+static void cpy_thread_##NAME##_sameshape(unsigned int nth, unsigned int ith, void * data) { \
+    struct htp_copy_context * ct = (struct htp_copy_context *) data;                         \
+    struct htp_ops_context * octx = ct->octx;                                                \
+    cpy_preamble;                                                                            \
+                                                                                             \
+    const uint32_t dr  = ct->src0_nrows_per_thread;                                          \
+    const uint32_t ir0 = ct->row_start + dr * ith;                                           \
+    const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);                           \
+    if (ir0 >= ir1) return;                                                                  \
+    const uint32_t nrows_thread = ir1 - ir0;                                                 \
+                                                                                             \
+    dma_queue * dma_q = octx->ctx->dma[ith];                                                 \
+    struct htp_thread_trace * tr = &octx->ctx->trace[ith];                                   \
+                                                                                             \
+    const struct htp_copy_convert_params * cvt = &ct->kparams->u.convert;                    \
+    const uint32_t src0_buf_size = cvt->src0_buf_size;                                       \
+    const uint32_t dst_buf_size  = cvt->dst_buf_size;                                        \
+    uint8_t * vtcm_src0_base = ct->vtcm_src0 + ith * cvt->spad0_size_per_thread;             \
+    uint8_t * vtcm_dst_base  = ct->vtcm_dst  + ith * cvt->spad1_size_per_thread;             \
+    const uint32_t src0_row_size = ne00 * ct->kparams->src0_type_size;                       \
+    const uint32_t dst_row_size  = ne00 * ct->kparams->dst_type_size;                        \
+                                                                                             \
+    const uint32_t ne02_ne01 = ne02 * ne01;                                                  \
+    uint32_t i03 = fastdiv(ir0, &cvt->div_ne02_ne01);                                        \
+    uint32_t rem = ir0 - i03 * ne02_ne01;                                                    \
+    uint32_t i02 = fastdiv(rem, &cvt->div_ne01);                                             \
+    uint32_t i01 = rem - i02 * ne01;                                                         \
+                                                                                             \
+    uint32_t f_i01 = i01, f_i02 = i02, f_i03 = i03;                                          \
+    dma_addr_t f_src0_addr = src0->data + f_i01*nb01 + f_i02*nb02 + f_i03*nb03;              \
+                                                                                             \
+    uint32_t c_i01 = i01, c_i02 = i02, c_i03 = i03;                                          \
+    dma_addr_t c_dst_addr = dst->data + c_i01*nb1 + c_i02*nb2 + c_i03*nb3;                   \
+                                                                                             \
+    for (uint32_t r = 0; r < nrows_thread && r < 2; ++r) {                                   \
+        uint8_t * src_spad = vtcm_src0_base + r * src0_buf_size;                             \
+        uint8_t * dst_spad = vtcm_dst_base  + r * dst_buf_size;                              \
+        dma_queue_push(dma_q, dma_make_data(dst->data, dst_spad),                            \
+                       dst_row_size, dst_buf_size, dst_row_size, 0);                         \
+        dma_queue_push(dma_q, dma_make_data(src_spad, f_src0_addr),                          \
+                       src0_buf_size, src0_row_size, src0_row_size, 1);                      \
+        f_src0_addr += nb01;                                                                 \
+        if (++f_i01 == ne01) {                                                               \
+            f_i01 = 0;                                                                       \
+            if (++f_i02 == ne02) {                                                           \
+                f_i02 = 0;                                                                   \
+                f_i03++;                                                                     \
+            }                                                                                \
+            f_src0_addr = src0->data + f_i02*nb02 + f_i03*nb03;                              \
+        }                                                                                    \
+    }                                                                                        \
+                                                                                             \
+    for (uint32_t r = 0; r < nrows_thread; ++r) {                                            \
+        uint8_t * dst_spad = (uint8_t *) (uintptr_t) dma_queue_pop(dma_q).src;               \
+        uint8_t * src_spad = (uint8_t *) (uintptr_t) dma_queue_pop(dma_q).dst;               \
+                                                                                             \
+        htp_trace_event_start(tr, HTP_TRACE_EVT_HVX_COMP, (uint16_t) r);                     \
+        CONV_FUNC(dst_spad, src_spad, ne00);                                                 \
+        htp_trace_event_stop(tr, HTP_TRACE_EVT_HVX_COMP, (uint16_t) r);                      \
+                                                                                             \
+        dma_queue_push(dma_q, dma_make_data(c_dst_addr, dst_spad),                           \
+                       dst_row_size, dst_buf_size, dst_row_size, 1);                         \
+        c_dst_addr += nb1;                                                                   \
+        if (++c_i01 == ne01) {                                                               \
+            c_i01 = 0;                                                                       \
+            if (++c_i02 == ne02) {                                                           \
+                c_i02 = 0;                                                                   \
+                c_i03++;                                                                     \
+            }                                                                                \
+            c_dst_addr = dst->data + c_i02*nb2 + c_i03*nb3;                                  \
+        }                                                                                    \
+                                                                                             \
+        if (r + 2 < nrows_thread) {                                                          \
+            dma_queue_push(dma_q, dma_make_data(src_spad, f_src0_addr),                      \
+                           src0_buf_size, src0_row_size, src0_row_size, 1);                  \
+            f_src0_addr += nb01;                                                             \
+            if (++f_i01 == ne01) {                                                           \
+                f_i01 = 0;                                                                   \
+                if (++f_i02 == ne02) {                                                       \
+                    f_i02 = 0;                                                               \
+                    f_i03++;                                                                 \
+                }                                                                            \
+                f_src0_addr = src0->data + f_i02*nb02 + f_i03*nb03;                          \
+            }                                                                                \
+        }                                                                                    \
+    }                                                                                        \
+    dma_queue_flush(dma_q);                                                                  \
 }

-static void cpy_thread_f32_f16_sameshape(unsigned int nth, unsigned int ith, void * data) {
-    struct htp_copy_context * ct = (struct htp_copy_context *) data;
-    struct htp_ops_context * octx = ct->octx;
-    cpy_preamble;
-
-    const uint32_t dr  = ct->src0_nrows_per_thread;
-    const uint32_t ir0 = ct->row_start + dr * ith;
-    const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);
-    if (ir0 >= ir1) return;
-
-    const uint32_t ne02_ne01 = ne02 * ne01;
-    uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);
-    uint32_t rem = ir0 - i03 * ne02_ne01;
-    uint32_t i02 = fastdiv(rem, &ct->div_ne01);
-    uint32_t i01 = rem - i02 * ne01;
-
-    uint8_t* dst_ptr  = (uint8_t*) dst->data  + i01*nb1  + i02*nb2  + i03*nb3;
-    uint8_t* src0_ptr = (uint8_t*) src0->data + i01*nb01 + i02*nb02 + i03*nb03;
-
-    for (uint32_t r = ir0; r < ir1; r++) {
-        hex_l2fetch(src0_ptr, ne00 * sizeof(__fp16), nb01, 2);
-        hvx_copy_f32_f16_uu(dst_ptr, src0_ptr, ne00);
-        dst_ptr  += nb1;
-        src0_ptr += nb01;
-        if (++i01 == ne01) {
-            i01 = 0;
-            if (++i02 == ne02) {
-                i02 = 0;
-                i03++;
-            }
-            dst_ptr  = (uint8_t*) dst->data  + i02*nb2  + i03*nb3;
-            src0_ptr = (uint8_t*) src0->data + i02*nb02 + i03*nb03;
-        }
+DEFINE_CPY_CONVERT_SAMESHAPE(f16_f32, hvx_copy_f16_f32_aa)
+DEFINE_CPY_CONVERT_SAMESHAPE(f32_f16, hvx_copy_f32_f16_aa)
+DEFINE_CPY_CONVERT_SAMESHAPE(i32_f32, hvx_copy_i32_f32_aa)
+DEFINE_CPY_CONVERT_SAMESHAPE(f32_i32, hvx_copy_f32_i32_aa)
+
+static int cpy_scalar(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+    if (octx->ctx->mdev.count > 1 && octx->ctx->mdev.idx > 0) {
+        return HTP_STATUS_OK;
     }
-}

-static void cpy_thread_i32_f32_sameshape(unsigned int nth, unsigned int ith, void * data) {
-    struct htp_copy_context * ct = (struct htp_copy_context *) data;
-    struct htp_ops_context * octx = ct->octx;
-    cpy_preamble;
-
-    const uint32_t dr  = ct->src0_nrows_per_thread;
-    const uint32_t ir0 = ct->row_start + dr * ith;
-    const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);
-    if (ir0 >= ir1) return;
-
-    const uint32_t ne02_ne01 = ne02 * ne01;
-    uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);
-    uint32_t rem = ir0 - i03 * ne02_ne01;
-    uint32_t i02 = fastdiv(rem, &ct->div_ne01);
-    uint32_t i01 = rem - i02 * ne01;
-
-    uint8_t* dst_ptr  = (uint8_t*) dst->data  + i01*nb1  + i02*nb2  + i03*nb3;
-    uint8_t* src0_ptr = (uint8_t*) src0->data + i01*nb01 + i02*nb02 + i03*nb03;
-
-    for (uint32_t r = ir0; r < ir1; r++) {
-        hex_l2fetch(src0_ptr, ne00 * sizeof(float), nb01, 2);
-        const float * restrict src_row = (const float *) src0_ptr;
-        int32_t * restrict dst_row = (int32_t *) dst_ptr;
-        for (uint32_t i = 0; i < ne00; i++) {
-            dst_row[i] = (int32_t) src_row[i];
-        }
-        dst_ptr  += nb1;
-        src0_ptr += nb01;
-        if (++i01 == ne01) {
-            i01 = 0;
-            if (++i02 == ne02) {
-                i02 = 0;
-                i03++;
-            }
-            dst_ptr  = (uint8_t*) dst->data  + i02*nb2  + i03*nb3;
-            src0_ptr = (uint8_t*) src0->data + i02*nb02 + i03*nb03;
-        }
+    const struct htp_tensor * src0 = octx->src[0];
+    const struct htp_tensor * dst  = octx->dst;
+
+    if (src0->type == dst->type) {
+        dma_cpy_sametype_reshape_contig(octx->ctx->dma[0], dst->data, src0->data, kparams->src0_type_size);
+        dma_queue_flush(octx->ctx->dma[0]);
+        return HTP_STATUS_OK;
     }
+
+    dma_queue * dma_q = octx->ctx->dma[0];
+    dma_addr_t s_vtcm = (dma_addr_t)(uintptr_t) octx->ctx->vtcm_base;
+    dma_addr_t d_vtcm = s_vtcm + VLEN;
+    const uint32_t s_size = kparams->src0_type_size;
+    const uint32_t d_size = kparams->dst_type_size;
+
+    dma_queue_push(dma_q, dma_make_data(s_vtcm, src0->data), s_size, s_size, s_size, 1);
+    dma_queue_pop(dma_q);
+
+    uint8_t * s_ptr = (uint8_t *) octx->ctx->vtcm_base;
+    uint8_t * d_ptr = s_ptr + VLEN;
+
+    const HVX_Vector v_src = hvx_vmem(s_ptr);
+    HVX_Vector v_dst;
+
+    if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_I32) {
+        v_dst = Q6_Vw_equals_Vsf(v_src);
+    } else if (src0->type == HTP_TYPE_I32 && dst->type == HTP_TYPE_F32) {
+        v_dst = Q6_Vsf_equals_Vw(v_src);
+    } else if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_F16) {
+        v_dst = hvx_vec_f32_to_f16(v_src, v_src);
+    } else if (src0->type == HTP_TYPE_F16 && dst->type == HTP_TYPE_F32) {
+        v_dst = Q6_V_lo_W(hvx_vec_f16_to_f32(v_src));
+    } else {
+        return HTP_STATUS_NO_SUPPORT;
+    }
+
+    hvx_vmem(d_ptr) = v_dst;
+
+    dma_queue_push(dma_q, dma_make_data(dst->data, d_vtcm), d_size, d_size, d_size, 1);
+    dma_queue_flush(dma_q);
+    return HTP_STATUS_OK;
 }

-static void cpy_thread_f32_i32_sameshape(unsigned int nth, unsigned int ith, void * data) {
-    struct htp_copy_context * ct = (struct htp_copy_context *) data;
-    struct htp_ops_context * octx = ct->octx;
-    cpy_preamble;
-
-    const uint32_t dr  = ct->src0_nrows_per_thread;
-    const uint32_t ir0 = ct->row_start + dr * ith;
-    const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);
-    if (ir0 >= ir1) return;
-
-    const uint32_t ne02_ne01 = ne02 * ne01;
-    uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);
-    uint32_t rem = ir0 - i03 * ne02_ne01;
-    uint32_t i02 = fastdiv(rem, &ct->div_ne01);
-    uint32_t i01 = rem - i02 * ne01;
-
-    uint8_t* dst_ptr  = (uint8_t*) dst->data  + i01*nb1  + i02*nb2  + i03*nb3;
-    uint8_t* src0_ptr = (uint8_t*) src0->data + i01*nb01 + i02*nb02 + i03*nb03;
-
-    for (uint32_t r = ir0; r < ir1; r++) {
-        hex_l2fetch(src0_ptr, ne00 * sizeof(int32_t), nb01, 2);
-        const int32_t * restrict src_row = (const int32_t *) src0_ptr;
-        float * restrict dst_row = (float *) dst_ptr;
-        for (uint32_t i = 0; i < ne00; i++) {
-            dst_row[i] = (float) src_row[i];
-        }
-        dst_ptr  += nb1;
-        src0_ptr += nb01;
-        if (++i01 == ne01) {
-            i01 = 0;
-            if (++i02 == ne02) {
-                i02 = 0;
-                i03++;
-            }
-            dst_ptr  = (uint8_t*) dst->data  + i02*nb2  + i03*nb3;
-            src0_ptr = (uint8_t*) src0->data + i02*nb02 + i03*nb03;
-        }
+static int cpy_1d_contig(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+    const struct htp_tensor * src0 = octx->src[0];
+    const struct htp_tensor * dst  = octx->dst;
+
+    uint32_t elem_start = 0;
+    uint32_t nelem      = kparams->total_elems;
+
+    if (octx->ctx->mdev.count > 1) {
+        const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(
+            kparams->total_elems, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+        elem_start = range.start;
+        nelem      = range.count;
     }
+
+    if (nelem > 0) {
+        dma_queue * q = octx->ctx->dma[0];
+        const uint32_t type_size = kparams->src0_type_size;
+        dma_addr_t dst_addr  = dst->data  + elem_start * type_size;
+        dma_addr_t src0_addr = src0->data + elem_start * type_size;
+        dma_cpy_sametype_reshape_contig(q, dst_addr, src0_addr, nelem * type_size);
+        dma_queue_flush(q);
+    }
+
+    return HTP_STATUS_OK;
 }

-static int exec_cpy(struct htp_ops_context * octx, bool * use_dma) {
-    cpy_preamble;
-    *use_dma = false;
+static int cpy_sameshape_sametype(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+    const struct htp_tensor * src0 = octx->src[0];
+    const struct htp_tensor * dst  = octx->dst;

-    const uint32_t total_elems_src = ne00 * ne01 * ne02 * ne03;
-    const uint32_t total_elems_dst = ne0 * ne1 * ne2 * ne3;
-    if (total_elems_src == 1 && total_elems_dst == 1) {
-        if (octx->ctx->mdev.count > 1 && octx->ctx->mdev.idx > 0) {
-            return HTP_STATUS_OK;
-        }
-        if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_I32) {
-            ((int32_t *) dst->data)[0] = (int32_t) (((const float *) src0->data)[0]);
-            return HTP_STATUS_OK;
-        }
-        if (src0->type == HTP_TYPE_I32 && dst->type == HTP_TYPE_F32) {
-            ((float *) dst->data)[0] = (float) (((const int32_t *) src0->data)[0]);
-            return HTP_STATUS_OK;
-        }
-        if (src0->type == HTP_TYPE_I32 && dst->type == HTP_TYPE_I32) {
-            ((int32_t *) dst->data)[0] = ((const int32_t *) src0->data)[0];
-            return HTP_STATUS_OK;
-        }
-        if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_F32) {
-            ((float *) dst->data)[0] = ((const float *) src0->data)[0];
-            return HTP_STATUS_OK;
-        }
-        if (src0->type == HTP_TYPE_F16 && dst->type == HTP_TYPE_F16) {
-            ((__fp16 *) dst->data)[0] = ((const __fp16 *) src0->data)[0];
-            return HTP_STATUS_OK;
-        }
-        if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_F16) {
-            ((__fp16 *) dst->data)[0] = (__fp16) (((const float *) src0->data)[0]);
-            return HTP_STATUS_OK;
-        }
-        if (src0->type == HTP_TYPE_F16 && dst->type == HTP_TYPE_F32) {
-            ((float *) dst->data)[0] = (float) (((const __fp16 *) src0->data)[0]);
-            return HTP_STATUS_OK;
-        }
+    uint32_t row_start = 0;
+    uint32_t nrows     = kparams->total_rows;
+
+    if (octx->ctx->mdev.count > 1) {
+        const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(
+            kparams->total_rows, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+        row_start = range.start;
+        nrows     = range.count;
     }

-    struct htp_copy_context ct;
-    ct.octx = octx;
+    if (nrows > 0) {
+        dma_queue * q = octx->ctx->dma[0];
+        dma_cpy_sametype_sameshape_range(q, dst, src0, kparams->src0_type_size, row_start, nrows);
+        dma_queue_flush(q);
+    }

-    switch (src0->type) {
-    case HTP_TYPE_F32: ct.src0_type_size = 4; ct.src0_block_size = 1; ct.src0_blocks_per_row = ne00 / 1; break;
-    case HTP_TYPE_F16: ct.src0_type_size = 2; ct.src0_block_size = 1; ct.src0_blocks_per_row = ne00 / 1; break;
-    case HTP_TYPE_I32: ct.src0_type_size = 4; ct.src0_block_size = 1; ct.src0_blocks_per_row = ne00 / 1; break;
-    default:
+    return HTP_STATUS_OK;
+}
+
+static int cpy_sameshape_convert(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+    const struct htp_tensor * src0 = octx->src[0];
+    const struct htp_tensor * dst  = octx->dst;
+
+    if (htp_tensor_is_extended(src0) || htp_tensor_is_extended(dst)) {
         return HTP_STATUS_NO_SUPPORT;
     }

-    switch (dst->type) {
-    case HTP_TYPE_F32: ct.dst_type_size = 4; ct.dst_block_size = 1; ct.dst_blocks_per_row = ne0 / 1; break;
-    case HTP_TYPE_F16: ct.dst_type_size = 2; ct.dst_block_size = 1; ct.dst_blocks_per_row = ne0 / 1; break;
-    case HTP_TYPE_I32: ct.dst_type_size = 4; ct.dst_block_size = 1; ct.dst_blocks_per_row = ne0 / 1; break;
-    default:
-        return HTP_STATUS_NO_SUPPORT;
+    if (!htp_ops_context_set_n_threads(octx, kparams->n_threads)) {
+        return HTP_STATUS_INVAL_PARAMS;
     }

-    const bool sametype   = (src0->type == dst->type);
-    const bool transposed = (nb00 > nb01) || (nb0 > nb1) ||
-                            (nb00 != ct.src0_type_size) || (nb0 != ct.dst_type_size) ||
-                            (nb01 < ne00 * ct.src0_type_size) || (nb1 < ne0 * ct.dst_type_size);
-    const bool sameshape  = !transposed && (ne00 == ne0 && ne01 == ne1 && ne02 == ne2 && ne03 == ne3);
+    uint32_t row_start = 0;
+    uint32_t nrows     = kparams->total_rows;

-    const uint32_t n_threads = octx->n_threads;
+    if (octx->ctx->mdev.count > 1) {
+        const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(
+            kparams->total_rows, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+        row_start = range.start;
+        nrows     = range.count;
+    }

-    const bool src_is_contiguous = htp_tensor_is_contiguous(src0, ct.src0_type_size);
-    const bool dst_is_contiguous = htp_tensor_is_contiguous(dst, ct.dst_type_size);
+    if (nrows == 0) {
+        return HTP_STATUS_OK;
+    }

-    if (htp_tensor_is_extended(src0) || htp_tensor_is_extended(dst)) {
-        if (!sametype) {
-            return HTP_STATUS_NO_SUPPORT;
-        }
-        if (!sameshape && !(src_is_contiguous && dst_is_contiguous && octx->ctx->mdev.count <= 1)) {
-            return HTP_STATUS_NO_SUPPORT;
-        }
+    if (kparams->vtcm_size > octx->ctx->vtcm_size) {
+        return HTP_STATUS_VTCM_TOO_SMALL;
     }

-    if (sameshape) {
-        const uint32_t total_rows = ne01 * ne02 * ne03;
-        const uint32_t row_size   = ne00 * ct.dst_type_size;
+    const uint32_t n_threads = octx->n_threads;
+    const struct htp_copy_convert_params * cvt = &kparams->u.convert;

-        ct.div_ne01      = init_fastdiv_values(ne01);
-        ct.div_ne02_ne01 = init_fastdiv_values(ne02 * ne01);
+    struct htp_copy_context ct;
+    ct.octx                  = octx;
+    ct.kparams               = kparams;
+    ct.row_start             = row_start;
+    ct.nrows                 = nrows;
+    ct.src0_nrows_per_thread = fastdiv(nrows + n_threads - 1, &octx->n_threads_div);
+
+    uint8_t * vtcm_base = (uint8_t *) octx->ctx->vtcm_base;
+    ct.vtcm_src0 = vtcm_base;
+    ct.vtcm_dst  = vtcm_base + (size_t) n_threads * cvt->spad0_size_per_thread;
+
+    work_queue_func_t copy_fun = NULL;
+    if (dst->type == HTP_TYPE_F16 && src0->type == HTP_TYPE_F32) {
+        copy_fun = cpy_thread_f16_f32_sameshape;
+    } else if (dst->type == HTP_TYPE_F32 && src0->type == HTP_TYPE_F16) {
+        copy_fun = cpy_thread_f32_f16_sameshape;
+    } else if (dst->type == HTP_TYPE_I32 && src0->type == HTP_TYPE_F32) {
+        copy_fun = cpy_thread_i32_f32_sameshape;
+    } else if (dst->type == HTP_TYPE_F32 && src0->type == HTP_TYPE_I32) {
+        copy_fun = cpy_thread_f32_i32_sameshape;
+    } else {
+        return HTP_STATUS_NO_SUPPORT;
+    }

-        uint32_t row_start = 0;
-        uint32_t nrows     = total_rows;
+    work_queue_run(octx->ctx->work_queue, copy_fun, &ct, n_threads);
+    return HTP_STATUS_OK;
+}

-        if (octx->ctx->mdev.count > 1) {
-            const uint32_t rows_per_chunk = (row_size > 0) ? (HEX_L2_LINE_SIZE / hex_gcd_u32(row_size, HEX_L2_LINE_SIZE)) : 1;
-            const bool can_split = htp_tensor_mdev_data_aligned(dst) && dst_is_contiguous;
-            const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_rows, can_split ? rows_per_chunk : 0,
-                                                               octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
-            row_start = range.start;
-            nrows     = range.count;
-        }
+static int cpy_reshape(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+    const struct htp_tensor * src0 = octx->src[0];
+    const struct htp_tensor * dst  = octx->dst;

-        if (nrows == 0) {
-            return HTP_STATUS_OK;
-        }
+    if (htp_tensor_is_extended(src0) || htp_tensor_is_extended(dst)) {
+        return HTP_STATUS_NO_SUPPORT;
+    }

-        ct.row_start = row_start;
-        ct.nrows     = nrows;
-        ct.src0_nrows_per_thread = fastdiv(nrows + n_threads - 1, &octx->n_threads_div);
-
-        if (sametype && (octx->ctx->mdev.count <= 1 || htp_tensor_is_extended(src0) || htp_tensor_is_extended(dst))) {
-            if (octx->ctx->mdev.idx == 0) {
-                *use_dma = true;
-                cpy_dma_sametype_sameshape(octx->ctx->dma[0], dst, src0, ct.src0_type_size);
-                dma_queue_flush(octx->ctx->dma[0]);
-            }
-        } else {
-            work_queue_func_t copy_fun = NULL;
-            if (sametype) {
-                switch (src0->type) {
-                    case HTP_TYPE_F32: copy_fun = cpy_thread_f32_sameshape; break;
-                    case HTP_TYPE_F16: copy_fun = cpy_thread_f16_sameshape; break;
-                    case HTP_TYPE_I32: copy_fun = cpy_thread_i32_sameshape; break;
-                    default: return HTP_STATUS_NO_SUPPORT;
-                }
-            } else if (dst->type == HTP_TYPE_F16 && src0->type == HTP_TYPE_F32) {
-                copy_fun = cpy_thread_f16_f32_sameshape;
-            } else if (dst->type == HTP_TYPE_F32 && src0->type == HTP_TYPE_F16) {
-                copy_fun = cpy_thread_f32_f16_sameshape;
-            } else if (dst->type == HTP_TYPE_I32 && src0->type == HTP_TYPE_F32) {
-                copy_fun = cpy_thread_i32_f32_sameshape;
-            } else if (dst->type == HTP_TYPE_F32 && src0->type == HTP_TYPE_I32) {
-                copy_fun = cpy_thread_f32_i32_sameshape;
-            } else {
-                return HTP_STATUS_NO_SUPPORT;
-            }
-            work_queue_run(octx->ctx->work_queue, copy_fun, &ct, n_threads);
-        }
-    } else if (sametype) {
-        const uint32_t total_elems = ne0 * ne1 * ne2 * ne3;
-        const uint32_t elems_per_line = (ct.dst_type_size == 4) ? 32 : 64;
-
-        if (octx->ctx->mdev.count <= 1 && dst_is_contiguous && src_is_contiguous) {
-            *use_dma = true;
-            cpy_dma_sametype_reshape_contig(octx->ctx->dma[0], dst->data, src0->data, total_elems * ct.dst_type_size);
-            dma_queue_flush(octx->ctx->dma[0]);
-            return HTP_STATUS_OK;
-        }
+    if (!htp_ops_context_set_n_threads(octx, kparams->n_threads)) {
+        return HTP_STATUS_INVAL_PARAMS;
+    }

-        ct.div_ne0            = init_fastdiv_values(ne0);
-        ct.div_ne1_ne0        = init_fastdiv_values(ne1 * ne0);
-        ct.div_ne2_ne1_ne0    = init_fastdiv_values(ne2 * ne1 * ne0);
-        ct.div_ne00           = init_fastdiv_values(ne00);
-        ct.div_ne01_ne00      = init_fastdiv_values(ne01 * ne00);
-        ct.div_ne02_ne01_ne00 = init_fastdiv_values(ne02 * ne01 * ne00);
-
-        uint32_t elem_start = 0;
-        uint32_t nelem      = total_elems;
-
-        if (octx->ctx->mdev.count > 1) {
-            const bool can_split = htp_tensor_mdev_data_aligned(dst) && dst_is_contiguous;
-            const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_elems, can_split ? elems_per_line : 0,
-                                                               octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
-            elem_start = range.start;
-            nelem      = range.count;
-        }
+    uint32_t elem_start = 0;
+    uint32_t nelem      = kparams->total_elems;

-        if (nelem == 0) {
-            return HTP_STATUS_OK;
-        }
+    if (octx->ctx->mdev.count > 1) {
+        const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(
+            kparams->total_elems, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+        elem_start = range.start;
+        nelem      = range.count;
+    }

-        ct.elem_start      = elem_start;
-        ct.nelem           = nelem;
-        ct.elem_per_thread = fastdiv(nelem + n_threads - 1, &octx->n_threads_div);
+    if (nelem == 0) {
+        return HTP_STATUS_OK;
+    }

-        work_queue_func_t copy_fun = NULL;
-        switch (src0->type) {
-            case HTP_TYPE_F32: copy_fun = cpy_thread_f32_reshape; break;
-            case HTP_TYPE_F16: copy_fun = cpy_thread_f16_reshape; break;
-            case HTP_TYPE_I32: copy_fun = cpy_thread_i32_reshape; break;
-            default: return HTP_STATUS_NO_SUPPORT;
-        }
-        work_queue_run(octx->ctx->work_queue, copy_fun, &ct, n_threads);
-    } else {
-        return HTP_STATUS_NO_SUPPORT;
+    const uint32_t n_threads = octx->n_threads;
+
+    struct htp_copy_context ct;
+    ct.octx            = octx;
+    ct.kparams         = kparams;
+    ct.elem_start      = elem_start;
+    ct.nelem           = nelem;
+    ct.elem_per_thread = fastdiv(nelem + n_threads - 1, &octx->n_threads_div);
+
+    work_queue_func_t copy_fun = NULL;
+    switch (src0->type) {
+        case HTP_TYPE_F32: copy_fun = cpy_thread_f32_reshape; break;
+        case HTP_TYPE_F16: copy_fun = cpy_thread_f16_reshape; break;
+        case HTP_TYPE_I32: copy_fun = cpy_thread_i32_reshape; break;
+        default: return HTP_STATUS_NO_SUPPORT;
     }

+    work_queue_run(octx->ctx->work_queue, copy_fun, &ct, n_threads);
     return HTP_STATUS_OK;
 }

 int op_cpy(struct htp_ops_context * octx) {
-    bool use_dma = false;
-    int status = exec_cpy(octx, &use_dma);
+    const struct htp_copy_kernel_params * kparams = (const struct htp_copy_kernel_params *) octx->kernel_params;
+    int status = HTP_STATUS_OK;
+
+    switch (kparams->kernel_type) {
+        case HTP_COPY_KERNEL_SCALAR:
+            status = cpy_scalar(octx, kparams);
+            break;
+        case HTP_COPY_KERNEL_1D_CONTIG:
+            status = cpy_1d_contig(octx, kparams);
+            break;
+        case HTP_COPY_KERNEL_SAMESHAPE_SAMETYPE:
+            status = cpy_sameshape_sametype(octx, kparams);
+            break;
+        case HTP_COPY_KERNEL_SAMESHAPE_CONVERT:
+            status = cpy_sameshape_convert(octx, kparams);
+            break;
+        case HTP_COPY_KERNEL_RESHAPE:
+            status = cpy_reshape(octx, kparams);
+            break;
+        default:
+            status = HTP_STATUS_NO_SUPPORT;
+            break;
+    }

     htp_ops_context_set_status(octx, status);

-    if (octx->op == HTP_OP_CPY_FENCE) {
-        if (!use_dma) {
-            htp_flush_dirty_ranges(octx->ctx);
-        }
-
+    if (octx->ctx->mdev.count > 1) {
         htp_mdev_group_barrier(octx);
+    }

+    if (octx->op == HTP_OP_CPY_FENCE) {
         if (octx->ctx->mdev.idx == 0) {
             const struct htp_tensor * sync = octx->src[1];
-            if (htp_tensor_is_extended(sync)) {
-                return HTP_STATUS_NO_SUPPORT;
-            }
             const uint32_t seq = (uint32_t) octx->op_params[0];
             atomic_uint * sync_fence = (atomic_uint *) (uintptr_t) sync->data;
             htp_fence_write(sync_fence, seq, octx->status);
diff --git a/ggml/src/ggml-hexagon/htp/cpy-ops.h b/ggml/src/ggml-hexagon/htp/cpy-ops.h
new file mode 100644
index 000000000..4cf7dc13b
--- /dev/null
+++ b/ggml/src/ggml-hexagon/htp/cpy-ops.h
@@ -0,0 +1,79 @@
+#ifndef HTP_CPY_OPS_H
+#define HTP_CPY_OPS_H
+
+#include "hex-common.h"
+#include "hex-fastdiv.h"
+#include <stdint.h>
+
+enum htp_copy_kernel_type {
+    HTP_COPY_KERNEL_UNSUPPORTED        = 0,
+    HTP_COPY_KERNEL_1D_CONTIG          = 1,
+    HTP_COPY_KERNEL_SAMESHAPE_SAMETYPE = 2,
+    HTP_COPY_KERNEL_SAMESHAPE_CONVERT  = 3,
+    HTP_COPY_KERNEL_RESHAPE            = 4,
+    HTP_COPY_KERNEL_SCALAR             = 5,
+};
+
+struct htp_copy_convert_params {
+    uint32_t              src0_buf_size;
+    uint32_t              dst_buf_size;
+    uint32_t              spad0_size_per_thread;
+    uint32_t              spad1_size_per_thread;
+    struct fastdiv_values div_ne01;
+    struct fastdiv_values div_ne02_ne01;
+};
+
+struct htp_copy_reshape_params {
+    struct fastdiv_values div_ne0;
+    struct fastdiv_values div_ne1_ne0;
+    struct fastdiv_values div_ne2_ne1_ne0;
+    struct fastdiv_values div_ne00;
+    struct fastdiv_values div_ne01_ne00;
+    struct fastdiv_values div_ne02_ne01_ne00;
+};
+
+struct htp_copy_kernel_params {
+    uint8_t  kernel_type;
+    uint8_t  src0_type_size;
+    uint8_t  dst_type_size;
+    uint8_t  n_threads;
+
+    uint32_t total_elems;
+    uint32_t total_rows;
+    uint32_t vtcm_size;
+
+    union {
+        struct htp_copy_convert_params convert;
+        struct htp_copy_reshape_params reshape;
+    } u;
+};
+
+struct htp_copy_convert_vtcm_layout {
+    uint32_t src0_buf_size;
+    uint32_t dst_buf_size;
+    uint32_t spad0_size_per_thread;
+    uint32_t spad1_size_per_thread;
+    uint32_t total_bytes;
+};
+
+static inline void htp_copy_convert_vtcm_layout_build(
+    struct htp_copy_convert_vtcm_layout * layout,
+    uint32_t ne00,
+    uint32_t src_type_size,
+    uint32_t dst_type_size,
+    uint32_t n_threads) {
+
+    layout->src0_buf_size = hex_round_up(ne00 * src_type_size, 256);
+    layout->dst_buf_size  = hex_round_up(ne00 * dst_type_size, 256);
+    layout->spad0_size_per_thread = 2 * layout->src0_buf_size;
+    layout->spad1_size_per_thread = 2 * layout->dst_buf_size;
+    layout->total_bytes = n_threads * (layout->spad0_size_per_thread + layout->spad1_size_per_thread);
+}
+
+#if defined(__cplusplus)
+static_assert(sizeof(struct htp_copy_kernel_params) <= 128, "htp_copy_kernel_params is too large for kernel_params blob");
+#else
+_Static_assert(sizeof(struct htp_copy_kernel_params) <= 128, "htp_copy_kernel_params is too large for kernel_params blob");
+#endif
+
+#endif // HTP_CPY_OPS_H
diff --git a/ggml/src/ggml-hexagon/htp/hex-cpy-dma.h b/ggml/src/ggml-hexagon/htp/dma-copy.h
similarity index 58%
rename from ggml/src/ggml-hexagon/htp/hex-cpy-dma.h
rename to ggml/src/ggml-hexagon/htp/dma-copy.h
index c87f495e2..b2c91a13c 100644
--- a/ggml/src/ggml-hexagon/htp/hex-cpy-dma.h
+++ b/ggml/src/ggml-hexagon/htp/dma-copy.h
@@ -1,9 +1,9 @@
-#ifndef HEX_CPY_DMA_H
-#define HEX_CPY_DMA_H
+#ifndef HTP_DMA_COPY_H
+#define HTP_DMA_COPY_H

 // DDR<->DDR DMA copies of same-type, same-shape tensors with arbitrary strides.
 // Used by CPY for the copy itself and by CONCAT, which is two such copies into
-// two views of its destination.  Every helper only pushes descriptors; the
+// two views of its destination. Every helper only pushes descriptors; the
 // caller flushes the queue when it needs the data.

 #include "dma-queue.h"
@@ -14,7 +14,7 @@
 #include <stdint.h>

 // Contiguous byte run, as 1d transfers of at most DMA_SAFE_CHUNK_SIZE each.
-static inline void cpy_dma_sametype_reshape_contig(dma_queue * dma_q,
+static inline void dma_cpy_sametype_reshape_contig(dma_queue * dma_q,
                                                    dma_addr_t  dst,
                                                    dma_addr_t  src0,
                                                    uint32_t    total_bytes) {
@@ -36,7 +36,7 @@ static inline void cpy_dma_sametype_reshape_contig(dma_queue * dma_q,
 }

 // One 2d transfer, split at the 16-bit nrows field.
-static inline void cpy_dma_push_2d_chunked(dma_queue * dma_q,
+static inline void dma_cpy_push_2d_chunked(dma_queue * dma_q,
                                            dma_addr_t  dst,
                                            dma_addr_t  src,
                                            size_t      dst_stride,
@@ -59,12 +59,18 @@ static inline void cpy_dma_push_2d_chunked(dma_queue * dma_q,
     }
 }

-// Copy src0 into dst: same type, same ne[], any nb[] above dim 0, dim 0 dense on
-// both sides (nb[0] == elem_size).
-static inline void cpy_dma_sametype_sameshape(dma_queue *               dma_q,
-                                              const struct htp_tensor * dst,
-                                              const struct htp_tensor * src0,
-                                              uint32_t                  elem_size) {
+// Copy a range of rows [row_start, row_start + nrows) from src0 into dst:
+// same type, same ne[], any nb[] above dim 0, dim 0 dense on both sides (nb[0] == elem_size).
+static inline void dma_cpy_sametype_sameshape_range(dma_queue *               dma_q,
+                                                    const struct htp_tensor * dst,
+                                                    const struct htp_tensor * src0,
+                                                    uint32_t                  elem_size,
+                                                    uint32_t                  row_start,
+                                                    uint32_t                  nrows) {
+    if (nrows == 0) {
+        return;
+    }
+
     const uint32_t ne00 = src0->ne[0];
     const uint32_t ne01 = src0->ne[1];
     const uint32_t ne02 = src0->ne[2];
@@ -85,13 +91,16 @@ static inline void cpy_dma_sametype_sameshape(dma_queue *               dma_q,
     const bool contiguous = htp_tensor_is_contiguous(src0, elem_size) && htp_tensor_is_contiguous(dst, elem_size);

     if (contiguous) {
-        cpy_dma_sametype_reshape_contig(dma_q, dst->data, src0->data, ne00 * elem_size * ne01 * ne02 * ne03);
+        dma_cpy_sametype_reshape_contig(dma_q,
+                                        dst->data  + (dma_addr_t) row_start * ne00 * elem_size,
+                                        src0->data + (dma_addr_t) row_start * ne00 * elem_size,
+                                        nrows * ne00 * elem_size);
         return;
     }

     // The single-descriptor path flattens (i01,i02,i03) into one row index, so every
     // row must sit at a constant stride: nb01 on the source, nb1 on the destination.
-    // Walk the outer dims and require each to continue that progression.  A dim of
+    // Walk the outer dims and require each to continue that progression. A dim of
     // extent 1 spans no rows, so it is skipped -- but its own stride must NOT then be
     // used to justify the next dim's stride, which is what comparing nb03 against
     // ne02*nb02 did: ggml leaves the stride of an extent-1 dim meaningless, so a view
@@ -109,18 +118,52 @@ static inline void cpy_dma_sametype_sameshape(dma_queue *               dma_q,
     }

     if (contiguous_outer) {
-        uint32_t total_rows = ne01 * ne02 * ne03;
-        cpy_dma_push_2d_chunked(dma_q, dst->data, src0->data, nb1, nb01, ne00 * elem_size, total_rows);
+        dma_cpy_push_2d_chunked(dma_q,
+                                dst->data  + (dma_addr_t) row_start * nb1,
+                                src0->data + (dma_addr_t) row_start * nb01,
+                                nb1, nb01, ne00 * elem_size, nrows);
         return;
     }

-    for (uint32_t i03 = 0; i03 < ne03; i03++) {
-        for (uint32_t i02 = 0; i02 < ne02; i02++) {
-            dma_addr_t dst_data  = dst->data + i02 * nb2 + i03 * nb3;
-            dma_addr_t src0_data = src0->data + i02 * nb02 + i03 * nb03;
-            cpy_dma_push_2d_chunked(dma_q, dst_data, src0_data, nb1, nb01, ne00 * elem_size, ne01);
+    const uint32_t ne02_ne01 = ne02 * ne01;
+    uint32_t i03 = row_start / ne02_ne01;
+    uint32_t rem = row_start - i03 * ne02_ne01;
+    uint32_t i02 = rem / ne01;
+    uint32_t i01 = rem - i02 * ne01;
+
+    dma_addr_t cur_dst  = dst->data  + (dma_addr_t) i01 * nb1  + (dma_addr_t) i02 * nb2  + (dma_addr_t) i03 * nb3;
+    dma_addr_t cur_src0 = src0->data + (dma_addr_t) i01 * nb01 + (dma_addr_t) i02 * nb02 + (dma_addr_t) i03 * nb03;
+
+    uint32_t r = row_start;
+    const uint32_t row_end = row_start + nrows;
+    while (r < row_end) {
+        uint32_t cur_rows = MIN(row_end - r, ne01 - i01);
+        dma_cpy_push_2d_chunked(dma_q, cur_dst, cur_src0, nb1, nb01, ne00 * elem_size, cur_rows);
+        r   += cur_rows;
+        i01 += cur_rows;
+        if (i01 == ne01) {
+            i01 = 0;
+            if (++i02 == ne02) {
+                i02 = 0;
+                i03++;
+            }
+            cur_dst  = dst->data  + (dma_addr_t) i02 * nb2  + (dma_addr_t) i03 * nb3;
+            cur_src0 = src0->data + (dma_addr_t) i02 * nb02 + (dma_addr_t) i03 * nb03;
+        } else {
+            cur_dst  += cur_rows * nb1;
+            cur_src0 += cur_rows * nb01;
         }
     }
 }

-#endif /* HEX_CPY_DMA_H */
+// Copy src0 into dst: same type, same ne[], any nb[] above dim 0, dim 0 dense on
+// both sides (nb[0] == elem_size).
+static inline void dma_cpy_sametype_sameshape(dma_queue *               dma_q,
+                                              const struct htp_tensor * dst,
+                                              const struct htp_tensor * src0,
+                                              uint32_t                  elem_size) {
+    const uint32_t total_rows = src0->ne[1] * src0->ne[2] * src0->ne[3];
+    dma_cpy_sametype_sameshape_range(dma_q, dst, src0, elem_size, 0, total_rows);
+}
+
+#endif /* HTP_DMA_COPY_H */
diff --git a/ggml/src/ggml-hexagon/htp/hex-utils.h b/ggml/src/ggml-hexagon/htp/hex-utils.h
index 853f1c1b2..abb44da25 100644
--- a/ggml/src/ggml-hexagon/htp/hex-utils.h
+++ b/ggml/src/ggml-hexagon/htp/hex-utils.h
@@ -38,21 +38,15 @@ static inline void hex_l2fetch_block(const void * addr, size_t size) {
 }

 #define HEX_L2_LINE_SIZE           128
-#define HEX_L2_BLOCK_SIZE          (HEX_L2_LINE_SIZE * 4) // flush granularity (lines per loop iteration)
+#define HEX_L2_BLOCK_SIZE          (HEX_L2_LINE_SIZE * 4) // flush granularity (chunks per thread)
 #define HEX_L2_FLUSH_WQ_THRESHOLD  (4 * 1024)
 #define HEX_L2_FLUSH_ALL_THRESHOLD (4 * 1024 * 1024)

 static inline void hex_l2flush(void * addr, size_t size) {
+    if (size == 0) return;
     const uint32_t s = ((uint32_t) addr) & ~(HEX_L2_LINE_SIZE - 1);
     const uint32_t e = (((uint32_t) addr) + size + HEX_L2_LINE_SIZE - 1) & ~(HEX_L2_LINE_SIZE - 1);
-    const uint32_t eb = s + ((e - s) & ~(HEX_L2_BLOCK_SIZE - 1));
-    for (uint32_t i = s; i < eb; i += HEX_L2_BLOCK_SIZE) {
-        Q6_dccleaninva_A((void *) (i + HEX_L2_LINE_SIZE * 0));
-        Q6_dccleaninva_A((void *) (i + HEX_L2_LINE_SIZE * 1));
-        Q6_dccleaninva_A((void *) (i + HEX_L2_LINE_SIZE * 2));
-        Q6_dccleaninva_A((void *) (i + HEX_L2_LINE_SIZE * 3));
-    }
-    for (uint32_t i = eb; i < e; i += HEX_L2_LINE_SIZE) {
+    for (uint32_t i = s; i < e; i += HEX_L2_LINE_SIZE) {
         Q6_dccleaninva_A((void *) i);
     }
 }
diff --git a/ggml/src/ggml-hexagon/htp/hvx-copy.h b/ggml/src/ggml-hexagon/htp/hvx-copy.h
index a3e33c3b3..08eb7cd61 100644
--- a/ggml/src/ggml-hexagon/htp/hvx-copy.h
+++ b/ggml/src/ggml-hexagon/htp/hvx-copy.h
@@ -44,7 +44,7 @@ static inline void hvx_splat_f32_u(void * restrict dst, float v, uint32_t n) {
 }

 static inline void hvx_splat_f16_a(void * restrict dst, _Float16 v, uint32_t n) {
-    hvx_splat_u(dst,  hvx_vec_splat_f16(v), n, sizeof(__fp16));
+    hvx_splat_a(dst,  hvx_vec_splat_f16(v), n, sizeof(__fp16));
 }

 static inline void hvx_splat_f16_u(void * restrict dst, _Float16 v, uint32_t n) {
@@ -106,42 +106,42 @@ static inline void hvx_copy_uu(uint8_t * restrict dst, const uint8_t * restrict
     hvx_copy_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
 }

-// copy n fp16 elements : source and destination are aligned to HVX Vector (128)
+// copy n fp16 elements : destination and source are aligned to HVX Vector (128)
 static inline void hvx_copy_f16_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_aa(dst, src, n, sizeof(__fp16));
 }

-// copy n fp16 elements : source is aligned, destination is potentially unaligned
+// copy n fp16 elements : destination is aligned, source is unaligned
 static inline void hvx_copy_f16_au(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_au(dst, src, n, sizeof(__fp16));
 }

-// copy n fp16 elements : source is aligned, destination is potentially unaligned
+// copy n fp16 elements : destination is unaligned, source is aligned
 static inline void hvx_copy_f16_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_ua(dst, src, n, sizeof(__fp16));
 }

-// copy n fp16 elements : source is aligned, destination is potentially unaligned
+// copy n fp16 elements : destination and source are unaligned
 static inline void hvx_copy_f16_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_uu(dst, src, n, sizeof(__fp16));
 }

-// copy n fp32 elements : source and destination are aligned to HVX Vector (128)
+// copy n fp32 elements : destination and source are aligned to HVX Vector (128)
 static inline void hvx_copy_f32_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_aa(dst, src, n, sizeof(float));
 }

-// copy n fp32 elements : source is aligned, destination is unaligned
-static inline void hvx_copy_f32_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
-    hvx_copy_ua(dst, src, n, sizeof(float));
-}
-
-// copy n fp32 elements : source is unaligned, destination is aligned
+// copy n fp32 elements : destination is aligned, source is unaligned
 static inline void hvx_copy_f32_au(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_au(dst, src, n, sizeof(float));
 }

-// copy n fp32 elements : source is unaligned, destination unaligned
+// copy n fp32 elements : destination is unaligned, source is aligned
+static inline void hvx_copy_f32_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+    hvx_copy_ua(dst, src, n, sizeof(float));
+}
+
+// copy n fp32 elements : destination and source are unaligned
 static inline void hvx_copy_f32_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_uu(dst, src, n, sizeof(float));
 }
@@ -170,26 +170,26 @@ static inline void hvx_copy_f32_uu(uint8_t * restrict dst, const uint8_t * restr
         }                                                                           \
     } while(0)

-// copy/convert n fp32 elements into n fp16 elements : source is aligned, destination is aligned
+// copy/convert n fp32 elements into n fp16 elements : destination and source are aligned
 static inline void hvx_copy_f16_f32_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     assert((unsigned long) dst % 128 == 0);
     assert((unsigned long) src % 128 == 0);
     hvx_copy_f16_f32_loop_body(HVX_Vector, HVX_Vector, hvx_vec_store_a);
 }

-// copy/convert n fp32 elements into n fp16 elements : source is unaligned, destination is aligned
+// copy/convert n fp32 elements into n fp16 elements : destination is aligned, source is unaligned
 static inline void hvx_copy_f16_f32_au(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     assert((unsigned long) dst % 128 == 0);
     hvx_copy_f16_f32_loop_body(HVX_Vector, HVX_UVector, hvx_vec_store_a);
 }

-// copy/convert n fp32 elements into n fp16 elements : source is aligned, destination is unaligned
+// copy/convert n fp32 elements into n fp16 elements : destination is unaligned, source is aligned
 static inline void hvx_copy_f16_f32_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     assert((unsigned long) src % 128 == 0);
     hvx_copy_f16_f32_loop_body(HVX_UVector, HVX_Vector, hvx_vec_store_u);
 }

-// copy/convert n fp32 elements into n fp16 elements : source is unaligned, destination is unaligned
+// copy/convert n fp32 elements into n fp16 elements : destination and source are unaligned
 static inline void hvx_copy_f16_f32_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_f16_f32_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
 }
@@ -235,28 +235,98 @@ static inline void hvx_copy_f16_f32_uu(uint8_t * restrict dst, const uint8_t * r
         }                                                                           \
     } while(0)

-// copy/convert n fp16 elements into n fp32 elements : source is aligned, destination is aligned
+// copy/convert n fp16 elements into n fp32 elements : destination and source are aligned
 static inline void hvx_copy_f32_f16_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     assert((unsigned long) dst % 128 == 0);
     assert((unsigned long) src % 128 == 0);
     hvx_copy_f32_f16_loop_body(HVX_Vector, HVX_Vector, hvx_vec_store_a);
 }

-// copy/convert n fp16 elements into n fp32 elements : source is unaligned, destination is aligned
+// copy/convert n fp16 elements into n fp32 elements : destination is aligned, source is unaligned
 static inline void hvx_copy_f32_f16_au(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     assert((unsigned long) dst % 128 == 0);
     hvx_copy_f32_f16_loop_body(HVX_Vector, HVX_UVector, hvx_vec_store_a);
 }

-// copy/convert n fp16 elements into n fp32 elements : source is aligned, destination is unaligned
+// copy/convert n fp16 elements into n fp32 elements : destination is unaligned, source is aligned
 static inline void hvx_copy_f32_f16_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     assert((unsigned long) src % 128 == 0);
     hvx_copy_f32_f16_loop_body(HVX_UVector, HVX_Vector, hvx_vec_store_u);
 }

-// copy/convert n fp16 elements into n fp32 elements : source is unaligned, destination is unaligned
+// copy/convert n fp16 elements into n fp32 elements : destination and source are unaligned
 static inline void hvx_copy_f32_f16_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
     hvx_copy_f32_f16_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
 }

+//// fp32 -> int32
+
+#define hvx_copy_i32_f32_loop_body(dst_type, src_type, vec_store) \
+    do {                                                          \
+        dst_type * restrict vdst = (dst_type *) dst;              \
+        src_type * restrict vsrc = (src_type *) src;              \
+                                                                  \
+        const uint32_t elem_size = sizeof(int32_t);               \
+        const uint32_t epv  = 128 / elem_size;                    \
+        const uint32_t nvec = n / epv;                            \
+        const uint32_t nloe = n % epv;                            \
+                                                                  \
+        uint32_t i = 0;                                           \
+        _Pragma("unroll(4)")                                      \
+        for (; i < nvec; i++) {                                   \
+            vdst[i] = Q6_Vw_equals_Vsf(vsrc[i]);                  \
+        }                                                         \
+        if (nloe) {                                               \
+            HVX_Vector v = Q6_Vw_equals_Vsf(vsrc[i]);             \
+            vec_store((void *) &vdst[i], nloe * elem_size, v);    \
+        }                                                         \
+    } while(0)
+
+// copy/convert n fp32 elements into n int32 elements : destination and source are aligned
+static inline void hvx_copy_i32_f32_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+    assert((unsigned long) dst % 128 == 0);
+    assert((unsigned long) src % 128 == 0);
+    hvx_copy_i32_f32_loop_body(HVX_Vector, HVX_Vector, hvx_vec_store_a);
+}
+
+// copy/convert n fp32 elements into n int32 elements : destination and source are unaligned
+static inline void hvx_copy_i32_f32_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+    hvx_copy_i32_f32_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
+}
+
+//// int32 -> fp32
+
+#define hvx_copy_f32_i32_loop_body(dst_type, src_type, vec_store) \
+    do {                                                          \
+        dst_type * restrict vdst = (dst_type *) dst;              \
+        src_type * restrict vsrc = (src_type *) src;              \
+                                                                  \
+        const uint32_t elem_size = sizeof(float);                 \
+        const uint32_t epv  = 128 / elem_size;                    \
+        const uint32_t nvec = n / epv;                            \
+        const uint32_t nloe = n % epv;                            \
+                                                                  \
+        uint32_t i = 0;                                           \
+        _Pragma("unroll(4)")                                      \
+        for (; i < nvec; i++) {                                   \
+            vdst[i] = Q6_Vsf_equals_Vw(vsrc[i]);                  \
+        }                                                         \
+        if (nloe) {                                               \
+            HVX_Vector v = Q6_Vsf_equals_Vw(vsrc[i]);             \
+            vec_store((void *) &vdst[i], nloe * elem_size, v);    \
+        }                                                         \
+    } while(0)
+
+// copy/convert n int32 elements into n fp32 elements : destination and source are aligned
+static inline void hvx_copy_f32_i32_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+    assert((unsigned long) dst % 128 == 0);
+    assert((unsigned long) src % 128 == 0);
+    hvx_copy_f32_i32_loop_body(HVX_Vector, HVX_Vector, hvx_vec_store_a);
+}
+
+// copy/convert n int32 elements into n fp32 elements : destination and source are unaligned
+static inline void hvx_copy_f32_i32_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+    hvx_copy_f32_i32_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
+}
+
 #endif // HVX_COPY_H