Commit 6c657a97 for libheif
commit 6c657a97cd4dd2d575d31078de743aebeb8d2802
Author: Dirk Farin <dirk.farin@gmail.com>
Date: Fri Oct 9 23:02:24 2026 +0200
Pad plane strides that are multiples of 512 bytes by one cache line
L1 data caches are set associative: the set that holds a cache line is
selected by the line's position within a 4 KB (x86) or 8/16 KB (ARM)
window of the address space, and a set holds only 4 to 12 lines. With a
stride that is a multiple of 512 bytes, 8 or more of any 64 consecutive
rows map to the same set, so a loop that walks down a column of rows,
such as the blocked rotation from #1934, evicts the rows before it comes
back to them and runs from L2 instead of L1.
Adding one cache line makes the stride an odd number of lines, which
spreads consecutive rows over all sets. Rotating a 4096x3072 4:2:0 image
(a common binned output size of 50 MP phone sensors) on a Ryzen 9 9950X,
median of 5 interleaved runs: 10 bit from 28.5 ms to 8.3 ms, 8 bit from
5.2 ms to 4.0 ms. Images whose strides are not multiples of 512, such as
4032x3024, are unchanged. The padding costs 64 bytes per row on the
affected planes. Decoded output is byte-identical.
diff --git a/libheif/image/pixelimage.cc b/libheif/image/pixelimage.cc
index e3735d19..7c57a1c3 100644
--- a/libheif/image/pixelimage.cc
+++ b/libheif/image/pixelimage.cc
@@ -519,6 +519,17 @@ Error HeifPixelImage::ComponentStorage::alloc(uint32_t width, uint32_t height, h
uint64_t stride_64 = static_cast<uint64_t>(m_mem_width) * bytes_per_pixel;
stride_64 = (stride_64 + alignment - 1U) & ~static_cast<uint64_t>(alignment - 1U);
+
+ // L1 data caches are set associative. The set that holds a cache line is selected by the
+ // line's position within a 4 KB (x86) or 8/16 KB (ARM) window of the address space, and a set
+ // holds only 4 to 12 lines. With a stride that is a multiple of 512 bytes, 8 or more of any 64
+ // consecutive rows fall into the same set, so a loop that walks down a column of rows, such as
+ // the blocked rotation, evicts the rows before it comes back to them. One extra cache line
+ // makes the stride an odd number of lines, which spreads consecutive rows over all sets.
+ if (stride_64 % 512 == 0) {
+ stride_64 += 64;
+ }
+
if (stride_64 > std::numeric_limits<size_t>::max()) {
return {heif_error_Memory_allocation_error,
heif_suberror_Security_limit_exceeded,