Commit 7a3c0289c3c8 for kernel

commit 7a3c0289c3c8eb4607dff448ae9ff9f902c813af
Author: Kumar Kartikeya Dwivedi <memxor@gmail.com>
Date:   Sun Aug 2 04:17:59 2026 +0200

    rqspinlock: Reset tail when preserving queue on deadlock

    Currently, the destruction of the waiter queue is suppressed for
    rqspinlock in cases where a deadlock is detected. Deadlock checks happen
    relatively frequently (on entry for AA, within 1ms for ABBA), and waiter
    threads may not be involved in locking scenarios involving deadlocks.
    Thus, it is useful to not flush the queue and let other waiters take a
    stab at acquiring the lock after we detect a deadlock and exit.

    However, we need to follow the same logic as what we did previously for
    the waitq_timeout label: reset the tail, and if we cannot, signal the
    next waiter appropriately. In case of deadlocks, this signal would just
    mark the MCS node as unlocked, and in case of timeouts, it would signal
    RES_TIMEOUT_VAL. The difference thus is in the value propagated, which
    decides whether the queue remains active or gets flushed.

    Not doing the tail reset, and waiting for the next waiter can lead to
    cases where we are the final waiter, and thus no next waiter arrives,
    leading to intermittent stalls in this path. Once the next waiter does
    join, we will be unblocked. In the theoretical case when the next waiter
    never joins, we risk stalling indefinitely.

    This can only happen for ABBA deadlocks, since entry into the wait queue
    is guarded with AA checks. A precise sequence of executions leading up
    to this scenario can be:

    CPU 0 holds lock A.
    CPU 1 holds lock B.
    CPU 2 attempts lock B, becomes the pending waiter for B.
    CPU 0 attempts lock B. B has locked+pending bits set, thus CPU 0 queues.
    CPU 1 attempts lock A.
    CPU 0 detects an ABBA deadlock.

    Once deadlock detection happens for CPU 0, it will sit waiting for the
    next waiter in the queue to populate node->next, which will experience
    delays until such a waiter arrives.

    Fix this by adjusting the logic for the check for deadlocks preceding
    the waitq_timeout label. It would make sense to consolidate code for
    both cases and use 'ret' to distinguish the value being propagated, but
    that is left as an exercise for a future refactoring task to avoid diff
    noise in this patch.

    Fixes: 7bd6e5ce5be6 ("rqspinlock: Disable queue destruction for deadlocks")
    Signed-off-by: Kumar Kartikeya Dwivedi <memxor@gmail.com>
    Link: https://patch.msgid.link/20260802021759.1139457-1-memxor@gmail.com
    Signed-off-by: Eduard Zingerman <eddyz87@gmail.com>

diff --git a/kernel/bpf/rqspinlock.c b/kernel/bpf/rqspinlock.c
index e4e338cdb437..2129defc4a9a 100644
--- a/kernel/bpf/rqspinlock.c
+++ b/kernel/bpf/rqspinlock.c
@@ -572,9 +572,10 @@ int __lockfunc resilient_queued_spin_lock_slowpath(rqspinlock_t *lock, u32 val)

 	/* Disable queue destruction when we detect deadlocks. */
 	if (ret == -EDEADLK) {
-		if (!next)
+		if (!try_cmpxchg_tail(lock, tail, 0)) {
 			next = smp_cond_load_relaxed(&node->next, (VAL));
-		arch_mcs_spin_unlock_contended(&next->locked);
+			arch_mcs_spin_unlock_contended(&next->locked);
+		}
 		goto err_release_node;
 	}