Re: [PATCH v6 06/18] nvme: Rapid Path Failure Recovery read controller identify fields

From: Achkinazi, Igor

Date: Tue Oct 06 2026 - 11:25:57 EST


Hi Randy, Mohamed,

We have tested the v6 patchset against our PowerFlex NVMe-oF target
implementation with TP8028 (Cross-Controller Reset) support.

Our testing covered:

- CCR for an impacted controller on the same node as the source
controller
- CCR for an impacted controller on a remote node in a distributed
system
- Fallback to CQT-based timeout recovery (2*KATO + CQT) when CCR
is not supported
- CCR limit exceeded and log page full error handling (tested on v3,
still applicable)

CCR dramatically reduces recovery time compared to the (2*KATO + CQT)
timeout path.

When PowerFlex returns CCR operation success (via IRS or the CCR log
page), the patched host resumes I/O immediately. When the target
returns success without the Validated/CLRI flags set, the host
correctly falls back to timeout-based recovery. On failure, the host
appropriately retries CCR via another controller.

We also confirmed that without this patchset, the Linux host can
retry commands too quickly during path recovery, resulting in
duplicate commands outstanding at storage simultaneously. This
patchset's fix for that retry behavior is long overdue.

We strongly support acceptance of this patchset. It makes a
significant and measurable improvement to Linux NVMe-oF usability
and correctness.

Tested-by: Igor Achkinazi <Igor.Achkinazi@xxxxxxxx>

Thank you,
Igor


Internal Use - Confidential