Re: [PATCH v6 1/2] md: Don't set MD_BROKEN for RAID1 and RAID10 when using FailFast
From: Martin Wilck
Date: Thu Sep 24 2026 - 14:40:31 EST
On Thu, 2026-09-24 at 16:27 +0000, Kenta Akagi wrote:
>
>
> On 2026/09/23 17:29, Martin Wilck wrote:
> > Hi Kenta,
> >
> > On Wed, 2026-09-23 at 04:08 +0000, Kenta Akagi wrote:
> > >
> > >
> > > Hi Martin,
> > >
> > > I have not given up on it, but I have not managed to post v7 yet.
> > > I still think failfast should be usable even in setups like that.
> > >
> > > It has been a while, but I intend to resume work on it.
> >
> > My thinking is that, in the fastfail case, code to prevent total
> > failure could be placed in the end-IO code code path, e.g. by
> > attempting a retry directly from the md layer when the last rdev
> > fails,
> > instead of setting failing the device. But I haven't thought it
> > through.
>
> Hi Martin,
>
> I may be misunderstanding your suggestion, but Neil's original
> failfast
> implementation already retries a failed failfast I/O to the last
> rdev.
> But a later change introduced a regression, which I intend to fix.
Ah OK, I wasn't aware of that.
> So the sequence should be:
>
> 1. The failfast bios to all mirrored rdevs fail.
> 2. The first rdev is marked faulty because its bio failed.
> 3. The other rdev is now the last, so it is not marked faulty.
> 4. Since the last rdev remains usable, its bio error handler
> retries the I/O without failfast.
Yes, that makes sense to me.
Martin
--
Dr. Martin Wilck <mwilck@xxxxxxxx>
SUSE Software Solutions Germany GmbH, Frankenstr. 146, 90461 Nürnberg,
Germany
Geschäftsführer: Stefan Gaiser, Jochen Jaser, Abhinav Puri (HRB
36809,AG Nürnberg)