Re: [RFC net-next 0/2] tipc: notify pending connects of peer node loss

From: shike liu

Date: Wed Oct 07 2026 - 21:52:12 EST


Hi,

I would like to clarify the next step for this RFC series.

The issue was observed on actual network devices operating in a
cluster. I have provided the discovery background in this thread.
The same selftest produced 8 passes and 10 failures on the unmodified
base, and all 18 tests passed with the series applied.

I have not received technical feedback on the approach yet. Since
this fixes an observed connection setup failure, would net be the
preferred target rather than net-next?

I have not yet reliably identified the original introducing commit
for a Fixes tag. I would appreciate guidance on that point and on
whether to proceed with a non-RFC submission.

Regards,
liushike

shike liu <liushike123@xxxxxxxxx> 于2026年9月30日周三 10:08写道:
>
> The issue was first observed on actual network devices operating in a
> cluster. The client was waiting in connect() for the server to accept the
> connection. When contact with the peer node was lost, the pending
> connect did not return promptly and instead continued waiting for its
> connection timeout.
>
> I subsequently investigated the TIPC connection setup and node-loss
> paths with LLM assistance. The analysis showed that pending connects
> were not registered in the peer node's conn_sks list, so
> node_lost_contact() did not notify them. The initial report came from
> the observed device failure, rather than an LLM or static-analysis scan.
>
> I then reproduced the behavior using the same final selftest on the
> unmodified base kernel and on that base with this series applied, in
> x86_64 QEMU/KVM guests. The tests use two network namespaces connected
> by veth pairs and disable TIPC Ethernet bearers to simulate loss of
> connectivity.
>
> On the unmodified kernel, 8 tests passed and 10 failed. With the series
> applied, all 18 tests passed, covering SOCK_STREAM and SOCK_SEQPACKET.
> The node-loss cases check completion within two seconds and the
> expected EHOSTUNREACH error.
>
> I will include this discovery background in the commit message if a
> revised series is needed. I am not reposting the series solely for this
> clarification.
>
> Regards,
> liushike
>
> shike liu <liushike123@xxxxxxxxx> 于2026年9月30日周三 10:01写道:
> >
> > Hi,
> >
> > > - How the issue was discovered, e.g. hit in production, hit during
> > > development, syzbot report, manual code inspection, LLM or static
> > > analysis tool scan.
> >
> > The issue was first observed on actual network devices operating in a
> > cluster. The client was waiting in connect() for the server to accept the
> > connection. When contact with the peer node was lost, the pending
> > connect did not return promptly and instead continued waiting for its
> > connection timeout.
> >
> > I subsequently investigated the TIPC connection setup and node-loss
> > paths with LLM assistance. The analysis showed that pending connects
> > were not registered in the peer node's conn_sks list, so
> > node_lost_contact() did not notify them. The initial report came from
> > the observed device failure, rather than an LLM or static-analysis scan.
> >
> > I then reproduced the behavior using the same final selftest on the
> > unmodified base kernel and on that base with this series applied, in
> > x86_64 QEMU/KVM guests. The tests use two network namespaces connected
> > by veth pairs and disable TIPC Ethernet bearers to simulate loss of
> > connectivity.
> >
> > On the unmodified kernel, 8 tests passed and 10 failed. With the series
> > applied, all 18 tests passed, covering SOCK_STREAM and SOCK_SEQPACKET.
> > The node-loss cases check completion within two seconds and the
> > expected EHOSTUNREACH error.
> >
> > I will include this discovery background in the commit message if a
> > revised series is needed. I am not reposting the series solely for this
> > clarification.
> >
> > Regards,
> > liushike
> >
> > <netdev-bot+sinfo@xxxxxxxxxx> 于2026年9月29日周二 14:54写道:
> >>
> >> Hi!
> >>
> >> This is an automated message. This series looks like a fix, but its
> >> commit messages seem to be missing some information:
> >>
> >> - How the issue was discovered, e.g. hit in production, hit during
> >> development, syzbot report, manual code inspection, LLM or static
> >> analysis tool scan.
> >>
> >> Please do not repost the series just to address the above. Instead,
> >> reply to this email with the missing information, so that reviewers
> >> can take it into account. If the series needs another revision for
> >> other reasons, please include the information in the commit messages
> >> then.
> >>
> >> The evaluation is done by an LLM so it may be wrong, if you think
> >> that is the case please reply and explain.