Re: [RFC net-next 0/2] tipc: notify pending connects of peer node loss

From: shike liu

Date: Tue Sep 29 2026 - 22:09:03 EST


The issue was first observed on actual network devices operating in a
cluster. The client was waiting in connect() for the server to accept the
connection. When contact with the peer node was lost, the pending
connect did not return promptly and instead continued waiting for its
connection timeout.

I subsequently investigated the TIPC connection setup and node-loss
paths with LLM assistance. The analysis showed that pending connects
were not registered in the peer node's conn_sks list, so
node_lost_contact() did not notify them. The initial report came from
the observed device failure, rather than an LLM or static-analysis scan.

I then reproduced the behavior using the same final selftest on the
unmodified base kernel and on that base with this series applied, in
x86_64 QEMU/KVM guests. The tests use two network namespaces connected
by veth pairs and disable TIPC Ethernet bearers to simulate loss of
connectivity.

On the unmodified kernel, 8 tests passed and 10 failed. With the series
applied, all 18 tests passed, covering SOCK_STREAM and SOCK_SEQPACKET.
The node-loss cases check completion within two seconds and the
expected EHOSTUNREACH error.

I will include this discovery background in the commit message if a
revised series is needed. I am not reposting the series solely for this
clarification.

Regards,
liushike

shike liu <liushike123@xxxxxxxxx> 于2026年9月30日周三 10:01写道:
>
> Hi,
>
> > - How the issue was discovered, e.g. hit in production, hit during
> > development, syzbot report, manual code inspection, LLM or static
> > analysis tool scan.
>
> The issue was first observed on actual network devices operating in a
> cluster. The client was waiting in connect() for the server to accept the
> connection. When contact with the peer node was lost, the pending
> connect did not return promptly and instead continued waiting for its
> connection timeout.
>
> I subsequently investigated the TIPC connection setup and node-loss
> paths with LLM assistance. The analysis showed that pending connects
> were not registered in the peer node's conn_sks list, so
> node_lost_contact() did not notify them. The initial report came from
> the observed device failure, rather than an LLM or static-analysis scan.
>
> I then reproduced the behavior using the same final selftest on the
> unmodified base kernel and on that base with this series applied, in
> x86_64 QEMU/KVM guests. The tests use two network namespaces connected
> by veth pairs and disable TIPC Ethernet bearers to simulate loss of
> connectivity.
>
> On the unmodified kernel, 8 tests passed and 10 failed. With the series
> applied, all 18 tests passed, covering SOCK_STREAM and SOCK_SEQPACKET.
> The node-loss cases check completion within two seconds and the
> expected EHOSTUNREACH error.
>
> I will include this discovery background in the commit message if a
> revised series is needed. I am not reposting the series solely for this
> clarification.
>
> Regards,
> liushike
>
> <netdev-bot+sinfo@xxxxxxxxxx> 于2026年9月29日周二 14:54写道:
>>
>> Hi!
>>
>> This is an automated message. This series looks like a fix, but its
>> commit messages seem to be missing some information:
>>
>> - How the issue was discovered, e.g. hit in production, hit during
>> development, syzbot report, manual code inspection, LLM or static
>> analysis tool scan.
>>
>> Please do not repost the series just to address the above. Instead,
>> reply to this email with the missing information, so that reviewers
>> can take it into account. If the series needs another revision for
>> other reasons, please include the information in the commit messages
>> then.
>>
>> The evaluation is done by an LLM so it may be wrong, if you think
>> that is the case please reply and explain.