RE: [PATCH net v2 2/2] tipc: serialize publication purging with name table updates
From: Tung Quang Nguyen
Date: Mon Oct 05 2026 - 00:09:33 EST
>Subject: [PATCH net v2 2/2] tipc: serialize publication purging with name table
>updates
>
>tipc_publ_notify() walks a failed node publication list after the node lock has
>been released. Its safe iterator is not protected by nametbl_lock, which is
>acquired only inside tipc_publ_purge().
>
>A concurrent withdrawal can unlink and schedule the saved next publication
>for freeing. The purge iterator then advances to that removed publication.
>It may access freed memory after the RCU grace period, or repeatedly follow
>the self-linked binding_node before then.
>
>The decoded causal stack is:
>
> tipc_nametbl_remove_publ net/tipc/name_table.c:543
> tipc_publ_purge net/tipc/name_distr.c:244
> tipc_publ_notify net/tipc/name_distr.c:261
> tipc_node_write_unlock net/tipc/node.c:425
> tipc_node_link_down net/tipc/node.c:1094
> tipc_node_delete_links net/tipc/node.c:1325
> bearer_disable net/tipc/bearer.c:414
> __tipc_nl_bearer_disable net/tipc/bearer.c:992
> tipc_nl_bearer_disable net/tipc/bearer.c:1002
>
>Move the failed node publications to a private list under nametbl_lock.
>Select, unlink and purge one publication during each lock acquisition, so no
>publication pointer is retained across an unlocked interval. Concurrent
>withdrawals can remove entries from the private list under the same lock.
>
>Holding the lock for the whole purge would keep bottom halves disabled while
>removing every publication. Releasing it after each entry avoids an excessive
>lock hold for nodes with many publications.
>
>node_lost_contact() purges queued name-table updates before scheduling the
>node-down notification. An update already dequeued by tipc_named_rcv()
>holds nametbl_lock until it updates the publication list, so it completes before
>the snapshot and is included. A publication accepted after the snapshot
>remains on the live node list for a later contact.
>
This does not fix the root cause of this issue: Not holding proper lock in tipc_node_write_unlock().
I will post fix for this issue. Thanks for your report.