Re: [PATCH net v2 2/2] tipc: serialize publication purging with name table updates

From: netdev-bot+sashiko

Date: Mon Oct 05 2026 - 14:31:22 EST


Thank you for your contribution! Sashiko AI review found 1 potential
issue(s) to consider.

Critical: 0 · High: 0 · Medium: 1 · Low: 0

If you believe none of the issues are legitimate, please reply to each
issue of Medium severity or higher and explain why the AI is wrong.
If any of them are legitimate and you plan to respin, please let patchwork
know by including "pw-bot: cr" as a separate line at the end of your reply
(one such reply per series is enough).

- [Medium] tipc_publ_notify() (net/tipc/name_distr.c) now calls
list_del_init(&publ->binding_node) on every selected publication…

--

tipc: serialize publication purging with name table updates

This patch moves a failed node's publications onto a private list while
holding nametbl_lock. It then unlinks and purges one publication per lock
acquisition, so no publication pointer is kept across an unlocked
interval.

Link: https://lore.kernel.org/netdev/20260927180806.1315902-1-nicoyip.dev@xxxxxxxxx/

> diff --git a/net/tipc/name_distr.c b/net/tipc/name_distr.c
> index acf96562608b9..9a400a1fa4d72 100644
> --- a/net/tipc/name_distr.c
> +++ b/net/tipc/name_distr.c
> @@ -230,20 +230,16 @@ void tipc_named_node_up(struct net *net, u32 dnode, u16 capabilities)

[ ... ]

> tipc_uaddr(&ua, TIPC_SERVICE_RANGE, p->scope, p->sr.type,
> p->sr.lower, p->sr.upper);
> - spin_lock_bh(&tn->nametbl_lock);
> _p = tipc_nametbl_remove_publ(net, &ua, &p->sk, p->key);
> - if (_p)
> - list_del_init(&_p->binding_node);
> - spin_unlock_bh(&tn->nametbl_lock);
> if (_p)
> kfree_rcu(_p, rcu);
> }
> @@ -254,10 +250,27 @@ void tipc_publ_notify(struct net *net, struct list_head *nsub_list,

[ ... ]

> + for (;;) {
> + spin_lock_bh(&tn->nametbl_lock);
> + if (list_empty(&purge_list)) {
> + spin_unlock_bh(&tn->nametbl_lock);
> + break;
> + }
> + publ = list_first_entry(&purge_list, struct publication,
> + binding_node);
> + list_del_init(&publ->binding_node);
> tipc_publ_purge(net, publ);

[Severity: Medium]
What happens to publ here if tipc_nametbl_remove_publ() returns NULL
inside tipc_publ_purge()?

The old code only unlinked binding_node when removal succeeded:

if (_p)
list_del_init(&_p->binding_node);

A failed removal left the publication on n->publ_list, so a later
node-down purge would try it again. With this change, list_del_init() runs
before the removal is attempted. tipc_publ_purge() also drops a NULL
result silently.

So a publication that fails removal stays in the name table but is no
longer on any node publication list. Can it then be orphaned for good?

A failed removal looks reachable through an existing problem with
duplicate ranges. This part predates the patch. In
tipc_service_create_range(), equal lower bounds always go left:

if (lower <= sr->lower)
n = &parent->rb_left;
else
n = &parent->rb_right;

After a rotation, an exact match in a right subtree is missed. For one
service type, with every range published by the same remote node:

1. Insert R1=[10,20] (pubA), then R2=[10,30] (pubB). R2 becomes the
left child of R1.
2. Insert R3=[10,40] (pubC). After the rebalance, R2 is the root, R3 is
on its left and R1 is on its right.
3. Insert pubD=[10,20]. The walk goes R2 -> R3 -> NULL and never
reaches R1. A duplicate [10,20] range R4 is created as the left
child of R3.
4. n->publ_list holds pubA, pubB, pubC, pubD in that order, because
tipc_node_subscribe() uses list_add_tail().

On node-down, the new loop picks pubA first and unlinks it. Then
tipc_service_find_range() returns the leftmost exact match, R4.
tipc_service_remove_publ(R4, skA, keyA) returns NULL because R4 only
holds pubD.

The loop goes on to purge pubB, pubC and pubD, and R4 is erased. pubA is
left in R1 and is no longer on any node list.

If the peer reconnects, tipc_service_insert_publ() rejects the
re-published pubA as a duplicate, so pubA never goes back on the node
list. If the peer then goes down without withdrawing it, pubA stays in
the name table until tipc_nametbl_stop().

Until then, lookups can still pick the dead node/port, and topology
subscribers never get TIPC_WITHDRAWN for it. The remote peer controls the
order through its bind() calls.

Could tipc_publ_purge() return whether the removal succeeded, so the
caller can keep or re-queue the publication when it fails?

> + spin_unlock_bh(&tn->nametbl_lock);
> + }
> +

[ ... ]

--
Sashiko AI review · https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20261001182924.3928331-1-nicoyip.dev%40gmail.com