[PATCH net v2 1/2] tcp: restore RACK list membership when undoing loss

From: nramaswamy

Date: Tue Oct 06 2026 - 01:17:43 EST


From: Neil Ramaswamy <nramaswamy@xxxxxxxxxx>

Partial undo can clear the TCPCB_LOST flag on segments already removed
from RACK's list, which prevents subsequent RACK loss detection, leading
to segments only being retransmitted after the RTO. Restoring them to
the RACK list as part of partial undo makes sure we can properly
reconsider them for fast retransmission in the future.

To do this, we first sort (by transmission time) the segments whose
TCPCB_LOST flag is being cleared and then linearly insert those back
into the RACK list, which is sorted by transmission time.

My investigation started from seeing repeated TCP stalls in prod and the
mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to
96 KiB. It also seems like others have seen similar symptoms before [1].

[1]
https://lore.kernel.org/netdev/35A4DDAA-7E8D-43CB-A1F5-D1E46A4ED42E@xxxxxxxxx/

Fixes: 043b87d7599e ("tcp: more efficient RACK loss detection")
Signed-off-by: Neil Ramaswamy <nramaswamy@xxxxxxxxxx>
Assisted-by: LLM sparse
---
net/ipv4/tcp_input.c | 35 +++++++++++++++++++++++++++++++++++
1 file changed, 35 insertions(+)

diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 92bc60716f33..51f04c7474dc 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -69,6 +69,7 @@
#include <linux/module.h>
#include <linux/sysctl.h>
#include <linux/kernel.h>
+#include <linux/list_sort.h>
#include <linux/prefetch.h>
#include <linux/bitops.h>
#include <net/dst.h>
@@ -2840,16 +2841,50 @@ static void DBGUNDO(struct sock *sk, const char *msg)
#endif
}

+static int tcp_rack_skb_cmp(void *priv, const struct list_head *a,
+ const struct list_head *b)
+{
+ const struct sk_buff *skb_a = list_entry(a, struct sk_buff,
+ tcp_tsorted_anchor);
+ const struct sk_buff *skb_b = list_entry(b, struct sk_buff,
+ tcp_tsorted_anchor);
+
+ return tcp_skb_sent_after(tcp_skb_timestamp_us(skb_a),
+ tcp_skb_timestamp_us(skb_b),
+ TCP_SKB_CB(skb_a)->end_seq,
+ TCP_SKB_CB(skb_b)->end_seq);
+}
+
static void tcp_undo_cwnd_reduction(struct sock *sk, bool unmark_loss)
{
struct tcp_sock *tp = tcp_sk(sk);

if (unmark_loss) {
+ LIST_HEAD(restored);
struct sk_buff *skb;

skb_rbtree_walk(skb, &sk->tcp_rtx_queue) {
+ if (TCP_SKB_CB(skb)->sacked & TCPCB_LOST)
+ list_move_tail(&skb->tcp_tsorted_anchor,
+ &restored);
TCP_SKB_CB(skb)->sacked &= ~TCPCB_LOST;
}
+ if (!list_empty(&restored)) {
+ struct list_head *pos = &tp->tsorted_sent_queue;
+
+ /* Ensure lost skbs are added in transmission order */
+ list_sort(NULL, &restored, tcp_rack_skb_cmp);
+ while (!list_empty(&restored)) {
+ struct list_head *entry = restored.next;
+
+ while (pos->next != &tp->tsorted_sent_queue &&
+ !tcp_rack_skb_cmp(NULL, pos->next,
+ entry))
+ pos = pos->next;
+ list_move(entry, pos);
+ pos = entry;
+ }
+ }
tp->lost_out = 0;
tcp_clear_all_retrans_hints(tp);
}
--
2.55.0