[PATCH net v3 1/3] tcp: restore RACK list membership when undoing loss
From: nramaswamy
Date: Fri Oct 09 2026 - 01:03:32 EST
From: Neil Ramaswamy <nramaswamy@xxxxxxxxxx>
Partial undo can clear the TCPCB_LOST flag on segments already removed
from RACK's list, which prevents subsequent RACK loss detection, leading
to segments only being retransmitted after the RTO.
This patch uses the implementation provided by Neal Cardwell to linearly
insert lost segments that have not been retransmitted back into the RACK
list. The core observation is that segments that are lost but not ever
retransmitted are already sorted by their transmission timestamp, which
allows us to insert them into RACK's list linearly, without needing to
pre-sort them.
Segments with TCPCB_EVER_RETRANS are left for RTO recovery. Their
transmission order can differ from their sequence order, so reinserting
them would require addiitonal work (e.g. sorting) before reinsertion into
the RACK list. We assume that this case is rare, and allow those segments
to be transmitted by the RTO.
My investigation started from seeing repeated TCP stalls in prod and the
mitigation that seemed to prevent these stalls was limiting SO_SNDBUF to
96 KiB. It also seems like others have seen similar symptoms before [1].
[1]
https://lore.kernel.org/netdev/35A4DDAA-7E8D-43CB-A1F5-D1E46A4ED42E@xxxxxxxxx/
Fixes: 043b87d7599e ("tcp: more efficient RACK loss detection")
Suggested-by: Neal Cardwell <ncardwell@xxxxxxxxxx>
Suggested-by: Yuchung Cheng <ycheng@xxxxxxxxxx>
Link: https://lore.kernel.org/netdev/20261006141300.1722466-1-ncardwell.sw@xxxxxxxxx/
Signed-off-by: Neil Ramaswamy <nramaswamy@xxxxxxxxxx>
Assisted-by: LLM sparse
---
net/ipv4/tcp_input.c | 50 +++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 49 insertions(+), 1 deletion(-)
diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
index 92bc60716f33..d0a1e8899c4b 100644
--- a/net/ipv4/tcp_input.c
+++ b/net/ipv4/tcp_input.c
@@ -2840,15 +2840,63 @@ static void DBGUNDO(struct sock *sk, const char *msg)
#endif
}
+/* Is skb @a after skb @b in tp->tsorted_sent_queue (send) order? */
+static bool tcp_tsorted_after(const struct list_head *a,
+ const struct list_head *b)
+{
+ const struct sk_buff *skb_a = list_entry(a, struct sk_buff,
+ tcp_tsorted_anchor);
+ const struct sk_buff *skb_b = list_entry(b, struct sk_buff,
+ tcp_tsorted_anchor);
+
+ return tcp_skb_sent_after(tcp_skb_timestamp_us(skb_a),
+ tcp_skb_timestamp_us(skb_b),
+ TCP_SKB_CB(skb_a)->end_seq,
+ TCP_SKB_CB(skb_b)->end_seq);
+}
+
+/* Link skb back into tp->tsorted_sent_queue in send order, at or after
+ * pos, and advance pos to it. Leave skb alone if it was sent before
+ * pos. During undo, all relink calls together traverse the RACK list
+ * at most once, as pos only moves forward.
+ */
+static void tcp_tsorted_relink_skb(struct tcp_sock *tp, struct sk_buff *skb,
+ struct list_head **pos_ptr)
+{
+ struct list_head *head = &tp->tsorted_sent_queue;
+ struct list_head *node = &skb->tcp_tsorted_anchor;
+ struct list_head *pos = *pos_ptr;
+
+ if (pos != head && !tcp_tsorted_after(node, pos))
+ return;
+ while (pos->next != head && !tcp_tsorted_after(pos->next, node))
+ pos = pos->next;
+ if (pos != node) /* not already linked in place */
+ list_move(node, pos);
+ *pos_ptr = node;
+}
+
static void tcp_undo_cwnd_reduction(struct sock *sk, bool unmark_loss)
{
struct tcp_sock *tp = tcp_sk(sk);
if (unmark_loss) {
+ struct list_head *pos = &tp->tsorted_sent_queue;
struct sk_buff *skb;
skb_rbtree_walk(skb, &sk->tcp_rtx_queue) {
- TCP_SKB_CB(skb)->sacked &= ~TCPCB_LOST;
+ u8 sacked = TCP_SKB_CB(skb)->sacked;
+
+ TCP_SKB_CB(skb)->sacked = sacked & ~TCPCB_LOST;
+ /* RACK unlinked the skbs it marked lost. Skbs never
+ * retransmitted keep their original send times, which
+ * increase with sequence, so in one forward pass we
+ * relink them all. For the rare case of undo after
+ * lost retransmissions, we will fall back to RTO.
+ */
+ if ((sacked & (TCPCB_LOST | TCPCB_EVER_RETRANS)) ==
+ TCPCB_LOST)
+ tcp_tsorted_relink_skb(tp, skb, &pos);
}
tp->lost_out = 0;
tcp_clear_all_retrans_hints(tp);
--
2.55.0