RE: [PATCH net-next v4] tcp: add TCP_ECN and TCP_ECN_OPTION socket options

From: Irlanki Sandeep

Date: Wed Sep 30 2026 - 08:30:47 EST


> On Mon, 21 Sep 2026 15:26:04 -0700, Jakub Kicinski <kuba@xxxxxxxxxx> wrote:
> > My gut reaction is that we shouldn't be adding a setsockopt here.
> > This is a routing property, really, IIUC you're adding the setsockopt
> > because you want to use BPF, in which case isn't it better to add
> > a kfunc for this? Why go thru all the sockopt plumbing?
> >
> > TCP maintainers, WDYT?

Hi Jakub,

Agreed. Moving to BPF kfuncs avoids UAPI clutter while giving us the exact per-connection granularity we need.

I adopted your suggestion and transitioned the implementation to two kfuncs:
- bpf_sock_ops_set_ecn_mode()
- bpf_sock_ops_set_accecn_option()

v5 with the kfuncs and selftests has been submitted:
https://lore.kernel.org/all/20260930102308.197808-1-irlanki.s@xxxxxxxxxxx/

> On Tue, 22 Sep 2026 02:48:10 +0200, Eric Dumazet <edumazet@xxxxxxxxxx> wrote:
> Speaking for myself, I also cannot see why we need to support so much
> flexibility.

Hi Eric,

Regarding the need for this flexibility:

In our testing on real-world networks across multiple interfaces (cellular, Wi-Fi):
1. Certain legacy middleboxes blackhole ECN flags or drop data packets carrying AccECN option headers. A global sysctl is all-or-nothing: disabling it loses L4S benefits on compliant paths, while enabling it globally causes connection failures on broken paths.
2. Under heavy CE markings, we observed throughput drops when comparing Prague (out-of-tree) against Cubic/BIC. We need the ability to selectively enable AccECN for latency-sensitive applications while avoiding it for bulk throughput-oriented traffic on paths showing throughput regressions.

This patch gives us this per-connection control dynamically without changing global sysctls.

> BTW:
>
> I have on my plate a complete walk through of recent fields additions
> in tcp_sock, mostly from the AccECN support.
>
> ecn_mode and ecn_options have nothing to do in tcp_sock_write_tx group
> in any case.
> This patch would hurt performance.

Thanks for pointing this out.

In v5, ecn_mode and ecn_option have been moved out of tcp_sock_write_tx and placed into the cold section of struct tcp_sock (immediately after accecn_fail_mode, outside hot fastpath cachelines). We have also updated Documentation/networking/net_cachelines/tcp_sock.rst accordingly.

Please review v5 when you get a chance.

Thanks,
Sandeep