Re: [RFC PATCH 0/12] drm/fabric: vendor-neutral topology infrastructure for scale-up accelerator interconnects
From: Rodrigo Vivi
Date: Fri Aug 28 2026 - 12:28:58 EST
On Thu, Aug 27, 2026 at 02:35:16PM +0200, Jiri Pirko wrote:
> Mon, Aug 24, 2026 at 10:09:28AM +0200, ksinyuk@xxxxxxxxxx wrote:
>
> [..]
>
> >Example queries using the in-tree YNL tool are:
> >
> > $ ./tools/net/ynl/pyynl/cli.py \
> > --spec Documentation/netlink/specs/drm_fabric.yaml \
> > --dump fabric-get
> > $ ./tools/net/ynl/pyynl/cli.py \
> > --spec Documentation/netlink/specs/drm_fabric.yaml \
> > --dump endpoint-get --json '{"fabric-id": <id>}'
> > $ ./tools/net/ynl/pyynl/cli.py \
> > --spec Documentation/netlink/specs/drm_fabric.yaml \
> > --do port-get --json '{"endpoint-id": <id>, "port-index": 0}'
>
> Using generic netlink instead of sysfs for this makes a lot of sense,
> but it may be a bit odd to use it outside the networking area.
> I've been struggling with the same in another non-networking use-case
> as well.
Please notice that this bubble was already broken. Netlink design
always had the dream to replace ioctl everywhere. And there are already
other usage in place that are not network related.
In our case we use the netlink API for GPU RAS error reporting: drm-ras.
Also there are other usages in netlink spec that apparently has nothing
to do with networking, like energy...
And in this particular case here, the drm-fabric is a 'network' of
GPU memory... At some point we even wondered if net/ was the right
place for this common API....
> I have been building a framework that keeps the benefits of
> generic netlink, extends it and is fd-based.
> I call it CTLV, here's a link to an early pre-RFC draft:
>
> https://github.com/jpirko/linux_mlxsw/commits/wip_ctlv_pre_rfc_draft1/
>
> What are the benefits over generic netlink:
>
> - No networking in the dependency chain. No netlink sockets,
> no CAP_NET_ADMIN, and no network namespace semantics to reason about
> for a device that has none.
>
> - One character device per registered instance, not one global family.
> Access control is the file: udev rules, ACLs, an fd passed into
> a container. The open mode decides what is allowed - actions need
> write, queries need read - so there is no permission model
> to reinvent per family.
>
> - Events per open file description, not a multicast group.
> Each descriptor has its own subscriptions and queue, and an overflow
> is reported in-band: the next read returns an event-overflow record
> naming the first and last sequence lost and how many.
>
> - Large payloads are referenced, not copied. A blob attribute carries
> a user VA, a memfd and offset, or a dma-buf fd. Nothing big travels
> through the message.
>
> - One YAML specification, everything generated from it: kernel metadata,
> UAPI headers, userspace bindings and the reference documentation.
> Introspectio. is answered out of the same metadata the validator
> enforces, so a device cannot advertise an operation it will refuse.
> A checked-in ABI snapshot makes any wire change a reviewable diff.
>
> - Family inheritance. A family inherits another's operations and fills
> declared extension points. The effective schema is resolved per
> device and the chain is exposed to user. A shared core with
> per-driver extensions is declared once, and both attributes and
> operations can be extended.
>
> - Introspection is per-device and live. The framework answers three
> queries on every device: the family chain, one entry per published
> operation, and any operation's full attribute tree with its bounds
> and limits.
>
> The answer is what this device accepts right now, including what
> a vendor extension added to an inherited operation and whether an op
> is currently disabled, and op-changed events carry the generation
> ops-dump reports, so a dump can be ordered against a change.
> GETFAMILY and GETPOLICY describe a family statically - there is
> no device in that model to ask.
>
> - Fragmented queries are built in, with a consistency check.
> A continuation carries the generation it started from, and if
> the answer changed underneath it is refused with ESTALE,
> reporting the expected and the current generation instead of
> reassembling a torn reply.
>
> - Attributes are 8-byte aligned. CTLV_ALIGNTO is 8, so
> a 64-bit payload is naturally aligned and there is no per-family
> padding to remember - netlink's 4-byte alignment is why
> nla_put_64bit() needs an explicit NOP pad attribute.
>
> - Cheaper round trip. A one-attribute query measured 2.6x cheaper
> than a genetlink one in the same guest - a debug kernel.
>
> - Async operations using io_uring are planned as a follow-up extension.
>
> The framework owns validation, schema resolution, blob acquisition,
> reply serialization and the event queues. All getters and putters
> are generated helpers. A family implements semantics only, and
> the aim is to make that hard to get wrong.
>
> Would this make sense to use for you?
>
> [..]