Re: [PATCH v3 06/15] Introduce structured tag value definition

From: David Gibson

Date: Sat Sep 19 2026 - 00:49:42 EST


On Fri, Sep 18, 2026 at 10:16:17AM +0200, Herve Codina wrote:
> Hi David,
>
> On Fri, 18 Sep 2026 14:41:11 +1000
> David Gibson <david@xxxxxxxxxxxxxxxxxxxxx> wrote:
>
> > On Thu, Sep 17, 2026 at 09:04:50AM +0200, Herve Codina wrote:
> > > Hi David,
> > >
> > > On Wed, 16 Sep 2026 15:21:15 +1000
> > > David Gibson <david@xxxxxxxxxxxxxxxxxxxxx> wrote:
> > >
> > > > On Mon, Sep 14, 2026 at 12:19:37PM +0200, Herve Codina wrote:
> > > > > Hi David,
> > > > >
> > > > > On Sat, 12 Sep 2026 12:34:24 +1000
> > > > > David Gibson <david@xxxxxxxxxxxxxxxxxxxxx> wrote:
> > > > >
> > > > > ...
> > > > >
> > > > > > > Do you mean that we should avoid the DATA_LEN_ENCODING and always have the
> > > > > > > 32-bit value right after the tag to give the size for all "skippable" tags?
> > > > > >
> > > > > > Yes.
> > > > > >
> > > > >
> > > > > I did a test using a dts file available in kernel sources. I used (arbitrary
> > > > > choice) juno.dts [0].
> > > > >
> > > > > Without any new tags, the size of the compiled dtb is 27067 bytes.
> > > > >
> > > > > With new metadata tags identifying phandles in properties (FDT_PROPDATA_PHANDLE),
> > > > > the size of the dtb becomes 29027 bytes and so 29027 - 27067 = 1960 bytes for
> > > > > those FDT_PROPDATA_PHANDLE tags (+7.2%).
> > > > >
> > > > > The tags used are composed of:
> > > > > 32-bit: FDT_PROPDATA_PHANDLE value encoding 1 x 32-bit for data
> > > > > 32-bit: offset in the property where a phandle is present.
> > > > >
> > > > > Removing the '1 x 32-bit' information from the tag value and adding a 32-bit
> > > > > 'length' in all cases will lead 3 x 32-bit values for a FDT_PROPDATA_PHANDLE
> > > > > tag (tag + length + offset) instead of the 2 x 32-bit (tag + offset).
> > > > >
> > > > > Back to juno.dts instead of 1960 bytes, the FDT_PROPDATA_PHANDLE will need
> > > > > 1960 * 3 / 2 = 2640 bytes (+9.7%). This leads to around +2.5% of the whole
> > > > > dtb just to have the 32-bit for length. This +2.5% can be easily avoided.
> > > > >
> > > > > Also, I will not be surprised to see more tags in the future adding some more
> > > > > metadata information and so increasing dtb sizes.
> > > > >
> > > > > Quite often you have mentioned memory constraints system where libfdt should
> > > > > be as small as possible. On those system, the dtb itself is embedded in the
> > > > > binary close to libfdt. The size of dtb should be taken into account.
> > > >
> > > > Yeah, those proportions are high enough that I think it's worth it.
> > > >
> > > > > If the SAFE_SKIP bit is removed, I even plan to use this now free bit in the
> > > > > length encoding part:
> > > > > 0b000: No data
> > > > > 0b001: 1 fdt32
> > > > > 0b010: 2 fdt32
> > > > > ...
> > > > > 0b110: 6 fdt32
> > > > > 0b111: On additional fdt32 to encode the length of data.
> > > > >
> > > > > IHMO, length encoding bits in tag value definition should be kept and used
> > > > > for all tags where the length is fixed and can be encoded using
> > > > > these bits.
> > > >
> > > > Well, I'm convinced we want some sort of compact encoding of the
> > > > length, but I think we can do better than the current proposal. It
> > > > seems implausible to me that we'll need 2^29 different metadata tags,
> > > > so I think we can spend some more of the tag bits on the length. How about:
> > > >
> > > > 0x80000000 structured tag bit
> > > > 0x7fff0000 tag type
> > > > 0x0000ffff tag length
> > > >
> > > > So we have up to 2^15 (32k) different structured tags each with a
> > > > length of [0..65534] bytes (length==65535 reserved for those that need
> > > > a full 32-bit length word).
> > > >
> > > > I believe that will avoid the extra length word for everything you
> > > > have currently drafted.
> > > >
> > >
> > > Yes, this will avoid the extra length field. The drawback is the that the
> > > tag value is no more a well fixed value. Each time we have to check the tag
> > > value we have to filter out the tag length.
> > >
> > > For instance:
> > > - FDT_PROPDATA_PHANDLE
> > > fixed data size 4 bytes for offset
> > > tag value: 0x80010004
> > >
> > > - FDT_PROPDATA_PHANDLE_REF
> > > data: 4 bytes for offset + N bytes for a string
> > > tag value 0x8002ssss with ssss for the size
> > >
> > > This will lead to code like this:
> > > tag = fdt_next_tag();
> > > if (tag == FDT_PROPDATA_PHANDLE)
> > > /* Do something */
> > >
> > > if (TAG_GET_ID(tag) == FDT_PROPDATA_PHANDLE_REF)
> > > /* Do something */
> >
> > True. But.. a similar problem kind of exists with the original
> > proposed encoding too: we *expect* a tag with fixed 4-byte contents to
> > use the "1 cell" flags, but we need to consider the case of encoding
> > it as VARLEN with a length field of 4. We could choose to make that
> > forbidden, but we'd still need to consider who's responsible for
> > enforcing that.
> >
> > Similarly, if a variable length metadata tag happens to have length 4
> > or 8 in a particular place, is it valid to encode it with the 1-cell
> > or 2-cell flag?
>
> My position was: if we expect a tag with "1-cell" flag, using varlen
> field encoding is considered as an other tag and so either an error
> or a skippable unknown tag.

So essentially tags of different (fixed) length live in different
namespaces. Ok, that makes good sense to me, but wasn't initially
obvious to me. So, I think we need to spell this out a bit better.

> The same apply for varlen defined tag. Even if the data is, let's
> say 8 bytes, the tag cannot be moved to a "2-cells" tag.
>
> The kind of data length encoding (1-cell, 2-cells, varlength) is done
> when the tag is defined and cannot be changed.
>
> Who is responsible for enforcing that ?
> I would say the documentation of the tag should clearly set the data
> encoding used for the tag and the documentation of the "skippable"
> format should say that data encoding is fixed for a given tag. It is
> set when the tag is defined and any changes at runtime should be
> considered as a different tag.
>
> Of course we can introduce dynamic length encoding to set the data length
> encoding in the tag according to the exact data length found at runtime.
>
> > As a variant on my proposal, I'd also be fine with dividing the
> > structed tags into several classes with bits indicating which is
> > which. Either:
> >
> > * "short" vs "long": "short" always has the length within the tag
> > word (and so cannot exceed 64k, or however many bits we set aside)
> > whereas long always has a length word
> > * "fixed" vs "variable", fixed length tag types always have the same
> > length, so the length can be considered part of the tag. Variable
> > would have a length word.
>
> Well, only strings, or more generally arrays, need a varlen. For those
> item, I would use the varlen word and so "long" in your definition.
>
> For all others where the sizeof(data) is well known when the tag is
> defined, I would use "fixed" and "long" only if sizeof(data) cannot
> be encoded by "fixed" (lengh > limit of dedicated bits).

Right, now understanding your thoughts on the originally proposed
encoding, the "fixed" versus "variable" distinction makes more sense I
think. I do see the advantage of never having a variable length
encoded within the tag word.

> Without any additional bits for any category, all of these fit with the
> following length encoding:
> 000...00: No data
> 000...01: 1 x 32-bit
> 111...10: N x 32-bit
> 111...11: varlen word

Right, so revising my suggestions in light of a better understanding
of what you had in mind originally, it comes down to:

* I think adding more bits to the size field would be worthwhile to
allow a wider variety of future fixed length metadata tags.

* Originally I was thinking that having a length in bytes rather than
just a length in 32-bit words would be worth it. But thinking
further, there's probably no benefit. The length rounded up to
32-bit words is all we need for skipping over it when unknown.
Even if we want a fixed length tag with, say, 3 bytes of data, we
can pad that out to a 1-word tag with a reserved byte.

Ok, so more length bits and better documentation of the fixed
vs. variable distinction are the only remaming suggestions.

> > Not sure if that makes things any easier, but they might, and I'd be
> > fine with either option (or some combination). In any of these cases
> > we do need to spell out what the requirements are: for dtb writers,
> > for dtb readers and for whoever defines a new tag.
> >
> > > Further more, in the code you will both a mix of both construction:
> > > while (tag == FDT_NOP || tag == FDT_BEGIN_NODE ||
> > > TAG_GET_ID(tag) == FDT_BEGIN_NODE_REF);
> > >
> > > with #define TAG_GET_ID(tag) ((tag) & 0xffff0000)
> > >
> > > Also when we write dtbs either in libfdt or dtc, tags value
> > > have to be built with the length when needed.
> > > #define TAG_VALUE(tag_id, length) (((tag_id) & 0xffff0000) || \
> > > ((length) & 0x0000ffff))
> > >
> > > I am totally fine with that but we need to have it in mind.
> >
> > Right, that's not a deal breaker for me - especially since I think
> > we'll need some similar stuff even with the original proposal.
> >
> > > Of course, I can encode FDT_PROPDATA_PHANDLE_REF with 0x8002ffff + length
> > > field but we lose all the benefits
> > >
> > > Ready to see TAG_GET_ID(tag) and TAG_VALUE(tag_id, length) when needed in
> > > the code?
> >
> > I think so, though that could change depending on what it ends up
> > looking like in practice.
> >
>
> Best regards,
> Hervé
>

--
David Gibson (he or they) | I'll have my music baroque, and my code
david AT gibson.dropbear.id.au | minimalist, thank you, not the other way
| around.
http://www.ozlabs.org/~dgibson

Attachment: signature.asc
Description: PGP signature