Re: [PATCH v2 21/21] kbuild: use pigz for gzip compression if available

From: Lorenzo Stoakes (ARM)

Date: Wed Sep 16 2026 - 11:10:56 EST


On Tue, Sep 15, 2026 at 10:31:20AM -0700, Kees Cook wrote:
> On Tue, Sep 15, 2026 at 03:30:43PM +0100, Lorenzo Stoakes (ARM) wrote:
> > On Mon, Sep 14, 2026 at 09:39:07AM -0700, Kees Cook wrote:
> > > On Mon, Sep 14, 2026 at 10:22:20AM +0100, Lorenzo Stoakes (ARM) wrote:
> > > > On a 128-core Threadripper, gzip -9 of a 36 MiB x86-64 vmlinux.bin takes
> > > > 1.6s, and with pigz it takes 0.09s, so the performance increase is
> > > > significant.
> > >
> > > Neat! I wanna go learn how pigz accomplishes this -- I thought the
> > > problem with gzip was a common lookup table. Anyway...
> >
> > Yeah not sure on the details :)
>
> Things I have now learned about a DEFLATE stream (RFC 1951):
>
> - The Huffman tables are per-block, so emitting multiple tables in a
> single stream is already required. (I only just knew about ZIP indexes
> being singular, but I was confused: that's about the container not
> the compression.)
> - Each pigz thread keeps a 32K back-reference window for its tables so
> it isn't starting a fresh table each block.
>
> Very cool! The first makes parallelism possible at all, and the second
> makes it work well: it's not just restarting the compression at every
> block boundary.

Nice!

>
> > > I think the more idiomatic way to do this is:
> > >
> > > KGZIP := $(call try-run,command -v pigz,pigz,gzip)
> >
> > try-run is defined in scripts/Makefile.compiler which is only included ~200
> > lines after KGZIP is set.
> >
> > That'll also set up and tear down a temp dir for a probe that doesn't need
> > that, so I think it's fine as it is.
>
> Yeah, good points. What you have is quite simple.
>
> > > However, parallelism needs to be set. We can't let it eat all CPUs: it
> > > needs to respect the -j make option (and make its CPU reservation known
> > > to "make"), which we already have a solution for in
> > > scripts/jobserver-exec.
> >
> > The only place where it's invoked is vmlinux.bin at the end of the serial
> > tail, where all the tokens would be free anyway.
> >
> > So I don't think it really buys anything at all?
>
> There are 2 things I'm thinking about:
>
> a)
> My main concern is the lack of respecting the -j make argument. In my
> mind, this is a blocker, because it means a build will now _always_ spin
> up max CPUs (not what -j has limited it to), and for CIs, shared compute
> systems, or whatever, this violates the requested parallelism level. For
> example, if I'm doing a long-running Coccinelle replacement in one tree
> (which uses half the CPUs), any builds I launch I'm asking for the other
> half of my CPUs to be used so they don't thrash my cache.
>
> This is the kind of "why are all the CPUs spinning up?" question I
> helped track down with commit 51e46c7a4007 ("docs, parallelism: Rearrange
> how jobserver reservations are made") forever ago.

Yeah that's fair enough and Arnd makes a good point about O(n^2) too.

>
> The next is kind of a special-case version of the above concern:
>
> b)
> There's nothing that ties KGZIP to only being used for final images
> (and in fact, it also does modules, as you show), and it's defined as
> part of "cmd_gzip", so it could be used at any moment in the build:
>
> $ git grep call.*,gzip | wc -l
> 23
>
> And it does have one use outside of the (presumed) final image build in
> the per-arch /boot/ rules besides modules, for config_data:
>
> $ git grep call.*,gzip | grep -v /boot/
> Documentation/kbuild/makefiles.rst: $(call if_changed,gzip)
> kernel/Makefile: $(call if_changed,gzip)
> scripts/Makefile.modinst: $(call cmd,gzip)
>
> So I'm nervous about a general-purpose tool and Kbuild infrastructure
> suddenly going max parallel in the middle of a build some day when
> another gzip use is added.

Hmm yeah. Splitting this off seems sensible then. As you suggest in your
other reply which I shall reply to (it seems the right way forward tbh!).

>
> > Using jobserver-exec would also put python3 and two wrapper scripts in
> > front of every gzip in the build including tar -I "$(KGZIP)" when packaging
> > and some arm and m68k scripts too.
>
> Does that impact wall-clock results meaningfully? I'd really like to
> avoid losing correctness in favor of speed here, especially when we
> already have a solution at hand for exactly this problem.

No (LLM says ~.01s per build for jobserve-exec, once) so actually that big
of a deal.

>
> > Overall I think it's less complexity and really no delta to just invoke it
> > as normal.
> >
> > It's designed as drop-in so it makes sense to use it as that.
>
> It is possible that it is so fast no one will notice, but I really worry
> that there are going to be many sysadmins driven to figuring out why
> their build systems suddenly spike the CPU use, and then waste their
> time tracking it down and reporting it to us, but we can solve it today.

Yeah that's a fair point also.

>
> > > Only RHEL appears a little glitchy, but likely they would trivially move
> > > it to base since it's already packaged, but off in EPEL.
> >
> > Definitely not something for this series, the fallback is one invocation of
> > command -v and I don't really want to break RHEL either :)
> >
> > If, once this has landed, we want to go that way then it's simple enough
> > for us to change it.
>
> Fair enough, though I do worry that this is vaguely "undiscoverable" in
> the sense that if you happen to have pigz installed, suddenly it gets
> used. (We have other such "try to use this other tool first" logic,

Yup that's a thing but I think this is a very worthwhile case!

> really; we have Kconfig stuff for detecting all sorts of capabilities,
> but I couldn't find examples like this one. Maybe I missed it.)
>
> So the "make it the default" suggestion is more about having better
> determinism in the build requirements. Perhaps add pigz to changes.rst
> and/or ver_linux so it is seen/recorded somewhere?

Ack!

(Will reply on your suggestion separately ofc)

>
> -Kees
>
> --
> Kees Cook

--
Cheers, Lorenzo