Re: [PATCH v2 0/4] isofs: bound the name conversion and return long Joliet names whole

From: Jan Kara

Date: Wed Sep 30 2026 - 07:23:15 EST


On Sat 26-09-26 12:29:12, Matthias Goergens wrote:
> Honza, this reworks v1 along the lines of your review, on top of your
> for_next: both converters take the buffer size, the callers pass it, and
> the buffer shrinks to what they need, without size asserts. But on long
> Joliet names it deliberately goes the other way: 1/4 returns them whole
> instead of cutting them at NAME_MAX, and 2/4 reports the longer limit in
> statfs. 1/4 also corrects the v1 history of the PAGE_SIZE bound.
>
> 1/4 is also a fix. get_joliet_filename() keeps the converted length in
> an unsigned char, so a Joliet name longer than 255 bytes of UTF-8 wraps
> and is listed under a garbled short name that cannot be looked up.
> PowerISO and UltraISO write such names by default (up to 110 UTF-16
> units, 330 bytes for CJK).
>
> Such names are rare: in the Joliet trees of a few thousand archive.org
> images, mostly CJK, Thai or Indic, none is longer than 255 bytes. The
> images I wrote with PowerISO and UltraISO (under wine), a crafted image
> with 111-unit names, and the survey scripts are at
>
> https://github.com/matthiasgoergens/linux/tree/isofs-joliet-long-names
>
> (or I can post them here).
>
> Why whole: as with empty directory blocks [1], I'd follow Windows.
> Joliet is Microsoft's extension, discs with such names are made with
> Windows tools for people who read them on Windows, and Windows returns
> the names whole, including 111-unit names in records that leave out the
> padding byte. Scripts and CI logs:
>
> https://github.com/matthiasgoergens/isofs-windows-probe
>
> On Linux, vfat and exfat already return names of up to 765 bytes and
> report a limit of 1530 in statfs, vfat since f68e542f3478 ("fat: Fix
> statfs->f_namelen"), and the VFS itself only limits a name to PATH_MAX
> (verify_dirent_name()). So what breaks is what breaks with vfat today:
> copying such a file to ext4 or tmpfs fails with ENAMETOOLONG, which cp,
> tar and rsync report before carrying on; readdir_r() and inotify readers
> sized by the man page fail; and fanotify reports the event without the
> name and hits the WARN_ON_ONCE() in fanotify_info_copy_name(), for which
> I have sent a fix [2]. 1/4 has the full list. Cutting at NAME_MAX
> avoids all of this, but lists names that Windows does not show and that
> can collide within a directory.
>
> If you still prefer NAME_MAX, I have that version ready and tested the
> same way: there 1/4 cuts long names at NAME_MAX on a character boundary,
> 2/4 is dropped, and 4/4 shrinks the buffer to NAME_MAX + 1.
>
> Testing: a KASAN and UBSAN kernel under qemu, 19 images mounted with no
> iocharset, utf8, iso8859-1, cp932 and euc-jp, each with and without
> norock. Only the four images with names over 255 bytes change, and only
> with no iocharset or utf8: their long names are now listed whole and
> open. Everything else lists, looks up and reads as before. Under a
> Debian userspace, ls, find, cp, tar, rsync, Python, readdir_r(), inotify
> and fanotify behave on the long-name images, including a crafted one
> with 111-unit names, as they do on a vfat with 765-byte names. This
> replaces v1 2/2's claim, whose test never ran the Joliet or zisofs code.
>
> v1: https://lore.kernel.org/all/20260922155524.1993425-1-matthias.goergens@xxxxxxxxx/
> [1] https://lore.kernel.org/all/cmlro2xzle2aa7ebflvxhbcv74mea7p6qvxhin6egpdvyzyuxh@s45tpgxdal3n/
> [2] https://lore.kernel.org/all/20260926020851.2938961-1-matthias.goergens@xxxxxxxxx/

Thanks! The patches look good and I've added them to my tree. I've just
added the following hunk to your last patch to document the situation in
the code:

+ /*
+ * Rockridge can produce names of NAME_MAX size, Acorn extensions at
+ * most 227 chars.
+ */
+ BUILD_BUG_ON(JOLIET_NAME_MAX < NAME_MAX || JOLIET_NAME_MAX < 227);

I've also noticed we are somewhat inconsistent with the terminating \0
character. In particular get_joliet_filename() -> utf16s_to_utf8s() doesn't
null-terminate the returned string. Similarly isofs_name_translate()
doesn't result in null-terminated string. Other functions do terminate it
(which isn't really necessary AFAICS). It would be nice to clean this up if
you're interested.

Honza
--
Jan Kara <jack@xxxxxxxx>
SUSE Labs, CR