Re: [PATCH] s390: Warn if kernel command line contains non-printable EBCDIC characters

From: Ilya Leoshkevich

Date: Thu Aug 27 2026 - 07:15:26 EST




On 8/27/26 10:32, David Laight wrote:
On Wed, 26 Aug 2026 16:08:50 +0200
Ilya Leoshkevich <iii@xxxxxxxxxxxxx> wrote:

On 8/26/26 15:51, David Laight wrote:
On Tue, 25 Aug 2026 17:08:08 +0200
Ilya Leoshkevich <iii@xxxxxxxxxxxxx> wrote:
Users may accidentally add multi-byte UTF-8 characters to zipl.conf
parmline, for example, by copying snippets containing non-breaking
spaces (\xC2\xA0) from web pages.

The kernel will then interpret the entire command line as EBCDIC,
making it unusable. Distinguish this situation from the legitimate
EBCDIC conversion by looking for non-printable characters and issue
a warning.

Would it be better to check for the entire line being printable ebcdic?
All of EBCDIC a-zA-Z0-9 have the 0x80 bit set and most of 0x20..0x7f
are invalid or control characters (or punctuation).

David

I actually started with that, but this required introducing a new
_ctype-like table (unfortunately it's not as simple as checking a
couple ranges), so I decided against that and took a shortcut via
ASCII.

Could you get the conversion function to return an error if it found
invalid EBCDIC characters?
If there is a single UTF8 character (eg non-breaking space) you really
want to treat the line as ASCII.
Actually you could count the number of characters with the 0x80 bit set.
If more than 1/2 assume EBCDIC (all of 0-9a-zA-Z have the bit set).

I also considered that, but setting any threshold feels arbitrary and
will probably fail for punctuation-heavy command lines.

I also didn't want to make it a hard fail, because I'm not certain that
I know all uses cases. Perhaps there are people who wants umlauts and
what not? It would be bad to break whatever they are doing.

But at the same time it was very painful to debug the &nbsp; issue, so
I settled for the compromise: add a warning that will be helpful to
99.9% users and will only mildly annoy the 0.1% umlaut users.


I just had an off-list discussion with Heiko and we think about going
with your first proposal for v3: a new _ctype table for EBCDIC for
determining whether characters are printable. The overhead from this is
not bad as I thought it would be.
(I didn't realise anyone still used EBCDIC.
I guess the unix implementation(s) use ASCII (otherwise too much code
is broken) but the old IBM OS uses EBCDIC.
I worked for ICL for a while, their old 1900 series (from the early
1970s) used 6bit characters (4 in a 24bit word) that were ACSII codes
32-95. The replacement 2900 series (very late 1970s) used EBCDIC internally
(I guess because IBM used it...) but all the peripherals were ASCII.)

Linux on s390 still uses it for interfacing with traditional IBM
hypervisors (z/VM and PR/SM), which are very much alive and used today.
David


[...]