Re: [PATCH v5 4/5] alpha: decode and acknowledge Tsunami system machine checks
From: Matt Turner
Date: Fri Oct 09 2026 - 21:39:58 EST
On Fri, Oct 9, 2026, Magnus Lindholm wrote:
> +#define TSUNAMI_PERROR_ERRMASK 0xdffUL
This leaves out bit 9, and errors[] has NULL in that slot, but
core_tsunami.h defines perror_m_rto as 0x200 and has a perror_v_rto
bitfield. Which one is right? If rev 4.0 of the HRM marks the bit
reserved, please fix the header in this series. If it is a real error
bit, it is neither reported nor acknowledged here.
> +static struct el_TSUNAMI_sysdata_mcheck *
> +tsunami_system_frame(struct el_common *header)
> +{
> + if ((header->code != 0x202 && header->code != 0x204) ||
> + !tsunami_frame_valid(header, sizeof(struct el_TSUNAMI_sysdata_mcheck)))
> + return NULL;
Do we know that DS10, DS20, XP1000 and the others use these two codes?
The vector already says it is a system frame, so I am not sure the code
check buys anything.
When this returns NULL the caller prints "Unrecognized system error"
followed by "Unknown or truncated Tsunami system frame", which is two
lines for one condition and none of the frame. Please dump it raw in
that case, as the EV6 code does for an error it cannot decode. The same
goes for the "layout unknown" environmental message on the non-Clipper
boards.
> + expected = vector != SCB_Q_SYSEVENT && mcheck_expected(smp_processor_id());
Nothing sets mcheck_expected on Tsunami. The only writer is in
core_tsunami.c under NXM_MACHINE_CHECKS_ON_TSUNAMI, which is never
defined. So this is always false, and neither the early return nor the
process_mcheck_info() call at the end can be reached.
I would drop both, along with the changelog paragraph about expected
probes. If you want to keep them for a probe path that might come back,
the changelog should say they are untested.
> + if (vector == SCB_Q_SYSMCHK || vector == SCB_Q_SYSERR) {
> + printk("%sSystem %s error (vector %lx, code %x) on CPU %d:\n",
> + err_print_prefix, vector == SCB_Q_SYSERR ? "correctable" :
> + "uncorrectable", vector, header->code, smp_processor_id());
The old handler went through process_mcheck_info(), which printed the PC
and called dik_show_regs(). System errors now get neither. For an NXM the
PC is the first thing I would want to see. Please add
dik_show_regs(get_irq_regs(), NULL) here, as ev6_machine_check() does.
> + if (fatal)
> + panic("Tsunami: uncorrectable ECC or PCI write parity error");
This is new policy. The old handler never panicked, and none of the
other err_*.c handlers do. I can see the case for UECC. I am less sure
about PERR, and fatal is computed from the live PERROR on every vector,
so a correctable CPU error would panic the machine if the bit happened
to be set.
Could the panic go in its own patch, with the reasoning in its
changelog? That would also make it easy to revert if it turns out to be
too eager on some board.
The decode and the write-one-to-clear acknowledgment look good to me.
Writing back what was read is a clear improvement on the old hardcoded
0x040.