Re: [PATCH v19 4/7] firmware: arm_rmm: Add support for SRO

From: Suzuki K Poulose

Date: Mon Sep 28 2026 - 16:45:31 EST


On 28/09/2026 18:28, Catalin Marinas wrote:
On Mon, Sep 28, 2026 at 11:13:39AM +0100, Suzuki K Poulose wrote:
On 28/09/2026 10:28, Catalin Marinas wrote:
On Thu, Sep 24, 2026 at 02:51:58PM +0100, Suzuki K Poulose wrote:
+long rmi_sro_memxfer_execute(struct rmi_sro_state *sro, gfp_t gfp)
+{
+ struct arm_smccc_1_2_regs *regs = &sro->regs;
+ bool cancelled = false;
+ unsigned long sro_handle;
+
+ rmi_smccc_invoke(regs);
+
+ sro_handle = regs->a1;
+ while (RMI_RESULT_STATUS(regs->a0) == RMI_INCOMPLETE) {
+ bool can_cancel = RMI_RESULT_CAN_CANCEL(regs->a0) == RMI_OP_CAN_CANCEL;
+ int ret = 0;
+
+ switch (RMI_RESULT_MEMREQ(regs->a0)) {
+ case RMI_OP_MEM_REQ_NONE:
+ rmi_op_continue(sro_handle, RMI_CONTINUE_KEEP_GOING,
+ regs);
+ break;
+ case RMI_OP_MEM_REQ_DONATE:
+ ret = rmi_sro_donate(sro, sro_handle, regs->a2, regs,
+ gfp);
+ break;
+ case RMI_OP_MEM_REQ_RECLAIM:
+ ret = rmi_sro_reclaim(sro, sro_handle, regs);
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ ret = -ENXIO;
+ break;
+ }

Another thing I came across while looking whether we can defer the
activation. It seems that the spec (I_JVYCH) lists some SROs as
PE-bound. Nothing here or in rmi_sro_execute() disables migration and
the memory allocation paths can even sleep with GFP_KERNEL.

No, this is not required. I agree this is confusing. I will get it
clarified.

So, there are two different sources for the SRO contexts. One is a global
pool and the other an Object.

e.g., For an RMI operation on an Object, SRO context can be the object
itself (e.g., REC_CREATE, REALM_ACTIVATE etc.)

However, when there is no reliable object for the command (e.g.,
RMI_GRANULE_RANGE_DELEGATE), the RMM must allocate a context from
the global pool. Now, the "PE" in there comes from a recommendation
to the RMM implementations, that the global pool size must depend on
the number of PEs on the system. This doesn't mean that the SRO
handles are only bound to those PEs. I will get this clarified
in the RMM spec.

This part of the spec needs rewriting, not clarifying. No matter how
hard you try, there's no way you can read it as a "global pool". For
example:

D_GZLMRA SRO context is bound to one of the following:
- A PE
- An RMM object

And take a random command:

B4.5.2 RMI_DPT_L0_CREATE command
Create a Level 0 DPT.
The RMI_DPT_L0_CREATE command may initiate a Stateful RMI
Operation whose context is bound to the current PE.

"bound to the current PE" pretty clearly shows the intention was to
disable preemption. It also doesn't say what happens when this pool is
exhausted (presumably it returns RMI_BLOCKED).

Agree. If there are not contexts available for the RMM to use, it
results in RMI_BLOCKED.

TBH, that's a pretty significant change for a bet3/4 release, though
arguably it can be seen as a relaxation. Code that relies on disabling
preemption should still work (somewhat, assuming the global pool is at
least the number of PEs and the host plays nicely to complete or cancel
all SROs).

The only case where the RMM mandates the execution continues on a "CPU"
(virtual PE rather) is Realm Attestation ABI exposed to the Realm (RSI_ATTEST_TOKEN_*).

Btw, SRO contexts are nothing but a "scratch" memory that the RMM uses
to hold the progress made in the SRO. e.g., the donated granules and
what is the next operation expected etc.



That said, such pool is a limited resource and we need some way to probe
its size if we want to do something smarter in the kernel, like a
semaphore to ensure we don't randomly fail because of an RMM limitation.
I don't really see how the number of PEs is relevant to this global
pool sizing, it's not that we limit the realms we can start to the
online CPUs.

The number of PEs is only a factor which can tell the RMM, how many
parallel requests it could get. That said, due to pre-emption there
could be multiple outstanding SRO operations on a single PE. But
also remember that not all SROs require a context from global pool.

Agreed that it would good to probe the number of contexts available.
Or may be even dynamically extend the pool for SRO contexts at
runtime. Will feed this back to the RMM spec, and hopefully we can
address this in the future versions without breaking the existing
ABI.

Thanks
Suzuki