Re: [PATCH v3 06/14] SUNRPC: report local rpcbind calls that get no answer

From: Jeff Layton

Date: Mon Aug 31 2026 - 08:12:11 EST


On Sun, 2026-08-30 at 11:51 -0400, Chuck Lever wrote:
>
> On Fri, Aug 28, 2026, at 12:37 PM, Jeff Layton wrote:
> > A caller that creates many listeners in one operation calls svc_register()
> > once for each of them. Every call waits for the local rpcbind on its own,
> > so a rpcbind that never answers costs the caller one timeout per listener.
> > The caller has no way to learn that the first call already failed.
> >
> > An rpcbind failure can occur one of two ways: either rpcbind fails to
> > respond, or it can respond with -EACCES to indicate that the user
> > doesn't own the current record.
> >
> > Give the first case its own errno. rpcb_register_call() returns -ENAVAIL
> > when the call got no answer, and the existing -EACCES continues to mean a
> > FALSE reply.
> >
> > svc_generic_rpcbind_set() has to let -ENAVAIL past vs_rpcb_optnl, since it
> > is not a refusal. svc_register() applies vs_rpcb_optnl to it instead, so a
> > v4-only server still creates its listeners, and then keeps a running total
> > in serv->sv_rpcb_failures. -ENAVAIL never escapes svc_register().
> >
> > svc_rpcb_failure_count() reports the total. A caller reads the count
> > before it starts and compares as it goes to determine if there have been
> > errors.
> >
> > The users of this infrastructure will be added in later patches.
> >
> > Assisted-by: LLM
> > Signed-off-by: Jeff Layton <jlayton@xxxxxxxxxx>
>
> > diff --git a/net/sunrpc/rpcb_clnt.c b/net/sunrpc/rpcb_clnt.c
> > index 0aa376b82a52..c680137f0fca 100644
> > --- a/net/sunrpc/rpcb_clnt.c
> > +++ b/net/sunrpc/rpcb_clnt.c
> > @@ -412,7 +412,8 @@ static struct rpc_clnt *rpcb_create(struct net
> > *net, const char *nodename,
> > return rpc_create(&args);
> > }
> >
> > -static int rpcb_register_call(struct sunrpc_net *sn, struct rpc_clnt
> > *clnt, struct rpc_message *msg, bool is_set)
> > +static int rpcb_register_call(struct sunrpc_net *sn, struct rpc_clnt
> > *clnt,
> > + struct rpc_message *msg, bool is_set)
> > {
> > int flags = RPC_TASK_NOCONNECT;
> > int error, result = 0;
> > @@ -422,8 +423,10 @@ static int rpcb_register_call(struct sunrpc_net
> > *sn, struct rpc_clnt *clnt, stru
> > msg->rpc_resp = &result;
> >
> > error = rpc_call_sync(clnt, msg, flags);
> > - if (error < 0)
> > + if (error == -EPROTONOSUPPORT)
> > return error;
> > + if (error < 0)
> > + return -ENAVAIL;
> >
> > if (!result)
> > return -EACCES;
>
> If I'm reading this correctly, rpcb_register_call() classifies
> every failure except -EPROTONOSUPPORT as "no answer".
>
> rpc_call_sync() returns negative errnos that are not "no answer":
> pre-dispatch local failures (-ENOMEM from rpc_new_task()), a fatal
> signal (-ERESTARTSYS), and reply-derived errors from
> rpc_verify_header(): -EPFNOSUPPORT, -EOPNOTSUPP, -EIO, -EACCES
> (auth error), -EKEYREJECTED. All of these show that rpcbind *did*
> answer.
>
> Now they become -ENAVAIL, get counted in sv_rpcb_failures, are
> silently converted to success for a vs_rpcb_optnl version, and
> reach userspace as a synthesized -ETIMEDOUT for mandatory versions.
>
> Consequences:
>
> * The commit message says "-EACCES continues to mean a FALSE
> reply," but the RPC layer's auth -EACCES is rewritten to
> -ENAVAIL before the two can be told apart. Its "one of two
> ways" failure taxonomy is not what the code implements.
>
> * "rpcbind not running" (-ECONNREFUSED/-ENOENT) and every
> other transport error reach nfsd's listener_set ack, the
> svc_register/svc_unregister tracepoints, and the printk
> as indistinguishable ETIMEDOUT/ENAVAIL, which IMO is an
> observability regression.
>
> * -ERESTARTSYS -> -ETIMEDOUT drops syscall-restart semantics
> on a fatal signal during registration.
>
> I'm probably missing something.
>

Fair point. I've mostly been trying to classify errors as to whether
there is any point is attempting further registrations. For most of the
errors besides -EACCES (aka rpc_auth_error), there is no point.

That said, I've gone in an cleaned up the error handling here, so v4
should be a bit more explicit about how errors are handled.
--
Jeff Layton <jlayton@xxxxxxxxxx>