Re: [PATCH 2/2] sched_ext: Sync common and compat headers from the scx repo

From: Tejun Heo

Date: Tue Aug 18 2026 - 17:33:32 EST


Hello,

On Tue, Aug 18, 2026 at 09:07:32PM +0000, sashiko-bot@xxxxxxxxxx wrote:
> > +/*
> > + * True if the non-sleepable BPF trampoline prolog (__bpf_prog_enter) calls
> > + * migrate_disable() for the current task. Recorded once by
> > + * scx_lib_init_probe, an fentry program on bpf_scx_reg() that fires during
> > + * the natural scheduler-attach call chain (auto-attached by scx_ops_attach!).
> > + *
> > + * Defaults to true (conservative). Over-reporting in is_migration_disabled()
> > + * causes local-only dispatch, which is safe. Under-reporting can crash the
> > + * scheduler, so we err high if the probe somehow fails to run.
> > + */
> > +bool __scx_prolog_disables_migration __weak = true;
>
> [Severity: High]
> Does this default value actually under-report if the probe fails to run?
>
> The documentation indicates that the fallback should err high (over-report)
> to avoid crashing the scheduler. However, looking at the slow path in
> is_migration_disabled() below, when __scx_prolog_disables_migration is true,
> the macro evaluates bpf_get_current_task_btf() != p. This causes it to
> return false for the current task, which means it incorrectly assumes
> migration is not disabled (under-reporting).
>
> Should the default value of __scx_prolog_disables_migration be false
> instead, so that it falls back to returning true?

The comment and the default do disagree. true matches the stock pre-v6.18
prolog behavior, so a failed probe still yields correct results on stock
kernels. However, if the probe fails on a kernel that cherry-picked
8e4f0b1ebcf2 ("bpf: use rcu_read_lock_dont_migrate() for trampoline.c"),
is_migration_disabled(current) under-reports, which is the direction the
comment claims to avoid. Erring high would be defaulting to false.

Changwoo, this code is synced from the scx repo. Can you update the scx
side, either flipping the default or fixing the comment? I'll re-sync
afterwards.

> > +static inline struct task_struct *__COMPAT_scx_bpf_cpu_curr(int cpu)
> > +{
> > + struct rq *rq;
> > +
> > + if (bpf_ksym_exists(scx_bpf_cpu_curr))
> > + return scx_bpf_cpu_curr(cpu);
> > +
> > + rq = scx_bpf_cpu_rq(cpu);
>
> [Severity: High]
> Will this unconditionally call scx_bpf_cpu_rq() on newer kernels?
>
> If both scx_bpf_cpu_curr and scx_bpf_cpu_rq have been dropped from the kernel
> tree, on modern kernels bpf_ksym_exists(scx_bpf_cpu_curr) evaluates to false
> and the fallback path is taken.
>
> Because the call to scx_bpf_cpu_rq() is not guarded by its own
> bpf_ksym_exists() check, won't libbpf poison the missing call and cause the
> BPF verifier to reject the program on newer kernels? Should we guard the
> fallback call as well?

scx_bpf_cpu_curr() wasn't dropped. It was added in v6.18 and only
scx_bpf_cpu_rq() is being removed, so there is no kernel where both are
missing. On kernels without scx_bpf_cpu_rq(),
bpf_ksym_exists(scx_bpf_cpu_curr) is constant true, the fallback is dead
code, and the poisoned call to the missing __weak ksym never reaches the
verifier.

Thanks.

--
tejun