When a sleeping task's affinity is changed, task_cpu(p) can be outside
of p->cpus_ptr until after select_task_rq() selects a new runqueue for
the task during wakeup.
Thus, the task's NUMA node determined by numa_select_cpu() can be
completely outside of the task's cpumask, leading to
scx_pick_{idle,any}_cpu_node() failing to find an eligible CPU and
returning -EBUSY. This leads to the numa.bpf.c scheduler abnormally
exiting with the following message in dmesg:
sched_ext: numa: invalid CPU -16
scx_bpf_cpu_node+0x120/0x190
bpf_prog_0a34b8e0f515771f_numa_select_cpu+0x108/0x14e
bpf__sched_ext_ops_select_cpu+0x4f/0xb4
select_task_rq_scx+0xb0/0x210
select_task_rq+0xa0/0xd0
__try_to_wake_up+0x196/0x650
complete_all+0x76/0x100
migration_cpu_stop+0x22b/0x300
cpu_stopper_thread+0xc1/0x180
smpboot_thread_fn+0x16b/0x230
kthread+0x2d7/0x350
ret_from_fork+0x1c2/0x350
ret_from_fork_asm+0x1a/0x30
Make numa_select_cpu() robust against this case by returning @prev_cpu
if no CPU could be found in the selected NUMA node _and_ we have reason
to believe that the task's affinity was changed while it was sleeping.
Fixes: 5ae5161820 ("selftests/sched_ext: Add NUMA-aware scheduler test")
Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
Signed-off-by: Tejun Heo <tj@kernel.org>