Merge tag 'cgroup-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup

Pull cgroup updates from Tejun Heo:

 - Attach path bug fixes: migrations spanning multiple source or
   destination cpusets were mishandled, most visibly leaving thread
   affinities stale when the controller is disabled in a threaded
   subtree. Configuration writes could also race an in-flight attach and
   apply stale state, and the deadline task count could get corrupted by
   concurrent updates, skewing SCHED_DEADLINE admission decisions.

 - Memory binding bug fixes: which node masks get applied differed
   between the binding update paths, and tasks cloned with
   CLONE_INTO_CGROUP skipped rebinding entirely. Rebinding also now runs
   once per process instead of repeating for every thread sharing the
   mm.

 - Overhead removals with no behavior change: CPU hotplug iterated tasks
   of cpusets that just inherit the parent's effective masks, and the
   slab-spreading task flag was still being maintained although the SLAB
   allocator that consumed it is long gone.

 - Data-race annotations for benign races so that KCSAN reports stay
   meaningful, selftest coverage for the fixes above along with
   flakiness and portability fixes, and documentation corrections.

* tag 'cgroup-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup: (34 commits)
  selftests/cgroup: Remove redundant chown in test_cgcore_lesser_ns_open
  selftests/cgroup: Preserve CPU hotplug write errors
  cgroup/cpuset: Add test for partition root invalidation returning wrong CPUs
  cgroup/cpuset: Remove obsolete PFA_SPREAD_SLAB task flag
  docs: cgroup-v2: fix stale "io" controller introduction
  selftests/cgroup: Avoid awk -e in cpuset tests
  cgroup/cpuset: Use WRITE_ONCE() for shared prs_err updates
  selftests/cgroup: add user_usec sanity check in test_cpucg_nice
  cgroup: drop unneeded semicolon
  docs: cgroup-v2: mark memory.pressure and io.pressure as read-write
  selftests/cgroup: Fix minor defects in test_cpuset
  Docs/admin-guide/cgroup-v2: fix delay_nsec unit in io.latency doc
  selftests/cgroup: Remove redundant cg_enter_current() call in test_core
  selftests/cgroup: Add test for cpuset affinity on controller disable
  cgroup/cpuset: Handle the special case of non-moving tasks in cpuset_can_attach()
  cgroup/cpuset: Support multiple destination cpusets for cpuset_*attach()
  selftests/cgroup: fix missing TAP output in test_hugetlb_memcg
  cgroup/cpuset: Support multiple source cpusets for cpuset_*attach()
  cgroup/cpuset: Move mpol_rebind_mm/cpuset_migrate_mm() calls inside cpuset_attach_task()
  cgroup/cpuset: Make attach_ctx.old_cs track task group leader
  ...
This commit is contained in:
Linus Torvalds
2026-08-20 10:37:42 -07:00
18 changed files with 749 additions and 305 deletions

View File

@@ -179,7 +179,7 @@ files describing that cpuset:
- cpuset.mem_hardwall flag: is memory allocation hardwalled
- cpuset.memory_pressure: measure of how much paging pressure in cpuset
- cpuset.memory_spread_page flag: if set, spread page cache evenly on allowed nodes
- cpuset.memory_spread_slab flag: OBSOLETE. Doesn't have any function.
- cpuset.memory_spread_slab flag: OBSOLETE. Has no effect on allocation behavior.
- cpuset.sched_load_balance flag: if set, load balance within CPUs on that cpuset
- cpuset.sched_relax_domain_level: the searching range when migrating tasks
@@ -318,26 +318,20 @@ times 1000.
1.6 What is memory spread ?
---------------------------
There are two boolean flag files per cpuset that control where the
kernel allocates pages for the file system buffers and related in
kernel data structures. They are called 'cpuset.memory_spread_page' and
'cpuset.memory_spread_slab'.
The 'cpuset.memory_spread_page' boolean flag file controls where the kernel
allocates page-cache pages.
The 'cpuset.memory_spread_slab' file is obsolete and has no effect on
allocation behavior, but is retained for compatibility.
If the per-cpuset boolean flag file 'cpuset.memory_spread_page' is set, then
the kernel will spread the file system buffers (page cache) evenly
over all the nodes that the faulting task is allowed to use, instead
of preferring to put those pages on the node where the task is running.
If the per-cpuset boolean flag file 'cpuset.memory_spread_slab' is set,
then the kernel will spread some file system related slab caches,
such as for inodes and dentries evenly over all the nodes that the
faulting task is allowed to use, instead of preferring to put those
pages on the node where the task is running.
The setting of these flags does not affect anonymous data segment or
The setting of this flag does not affect anonymous data segment or
stack segment pages of a task.
By default, both kinds of memory spreading are off, and memory
By default, page cache memory spreading is off, and memory
pages are allocated on the node local to where the task is running,
except perhaps as modified by the task's NUMA mempolicy or cpuset
configuration, so long as sufficient free memory pages are available.
@@ -345,18 +339,18 @@ configuration, so long as sufficient free memory pages are available.
When new cpusets are created, they inherit the memory spread settings
of their parent.
Setting memory spreading causes allocations for the affected page
or slab caches to ignore the task's NUMA mempolicy and be spread
instead. Tasks using mbind() or set_mempolicy() calls to set NUMA
mempolicies will not notice any change in these calls as a result of
their containing task's memory spread settings. If memory spreading
Setting page cache memory spreading causes affected allocations to ignore the
task's NUMA mempolicy and be spread instead. Tasks using mbind() or
set_mempolicy() to set NUMA mempolicies will not notice any change as a
result of their containing task's memory spread settings. If memory spreading
is turned off, then the currently specified NUMA mempolicy once again
applies to memory page allocations.
Both 'cpuset.memory_spread_page' and 'cpuset.memory_spread_slab' are boolean flag
files. By default they contain "0", meaning that the feature is off
for that cpuset. If a "1" is written to that file, then that turns
the named feature on.
Both 'cpuset.memory_spread_page' and 'cpuset.memory_spread_slab' are boolean
flag files. In the root cpuset, both files initially contain "0". Writing "1"
or "0" to 'cpuset.memory_spread_page' enables or disables page-cache spreading,
respectively. The value of 'cpuset.memory_spread_slab' is retained, can be read
back and inherited, but it does not affect allocation behavior.
The implementation is simple.
@@ -367,10 +361,6 @@ is modified to perform an inline check for this PFA_SPREAD_PAGE task
flag, and if set, a call to a new routine cpuset_mem_spread_node()
returns the node to prefer for the allocation.
Similarly, setting 'cpuset.memory_spread_slab' turns on the flag
PFA_SPREAD_SLAB, and appropriately marked slab caches will allocate
pages from the node returned by cpuset_mem_spread_node().
The cpuset_mem_spread_node() routine is also simple. It uses the
value of a per-task rotor cpuset_mem_spread_rotor to select the next
node in the current task's mems_allowed to prefer for the allocation.

View File

@@ -10,7 +10,7 @@ Because VM is getting complex (one of reasons is memcg...), memcg's behavior
is complex. This is a document for memcg's internal behavior.
Please note that implementation details can be changed.
(*) Topics on API should be in Documentation/admin-guide/cgroup-v1/memory.rst)
(*) Topics on API should be in Documentation/admin-guide/cgroup-v1/memory.rst
0. How to record usage ?
========================

View File

@@ -1145,7 +1145,7 @@ will be referred to. All time durations are in microseconds.
This file exists whether the controller is enabled or not.
It always reports the following three stats, which account for all the
processes in the cgroup:
processes in the cgroup (including those in descendant cgroups):
- usage_usec
- user_usec
@@ -1160,6 +1160,27 @@ will be referred to. All time durations are in microseconds.
- nr_bursts
- burst_usec
Note that the above five CFS bandwidth stats are non-hierarchical;
they only account for throttling caused by this cgroup's own bandwidth
limit, not including throttling inherited from ancestor cgroups.
cpu.stat.local
A read-only flat-keyed file.
This file exists whether the controller is enabled or not.
It reports the following stat when the controller is enabled:
- throttled_usec
Unlike the ``throttled_usec`` reported by ``cpu.stat`` which
accounts for throttling caused by this cgroup's own CFS
bandwidth limit, ``cpu.stat.local`` reports the actual
throttling time incurred by this cgroup's own runqueues,
which may include throttling inherited from ancestor
cgroup bandwidth limits.
When the controller is not enabled, this stat is not reported.
cpu.weight
A read-write single value file which exists on non-root
cgroups. The default is "100".
@@ -1909,7 +1930,7 @@ The following nested keys are defined.
is allowed unless memory.swap.max is set to 0.
memory.pressure
A read-only nested-keyed file.
A read-write nested-keyed file.
Shows pressure stall information for memory. See
:ref:`Documentation/accounting/psi.rst <psi>` for details.
@@ -1984,9 +2005,13 @@ IO
The "io" controller regulates the distribution of IO resources. This
controller implements both weight based and absolute bandwidth or IOPS
limit distribution; however, weight based distribution is available
only if cfq-iosched is in use and neither scheme is available for
blk-mq devices.
limit distribution. Absolute BPS and IOPS limits are enforced by
blk-throttle and apply to all devices, while weight based proportional
distribution is provided by the iocost cost model controller
(CONFIG_BLK_CGROUP_IOCOST) and, when the BFQ I/O scheduler is in use
for a device, by BFQ's own cgroup support. Latency-based protection
(CONFIG_BLK_CGROUP_IOLATENCY) and I/O priority assignment
(CONFIG_BLK_CGROUP_IOPRIO) are also available.
IO Interface Files
@@ -2169,7 +2194,7 @@ IO Interface Files
8:16 rbps=2097152 wbps=max riops=max wiops=max
io.pressure
A read-only nested-keyed file.
A read-write nested-keyed file.
Shows pressure stall information for IO. See
:ref:`Documentation/accounting/psi.rst <psi>` for details.
@@ -2284,9 +2309,9 @@ This throttling takes 2 forms:
throttled without possibly adversely affecting higher priority groups. This
includes swapping and metadata IO. These types of IO are allowed to occur
normally, however they are "charged" to the originating group. If the
originating group is being throttled you will see the use_delay and delay
fields in io.stat increase. The delay value is how many microseconds that are
being added to any process that runs in this group. Because this number can
originating group is being throttled you will see the use_delay and delay_nsec
fields in io.stat increase. The delay_nsec value is how many nanoseconds that
are being added to any process that runs in this group. Because this number can
grow quite large if there is a lot of swapping or metadata IO occurring we
limit the individual delay events to 1 second at a time.
@@ -2552,6 +2577,13 @@ Cpuset Interface Files
a need to change "cpuset.mems" with active tasks, it shouldn't
be done frequently.
For a multithreaded process, the threadgroup leader is
considered the owner of the group's memory. Memory policy
rebinding and migration will only happen with respect to the
threadgroup leader. To avoid unexpected results, non-leading
threads shouldn't be put into another cgroup whose "cpuset.mems"
doesn't fully overlap that of the threadgroup leader.
cpuset.mems.effective
A read-only multiple values file which exists on all
cpuset-enabled cgroups.

View File

@@ -896,7 +896,7 @@ static inline void cgroup_threadgroup_change_begin(struct task_struct *tsk)
* cgroup_threadgroup_change_end - threadgroup exclusion for cgroups
* @tsk: target task
*
* Counterpart of cgroup_threadcgroup_change_begin().
* Counterpart of cgroup_threadgroup_change_begin().
*/
static inline void cgroup_threadgroup_change_end(struct task_struct *tsk)
{

View File

@@ -480,7 +480,7 @@ static inline void cgroup_unlock(void)
rcu_read_lock_sched_held() || \
lockdep_is_held(&cgroup_mutex) || \
lockdep_is_held(&css_set_lock) || \
((task)->flags & PF_EXITING) || (__c))
(data_race((task)->flags) & PF_EXITING) || (__c))
#else
#define task_css_set_check(task, __c) \
rcu_dereference((task)->cgroups)

View File

@@ -1860,7 +1860,6 @@ static __always_inline bool is_user_task(struct task_struct *task)
/* Per-process atomic flags. */
#define PFA_NO_NEW_PRIVS 0 /* May not gain new privileges. */
#define PFA_SPREAD_PAGE 1 /* Spread page cache over cpuset */
#define PFA_SPREAD_SLAB 2 /* Spread some slab caches over cpuset */
#define PFA_SPEC_SSB_DISABLE 3 /* Speculative Store Bypass disabled */
#define PFA_SPEC_SSB_FORCE_DISABLE 4 /* Speculative Store Bypass force disabled*/
#define PFA_SPEC_IB_DISABLE 5 /* Indirect branch speculation restricted */
@@ -1886,10 +1885,6 @@ TASK_PFA_TEST(SPREAD_PAGE, spread_page)
TASK_PFA_SET(SPREAD_PAGE, spread_page)
TASK_PFA_CLEAR(SPREAD_PAGE, spread_page)
TASK_PFA_TEST(SPREAD_SLAB, spread_slab)
TASK_PFA_SET(SPREAD_SLAB, spread_slab)
TASK_PFA_CLEAR(SPREAD_SLAB, spread_slab)
TASK_PFA_TEST(SPEC_SSB_DISABLE, spec_ssb_disable)
TASK_PFA_SET(SPEC_SSB_DISABLE, spec_ssb_disable)
TASK_PFA_CLEAR(SPEC_SSB_DISABLE, spec_ssb_disable)

View File

@@ -104,7 +104,7 @@ DEFINE_PERCPU_RWSEM(cgroup_threadgroup_rwsem);
#define cgroup_assert_mutex_or_rcu_locked() \
RCU_LOCKDEP_WARN(!rcu_read_lock_held() && \
!lockdep_is_held(&cgroup_mutex), \
"cgroup_mutex or RCU read lock required");
"cgroup_mutex or RCU read lock required")
/*
* cgroup destruction makes heavy use of work items and there can be a lot

View File

@@ -146,10 +146,9 @@ struct cpuset {
nodemask_t old_mems_allowed;
/*
* Tasks are being attached to this cpuset. Used to prevent
* zeroing cpus/mems_allowed between ->can_attach() and ->attach().
* For linking impacted cpusets during an attach operation.
*/
int attach_in_progress;
struct llist_node attach_node;
/* partition root state */
int partition_root_state;
@@ -165,7 +164,7 @@ struct cpuset {
* number of SCHED_DEADLINE tasks attached to this cpuset, so that we
* know when to rebuild associated root domain bandwidth information.
*/
int nr_deadline_tasks;
atomic_t nr_deadline_tasks;
int nr_migrate_dl_tasks;
/* DL bandwidth that needs destination reservation for this attach. */
u64 sum_migrate_dl_bw;
@@ -269,10 +268,7 @@ static inline int nr_cpusets(void)
static inline bool cpuset_is_populated(struct cpuset *cs)
{
lockdep_assert_cpuset_lock_held();
/* Cpusets in the process of attaching should be considered as populated */
return cgroup_is_populated(cs->css.cgroup) ||
cs->attach_in_progress;
return cgroup_is_populated(cs->css.cgroup);
}
/**

View File

@@ -204,7 +204,7 @@ static s64 cpuset_read_s64(struct cgroup_subsys_state *css, struct cftype *cft)
}
/*
* update task's spread flag if cpuset's page/slab spread flag is set
* Update a task's spread flag if the cpuset's page spread flag is set.
*
* Call with callback_lock or cpuset_mutex held. The check can be skipped
* if on default hierarchy.
@@ -219,18 +219,13 @@ void cpuset1_update_task_spread_flags(struct cpuset *cs,
task_set_spread_page(tsk);
else
task_clear_spread_page(tsk);
if (is_spread_slab(cs))
task_set_spread_slab(tsk);
else
task_clear_spread_slab(tsk);
}
/**
* cpuset1_update_tasks_flags - update the spread flags of tasks in the cpuset.
* @cs: the cpuset in which each task's spread flags needs to be changed
* cpuset1_update_tasks_flags - update the page spread flag of cpuset tasks
* @cs: the cpuset whose tasks need their page spread flag updated
*
* Iterate through each task of @cs updating its spread flags. As this
* Iterate through each task of @cs updating its page spread flag. As this
* function is called with cpuset_mutex held, cpuset membership stays
* stable.
*/

View File

@@ -37,6 +37,7 @@
#include <linux/wait.h>
#include <linux/workqueue.h>
#include <linux/task_work.h>
#include <linux/llist.h>
DEFINE_STATIC_KEY_FALSE(cpusets_pre_enable_key);
DEFINE_STATIC_KEY_FALSE(cpusets_enabled_key);
@@ -222,14 +223,14 @@ void inc_dl_tasks_cs(struct task_struct *p)
{
struct cpuset *cs = task_cs(p);
cs->nr_deadline_tasks++;
atomic_inc(&cs->nr_deadline_tasks);
}
void dec_dl_tasks_cs(struct task_struct *p)
{
struct cpuset *cs = task_cs(p);
cs->nr_deadline_tasks--;
atomic_dec(&cs->nr_deadline_tasks);
}
static inline bool is_partition_valid(const struct cpuset *cs)
@@ -356,6 +357,41 @@ static struct workqueue_struct *cpuset_migrate_mm_wq;
static DECLARE_WAIT_QUEUE_HEAD(cpuset_attach_wq);
/*
* Cpuset task attach context
* Protected by cpuset_mutex
*/
static struct {
int in_progress;
bool cpus_updated;
bool mems_updated;
bool task_work_queued;
bool many_dest_cs; /* Have many destination cpusets */
struct cpuset *old_cs; /* Source cpuset */
nodemask_t nodemask_to;
} attach_ctx;
static LLIST_HEAD(src_cs_head);
static LLIST_HEAD(dst_cs_head);
/*
* Wait if task attach is in progress until it is done and then acquire
* cpuset_mutex before returning.
*/
static void wait_attach_done_lock(void)
__acquires(&cpuset_mutex)
{
for (;;) {
mutex_lock(&cpuset_mutex);
if (!attach_ctx.in_progress)
return;
mutex_unlock(&cpuset_mutex);
/* Wait until attach operation is done to prevent racing */
wait_event(cpuset_attach_wq, attach_ctx.in_progress == 0);
}
}
static inline void check_insane_mems_config(nodemask_t *nodes)
{
if (!cpusets_insane_config() &&
@@ -368,22 +404,22 @@ static inline void check_insane_mems_config(nodemask_t *nodes)
}
/*
* decrease cs->attach_in_progress.
* wake_up cpuset_attach_wq if cs->attach_in_progress==0.
* decrease attach_ctx.in_progress.
* wake_up cpuset_attach_wq if attach_ctx.in_progress==0.
*/
static inline void dec_attach_in_progress_locked(struct cpuset *cs)
static inline void dec_attach_in_progress_locked(void)
{
lockdep_assert_cpuset_lock_held();
cs->attach_in_progress--;
if (!cs->attach_in_progress)
attach_ctx.in_progress--;
if (!attach_ctx.in_progress)
wake_up(&cpuset_attach_wq);
}
static inline void dec_attach_in_progress(struct cpuset *cs)
static inline void dec_attach_in_progress(void)
{
mutex_lock(&cpuset_mutex);
dec_attach_in_progress_locked(cs);
dec_attach_in_progress_locked();
mutex_unlock(&cpuset_mutex);
}
@@ -432,8 +468,7 @@ static inline bool partition_is_populated(struct cpuset *cs,
* nr_populated_domain_children may include populated
* csets from descendants that are partitions.
*/
if (cgroup_has_tasks(cs->css.cgroup) ||
cs->attach_in_progress)
if (cgroup_has_tasks(cs->css.cgroup))
return true;
rcu_read_lock();
@@ -489,7 +524,10 @@ static void guarantee_active_cpus(struct task_struct *tsk,
* Return in *pmask the portion of a cpusets's mems_allowed that
* are online, with memory. If none are online with memory, walk
* up the cpuset hierarchy until we find one that does have some
* online mems. The top cpuset always has some mems online.
* online mems. The top cpuset always has some mems online. With v2,
* effective_mems should always contain online memory nodes except
* during the transition period where a memory node hotunplug operation
* is in progress.
*
* One way or another, we guarantee to return some non-empty subset
* of node_states[N_MEMORY].
@@ -581,6 +619,7 @@ static struct cpuset *dup_or_alloc_cpuset(struct cpuset *cs)
return NULL;
trial->dl_bw_cpu = -1;
init_llist_node(&trial->attach_node);
/* Setup cpumask pointer array */
cpumask_var_t *pmask[4] = {
@@ -918,7 +957,7 @@ static void dl_update_tasks_root_domain(struct cpuset *cs)
struct css_task_iter it;
struct task_struct *task;
if (cs->nr_deadline_tasks == 0)
if (atomic_read(&cs->nr_deadline_tasks) == 0)
return;
css_task_iter_start(&cs->css, 0, &it);
@@ -1089,12 +1128,35 @@ void cpuset_update_tasks_cpumask(struct cpuset *cs, struct cpumask *new_cpus)
* @cs: the cpuset the need to recompute the new effective_cpus mask
* @parent: the parent cpuset
*
* For v2, the parent's effective_cpus is inherited if cpumask is empty.
* The result is valid only if the given cpuset isn't a partition root.
*/
static void compute_effective_cpumask(struct cpumask *new_cpus,
struct cpuset *cs, struct cpuset *parent)
{
cpumask_and(new_cpus, cs->cpus_allowed, parent->effective_cpus);
bool has_cpus;
has_cpus = cpumask_and(new_cpus, cs->cpus_allowed, parent->effective_cpus);
if (!has_cpus && is_in_v2_mode())
cpumask_copy(new_cpus, parent->effective_cpus);
}
/**
* compute_effective_nodemask - Compute the effective nodemask of the cpuset
* @new_mems: the temp variable for the new effective_mems mask
* @cs: the cpuset the need to recompute the new effective_mems mask
* @parent: the parent cpuset
*
* For v2, the parent's effective_mems is inherited if nodemask is empty.
*/
static void compute_effective_nodemask(nodemask_t *new_mems,
struct cpuset *cs, struct cpuset *parent)
{
bool has_mems;
has_mems = nodes_and(*new_mems, cs->mems_allowed, parent->effective_mems);
if (!has_mems && is_in_v2_mode())
nodes_copy(*new_mems, parent->effective_mems);
}
/*
@@ -1525,7 +1587,7 @@ static int remote_partition_enable(struct cpuset *cs, int new_prs,
cpumask_copy(cs->effective_xcpus, tmp->new_cpus);
spin_unlock_irq(&callback_lock);
cpuset_force_rebuild();
cs->prs_err = 0;
WRITE_ONCE(cs->prs_err, 0);
/*
* Propagate changes in top_cpuset's effective_cpus down the hierarchy.
@@ -1599,7 +1661,7 @@ static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
WARN_ON_ONCE(!cpumask_subset(cs->effective_xcpus, subpartitions_cpus));
if (cpumask_empty(excpus)) {
cs->prs_err = PERR_CPUSEMPTY;
WRITE_ONCE(cs->prs_err, PERR_CPUSEMPTY);
goto invalidate;
}
@@ -1614,13 +1676,13 @@ static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
if (adding) {
WARN_ON_ONCE(cpumask_intersects(tmp->addmask, subpartitions_cpus));
if (!capable(CAP_SYS_ADMIN))
cs->prs_err = PERR_ACCESS;
WRITE_ONCE(cs->prs_err, PERR_ACCESS);
else if (cpumask_intersects(tmp->addmask, subpartitions_cpus) ||
cpumask_subset(top_cpuset.effective_cpus, tmp->addmask))
cs->prs_err = PERR_NOCPUS;
WRITE_ONCE(cs->prs_err, PERR_NOCPUS);
else if ((prs == PRS_ISOLATED) &&
!isolated_cpus_can_update(tmp->addmask, tmp->delmask))
cs->prs_err = PERR_HKEEPING;
WRITE_ONCE(cs->prs_err, PERR_HKEEPING);
if (cs->prs_err)
goto invalidate;
}
@@ -2048,13 +2110,13 @@ static void compute_partition_effective_cpumask(struct cpuset *cs,
* partition root.
*/
WARN_ON_ONCE(is_remote_partition(child));
child->prs_err = 0;
WRITE_ONCE(child->prs_err, 0);
if (!cpumask_subset(child->effective_xcpus,
cs->effective_xcpus))
child->prs_err = PERR_INVCPUS;
WRITE_ONCE(child->prs_err, PERR_INVCPUS);
else if (populated &&
cpumask_subset(new_ecpus, child->effective_xcpus))
child->prs_err = PERR_NOCPUS;
WRITE_ONCE(child->prs_err, PERR_NOCPUS);
if (child->prs_err) {
int old_prs = child->partition_root_state;
@@ -2143,15 +2205,6 @@ static void update_cpumasks_hier(struct cpuset *cs, struct tmpmasks *tmp,
goto update_parent_effective;
}
/*
* If it becomes empty, inherit the effective mask of the
* parent, which is guaranteed to have some CPUs unless
* it is a partition root that has explicitly distributed
* out all its CPUs.
*/
if (is_in_v2_mode() && !remote && cpumask_empty(tmp->new_cpus))
cpumask_copy(tmp->new_cpus, parent->effective_cpus);
/*
* Skip the whole subtree if
* 1) the cpumask remains the same,
@@ -2367,8 +2420,10 @@ static void partition_cpus_change(struct cpuset *cs, struct cpuset *trialcs,
return;
prs_err = validate_partition(cs, trialcs);
if (prs_err)
trialcs->prs_err = cs->prs_err = prs_err;
if (prs_err) {
WRITE_ONCE(cs->prs_err, prs_err);
trialcs->prs_err = prs_err;
}
if (is_remote_partition(cs)) {
if (trialcs->prs_err)
@@ -2619,6 +2674,14 @@ static void *cpuset_being_rebound;
* Iterate through each task of @cs updating its mems_allowed to the
* effective cpuset's. As this function is called with cpuset_mutex held,
* cpuset membership stays stable.
*
* - cpuset_change_task_nodemask(): guarantee_online_mems()
* - mpol_rebind_mm(): effective_mems
* - cpuset_migrate_mm(): guarantee_online_mems()
* - old_mems_allowed: guarantee_online_mems()
*
* For v2, guarantee_online_mems() should return a node mask that is the same
* as the effective_mems of current cpuset.
*/
void cpuset_update_tasks_nodemask(struct cpuset *cs)
{
@@ -2627,7 +2690,6 @@ void cpuset_update_tasks_nodemask(struct cpuset *cs)
struct task_struct *task;
cpuset_being_rebound = cs; /* causes mpol_dup() rebind */
guarantee_online_mems(cs, &newmems);
/*
@@ -2647,6 +2709,10 @@ void cpuset_update_tasks_nodemask(struct cpuset *cs)
cpuset_change_task_nodemask(task, &newmems);
/* Rebind and migrate mm only for thread group leader */
if (!thread_group_leader(task))
continue;
mm = get_task_mm(task);
if (!mm)
continue;
@@ -2697,14 +2763,7 @@ static void update_nodemasks_hier(struct cpuset *cs, nodemask_t *new_mems)
cpuset_for_each_descendant_pre(cp, pos_css, cs) {
struct cpuset *parent = parent_cs(cp);
bool has_mems = nodes_and(*new_mems, cp->mems_allowed, parent->effective_mems);
/*
* If it becomes empty, inherit the effective mask of the
* parent, which is guaranteed to have some MEMs.
*/
if (is_in_v2_mode() && !has_mems)
*new_mems = parent->effective_mems;
compute_effective_nodemask(new_mems, cp, parent);
/* Skip the whole subtree if the nodemask remains the same. */
if (nodes_equal(*new_mems, cp->effective_mems)) {
@@ -2805,7 +2864,7 @@ int cpuset_update_flag(cpuset_flagbits_t bit, struct cpuset *cs,
{
struct cpuset *trialcs;
int balance_flag_changed;
int spread_flag_changed;
int spread_page_changed;
int err;
trialcs = dup_or_alloc_cpuset(cs);
@@ -2824,8 +2883,7 @@ int cpuset_update_flag(cpuset_flagbits_t bit, struct cpuset *cs,
balance_flag_changed = (is_sched_load_balance(cs) !=
is_sched_load_balance(trialcs));
spread_flag_changed = ((is_spread_slab(cs) != is_spread_slab(trialcs))
|| (is_spread_page(cs) != is_spread_page(trialcs)));
spread_page_changed = is_spread_page(cs) != is_spread_page(trialcs);
spin_lock_irq(&callback_lock);
cs->flags = trialcs->flags;
@@ -2838,7 +2896,7 @@ int cpuset_update_flag(cpuset_flagbits_t bit, struct cpuset *cs,
rebuild_sched_domains_locked();
}
if (spread_flag_changed)
if (spread_page_changed)
cpuset1_update_tasks_flags(cs);
out:
free_cpuset(trialcs);
@@ -2973,58 +3031,47 @@ static int update_prstate(struct cpuset *cs, int new_prs)
return 0;
}
static struct cpuset *cpuset_attach_old_cs;
/*
* Check to see if a cpuset can accept a new task
* For v1, cpus_allowed and mems_allowed can't be empty.
* For v2, effective_cpus can't be empty.
* Note that in v1, effective_cpus = cpus_allowed.
*
* Also set the boolean flag passed in by @psetsched depending on if
* security_task_setscheduler() call is needed and @oldcs is not NULL.
*/
static int cpuset_can_attach_check(struct cpuset *cs)
static int cpuset_can_attach_check(struct cpuset *cs, struct cpuset *oldcs,
bool *psetsched)
{
bool cpus_updated, mems_updated;
if (cpumask_empty(cs->effective_cpus) ||
(!is_in_v2_mode() && nodes_empty(cs->mems_allowed)))
return -ENOSPC;
return 0;
}
static void reset_migrate_dl_data(struct cpuset *cs)
{
cs->nr_migrate_dl_tasks = 0;
cs->sum_migrate_dl_bw = 0;
cs->dl_bw_cpu = -1;
}
if (!oldcs)
return 0;
/* Called by cgroups to determine if a cpuset is usable; cpuset_mutex held */
static int cpuset_can_attach(struct cgroup_taskset *tset)
{
struct cgroup_subsys_state *css;
struct cpuset *cs, *oldcs;
struct task_struct *task;
bool setsched_check;
int cpu, ret;
if (!llist_on_list(&oldcs->attach_node))
llist_add(&oldcs->attach_node, &src_cs_head);
/* used later by cpuset_attach() */
cpuset_attach_old_cs = task_cs(cgroup_taskset_first(tset, &css));
oldcs = cpuset_attach_old_cs;
cs = css_cs(css);
if (!llist_on_list(&cs->attach_node))
llist_add(&cs->attach_node, &dst_cs_head);
mutex_lock(&cpuset_mutex);
cpus_updated = !cpumask_equal(cs->effective_cpus, oldcs->effective_cpus);
mems_updated = !nodes_equal(cs->effective_mems, oldcs->effective_mems);
/* Check to see if task is allowed in the cpuset */
ret = cpuset_can_attach_check(cs);
if (ret)
goto out_unlock;
if (cpus_updated)
attach_ctx.cpus_updated = true;
if (mems_updated)
attach_ctx.mems_updated = true;
/*
* Skip rights over task setsched check in v2 when nothing changes,
* migration permission derives from hierarchy ownership in
* cgroup_procs_write_permission()).
* Skip rights over task setsched check in v2 when nothing changes for
* the current oldcs/cs pair, migration permission derives from
* hierarchy ownership in cgroup_procs_write_permission()).
*/
setsched_check = !cpuset_v2() ||
!cpumask_equal(cs->effective_cpus, oldcs->effective_cpus) ||
!nodes_equal(cs->effective_mems, oldcs->effective_mems);
*psetsched = !cpuset_v2() || cpus_updated || mems_updated;
/*
* A v1 cpuset with tasks will have no CPU left only when CPU hotplug
@@ -3034,13 +3081,123 @@ static int cpuset_can_attach(struct cgroup_taskset *tset)
* sure they will be able to run after migration.
*/
if (!is_in_v2_mode() && cpumask_empty(oldcs->effective_cpus))
setsched_check = false;
*psetsched = false;
return 0;
}
static int cpuset_reserve_dl_bw(void)
{
struct cpuset *cs;
int cpu, ret;
llist_for_each_entry(cs, dst_cs_head.first, attach_node) {
if (!cs->sum_migrate_dl_bw)
continue;
cpu = cpumask_any_and(cpu_active_mask, cs->effective_cpus);
if (unlikely(cpu >= nr_cpu_ids))
return -EINVAL;
ret = dl_bw_alloc(cpu, cs->sum_migrate_dl_bw);
if (ret)
return ret;
cs->dl_bw_cpu = cpu;
}
return 0;
}
/*
* Clear and optionally apply (@cancel is false) the attach related data in the
* source or destination cpuset.
*/
static void clear_attach_data(struct llist_head *head, bool cancel)
{
struct cpuset *cs, *next;
struct llist_node *lnode = __llist_del_all(head);
llist_for_each_entry_safe(cs, next, lnode, attach_node) {
init_llist_node(&cs->attach_node);
if (cs->nr_migrate_dl_tasks) {
if (!cancel)
atomic_add(cs->nr_migrate_dl_tasks, &cs->nr_deadline_tasks);
else if (cs->dl_bw_cpu >= 0) /* && cancel */
dl_bw_free(cs->dl_bw_cpu, cs->sum_migrate_dl_bw);
cs->nr_migrate_dl_tasks = 0;
cs->sum_migrate_dl_bw = 0;
cs->dl_bw_cpu = -1;
}
}
}
/* Called by cgroups to determine if a cpuset is usable; cpuset_mutex held */
static int cpuset_can_attach(struct cgroup_taskset *tset)
{
struct cgroup_subsys_state *css;
struct cpuset *cs, *oldcs;
struct task_struct *task;
bool setsched_check;
int ret;
cs = oldcs = NULL;
mutex_lock(&cpuset_mutex);
attach_ctx.old_cs = NULL; /* Used later in cpuset_attach_task() */
attach_ctx.cpus_updated = false;
attach_ctx.mems_updated = false;
attach_ctx.many_dest_cs = false;
/*
* The attach_ctx.old_cs is used mainly by cpuset_migrate_mm() to get
* the old_mems_allowed value. There are two ways that many-to-one
* cpuset migration can happen:
* 1) A multithread application with threads in different cpusets is
* wholely migrated to a new cpuset.
* 2) Disabling v2 cpuset controller will move all the tasks in child
* cpusets to the parent cpuset.
*
* In the former case, it is the mm setting of the group leader that
* really matters. So attach_ctx.old_cs should track the oldcs of the
* group leader. It falls back to the oldcs of the first task if there
* is no group leader in the taskset. In the latter case, effective_mems
* of child cpusets must always be a subset of the parent. So no real
* page migration will be necessary no matter which child cpuset is
* selected as attach_ctx.old_cs.
*
* For a v2 threaded subtree where cpuset isn't enabled in some of the
* cgroups, it is possible that oldcs == cs for some of the tasks.
* In this case, we can skip checking on those tasks as there is no
* actual migration wrt cpuset.
*/
cgroup_taskset_for_each(task, css, tset) {
struct cpuset *new_cs = css_cs(css);
struct cpuset *new_oldcs = task_cs(task);
if ((new_oldcs != oldcs) || (new_cs != cs)) {
if (cs && (new_cs != cs))
attach_ctx.many_dest_cs = true;
cs = new_cs;
oldcs = new_oldcs;
if (oldcs == cs)
continue;
if (!attach_ctx.old_cs)
attach_ctx.old_cs = oldcs;
ret = cpuset_can_attach_check(cs, oldcs, &setsched_check);
if (ret)
goto out_unlock;
}
if (oldcs == cs)
continue;
ret = task_can_attach(task);
if (ret)
goto out_unlock;
/* Update attach_ctx.old_cs to the latest group leader */
if (task == task->group_leader)
attach_ctx.old_cs = task_cs(task);
if (setsched_check) {
ret = security_task_setscheduler(task);
if (ret)
@@ -3054,57 +3211,48 @@ static int cpuset_can_attach(struct cgroup_taskset *tset)
* contribute to sum_migrate_dl_bw.
*/
cs->nr_migrate_dl_tasks++;
oldcs->nr_migrate_dl_tasks--;
if (dl_task_needs_bw_move(task, cs->effective_cpus))
cs->sum_migrate_dl_bw += task->dl.dl_bw;
}
}
if (!cs->sum_migrate_dl_bw)
goto out_success;
cpu = cpumask_any_and(cpu_active_mask, cs->effective_cpus);
if (unlikely(cpu >= nr_cpu_ids)) {
/*
* The only case where there are multiple destination cpusets for
* task migration is when enabling a v2 cpuset controllers where
* tasks will be migrated to multiple child cpusets from a parent
* cpuset with the same effective CPUs and memory nodes. IOW,
* both attach_cpus_updated and attach_mems_updated should be false.
* If not, it is a condition that the current code cannot handle.
* Print a warning and abort the attach operation as further code
* change may be needed.
*/
if (WARN_ON_ONCE(attach_ctx.many_dest_cs && (!cpuset_v2() ||
attach_ctx.cpus_updated || attach_ctx.mems_updated))) {
ret = -EINVAL;
goto out_unlock;
}
ret = dl_bw_alloc(cpu, cs->sum_migrate_dl_bw);
if (ret)
goto out_unlock;
cs->dl_bw_cpu = cpu;
out_success:
/*
* Mark attach is in progress. This makes validate_change() fail
* changes which zero cpus/mems_allowed.
*/
cs->attach_in_progress++;
ret = cpuset_reserve_dl_bw();
out_unlock:
if (ret)
reset_migrate_dl_data(cs);
if (ret) {
clear_attach_data(&src_cs_head, true);
clear_attach_data(&dst_cs_head, true);
} else {
attach_ctx.in_progress++;
}
mutex_unlock(&cpuset_mutex);
return ret;
}
static void cpuset_cancel_attach(struct cgroup_taskset *tset)
{
struct cgroup_subsys_state *css;
struct cpuset *cs;
cgroup_taskset_first(tset, &css);
cs = css_cs(css);
mutex_lock(&cpuset_mutex);
dec_attach_in_progress_locked(cs);
if (cs->dl_bw_cpu >= 0)
dl_bw_free(cs->dl_bw_cpu, cs->sum_migrate_dl_bw);
if (cs->nr_migrate_dl_tasks)
reset_migrate_dl_data(cs);
dec_attach_in_progress_locked();
clear_attach_data(&src_cs_head, true);
clear_attach_data(&dst_cs_head, true);
mutex_unlock(&cpuset_mutex);
}
@@ -3114,10 +3262,11 @@ static void cpuset_cancel_attach(struct cgroup_taskset *tset)
* allocate from cpuset_init().
*/
static cpumask_var_t cpus_attach;
static nodemask_t cpuset_attach_nodemask_to;
static void cpuset_attach_task(struct cpuset *cs, struct task_struct *task)
{
struct mm_struct *mm;
lockdep_assert_cpuset_lock_held();
if (cs != &top_cpuset)
@@ -3131,90 +3280,88 @@ static void cpuset_attach_task(struct cpuset *cs, struct task_struct *task)
*/
WARN_ON_ONCE(set_cpus_allowed_ptr(task, cpus_attach));
cpuset_change_task_nodemask(task, &cpuset_attach_nodemask_to);
if (cpuset_v2() && !attach_ctx.mems_updated)
return;
cpuset_change_task_nodemask(task, &attach_ctx.nodemask_to);
cpuset1_update_task_spread_flags(cs, task);
if ((task != task->group_leader) || !attach_ctx.mems_updated)
return;
/*
* Change mm for threadgroup leader. This is expensive and may
* sleep and should be moved outside migration path proper.
*/
mm = get_task_mm(task);
if (mm) {
struct cpuset *oldcs = attach_ctx.old_cs;
mpol_rebind_mm(mm, &cs->effective_mems);
/*
* old_mems_allowed is the same with mems_allowed
* here, except if this task is being moved
* automatically due to hotplug. In that case
* @mems_allowed has been updated and is empty, so
* @old_mems_allowed is the right nodesets that we
* migrate mm from.
*/
if (is_memory_migrate(cs)) {
cpuset_migrate_mm(mm, &oldcs->old_mems_allowed,
&attach_ctx.nodemask_to);
attach_ctx.task_work_queued = true;
} else {
mmput(mm);
}
}
}
static void cpuset_attach(struct cgroup_taskset *tset)
{
struct task_struct *task;
struct task_struct *leader;
struct cgroup_subsys_state *css;
struct cpuset *cs;
struct cpuset *oldcs = cpuset_attach_old_cs;
bool cpus_updated, mems_updated;
bool queue_task_work = false;
cgroup_taskset_first(tset, &css);
cs = css_cs(css);
lockdep_assert_cpus_held(); /* see cgroup_attach_lock() */
mutex_lock(&cpuset_mutex);
cpus_updated = !cpumask_equal(cs->effective_cpus,
oldcs->effective_cpus);
mems_updated = !nodes_equal(cs->effective_mems, oldcs->effective_mems);
attach_ctx.task_work_queued = false;
guarantee_online_mems(cs, &attach_ctx.nodemask_to);
/*
* attach_ctx.old_cs can only be NULL if no task is actually migrating.
* This is highly unlikely. If it happens at all, we can skip task
* iteration and setting old_mems_allowed.
*/
if (unlikely(!attach_ctx.old_cs))
goto out;
/*
* In the default hierarchy, enabling cpuset in the child cgroups
* will trigger a number of cpuset_attach() calls with no change
* in effective cpus and mems. In that case, we can optimize out
* by skipping the task iteration and update.
* will trigger a cpuset_attach() call with no change in effective cpus
* and mems. In that case, we can optimize out by skipping the task
* iteration and the destination cpuset list is iterated to set
* old_mems_allowed.
*/
if (cpuset_v2() && !cpus_updated && !mems_updated) {
cpuset_attach_nodemask_to = cs->effective_mems;
if (cpuset_v2() && !attach_ctx.cpus_updated && !attach_ctx.mems_updated) {
llist_for_each_entry(cs, dst_cs_head.first, attach_node)
cs->old_mems_allowed = attach_ctx.nodemask_to;
goto out;
}
guarantee_online_mems(cs, &cpuset_attach_nodemask_to);
cgroup_taskset_for_each(task, css, tset)
cpuset_attach_task(cs, task);
/*
* Change mm for all threadgroup leaders. This is expensive and may
* sleep and should be moved outside migration path proper. Skip it
* if there is no change in effective_mems and CS_MEMORY_MIGRATE is
* not set.
*/
cpuset_attach_nodemask_to = cs->effective_mems;
if (!is_memory_migrate(cs) && !mems_updated)
goto out;
cgroup_taskset_for_each_leader(leader, css, tset) {
struct mm_struct *mm = get_task_mm(leader);
if (mm) {
mpol_rebind_mm(mm, &cpuset_attach_nodemask_to);
/*
* old_mems_allowed is the same with mems_allowed
* here, except if this task is being moved
* automatically due to hotplug. In that case
* @mems_allowed has been updated and is empty, so
* @old_mems_allowed is the right nodesets that we
* migrate mm from.
*/
if (is_memory_migrate(cs)) {
cpuset_migrate_mm(mm, &oldcs->old_mems_allowed,
&cpuset_attach_nodemask_to);
queue_task_work = true;
} else
mmput(mm);
}
}
out:
if (queue_task_work)
if (attach_ctx.task_work_queued)
schedule_flush_migrate_mm();
cs->old_mems_allowed = cpuset_attach_nodemask_to;
if (cs->nr_migrate_dl_tasks) {
cs->nr_deadline_tasks += cs->nr_migrate_dl_tasks;
oldcs->nr_deadline_tasks -= cs->nr_migrate_dl_tasks;
reset_migrate_dl_data(cs);
}
dec_attach_in_progress_locked(cs);
cs->old_mems_allowed = attach_ctx.nodemask_to;
out:
clear_attach_data(&src_cs_head, false);
clear_attach_data(&dst_cs_head, false);
dec_attach_in_progress_locked();
mutex_unlock(&cpuset_mutex);
}
@@ -3234,7 +3381,12 @@ ssize_t cpuset_write_resmask(struct kernfs_open_file *of,
return -EACCES;
buf = strstrip(buf);
cpuset_full_lock();
/* cpuset_mutex acquired in wait_attach_done_lock() */
mutex_lock(&cpuset_top_mutex);
cpus_read_lock();
wait_attach_done_lock();
if (!is_cpuset_online(cs))
goto out_unlock;
@@ -3365,7 +3517,10 @@ static ssize_t cpuset_partition_write(struct kernfs_open_file *of, char *buf,
else
return -EINVAL;
cpuset_full_lock();
mutex_lock(&cpuset_top_mutex);
cpus_read_lock();
wait_attach_done_lock();
if (is_cpuset_online(cs))
retval = update_prstate(cs, val);
cpuset_update_sd_hk_unlock();
@@ -3592,7 +3747,7 @@ static int cpuset_can_fork(struct task_struct *task, struct css_set *cset)
mutex_lock(&cpuset_mutex);
/* Check to see if task is allowed in the cpuset */
ret = cpuset_can_attach_check(cs);
ret = cpuset_can_attach_check(cs, NULL, NULL);
if (ret)
goto out_unlock;
@@ -3604,11 +3759,7 @@ static int cpuset_can_fork(struct task_struct *task, struct css_set *cset)
if (ret)
goto out_unlock;
/*
* Mark attach is in progress. This makes validate_change() fail
* changes which zero cpus/mems_allowed.
*/
cs->attach_in_progress++;
attach_ctx.in_progress++;
out_unlock:
mutex_unlock(&cpuset_mutex);
return ret;
@@ -3626,7 +3777,7 @@ static void cpuset_cancel_fork(struct task_struct *task, struct css_set *cset)
if (same_cs)
return;
dec_attach_in_progress(cs);
dec_attach_in_progress();
}
/*
@@ -3636,15 +3787,14 @@ static void cpuset_cancel_fork(struct task_struct *task, struct css_set *cset)
*/
static void cpuset_fork(struct task_struct *task)
{
struct cpuset *cs;
bool same_cs;
struct cpuset *cs, *oldcs;
rcu_read_lock();
cs = task_cs(task);
same_cs = (cs == task_cs(current));
oldcs = task_cs(current);
rcu_read_unlock();
if (same_cs) {
if (cs == oldcs) {
if (cs == &top_cpuset)
return;
@@ -3655,10 +3805,22 @@ static void cpuset_fork(struct task_struct *task)
/* CLONE_INTO_CGROUP */
mutex_lock(&cpuset_mutex);
guarantee_online_mems(cs, &cpuset_attach_nodemask_to);
cpuset_attach_task(cs, task);
guarantee_online_mems(cs, &attach_ctx.nodemask_to);
cs->old_mems_allowed = attach_ctx.nodemask_to;
dec_attach_in_progress_locked(cs);
/*
* Assume CPUs and memory nodes are updated
* A CLONE_INTO_CGROUP operation should have taken the cgroup mutex
* and so there shouldn't be a competing cpuset_attach() operation.
*/
attach_ctx.cpus_updated = attach_ctx.mems_updated = true;
attach_ctx.task_work_queued = false;
attach_ctx.old_cs = oldcs;
cpuset_attach_task(cs, task);
if (attach_ctx.task_work_queued)
schedule_flush_migrate_mm();
dec_attach_in_progress_locked();
mutex_unlock(&cpuset_mutex);
}
@@ -3705,6 +3867,7 @@ int __init cpuset_init(void)
cpumask_setall(top_cpuset.effective_xcpus);
cpumask_setall(top_cpuset.exclusive_cpus);
nodes_setall(top_cpuset.effective_mems);
init_llist_node(&top_cpuset.attach_node);
cpuset1_init(&top_cpuset);
@@ -3762,23 +3925,11 @@ static void cpuset_hotplug_update_tasks(struct cpuset *cs, struct tmpmasks *tmp)
bool remote;
int partcmd = -1;
struct cpuset *parent;
retry:
wait_event(cpuset_attach_wq, cs->attach_in_progress == 0);
mutex_lock(&cpuset_mutex);
/*
* We have raced with task attaching. We wait until attaching
* is finished, so we won't attach a task to an empty cpuset.
*/
if (cs->attach_in_progress) {
mutex_unlock(&cpuset_mutex);
goto retry;
}
wait_attach_done_lock();
parent = parent_cs(cs);
compute_effective_cpumask(&new_cpus, cs, parent);
nodes_and(new_mems, cs->mems_allowed, parent->effective_mems);
compute_effective_nodemask(&new_mems, cs, parent);
if (!tmp || !cs->partition_root_state)
goto update_tasks;
@@ -3794,7 +3945,7 @@ static void cpuset_hotplug_update_tasks(struct cpuset *cs, struct tmpmasks *tmp)
if (remote && (cpumask_empty(subpartitions_cpus) ||
(cpumask_empty(&new_cpus) &&
partition_is_populated(cs, NULL)))) {
cs->prs_err = PERR_HOTPLUG;
WRITE_ONCE(cs->prs_err, PERR_HOTPLUG);
remote_partition_disable(cs, tmp);
compute_effective_cpumask(&new_cpus, cs, parent);
remote = false;
@@ -4337,14 +4488,10 @@ void cpuset_nodes_allowed(struct cgroup *cgroup, nodemask_t *mask)
* cpuset_spread_node() - On which node to begin search for a page
* @rotor: round robin rotor
*
* If a task is marked PF_SPREAD_PAGE or PF_SPREAD_SLAB (as for
* tasks in a cpuset with is_spread_page or is_spread_slab set),
* and if the memory allocation used cpuset_mem_spread_node()
* to determine on which node to start looking, as it will for
* certain page cache or slab cache pages such as used for file
* system buffers and inode caches, then instead of starting on the
* local node to look for a free page, rather spread the starting
* node around the tasks mems_allowed nodes.
* If a task is marked PFA_SPREAD_PAGE and a page cache allocation uses
* cpuset_mem_spread_node() to determine where to start looking, spread the
* starting node around the task's mems_allowed nodes instead of starting on
* the local node.
*
* We don't have to worry about the returned node being offline
* because "it can't happen", and even if it did, it would be ok.

View File

@@ -15,11 +15,6 @@ import time
import json
import math
import drgn
from drgn import container_of
from drgn.helpers.linux.list import list_for_each_entry,list_empty
from drgn.helpers.linux.radixtree import radix_tree_for_each,radix_tree_lookup
import argparse
parser = argparse.ArgumentParser(description=desc,
formatter_class=argparse.RawTextHelpFormatter)
@@ -34,6 +29,11 @@ parser.add_argument('--json', action='store_true',
help='Output in json')
args = parser.parse_args()
import drgn
from drgn import container_of
from drgn.helpers.linux.list import list_for_each_entry,list_empty
from drgn.helpers.linux.radixtree import radix_tree_for_each,radix_tree_lookup
def err(s):
print(s, file=sys.stderr, flush=True)
sys.exit(1)

View File

@@ -8,6 +8,7 @@
#define MB(x) (x << 20)
#define NSEC_PER_USEC 1000L
#define USEC_PER_SEC 1000000L
#define NSEC_PER_SEC 1000000000L

View File

@@ -426,7 +426,6 @@ static int test_cgcore_no_internal_process_constraint_on_threads(const char *roo
ret = KSFT_PASS;
cleanup:
cg_enter_current(root);
cg_enter_current(root);
if (child)
cg_destroy(child);
@@ -795,10 +794,9 @@ static int lesser_ns_open_thread_fn(void *arg)
static int test_cgcore_lesser_ns_open(const char *root)
{
static char stack[65536];
const uid_t test_euid = 65534; /* usually nobody, any !root is fine */
int ret = KSFT_FAIL;
char *cg_test_a = NULL, *cg_test_b = NULL;
char *cg_test_a_procs = NULL, *cg_test_b_procs = NULL;
char *cg_test_b_procs = NULL;
int cg_test_b_procs_fd = -1;
struct lesser_ns_open_thread_arg targ = { .fd = -1 };
pid_t pid;
@@ -813,10 +811,9 @@ static int test_cgcore_lesser_ns_open(const char *root)
if (!cg_test_a || !cg_test_b)
goto cleanup;
cg_test_a_procs = cg_name(cg_test_a, "cgroup.procs");
cg_test_b_procs = cg_name(cg_test_b, "cgroup.procs");
if (!cg_test_a_procs || !cg_test_b_procs)
if (!cg_test_b_procs)
goto cleanup;
if (cg_create(cg_test_a) || cg_create(cg_test_b))
@@ -825,10 +822,6 @@ static int test_cgcore_lesser_ns_open(const char *root)
if (cg_enter_current(cg_test_b))
goto cleanup;
if (chown(cg_test_a_procs, test_euid, -1) ||
chown(cg_test_b_procs, test_euid, -1))
goto cleanup;
targ.path = cg_test_b_procs;
pid = clone(lesser_ns_open_thread_fn, stack + sizeof(stack),
CLONE_NEWCGROUP | CLONE_FILES | CLONE_VM | SIGCHLD,
@@ -863,7 +856,6 @@ static int test_cgcore_lesser_ns_open(const char *root)
if (cg_test_a)
cg_destroy(cg_test_a);
free(cg_test_b_procs);
free(cg_test_a_procs);
free(cg_test_b);
free(cg_test_a);
return ret;

View File

@@ -291,6 +291,8 @@ static int test_cpucg_nice(const char *root)
user_usec = cg_read_key_long(cpucg, "cpu.stat", "user_usec");
nice_usec = cg_read_key_long(cpucg, "cpu.stat", "nice_usec");
if (user_usec <= 0)
goto cleanup;
if (!values_close_report(nice_usec, expected_nice_usec, 1))
goto cleanup;
@@ -639,6 +641,31 @@ test_cpucg_nested_weight_underprovisioned(const char *root)
return run_cpucg_nested_weight_test(root, false);
}
/*
* Best effort attempt to get the kernel's HZ value from the config.
* Return the HZ value if found otherwise return 1000 (the default) to
* indicate failure.
*/
static long
get_config_hz(void)
{
long hz = 1000;
FILE *f;
char cmd[256] = "zcat /proc/config.gz 2>/dev/null | grep '^CONFIG_HZ='";
f = popen(cmd, "r");
if (!f)
return hz;
if (fscanf(f, "CONFIG_HZ=%ld", &hz) == EOF)
goto out;
out:
pclose(f);
return hz;
}
/*
* This test creates a cgroup with some maximum value within a period, and
* verifies that a process in the cgroup is not overscheduled.
@@ -646,15 +673,18 @@ test_cpucg_nested_weight_underprovisioned(const char *root)
static int test_cpucg_max(const char *root)
{
int ret = KSFT_FAIL;
long hz = get_config_hz();
long quota_usec = 1000;
long default_period_usec = 100000; /* cpu.max's default period */
long duration_seconds = 1;
long duration_usec = duration_seconds * USEC_PER_SEC;
long duration_usec;
long usage_usec, n_periods, remainder_usec, expected_usage_usec;
char *cpucg;
char quota_buf[32];
duration_usec = duration_seconds * USEC_PER_SEC * 1000 / hz;
snprintf(quota_buf, sizeof(quota_buf), "%ld", quota_usec);
cpucg = cg_name(root, "cpucg_test");
@@ -670,8 +700,8 @@ static int test_cpucg_max(const char *root)
struct cpu_hog_func_param param = {
.nprocs = 1,
.ts = {
.tv_sec = duration_seconds,
.tv_nsec = 0,
.tv_sec = duration_usec / USEC_PER_SEC,
.tv_nsec = duration_usec % USEC_PER_SEC * NSEC_PER_USEC,
},
.clock_type = CPU_HOG_CLOCK_WALL,
};
@@ -710,15 +740,18 @@ static int test_cpucg_max(const char *root)
static int test_cpucg_max_nested(const char *root)
{
int ret = KSFT_FAIL;
long hz = get_config_hz();
long quota_usec = 1000;
long default_period_usec = 100000; /* cpu.max's default period */
long duration_seconds = 1;
long duration_usec = duration_seconds * USEC_PER_SEC;
long duration_usec;
long usage_usec, n_periods, remainder_usec, expected_usage_usec;
char *parent, *child;
char quota_buf[32];
duration_usec = duration_seconds * USEC_PER_SEC * 1000 / hz;
snprintf(quota_buf, sizeof(quota_buf), "%ld", quota_usec);
parent = cg_name(root, "cpucg_parent");
@@ -741,8 +774,8 @@ static int test_cpucg_max_nested(const char *root)
struct cpu_hog_func_param param = {
.nprocs = 1,
.ts = {
.tv_sec = duration_seconds,
.tv_nsec = 0,
.tv_sec = duration_usec / USEC_PER_SEC,
.tv_nsec = duration_usec % USEC_PER_SEC * NSEC_PER_USEC,
},
.clock_type = CPU_HOG_CLOCK_WALL,
};

View File

@@ -1,7 +1,13 @@
// SPDX-License-Identifier: GPL-2.0
#define _GNU_SOURCE
#include <assert.h>
#include <linux/limits.h>
#include <pthread.h>
#include <sched.h>
#include <signal.h>
#include <sys/syscall.h>
#include <unistd.h>
#include "kselftest.h"
#include "cgroup_util.h"
@@ -232,6 +238,246 @@ static int test_cpuset_perms_subtree(const char *root)
return ret;
}
static int get_cpu_affinity(cpu_set_t *mask)
{
CPU_ZERO(mask);
return sched_getaffinity(0, sizeof(*mask), mask);
}
static int cpu_set_equal(cpu_set_t *dst, unsigned long mask)
{
cpu_set_t expected;
CPU_ZERO(&expected);
assert(sizeof(mask) < CPU_SETSIZE);
for (int cpu = 0; cpu < sizeof(mask) * 8; ++cpu)
if ((1UL << cpu) & mask)
CPU_SET(cpu, &expected);
return CPU_EQUAL(&expected, dst);
}
enum test_phase {
AFFINITY_SETUP,
AFFINITY_CONTROLLER_DISABLED,
AFFINITY_COMPLETE,
AFFINITY_ERROR
};
struct thread_args {
const char *cgroup;
cpu_set_t *affinity_before;
cpu_set_t *affinity_after;
int affinity_before_ready;
};
static pthread_mutex_t test_mutex = PTHREAD_MUTEX_INITIALIZER;
static pthread_cond_t test_cond = PTHREAD_COND_INITIALIZER;
static enum test_phase test_phase;
static void *affinity_thread_fn(void *arg)
{
struct thread_args *args = (struct thread_args *)arg;
if (cg_enter_current_thread(args->cgroup))
goto fail;
if (get_cpu_affinity(args->affinity_before) != 0)
goto fail;
pthread_mutex_lock(&test_mutex);
args->affinity_before_ready = 1;
pthread_cond_broadcast(&test_cond);
while (test_phase < AFFINITY_CONTROLLER_DISABLED)
pthread_cond_wait(&test_cond, &test_mutex);
pthread_mutex_unlock(&test_mutex);
if (get_cpu_affinity(args->affinity_after) != 0)
goto fail;
return NULL;
fail:
pthread_mutex_lock(&test_mutex);
test_phase = AFFINITY_ERROR;
pthread_cond_broadcast(&test_cond);
pthread_mutex_unlock(&test_mutex);
return NULL;
}
/*
* Test that disabling cpuset controller properly updates thread affinity.
*
* This test exposes a bug in cpuset_attach() where threads in child cgroups
* don't get their affinity updated when the cpuset controller is disabled.
*
* Setup:
* - Create parent cgroup with cpuset.cpus=0-1
* - Create child A with cpuset.cpus=0-1
* - Create child B with cpuset.cpus=1
* - Place multithreaded process: group leader + thread_a in A, thread_b in B
* - Disable cpuset controller on parent
*
* Expected: thread_b's affinity should expand from {1} to {0-1}
* Buggy: thread_b's affinity remains {1}
*/
static int test_cpuset_affinity_on_controller_disable(const char *root)
{
char *parent = NULL, *child_a = NULL, *child_b = NULL;
pthread_t thread_a, thread_b;
int thread_a_created = 0, thread_b_created = 0;
cpu_set_t affinity_a_before, affinity_a_after;
cpu_set_t affinity_b_before, affinity_b_after;
int ret = KSFT_FAIL;
parent = cg_name(root, "cpuset_affinity_test");
if (!parent)
goto cleanup;
if (cg_create(parent))
goto cleanup;
if (cg_write(parent, "cgroup.type", "threaded"))
goto cleanup;
child_a = cg_name(parent, "A");
if (!child_a)
goto cleanup;
if (cg_create(child_a))
goto cleanup;
if (cg_write(child_a, "cgroup.type", "threaded"))
goto cleanup;
child_b = cg_name(parent, "B");
if (!child_b)
goto cleanup;
if (cg_create(child_b))
goto cleanup;
if (cg_write(child_b, "cgroup.type", "threaded"))
goto cleanup;
/* Now enable cpuset controller in parent */
if (cg_write(parent, "cgroup.subtree_control", "+cpuset"))
goto skip;
/*
* Set CPU affinity constraints
* Skip the test if the setting of "cpuset.cpus" fails as the test
* system may not have CPU 1.
*/
if (cg_write(parent, "cpuset.cpus", "0-1"))
goto skip;
if (cg_write(child_a, "cpuset.cpus", "0-1"))
goto skip;
if (cg_write(child_b, "cpuset.cpus", "1"))
goto skip;
/* Move group leader (main thread) to child A */
if (cg_enter_current(child_a))
goto cleanup;
/* Create threads - they will move themselves to their respective cgroups */
test_phase = AFFINITY_SETUP;
struct thread_args args_a = {
.cgroup = child_a,
.affinity_before = &affinity_a_before,
.affinity_after = &affinity_a_after,
.affinity_before_ready = 0,
};
if (pthread_create(&thread_a, NULL, affinity_thread_fn, &args_a))
goto cleanup;
thread_a_created = 1;
struct thread_args args_b = {
.cgroup = child_b,
.affinity_before = &affinity_b_before,
.affinity_after = &affinity_b_after,
.affinity_before_ready = 0,
};
if (pthread_create(&thread_b, NULL, affinity_thread_fn, &args_b))
goto cleanup_threads;
thread_b_created = 1;
pthread_mutex_lock(&test_mutex);
while ((test_phase < AFFINITY_ERROR) &&
(args_a.affinity_before_ready + args_b.affinity_before_ready < 2))
pthread_cond_wait(&test_cond, &test_mutex);
/* If a thread failed during setup, bail out */
if (test_phase == AFFINITY_ERROR) {
pthread_mutex_unlock(&test_mutex);
goto cleanup_threads;
}
pthread_mutex_unlock(&test_mutex);
if (!cpu_set_equal(&affinity_a_before, 0x3)) {
ksft_print_msg("FAIL: thread_a initial affinity incorrect\n");
goto cleanup_threads;
}
if (!cpu_set_equal(&affinity_b_before, 0x2)) {
ksft_print_msg("FAIL: thread_b initial affinity incorrect\n");
goto cleanup_threads;
}
/* Disable cpuset controller - this should trigger affinity update */
if (cg_write(parent, "cgroup.subtree_control", "-cpuset"))
goto cleanup_threads;
/* Signal threads to save their final affinity and exit */
pthread_mutex_lock(&test_mutex);
test_phase = AFFINITY_CONTROLLER_DISABLED;
pthread_cond_broadcast(&test_cond);
pthread_mutex_unlock(&test_mutex);
pthread_join(thread_a, NULL);
pthread_join(thread_b, NULL);
/* Verify thread affinities AFTER disabling controller */
if (!cpu_set_equal(&affinity_a_after, 0x3)) {
ksft_print_msg("FAIL: thread_a final affinity incorrect\n");
goto cleanup;
}
if (!cpu_set_equal(&affinity_b_after, 0x3)) {
ksft_print_msg("FAIL: thread_b affinity did not expand to {0-1}\n");
goto cleanup;
}
ret = KSFT_PASS;
goto cleanup;
skip:
ret = KSFT_SKIP;
goto cleanup;
cleanup_threads:
pthread_mutex_lock(&test_mutex);
test_phase = AFFINITY_COMPLETE;
pthread_cond_broadcast(&test_cond);
pthread_mutex_unlock(&test_mutex);
if (thread_a_created)
pthread_join(thread_a, NULL);
if (thread_b_created)
pthread_join(thread_b, NULL);
cleanup:
/* Move back to root before cleanup */
cg_enter_current(root);
cg_destroy(child_b);
free(child_b);
cg_destroy(child_a);
free(child_a);
cg_destroy(parent);
free(parent);
return ret;
}
#define T(x) { x, #x }
struct cpuset_test {
@@ -241,6 +487,7 @@ struct cpuset_test {
T(test_cpuset_perms_object_allow),
T(test_cpuset_perms_object_deny),
T(test_cpuset_perms_subtree),
T(test_cpuset_affinity_on_controller_disable),
};
#undef T

View File

@@ -20,7 +20,7 @@ skip_test() {
WAIT_INOTIFY=$(cd $(dirname $0); pwd)/wait_inotify
# Find cgroup v2 mount point
CGROUP2=$(mount -t cgroup2 | head -1 | awk -e '{print $3}')
CGROUP2=$(mount -t cgroup2 | head -1 | awk '{print $3}')
[[ -n "$CGROUP2" ]] || skip_test "Cgroup v2 mount point not found!"
SUBPARTS_CPUS=$CGROUP2/.__DEBUG__.cpuset.cpus.subpartitions
CPULIST=$(cat $CGROUP2/cpuset.cpus.effective)
@@ -495,13 +495,26 @@ REMOTE_TEST_MATRIX=(
# Narrowing cpuset.cpus to previously sibling-excluded CPUs should
# not return CPUs that were never actually owned.
" C1-4:P1 . C1-2:P1 C1-3:P2 . . \
. . . C3 . . p1:4|c11:1-2|c12:3 \
. . . C3 . . p1:4|c11:1-2|c12:3 \
p1:P1|c11:P1|c12:P2 3"
# Expanding cpuset.cpus to include a previously sibling-excluded CPU
# after the sibling has become a member should correctly request it.
" C1-4:P1 . C1-2:P1 C1-3:P2 . . \
. . P0 C2-3 . . p1:1,4|c11:1|c12:2-3 \
. . P0 C2-3 . . p1:1,4|c11:1|c12:2-3 \
p1:P1|c11:P0|c12:P2 2-3"
# Changing a sibling partition's cpuset.cpus to overlap with another
# sibling partition should invalidate itself and return only actually
# allocated CPUs (effective_xcpus) to the parent.
" C1-4:P1 . C1-2:P1 C2-4:P2 . . \
. . . C1-2 . . p1:3-4|c11:1-2|c12:3-4 \
p1:P1|c11:P1|c12:P-2"
# Cpusets with empty cpuset.cpus should inherit parent's effective_cpus
" C1-4:P1 C5-6 C1-2 . C5 . \
. P1 P1 . . . p1:3-4|p2:5-6|c11:1-2|c12:3-4|c21:5|c22:5-6 \
p1:P1|p2:P1|c11:P1"
" C1-4:P1 C5-6 C1-2 . C5 . \
. P1 P1 . O5=0 . p1:3-4|p2:6|c11:1-2|c12:3-4|c21:6|c22:6 \
p1:P1|p2:P1|c11:P1"
)
#
@@ -513,6 +526,7 @@ write_cpu_online()
CPU=${1%=*}
VAL=${1#*=}
CPUFILE=//sys/devices/system/cpu/cpu${CPU}/online
echo $VAL > $CPUFILE || return 1
if [[ $VAL -eq 0 ]]
then
OFFLINE_CPUS="$OFFLINE_CPUS $CPU"
@@ -522,7 +536,6 @@ write_cpu_online()
sort | uniq -u)
}
fi
echo $VAL > $CPUFILE
pause 0.05
}
@@ -590,7 +603,8 @@ set_ctrl_state()
eval $COMM $REDIRECT
;;
O*) VAL=${CMD#?}
write_cpu_online $VAL
COMM="write_cpu_online $VAL"
eval $COMM $REDIRECT
;;
T*) COMM="echo 0 > $TFILE"
eval $COMM $REDIRECT

View File

@@ -14,7 +14,7 @@ skip_test() {
[[ $(id -u) -eq 0 ]] || skip_test "Test must be run as root!"
# Find cpuset v1 mount point
CPUSET=$(mount -t cgroup | grep cpuset | head -1 | awk -e '{print $3}')
CPUSET=$(mount -t cgroup | grep cpuset | head -1 | awk '{print $3}')
[[ -n "$CPUSET" ]] || skip_test "cpuset v1 mount point not found!"
#

View File

@@ -199,7 +199,10 @@ static int test_hugetlb_memcg(char *root)
int main(int argc, char **argv)
{
char root[PATH_MAX];
int ret = EXIT_SUCCESS, has_memory_hugetlb_acc;
int has_memory_hugetlb_acc;
ksft_print_header();
ksft_set_plan(1);
has_memory_hugetlb_acc = proc_mount_contains("memory_hugetlb_accounting");
if (has_memory_hugetlb_acc < 0)
@@ -211,7 +214,7 @@ int main(int argc, char **argv)
if (get_hugepage_size() != 2048) {
ksft_print_msg("test_hugetlb_memcg requires 2MB hugepages\n");
ksft_test_result_skip("test_hugetlb_memcg\n");
return ret;
ksft_finished();
}
if (cg_find_unified_root(root, sizeof(root), NULL))
@@ -233,10 +236,9 @@ int main(int argc, char **argv)
ksft_test_result_skip("test_hugetlb_memcg\n");
break;
default:
ret = EXIT_FAILURE;
ksft_test_result_fail("test_hugetlb_memcg\n");
break;
}
return ret;
ksft_finished();
}