mirror of
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
synced 2026-08-31 14:04:27 -04:00
Merge tag 'cgroup-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup
Pull cgroup updates from Tejun Heo: - Attach path bug fixes: migrations spanning multiple source or destination cpusets were mishandled, most visibly leaving thread affinities stale when the controller is disabled in a threaded subtree. Configuration writes could also race an in-flight attach and apply stale state, and the deadline task count could get corrupted by concurrent updates, skewing SCHED_DEADLINE admission decisions. - Memory binding bug fixes: which node masks get applied differed between the binding update paths, and tasks cloned with CLONE_INTO_CGROUP skipped rebinding entirely. Rebinding also now runs once per process instead of repeating for every thread sharing the mm. - Overhead removals with no behavior change: CPU hotplug iterated tasks of cpusets that just inherit the parent's effective masks, and the slab-spreading task flag was still being maintained although the SLAB allocator that consumed it is long gone. - Data-race annotations for benign races so that KCSAN reports stay meaningful, selftest coverage for the fixes above along with flakiness and portability fixes, and documentation corrections. * tag 'cgroup-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup: (34 commits) selftests/cgroup: Remove redundant chown in test_cgcore_lesser_ns_open selftests/cgroup: Preserve CPU hotplug write errors cgroup/cpuset: Add test for partition root invalidation returning wrong CPUs cgroup/cpuset: Remove obsolete PFA_SPREAD_SLAB task flag docs: cgroup-v2: fix stale "io" controller introduction selftests/cgroup: Avoid awk -e in cpuset tests cgroup/cpuset: Use WRITE_ONCE() for shared prs_err updates selftests/cgroup: add user_usec sanity check in test_cpucg_nice cgroup: drop unneeded semicolon docs: cgroup-v2: mark memory.pressure and io.pressure as read-write selftests/cgroup: Fix minor defects in test_cpuset Docs/admin-guide/cgroup-v2: fix delay_nsec unit in io.latency doc selftests/cgroup: Remove redundant cg_enter_current() call in test_core selftests/cgroup: Add test for cpuset affinity on controller disable cgroup/cpuset: Handle the special case of non-moving tasks in cpuset_can_attach() cgroup/cpuset: Support multiple destination cpusets for cpuset_*attach() selftests/cgroup: fix missing TAP output in test_hugetlb_memcg cgroup/cpuset: Support multiple source cpusets for cpuset_*attach() cgroup/cpuset: Move mpol_rebind_mm/cpuset_migrate_mm() calls inside cpuset_attach_task() cgroup/cpuset: Make attach_ctx.old_cs track task group leader ...
This commit is contained in:
@@ -179,7 +179,7 @@ files describing that cpuset:
|
||||
- cpuset.mem_hardwall flag: is memory allocation hardwalled
|
||||
- cpuset.memory_pressure: measure of how much paging pressure in cpuset
|
||||
- cpuset.memory_spread_page flag: if set, spread page cache evenly on allowed nodes
|
||||
- cpuset.memory_spread_slab flag: OBSOLETE. Doesn't have any function.
|
||||
- cpuset.memory_spread_slab flag: OBSOLETE. Has no effect on allocation behavior.
|
||||
- cpuset.sched_load_balance flag: if set, load balance within CPUs on that cpuset
|
||||
- cpuset.sched_relax_domain_level: the searching range when migrating tasks
|
||||
|
||||
@@ -318,26 +318,20 @@ times 1000.
|
||||
|
||||
1.6 What is memory spread ?
|
||||
---------------------------
|
||||
There are two boolean flag files per cpuset that control where the
|
||||
kernel allocates pages for the file system buffers and related in
|
||||
kernel data structures. They are called 'cpuset.memory_spread_page' and
|
||||
'cpuset.memory_spread_slab'.
|
||||
The 'cpuset.memory_spread_page' boolean flag file controls where the kernel
|
||||
allocates page-cache pages.
|
||||
The 'cpuset.memory_spread_slab' file is obsolete and has no effect on
|
||||
allocation behavior, but is retained for compatibility.
|
||||
|
||||
If the per-cpuset boolean flag file 'cpuset.memory_spread_page' is set, then
|
||||
the kernel will spread the file system buffers (page cache) evenly
|
||||
over all the nodes that the faulting task is allowed to use, instead
|
||||
of preferring to put those pages on the node where the task is running.
|
||||
|
||||
If the per-cpuset boolean flag file 'cpuset.memory_spread_slab' is set,
|
||||
then the kernel will spread some file system related slab caches,
|
||||
such as for inodes and dentries evenly over all the nodes that the
|
||||
faulting task is allowed to use, instead of preferring to put those
|
||||
pages on the node where the task is running.
|
||||
|
||||
The setting of these flags does not affect anonymous data segment or
|
||||
The setting of this flag does not affect anonymous data segment or
|
||||
stack segment pages of a task.
|
||||
|
||||
By default, both kinds of memory spreading are off, and memory
|
||||
By default, page cache memory spreading is off, and memory
|
||||
pages are allocated on the node local to where the task is running,
|
||||
except perhaps as modified by the task's NUMA mempolicy or cpuset
|
||||
configuration, so long as sufficient free memory pages are available.
|
||||
@@ -345,18 +339,18 @@ configuration, so long as sufficient free memory pages are available.
|
||||
When new cpusets are created, they inherit the memory spread settings
|
||||
of their parent.
|
||||
|
||||
Setting memory spreading causes allocations for the affected page
|
||||
or slab caches to ignore the task's NUMA mempolicy and be spread
|
||||
instead. Tasks using mbind() or set_mempolicy() calls to set NUMA
|
||||
mempolicies will not notice any change in these calls as a result of
|
||||
their containing task's memory spread settings. If memory spreading
|
||||
Setting page cache memory spreading causes affected allocations to ignore the
|
||||
task's NUMA mempolicy and be spread instead. Tasks using mbind() or
|
||||
set_mempolicy() to set NUMA mempolicies will not notice any change as a
|
||||
result of their containing task's memory spread settings. If memory spreading
|
||||
is turned off, then the currently specified NUMA mempolicy once again
|
||||
applies to memory page allocations.
|
||||
|
||||
Both 'cpuset.memory_spread_page' and 'cpuset.memory_spread_slab' are boolean flag
|
||||
files. By default they contain "0", meaning that the feature is off
|
||||
for that cpuset. If a "1" is written to that file, then that turns
|
||||
the named feature on.
|
||||
Both 'cpuset.memory_spread_page' and 'cpuset.memory_spread_slab' are boolean
|
||||
flag files. In the root cpuset, both files initially contain "0". Writing "1"
|
||||
or "0" to 'cpuset.memory_spread_page' enables or disables page-cache spreading,
|
||||
respectively. The value of 'cpuset.memory_spread_slab' is retained, can be read
|
||||
back and inherited, but it does not affect allocation behavior.
|
||||
|
||||
The implementation is simple.
|
||||
|
||||
@@ -367,10 +361,6 @@ is modified to perform an inline check for this PFA_SPREAD_PAGE task
|
||||
flag, and if set, a call to a new routine cpuset_mem_spread_node()
|
||||
returns the node to prefer for the allocation.
|
||||
|
||||
Similarly, setting 'cpuset.memory_spread_slab' turns on the flag
|
||||
PFA_SPREAD_SLAB, and appropriately marked slab caches will allocate
|
||||
pages from the node returned by cpuset_mem_spread_node().
|
||||
|
||||
The cpuset_mem_spread_node() routine is also simple. It uses the
|
||||
value of a per-task rotor cpuset_mem_spread_rotor to select the next
|
||||
node in the current task's mems_allowed to prefer for the allocation.
|
||||
|
||||
@@ -10,7 +10,7 @@ Because VM is getting complex (one of reasons is memcg...), memcg's behavior
|
||||
is complex. This is a document for memcg's internal behavior.
|
||||
Please note that implementation details can be changed.
|
||||
|
||||
(*) Topics on API should be in Documentation/admin-guide/cgroup-v1/memory.rst)
|
||||
(*) Topics on API should be in Documentation/admin-guide/cgroup-v1/memory.rst
|
||||
|
||||
0. How to record usage ?
|
||||
========================
|
||||
|
||||
@@ -1145,7 +1145,7 @@ will be referred to. All time durations are in microseconds.
|
||||
This file exists whether the controller is enabled or not.
|
||||
|
||||
It always reports the following three stats, which account for all the
|
||||
processes in the cgroup:
|
||||
processes in the cgroup (including those in descendant cgroups):
|
||||
|
||||
- usage_usec
|
||||
- user_usec
|
||||
@@ -1160,6 +1160,27 @@ will be referred to. All time durations are in microseconds.
|
||||
- nr_bursts
|
||||
- burst_usec
|
||||
|
||||
Note that the above five CFS bandwidth stats are non-hierarchical;
|
||||
they only account for throttling caused by this cgroup's own bandwidth
|
||||
limit, not including throttling inherited from ancestor cgroups.
|
||||
|
||||
cpu.stat.local
|
||||
A read-only flat-keyed file.
|
||||
This file exists whether the controller is enabled or not.
|
||||
|
||||
It reports the following stat when the controller is enabled:
|
||||
|
||||
- throttled_usec
|
||||
|
||||
Unlike the ``throttled_usec`` reported by ``cpu.stat`` which
|
||||
accounts for throttling caused by this cgroup's own CFS
|
||||
bandwidth limit, ``cpu.stat.local`` reports the actual
|
||||
throttling time incurred by this cgroup's own runqueues,
|
||||
which may include throttling inherited from ancestor
|
||||
cgroup bandwidth limits.
|
||||
|
||||
When the controller is not enabled, this stat is not reported.
|
||||
|
||||
cpu.weight
|
||||
A read-write single value file which exists on non-root
|
||||
cgroups. The default is "100".
|
||||
@@ -1909,7 +1930,7 @@ The following nested keys are defined.
|
||||
is allowed unless memory.swap.max is set to 0.
|
||||
|
||||
memory.pressure
|
||||
A read-only nested-keyed file.
|
||||
A read-write nested-keyed file.
|
||||
|
||||
Shows pressure stall information for memory. See
|
||||
:ref:`Documentation/accounting/psi.rst <psi>` for details.
|
||||
@@ -1984,9 +2005,13 @@ IO
|
||||
|
||||
The "io" controller regulates the distribution of IO resources. This
|
||||
controller implements both weight based and absolute bandwidth or IOPS
|
||||
limit distribution; however, weight based distribution is available
|
||||
only if cfq-iosched is in use and neither scheme is available for
|
||||
blk-mq devices.
|
||||
limit distribution. Absolute BPS and IOPS limits are enforced by
|
||||
blk-throttle and apply to all devices, while weight based proportional
|
||||
distribution is provided by the iocost cost model controller
|
||||
(CONFIG_BLK_CGROUP_IOCOST) and, when the BFQ I/O scheduler is in use
|
||||
for a device, by BFQ's own cgroup support. Latency-based protection
|
||||
(CONFIG_BLK_CGROUP_IOLATENCY) and I/O priority assignment
|
||||
(CONFIG_BLK_CGROUP_IOPRIO) are also available.
|
||||
|
||||
|
||||
IO Interface Files
|
||||
@@ -2169,7 +2194,7 @@ IO Interface Files
|
||||
8:16 rbps=2097152 wbps=max riops=max wiops=max
|
||||
|
||||
io.pressure
|
||||
A read-only nested-keyed file.
|
||||
A read-write nested-keyed file.
|
||||
|
||||
Shows pressure stall information for IO. See
|
||||
:ref:`Documentation/accounting/psi.rst <psi>` for details.
|
||||
@@ -2284,9 +2309,9 @@ This throttling takes 2 forms:
|
||||
throttled without possibly adversely affecting higher priority groups. This
|
||||
includes swapping and metadata IO. These types of IO are allowed to occur
|
||||
normally, however they are "charged" to the originating group. If the
|
||||
originating group is being throttled you will see the use_delay and delay
|
||||
fields in io.stat increase. The delay value is how many microseconds that are
|
||||
being added to any process that runs in this group. Because this number can
|
||||
originating group is being throttled you will see the use_delay and delay_nsec
|
||||
fields in io.stat increase. The delay_nsec value is how many nanoseconds that
|
||||
are being added to any process that runs in this group. Because this number can
|
||||
grow quite large if there is a lot of swapping or metadata IO occurring we
|
||||
limit the individual delay events to 1 second at a time.
|
||||
|
||||
@@ -2552,6 +2577,13 @@ Cpuset Interface Files
|
||||
a need to change "cpuset.mems" with active tasks, it shouldn't
|
||||
be done frequently.
|
||||
|
||||
For a multithreaded process, the threadgroup leader is
|
||||
considered the owner of the group's memory. Memory policy
|
||||
rebinding and migration will only happen with respect to the
|
||||
threadgroup leader. To avoid unexpected results, non-leading
|
||||
threads shouldn't be put into another cgroup whose "cpuset.mems"
|
||||
doesn't fully overlap that of the threadgroup leader.
|
||||
|
||||
cpuset.mems.effective
|
||||
A read-only multiple values file which exists on all
|
||||
cpuset-enabled cgroups.
|
||||
|
||||
@@ -896,7 +896,7 @@ static inline void cgroup_threadgroup_change_begin(struct task_struct *tsk)
|
||||
* cgroup_threadgroup_change_end - threadgroup exclusion for cgroups
|
||||
* @tsk: target task
|
||||
*
|
||||
* Counterpart of cgroup_threadcgroup_change_begin().
|
||||
* Counterpart of cgroup_threadgroup_change_begin().
|
||||
*/
|
||||
static inline void cgroup_threadgroup_change_end(struct task_struct *tsk)
|
||||
{
|
||||
|
||||
@@ -480,7 +480,7 @@ static inline void cgroup_unlock(void)
|
||||
rcu_read_lock_sched_held() || \
|
||||
lockdep_is_held(&cgroup_mutex) || \
|
||||
lockdep_is_held(&css_set_lock) || \
|
||||
((task)->flags & PF_EXITING) || (__c))
|
||||
(data_race((task)->flags) & PF_EXITING) || (__c))
|
||||
#else
|
||||
#define task_css_set_check(task, __c) \
|
||||
rcu_dereference((task)->cgroups)
|
||||
|
||||
@@ -1860,7 +1860,6 @@ static __always_inline bool is_user_task(struct task_struct *task)
|
||||
/* Per-process atomic flags. */
|
||||
#define PFA_NO_NEW_PRIVS 0 /* May not gain new privileges. */
|
||||
#define PFA_SPREAD_PAGE 1 /* Spread page cache over cpuset */
|
||||
#define PFA_SPREAD_SLAB 2 /* Spread some slab caches over cpuset */
|
||||
#define PFA_SPEC_SSB_DISABLE 3 /* Speculative Store Bypass disabled */
|
||||
#define PFA_SPEC_SSB_FORCE_DISABLE 4 /* Speculative Store Bypass force disabled*/
|
||||
#define PFA_SPEC_IB_DISABLE 5 /* Indirect branch speculation restricted */
|
||||
@@ -1886,10 +1885,6 @@ TASK_PFA_TEST(SPREAD_PAGE, spread_page)
|
||||
TASK_PFA_SET(SPREAD_PAGE, spread_page)
|
||||
TASK_PFA_CLEAR(SPREAD_PAGE, spread_page)
|
||||
|
||||
TASK_PFA_TEST(SPREAD_SLAB, spread_slab)
|
||||
TASK_PFA_SET(SPREAD_SLAB, spread_slab)
|
||||
TASK_PFA_CLEAR(SPREAD_SLAB, spread_slab)
|
||||
|
||||
TASK_PFA_TEST(SPEC_SSB_DISABLE, spec_ssb_disable)
|
||||
TASK_PFA_SET(SPEC_SSB_DISABLE, spec_ssb_disable)
|
||||
TASK_PFA_CLEAR(SPEC_SSB_DISABLE, spec_ssb_disable)
|
||||
|
||||
@@ -104,7 +104,7 @@ DEFINE_PERCPU_RWSEM(cgroup_threadgroup_rwsem);
|
||||
#define cgroup_assert_mutex_or_rcu_locked() \
|
||||
RCU_LOCKDEP_WARN(!rcu_read_lock_held() && \
|
||||
!lockdep_is_held(&cgroup_mutex), \
|
||||
"cgroup_mutex or RCU read lock required");
|
||||
"cgroup_mutex or RCU read lock required")
|
||||
|
||||
/*
|
||||
* cgroup destruction makes heavy use of work items and there can be a lot
|
||||
|
||||
@@ -146,10 +146,9 @@ struct cpuset {
|
||||
nodemask_t old_mems_allowed;
|
||||
|
||||
/*
|
||||
* Tasks are being attached to this cpuset. Used to prevent
|
||||
* zeroing cpus/mems_allowed between ->can_attach() and ->attach().
|
||||
* For linking impacted cpusets during an attach operation.
|
||||
*/
|
||||
int attach_in_progress;
|
||||
struct llist_node attach_node;
|
||||
|
||||
/* partition root state */
|
||||
int partition_root_state;
|
||||
@@ -165,7 +164,7 @@ struct cpuset {
|
||||
* number of SCHED_DEADLINE tasks attached to this cpuset, so that we
|
||||
* know when to rebuild associated root domain bandwidth information.
|
||||
*/
|
||||
int nr_deadline_tasks;
|
||||
atomic_t nr_deadline_tasks;
|
||||
int nr_migrate_dl_tasks;
|
||||
/* DL bandwidth that needs destination reservation for this attach. */
|
||||
u64 sum_migrate_dl_bw;
|
||||
@@ -269,10 +268,7 @@ static inline int nr_cpusets(void)
|
||||
static inline bool cpuset_is_populated(struct cpuset *cs)
|
||||
{
|
||||
lockdep_assert_cpuset_lock_held();
|
||||
|
||||
/* Cpusets in the process of attaching should be considered as populated */
|
||||
return cgroup_is_populated(cs->css.cgroup) ||
|
||||
cs->attach_in_progress;
|
||||
return cgroup_is_populated(cs->css.cgroup);
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -204,7 +204,7 @@ static s64 cpuset_read_s64(struct cgroup_subsys_state *css, struct cftype *cft)
|
||||
}
|
||||
|
||||
/*
|
||||
* update task's spread flag if cpuset's page/slab spread flag is set
|
||||
* Update a task's spread flag if the cpuset's page spread flag is set.
|
||||
*
|
||||
* Call with callback_lock or cpuset_mutex held. The check can be skipped
|
||||
* if on default hierarchy.
|
||||
@@ -219,18 +219,13 @@ void cpuset1_update_task_spread_flags(struct cpuset *cs,
|
||||
task_set_spread_page(tsk);
|
||||
else
|
||||
task_clear_spread_page(tsk);
|
||||
|
||||
if (is_spread_slab(cs))
|
||||
task_set_spread_slab(tsk);
|
||||
else
|
||||
task_clear_spread_slab(tsk);
|
||||
}
|
||||
|
||||
/**
|
||||
* cpuset1_update_tasks_flags - update the spread flags of tasks in the cpuset.
|
||||
* @cs: the cpuset in which each task's spread flags needs to be changed
|
||||
* cpuset1_update_tasks_flags - update the page spread flag of cpuset tasks
|
||||
* @cs: the cpuset whose tasks need their page spread flag updated
|
||||
*
|
||||
* Iterate through each task of @cs updating its spread flags. As this
|
||||
* Iterate through each task of @cs updating its page spread flag. As this
|
||||
* function is called with cpuset_mutex held, cpuset membership stays
|
||||
* stable.
|
||||
*/
|
||||
|
||||
@@ -37,6 +37,7 @@
|
||||
#include <linux/wait.h>
|
||||
#include <linux/workqueue.h>
|
||||
#include <linux/task_work.h>
|
||||
#include <linux/llist.h>
|
||||
|
||||
DEFINE_STATIC_KEY_FALSE(cpusets_pre_enable_key);
|
||||
DEFINE_STATIC_KEY_FALSE(cpusets_enabled_key);
|
||||
@@ -222,14 +223,14 @@ void inc_dl_tasks_cs(struct task_struct *p)
|
||||
{
|
||||
struct cpuset *cs = task_cs(p);
|
||||
|
||||
cs->nr_deadline_tasks++;
|
||||
atomic_inc(&cs->nr_deadline_tasks);
|
||||
}
|
||||
|
||||
void dec_dl_tasks_cs(struct task_struct *p)
|
||||
{
|
||||
struct cpuset *cs = task_cs(p);
|
||||
|
||||
cs->nr_deadline_tasks--;
|
||||
atomic_dec(&cs->nr_deadline_tasks);
|
||||
}
|
||||
|
||||
static inline bool is_partition_valid(const struct cpuset *cs)
|
||||
@@ -356,6 +357,41 @@ static struct workqueue_struct *cpuset_migrate_mm_wq;
|
||||
|
||||
static DECLARE_WAIT_QUEUE_HEAD(cpuset_attach_wq);
|
||||
|
||||
/*
|
||||
* Cpuset task attach context
|
||||
* Protected by cpuset_mutex
|
||||
*/
|
||||
static struct {
|
||||
int in_progress;
|
||||
bool cpus_updated;
|
||||
bool mems_updated;
|
||||
bool task_work_queued;
|
||||
bool many_dest_cs; /* Have many destination cpusets */
|
||||
struct cpuset *old_cs; /* Source cpuset */
|
||||
nodemask_t nodemask_to;
|
||||
} attach_ctx;
|
||||
static LLIST_HEAD(src_cs_head);
|
||||
static LLIST_HEAD(dst_cs_head);
|
||||
|
||||
/*
|
||||
* Wait if task attach is in progress until it is done and then acquire
|
||||
* cpuset_mutex before returning.
|
||||
*/
|
||||
static void wait_attach_done_lock(void)
|
||||
__acquires(&cpuset_mutex)
|
||||
{
|
||||
for (;;) {
|
||||
mutex_lock(&cpuset_mutex);
|
||||
if (!attach_ctx.in_progress)
|
||||
return;
|
||||
|
||||
mutex_unlock(&cpuset_mutex);
|
||||
|
||||
/* Wait until attach operation is done to prevent racing */
|
||||
wait_event(cpuset_attach_wq, attach_ctx.in_progress == 0);
|
||||
}
|
||||
}
|
||||
|
||||
static inline void check_insane_mems_config(nodemask_t *nodes)
|
||||
{
|
||||
if (!cpusets_insane_config() &&
|
||||
@@ -368,22 +404,22 @@ static inline void check_insane_mems_config(nodemask_t *nodes)
|
||||
}
|
||||
|
||||
/*
|
||||
* decrease cs->attach_in_progress.
|
||||
* wake_up cpuset_attach_wq if cs->attach_in_progress==0.
|
||||
* decrease attach_ctx.in_progress.
|
||||
* wake_up cpuset_attach_wq if attach_ctx.in_progress==0.
|
||||
*/
|
||||
static inline void dec_attach_in_progress_locked(struct cpuset *cs)
|
||||
static inline void dec_attach_in_progress_locked(void)
|
||||
{
|
||||
lockdep_assert_cpuset_lock_held();
|
||||
|
||||
cs->attach_in_progress--;
|
||||
if (!cs->attach_in_progress)
|
||||
attach_ctx.in_progress--;
|
||||
if (!attach_ctx.in_progress)
|
||||
wake_up(&cpuset_attach_wq);
|
||||
}
|
||||
|
||||
static inline void dec_attach_in_progress(struct cpuset *cs)
|
||||
static inline void dec_attach_in_progress(void)
|
||||
{
|
||||
mutex_lock(&cpuset_mutex);
|
||||
dec_attach_in_progress_locked(cs);
|
||||
dec_attach_in_progress_locked();
|
||||
mutex_unlock(&cpuset_mutex);
|
||||
}
|
||||
|
||||
@@ -432,8 +468,7 @@ static inline bool partition_is_populated(struct cpuset *cs,
|
||||
* nr_populated_domain_children may include populated
|
||||
* csets from descendants that are partitions.
|
||||
*/
|
||||
if (cgroup_has_tasks(cs->css.cgroup) ||
|
||||
cs->attach_in_progress)
|
||||
if (cgroup_has_tasks(cs->css.cgroup))
|
||||
return true;
|
||||
|
||||
rcu_read_lock();
|
||||
@@ -489,7 +524,10 @@ static void guarantee_active_cpus(struct task_struct *tsk,
|
||||
* Return in *pmask the portion of a cpusets's mems_allowed that
|
||||
* are online, with memory. If none are online with memory, walk
|
||||
* up the cpuset hierarchy until we find one that does have some
|
||||
* online mems. The top cpuset always has some mems online.
|
||||
* online mems. The top cpuset always has some mems online. With v2,
|
||||
* effective_mems should always contain online memory nodes except
|
||||
* during the transition period where a memory node hotunplug operation
|
||||
* is in progress.
|
||||
*
|
||||
* One way or another, we guarantee to return some non-empty subset
|
||||
* of node_states[N_MEMORY].
|
||||
@@ -581,6 +619,7 @@ static struct cpuset *dup_or_alloc_cpuset(struct cpuset *cs)
|
||||
return NULL;
|
||||
|
||||
trial->dl_bw_cpu = -1;
|
||||
init_llist_node(&trial->attach_node);
|
||||
|
||||
/* Setup cpumask pointer array */
|
||||
cpumask_var_t *pmask[4] = {
|
||||
@@ -918,7 +957,7 @@ static void dl_update_tasks_root_domain(struct cpuset *cs)
|
||||
struct css_task_iter it;
|
||||
struct task_struct *task;
|
||||
|
||||
if (cs->nr_deadline_tasks == 0)
|
||||
if (atomic_read(&cs->nr_deadline_tasks) == 0)
|
||||
return;
|
||||
|
||||
css_task_iter_start(&cs->css, 0, &it);
|
||||
@@ -1089,12 +1128,35 @@ void cpuset_update_tasks_cpumask(struct cpuset *cs, struct cpumask *new_cpus)
|
||||
* @cs: the cpuset the need to recompute the new effective_cpus mask
|
||||
* @parent: the parent cpuset
|
||||
*
|
||||
* For v2, the parent's effective_cpus is inherited if cpumask is empty.
|
||||
* The result is valid only if the given cpuset isn't a partition root.
|
||||
*/
|
||||
static void compute_effective_cpumask(struct cpumask *new_cpus,
|
||||
struct cpuset *cs, struct cpuset *parent)
|
||||
{
|
||||
cpumask_and(new_cpus, cs->cpus_allowed, parent->effective_cpus);
|
||||
bool has_cpus;
|
||||
|
||||
has_cpus = cpumask_and(new_cpus, cs->cpus_allowed, parent->effective_cpus);
|
||||
if (!has_cpus && is_in_v2_mode())
|
||||
cpumask_copy(new_cpus, parent->effective_cpus);
|
||||
}
|
||||
|
||||
/**
|
||||
* compute_effective_nodemask - Compute the effective nodemask of the cpuset
|
||||
* @new_mems: the temp variable for the new effective_mems mask
|
||||
* @cs: the cpuset the need to recompute the new effective_mems mask
|
||||
* @parent: the parent cpuset
|
||||
*
|
||||
* For v2, the parent's effective_mems is inherited if nodemask is empty.
|
||||
*/
|
||||
static void compute_effective_nodemask(nodemask_t *new_mems,
|
||||
struct cpuset *cs, struct cpuset *parent)
|
||||
{
|
||||
bool has_mems;
|
||||
|
||||
has_mems = nodes_and(*new_mems, cs->mems_allowed, parent->effective_mems);
|
||||
if (!has_mems && is_in_v2_mode())
|
||||
nodes_copy(*new_mems, parent->effective_mems);
|
||||
}
|
||||
|
||||
/*
|
||||
@@ -1525,7 +1587,7 @@ static int remote_partition_enable(struct cpuset *cs, int new_prs,
|
||||
cpumask_copy(cs->effective_xcpus, tmp->new_cpus);
|
||||
spin_unlock_irq(&callback_lock);
|
||||
cpuset_force_rebuild();
|
||||
cs->prs_err = 0;
|
||||
WRITE_ONCE(cs->prs_err, 0);
|
||||
|
||||
/*
|
||||
* Propagate changes in top_cpuset's effective_cpus down the hierarchy.
|
||||
@@ -1599,7 +1661,7 @@ static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
|
||||
WARN_ON_ONCE(!cpumask_subset(cs->effective_xcpus, subpartitions_cpus));
|
||||
|
||||
if (cpumask_empty(excpus)) {
|
||||
cs->prs_err = PERR_CPUSEMPTY;
|
||||
WRITE_ONCE(cs->prs_err, PERR_CPUSEMPTY);
|
||||
goto invalidate;
|
||||
}
|
||||
|
||||
@@ -1614,13 +1676,13 @@ static void remote_cpus_update(struct cpuset *cs, struct cpumask *xcpus,
|
||||
if (adding) {
|
||||
WARN_ON_ONCE(cpumask_intersects(tmp->addmask, subpartitions_cpus));
|
||||
if (!capable(CAP_SYS_ADMIN))
|
||||
cs->prs_err = PERR_ACCESS;
|
||||
WRITE_ONCE(cs->prs_err, PERR_ACCESS);
|
||||
else if (cpumask_intersects(tmp->addmask, subpartitions_cpus) ||
|
||||
cpumask_subset(top_cpuset.effective_cpus, tmp->addmask))
|
||||
cs->prs_err = PERR_NOCPUS;
|
||||
WRITE_ONCE(cs->prs_err, PERR_NOCPUS);
|
||||
else if ((prs == PRS_ISOLATED) &&
|
||||
!isolated_cpus_can_update(tmp->addmask, tmp->delmask))
|
||||
cs->prs_err = PERR_HKEEPING;
|
||||
WRITE_ONCE(cs->prs_err, PERR_HKEEPING);
|
||||
if (cs->prs_err)
|
||||
goto invalidate;
|
||||
}
|
||||
@@ -2048,13 +2110,13 @@ static void compute_partition_effective_cpumask(struct cpuset *cs,
|
||||
* partition root.
|
||||
*/
|
||||
WARN_ON_ONCE(is_remote_partition(child));
|
||||
child->prs_err = 0;
|
||||
WRITE_ONCE(child->prs_err, 0);
|
||||
if (!cpumask_subset(child->effective_xcpus,
|
||||
cs->effective_xcpus))
|
||||
child->prs_err = PERR_INVCPUS;
|
||||
WRITE_ONCE(child->prs_err, PERR_INVCPUS);
|
||||
else if (populated &&
|
||||
cpumask_subset(new_ecpus, child->effective_xcpus))
|
||||
child->prs_err = PERR_NOCPUS;
|
||||
WRITE_ONCE(child->prs_err, PERR_NOCPUS);
|
||||
|
||||
if (child->prs_err) {
|
||||
int old_prs = child->partition_root_state;
|
||||
@@ -2143,15 +2205,6 @@ static void update_cpumasks_hier(struct cpuset *cs, struct tmpmasks *tmp,
|
||||
goto update_parent_effective;
|
||||
}
|
||||
|
||||
/*
|
||||
* If it becomes empty, inherit the effective mask of the
|
||||
* parent, which is guaranteed to have some CPUs unless
|
||||
* it is a partition root that has explicitly distributed
|
||||
* out all its CPUs.
|
||||
*/
|
||||
if (is_in_v2_mode() && !remote && cpumask_empty(tmp->new_cpus))
|
||||
cpumask_copy(tmp->new_cpus, parent->effective_cpus);
|
||||
|
||||
/*
|
||||
* Skip the whole subtree if
|
||||
* 1) the cpumask remains the same,
|
||||
@@ -2367,8 +2420,10 @@ static void partition_cpus_change(struct cpuset *cs, struct cpuset *trialcs,
|
||||
return;
|
||||
|
||||
prs_err = validate_partition(cs, trialcs);
|
||||
if (prs_err)
|
||||
trialcs->prs_err = cs->prs_err = prs_err;
|
||||
if (prs_err) {
|
||||
WRITE_ONCE(cs->prs_err, prs_err);
|
||||
trialcs->prs_err = prs_err;
|
||||
}
|
||||
|
||||
if (is_remote_partition(cs)) {
|
||||
if (trialcs->prs_err)
|
||||
@@ -2619,6 +2674,14 @@ static void *cpuset_being_rebound;
|
||||
* Iterate through each task of @cs updating its mems_allowed to the
|
||||
* effective cpuset's. As this function is called with cpuset_mutex held,
|
||||
* cpuset membership stays stable.
|
||||
*
|
||||
* - cpuset_change_task_nodemask(): guarantee_online_mems()
|
||||
* - mpol_rebind_mm(): effective_mems
|
||||
* - cpuset_migrate_mm(): guarantee_online_mems()
|
||||
* - old_mems_allowed: guarantee_online_mems()
|
||||
*
|
||||
* For v2, guarantee_online_mems() should return a node mask that is the same
|
||||
* as the effective_mems of current cpuset.
|
||||
*/
|
||||
void cpuset_update_tasks_nodemask(struct cpuset *cs)
|
||||
{
|
||||
@@ -2627,7 +2690,6 @@ void cpuset_update_tasks_nodemask(struct cpuset *cs)
|
||||
struct task_struct *task;
|
||||
|
||||
cpuset_being_rebound = cs; /* causes mpol_dup() rebind */
|
||||
|
||||
guarantee_online_mems(cs, &newmems);
|
||||
|
||||
/*
|
||||
@@ -2647,6 +2709,10 @@ void cpuset_update_tasks_nodemask(struct cpuset *cs)
|
||||
|
||||
cpuset_change_task_nodemask(task, &newmems);
|
||||
|
||||
/* Rebind and migrate mm only for thread group leader */
|
||||
if (!thread_group_leader(task))
|
||||
continue;
|
||||
|
||||
mm = get_task_mm(task);
|
||||
if (!mm)
|
||||
continue;
|
||||
@@ -2697,14 +2763,7 @@ static void update_nodemasks_hier(struct cpuset *cs, nodemask_t *new_mems)
|
||||
cpuset_for_each_descendant_pre(cp, pos_css, cs) {
|
||||
struct cpuset *parent = parent_cs(cp);
|
||||
|
||||
bool has_mems = nodes_and(*new_mems, cp->mems_allowed, parent->effective_mems);
|
||||
|
||||
/*
|
||||
* If it becomes empty, inherit the effective mask of the
|
||||
* parent, which is guaranteed to have some MEMs.
|
||||
*/
|
||||
if (is_in_v2_mode() && !has_mems)
|
||||
*new_mems = parent->effective_mems;
|
||||
compute_effective_nodemask(new_mems, cp, parent);
|
||||
|
||||
/* Skip the whole subtree if the nodemask remains the same. */
|
||||
if (nodes_equal(*new_mems, cp->effective_mems)) {
|
||||
@@ -2805,7 +2864,7 @@ int cpuset_update_flag(cpuset_flagbits_t bit, struct cpuset *cs,
|
||||
{
|
||||
struct cpuset *trialcs;
|
||||
int balance_flag_changed;
|
||||
int spread_flag_changed;
|
||||
int spread_page_changed;
|
||||
int err;
|
||||
|
||||
trialcs = dup_or_alloc_cpuset(cs);
|
||||
@@ -2824,8 +2883,7 @@ int cpuset_update_flag(cpuset_flagbits_t bit, struct cpuset *cs,
|
||||
balance_flag_changed = (is_sched_load_balance(cs) !=
|
||||
is_sched_load_balance(trialcs));
|
||||
|
||||
spread_flag_changed = ((is_spread_slab(cs) != is_spread_slab(trialcs))
|
||||
|| (is_spread_page(cs) != is_spread_page(trialcs)));
|
||||
spread_page_changed = is_spread_page(cs) != is_spread_page(trialcs);
|
||||
|
||||
spin_lock_irq(&callback_lock);
|
||||
cs->flags = trialcs->flags;
|
||||
@@ -2838,7 +2896,7 @@ int cpuset_update_flag(cpuset_flagbits_t bit, struct cpuset *cs,
|
||||
rebuild_sched_domains_locked();
|
||||
}
|
||||
|
||||
if (spread_flag_changed)
|
||||
if (spread_page_changed)
|
||||
cpuset1_update_tasks_flags(cs);
|
||||
out:
|
||||
free_cpuset(trialcs);
|
||||
@@ -2973,58 +3031,47 @@ static int update_prstate(struct cpuset *cs, int new_prs)
|
||||
return 0;
|
||||
}
|
||||
|
||||
static struct cpuset *cpuset_attach_old_cs;
|
||||
|
||||
/*
|
||||
* Check to see if a cpuset can accept a new task
|
||||
* For v1, cpus_allowed and mems_allowed can't be empty.
|
||||
* For v2, effective_cpus can't be empty.
|
||||
* Note that in v1, effective_cpus = cpus_allowed.
|
||||
*
|
||||
* Also set the boolean flag passed in by @psetsched depending on if
|
||||
* security_task_setscheduler() call is needed and @oldcs is not NULL.
|
||||
*/
|
||||
static int cpuset_can_attach_check(struct cpuset *cs)
|
||||
static int cpuset_can_attach_check(struct cpuset *cs, struct cpuset *oldcs,
|
||||
bool *psetsched)
|
||||
{
|
||||
bool cpus_updated, mems_updated;
|
||||
|
||||
if (cpumask_empty(cs->effective_cpus) ||
|
||||
(!is_in_v2_mode() && nodes_empty(cs->mems_allowed)))
|
||||
return -ENOSPC;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static void reset_migrate_dl_data(struct cpuset *cs)
|
||||
{
|
||||
cs->nr_migrate_dl_tasks = 0;
|
||||
cs->sum_migrate_dl_bw = 0;
|
||||
cs->dl_bw_cpu = -1;
|
||||
}
|
||||
if (!oldcs)
|
||||
return 0;
|
||||
|
||||
/* Called by cgroups to determine if a cpuset is usable; cpuset_mutex held */
|
||||
static int cpuset_can_attach(struct cgroup_taskset *tset)
|
||||
{
|
||||
struct cgroup_subsys_state *css;
|
||||
struct cpuset *cs, *oldcs;
|
||||
struct task_struct *task;
|
||||
bool setsched_check;
|
||||
int cpu, ret;
|
||||
if (!llist_on_list(&oldcs->attach_node))
|
||||
llist_add(&oldcs->attach_node, &src_cs_head);
|
||||
|
||||
/* used later by cpuset_attach() */
|
||||
cpuset_attach_old_cs = task_cs(cgroup_taskset_first(tset, &css));
|
||||
oldcs = cpuset_attach_old_cs;
|
||||
cs = css_cs(css);
|
||||
if (!llist_on_list(&cs->attach_node))
|
||||
llist_add(&cs->attach_node, &dst_cs_head);
|
||||
|
||||
mutex_lock(&cpuset_mutex);
|
||||
cpus_updated = !cpumask_equal(cs->effective_cpus, oldcs->effective_cpus);
|
||||
mems_updated = !nodes_equal(cs->effective_mems, oldcs->effective_mems);
|
||||
|
||||
/* Check to see if task is allowed in the cpuset */
|
||||
ret = cpuset_can_attach_check(cs);
|
||||
if (ret)
|
||||
goto out_unlock;
|
||||
if (cpus_updated)
|
||||
attach_ctx.cpus_updated = true;
|
||||
if (mems_updated)
|
||||
attach_ctx.mems_updated = true;
|
||||
|
||||
/*
|
||||
* Skip rights over task setsched check in v2 when nothing changes,
|
||||
* migration permission derives from hierarchy ownership in
|
||||
* cgroup_procs_write_permission()).
|
||||
* Skip rights over task setsched check in v2 when nothing changes for
|
||||
* the current oldcs/cs pair, migration permission derives from
|
||||
* hierarchy ownership in cgroup_procs_write_permission()).
|
||||
*/
|
||||
setsched_check = !cpuset_v2() ||
|
||||
!cpumask_equal(cs->effective_cpus, oldcs->effective_cpus) ||
|
||||
!nodes_equal(cs->effective_mems, oldcs->effective_mems);
|
||||
*psetsched = !cpuset_v2() || cpus_updated || mems_updated;
|
||||
|
||||
/*
|
||||
* A v1 cpuset with tasks will have no CPU left only when CPU hotplug
|
||||
@@ -3034,13 +3081,123 @@ static int cpuset_can_attach(struct cgroup_taskset *tset)
|
||||
* sure they will be able to run after migration.
|
||||
*/
|
||||
if (!is_in_v2_mode() && cpumask_empty(oldcs->effective_cpus))
|
||||
setsched_check = false;
|
||||
*psetsched = false;
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int cpuset_reserve_dl_bw(void)
|
||||
{
|
||||
struct cpuset *cs;
|
||||
int cpu, ret;
|
||||
|
||||
llist_for_each_entry(cs, dst_cs_head.first, attach_node) {
|
||||
if (!cs->sum_migrate_dl_bw)
|
||||
continue;
|
||||
|
||||
cpu = cpumask_any_and(cpu_active_mask, cs->effective_cpus);
|
||||
if (unlikely(cpu >= nr_cpu_ids))
|
||||
return -EINVAL;
|
||||
|
||||
ret = dl_bw_alloc(cpu, cs->sum_migrate_dl_bw);
|
||||
if (ret)
|
||||
return ret;
|
||||
|
||||
cs->dl_bw_cpu = cpu;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
/*
|
||||
* Clear and optionally apply (@cancel is false) the attach related data in the
|
||||
* source or destination cpuset.
|
||||
*/
|
||||
static void clear_attach_data(struct llist_head *head, bool cancel)
|
||||
{
|
||||
struct cpuset *cs, *next;
|
||||
struct llist_node *lnode = __llist_del_all(head);
|
||||
|
||||
llist_for_each_entry_safe(cs, next, lnode, attach_node) {
|
||||
init_llist_node(&cs->attach_node);
|
||||
if (cs->nr_migrate_dl_tasks) {
|
||||
if (!cancel)
|
||||
atomic_add(cs->nr_migrate_dl_tasks, &cs->nr_deadline_tasks);
|
||||
else if (cs->dl_bw_cpu >= 0) /* && cancel */
|
||||
dl_bw_free(cs->dl_bw_cpu, cs->sum_migrate_dl_bw);
|
||||
cs->nr_migrate_dl_tasks = 0;
|
||||
cs->sum_migrate_dl_bw = 0;
|
||||
cs->dl_bw_cpu = -1;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/* Called by cgroups to determine if a cpuset is usable; cpuset_mutex held */
|
||||
static int cpuset_can_attach(struct cgroup_taskset *tset)
|
||||
{
|
||||
struct cgroup_subsys_state *css;
|
||||
struct cpuset *cs, *oldcs;
|
||||
struct task_struct *task;
|
||||
bool setsched_check;
|
||||
int ret;
|
||||
|
||||
cs = oldcs = NULL;
|
||||
mutex_lock(&cpuset_mutex);
|
||||
attach_ctx.old_cs = NULL; /* Used later in cpuset_attach_task() */
|
||||
attach_ctx.cpus_updated = false;
|
||||
attach_ctx.mems_updated = false;
|
||||
attach_ctx.many_dest_cs = false;
|
||||
|
||||
/*
|
||||
* The attach_ctx.old_cs is used mainly by cpuset_migrate_mm() to get
|
||||
* the old_mems_allowed value. There are two ways that many-to-one
|
||||
* cpuset migration can happen:
|
||||
* 1) A multithread application with threads in different cpusets is
|
||||
* wholely migrated to a new cpuset.
|
||||
* 2) Disabling v2 cpuset controller will move all the tasks in child
|
||||
* cpusets to the parent cpuset.
|
||||
*
|
||||
* In the former case, it is the mm setting of the group leader that
|
||||
* really matters. So attach_ctx.old_cs should track the oldcs of the
|
||||
* group leader. It falls back to the oldcs of the first task if there
|
||||
* is no group leader in the taskset. In the latter case, effective_mems
|
||||
* of child cpusets must always be a subset of the parent. So no real
|
||||
* page migration will be necessary no matter which child cpuset is
|
||||
* selected as attach_ctx.old_cs.
|
||||
*
|
||||
* For a v2 threaded subtree where cpuset isn't enabled in some of the
|
||||
* cgroups, it is possible that oldcs == cs for some of the tasks.
|
||||
* In this case, we can skip checking on those tasks as there is no
|
||||
* actual migration wrt cpuset.
|
||||
*/
|
||||
cgroup_taskset_for_each(task, css, tset) {
|
||||
struct cpuset *new_cs = css_cs(css);
|
||||
struct cpuset *new_oldcs = task_cs(task);
|
||||
|
||||
if ((new_oldcs != oldcs) || (new_cs != cs)) {
|
||||
if (cs && (new_cs != cs))
|
||||
attach_ctx.many_dest_cs = true;
|
||||
cs = new_cs;
|
||||
oldcs = new_oldcs;
|
||||
if (oldcs == cs)
|
||||
continue;
|
||||
if (!attach_ctx.old_cs)
|
||||
attach_ctx.old_cs = oldcs;
|
||||
ret = cpuset_can_attach_check(cs, oldcs, &setsched_check);
|
||||
if (ret)
|
||||
goto out_unlock;
|
||||
}
|
||||
|
||||
if (oldcs == cs)
|
||||
continue;
|
||||
|
||||
ret = task_can_attach(task);
|
||||
if (ret)
|
||||
goto out_unlock;
|
||||
|
||||
/* Update attach_ctx.old_cs to the latest group leader */
|
||||
if (task == task->group_leader)
|
||||
attach_ctx.old_cs = task_cs(task);
|
||||
|
||||
if (setsched_check) {
|
||||
ret = security_task_setscheduler(task);
|
||||
if (ret)
|
||||
@@ -3054,57 +3211,48 @@ static int cpuset_can_attach(struct cgroup_taskset *tset)
|
||||
* contribute to sum_migrate_dl_bw.
|
||||
*/
|
||||
cs->nr_migrate_dl_tasks++;
|
||||
oldcs->nr_migrate_dl_tasks--;
|
||||
if (dl_task_needs_bw_move(task, cs->effective_cpus))
|
||||
cs->sum_migrate_dl_bw += task->dl.dl_bw;
|
||||
}
|
||||
}
|
||||
|
||||
if (!cs->sum_migrate_dl_bw)
|
||||
goto out_success;
|
||||
|
||||
cpu = cpumask_any_and(cpu_active_mask, cs->effective_cpus);
|
||||
if (unlikely(cpu >= nr_cpu_ids)) {
|
||||
/*
|
||||
* The only case where there are multiple destination cpusets for
|
||||
* task migration is when enabling a v2 cpuset controllers where
|
||||
* tasks will be migrated to multiple child cpusets from a parent
|
||||
* cpuset with the same effective CPUs and memory nodes. IOW,
|
||||
* both attach_cpus_updated and attach_mems_updated should be false.
|
||||
* If not, it is a condition that the current code cannot handle.
|
||||
* Print a warning and abort the attach operation as further code
|
||||
* change may be needed.
|
||||
*/
|
||||
if (WARN_ON_ONCE(attach_ctx.many_dest_cs && (!cpuset_v2() ||
|
||||
attach_ctx.cpus_updated || attach_ctx.mems_updated))) {
|
||||
ret = -EINVAL;
|
||||
goto out_unlock;
|
||||
}
|
||||
|
||||
ret = dl_bw_alloc(cpu, cs->sum_migrate_dl_bw);
|
||||
if (ret)
|
||||
goto out_unlock;
|
||||
|
||||
cs->dl_bw_cpu = cpu;
|
||||
|
||||
out_success:
|
||||
/*
|
||||
* Mark attach is in progress. This makes validate_change() fail
|
||||
* changes which zero cpus/mems_allowed.
|
||||
*/
|
||||
cs->attach_in_progress++;
|
||||
ret = cpuset_reserve_dl_bw();
|
||||
|
||||
out_unlock:
|
||||
if (ret)
|
||||
reset_migrate_dl_data(cs);
|
||||
if (ret) {
|
||||
clear_attach_data(&src_cs_head, true);
|
||||
clear_attach_data(&dst_cs_head, true);
|
||||
} else {
|
||||
attach_ctx.in_progress++;
|
||||
}
|
||||
|
||||
mutex_unlock(&cpuset_mutex);
|
||||
return ret;
|
||||
}
|
||||
|
||||
static void cpuset_cancel_attach(struct cgroup_taskset *tset)
|
||||
{
|
||||
struct cgroup_subsys_state *css;
|
||||
struct cpuset *cs;
|
||||
|
||||
cgroup_taskset_first(tset, &css);
|
||||
cs = css_cs(css);
|
||||
|
||||
mutex_lock(&cpuset_mutex);
|
||||
dec_attach_in_progress_locked(cs);
|
||||
|
||||
if (cs->dl_bw_cpu >= 0)
|
||||
dl_bw_free(cs->dl_bw_cpu, cs->sum_migrate_dl_bw);
|
||||
|
||||
if (cs->nr_migrate_dl_tasks)
|
||||
reset_migrate_dl_data(cs);
|
||||
|
||||
dec_attach_in_progress_locked();
|
||||
clear_attach_data(&src_cs_head, true);
|
||||
clear_attach_data(&dst_cs_head, true);
|
||||
mutex_unlock(&cpuset_mutex);
|
||||
}
|
||||
|
||||
@@ -3114,10 +3262,11 @@ static void cpuset_cancel_attach(struct cgroup_taskset *tset)
|
||||
* allocate from cpuset_init().
|
||||
*/
|
||||
static cpumask_var_t cpus_attach;
|
||||
static nodemask_t cpuset_attach_nodemask_to;
|
||||
|
||||
static void cpuset_attach_task(struct cpuset *cs, struct task_struct *task)
|
||||
{
|
||||
struct mm_struct *mm;
|
||||
|
||||
lockdep_assert_cpuset_lock_held();
|
||||
|
||||
if (cs != &top_cpuset)
|
||||
@@ -3131,90 +3280,88 @@ static void cpuset_attach_task(struct cpuset *cs, struct task_struct *task)
|
||||
*/
|
||||
WARN_ON_ONCE(set_cpus_allowed_ptr(task, cpus_attach));
|
||||
|
||||
cpuset_change_task_nodemask(task, &cpuset_attach_nodemask_to);
|
||||
if (cpuset_v2() && !attach_ctx.mems_updated)
|
||||
return;
|
||||
|
||||
cpuset_change_task_nodemask(task, &attach_ctx.nodemask_to);
|
||||
cpuset1_update_task_spread_flags(cs, task);
|
||||
|
||||
if ((task != task->group_leader) || !attach_ctx.mems_updated)
|
||||
return;
|
||||
|
||||
/*
|
||||
* Change mm for threadgroup leader. This is expensive and may
|
||||
* sleep and should be moved outside migration path proper.
|
||||
*/
|
||||
mm = get_task_mm(task);
|
||||
if (mm) {
|
||||
struct cpuset *oldcs = attach_ctx.old_cs;
|
||||
|
||||
mpol_rebind_mm(mm, &cs->effective_mems);
|
||||
|
||||
/*
|
||||
* old_mems_allowed is the same with mems_allowed
|
||||
* here, except if this task is being moved
|
||||
* automatically due to hotplug. In that case
|
||||
* @mems_allowed has been updated and is empty, so
|
||||
* @old_mems_allowed is the right nodesets that we
|
||||
* migrate mm from.
|
||||
*/
|
||||
if (is_memory_migrate(cs)) {
|
||||
cpuset_migrate_mm(mm, &oldcs->old_mems_allowed,
|
||||
&attach_ctx.nodemask_to);
|
||||
attach_ctx.task_work_queued = true;
|
||||
} else {
|
||||
mmput(mm);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
static void cpuset_attach(struct cgroup_taskset *tset)
|
||||
{
|
||||
struct task_struct *task;
|
||||
struct task_struct *leader;
|
||||
struct cgroup_subsys_state *css;
|
||||
struct cpuset *cs;
|
||||
struct cpuset *oldcs = cpuset_attach_old_cs;
|
||||
bool cpus_updated, mems_updated;
|
||||
bool queue_task_work = false;
|
||||
|
||||
cgroup_taskset_first(tset, &css);
|
||||
cs = css_cs(css);
|
||||
|
||||
lockdep_assert_cpus_held(); /* see cgroup_attach_lock() */
|
||||
mutex_lock(&cpuset_mutex);
|
||||
cpus_updated = !cpumask_equal(cs->effective_cpus,
|
||||
oldcs->effective_cpus);
|
||||
mems_updated = !nodes_equal(cs->effective_mems, oldcs->effective_mems);
|
||||
attach_ctx.task_work_queued = false;
|
||||
guarantee_online_mems(cs, &attach_ctx.nodemask_to);
|
||||
|
||||
/*
|
||||
* attach_ctx.old_cs can only be NULL if no task is actually migrating.
|
||||
* This is highly unlikely. If it happens at all, we can skip task
|
||||
* iteration and setting old_mems_allowed.
|
||||
*/
|
||||
if (unlikely(!attach_ctx.old_cs))
|
||||
goto out;
|
||||
|
||||
/*
|
||||
* In the default hierarchy, enabling cpuset in the child cgroups
|
||||
* will trigger a number of cpuset_attach() calls with no change
|
||||
* in effective cpus and mems. In that case, we can optimize out
|
||||
* by skipping the task iteration and update.
|
||||
* will trigger a cpuset_attach() call with no change in effective cpus
|
||||
* and mems. In that case, we can optimize out by skipping the task
|
||||
* iteration and the destination cpuset list is iterated to set
|
||||
* old_mems_allowed.
|
||||
*/
|
||||
if (cpuset_v2() && !cpus_updated && !mems_updated) {
|
||||
cpuset_attach_nodemask_to = cs->effective_mems;
|
||||
if (cpuset_v2() && !attach_ctx.cpus_updated && !attach_ctx.mems_updated) {
|
||||
llist_for_each_entry(cs, dst_cs_head.first, attach_node)
|
||||
cs->old_mems_allowed = attach_ctx.nodemask_to;
|
||||
goto out;
|
||||
}
|
||||
|
||||
guarantee_online_mems(cs, &cpuset_attach_nodemask_to);
|
||||
|
||||
cgroup_taskset_for_each(task, css, tset)
|
||||
cpuset_attach_task(cs, task);
|
||||
|
||||
/*
|
||||
* Change mm for all threadgroup leaders. This is expensive and may
|
||||
* sleep and should be moved outside migration path proper. Skip it
|
||||
* if there is no change in effective_mems and CS_MEMORY_MIGRATE is
|
||||
* not set.
|
||||
*/
|
||||
cpuset_attach_nodemask_to = cs->effective_mems;
|
||||
if (!is_memory_migrate(cs) && !mems_updated)
|
||||
goto out;
|
||||
|
||||
cgroup_taskset_for_each_leader(leader, css, tset) {
|
||||
struct mm_struct *mm = get_task_mm(leader);
|
||||
|
||||
if (mm) {
|
||||
mpol_rebind_mm(mm, &cpuset_attach_nodemask_to);
|
||||
|
||||
/*
|
||||
* old_mems_allowed is the same with mems_allowed
|
||||
* here, except if this task is being moved
|
||||
* automatically due to hotplug. In that case
|
||||
* @mems_allowed has been updated and is empty, so
|
||||
* @old_mems_allowed is the right nodesets that we
|
||||
* migrate mm from.
|
||||
*/
|
||||
if (is_memory_migrate(cs)) {
|
||||
cpuset_migrate_mm(mm, &oldcs->old_mems_allowed,
|
||||
&cpuset_attach_nodemask_to);
|
||||
queue_task_work = true;
|
||||
} else
|
||||
mmput(mm);
|
||||
}
|
||||
}
|
||||
|
||||
out:
|
||||
if (queue_task_work)
|
||||
if (attach_ctx.task_work_queued)
|
||||
schedule_flush_migrate_mm();
|
||||
cs->old_mems_allowed = cpuset_attach_nodemask_to;
|
||||
|
||||
if (cs->nr_migrate_dl_tasks) {
|
||||
cs->nr_deadline_tasks += cs->nr_migrate_dl_tasks;
|
||||
oldcs->nr_deadline_tasks -= cs->nr_migrate_dl_tasks;
|
||||
reset_migrate_dl_data(cs);
|
||||
}
|
||||
|
||||
dec_attach_in_progress_locked(cs);
|
||||
cs->old_mems_allowed = attach_ctx.nodemask_to;
|
||||
out:
|
||||
clear_attach_data(&src_cs_head, false);
|
||||
clear_attach_data(&dst_cs_head, false);
|
||||
dec_attach_in_progress_locked();
|
||||
|
||||
mutex_unlock(&cpuset_mutex);
|
||||
}
|
||||
@@ -3234,7 +3381,12 @@ ssize_t cpuset_write_resmask(struct kernfs_open_file *of,
|
||||
return -EACCES;
|
||||
|
||||
buf = strstrip(buf);
|
||||
cpuset_full_lock();
|
||||
|
||||
/* cpuset_mutex acquired in wait_attach_done_lock() */
|
||||
mutex_lock(&cpuset_top_mutex);
|
||||
cpus_read_lock();
|
||||
wait_attach_done_lock();
|
||||
|
||||
if (!is_cpuset_online(cs))
|
||||
goto out_unlock;
|
||||
|
||||
@@ -3365,7 +3517,10 @@ static ssize_t cpuset_partition_write(struct kernfs_open_file *of, char *buf,
|
||||
else
|
||||
return -EINVAL;
|
||||
|
||||
cpuset_full_lock();
|
||||
mutex_lock(&cpuset_top_mutex);
|
||||
cpus_read_lock();
|
||||
wait_attach_done_lock();
|
||||
|
||||
if (is_cpuset_online(cs))
|
||||
retval = update_prstate(cs, val);
|
||||
cpuset_update_sd_hk_unlock();
|
||||
@@ -3592,7 +3747,7 @@ static int cpuset_can_fork(struct task_struct *task, struct css_set *cset)
|
||||
mutex_lock(&cpuset_mutex);
|
||||
|
||||
/* Check to see if task is allowed in the cpuset */
|
||||
ret = cpuset_can_attach_check(cs);
|
||||
ret = cpuset_can_attach_check(cs, NULL, NULL);
|
||||
if (ret)
|
||||
goto out_unlock;
|
||||
|
||||
@@ -3604,11 +3759,7 @@ static int cpuset_can_fork(struct task_struct *task, struct css_set *cset)
|
||||
if (ret)
|
||||
goto out_unlock;
|
||||
|
||||
/*
|
||||
* Mark attach is in progress. This makes validate_change() fail
|
||||
* changes which zero cpus/mems_allowed.
|
||||
*/
|
||||
cs->attach_in_progress++;
|
||||
attach_ctx.in_progress++;
|
||||
out_unlock:
|
||||
mutex_unlock(&cpuset_mutex);
|
||||
return ret;
|
||||
@@ -3626,7 +3777,7 @@ static void cpuset_cancel_fork(struct task_struct *task, struct css_set *cset)
|
||||
if (same_cs)
|
||||
return;
|
||||
|
||||
dec_attach_in_progress(cs);
|
||||
dec_attach_in_progress();
|
||||
}
|
||||
|
||||
/*
|
||||
@@ -3636,15 +3787,14 @@ static void cpuset_cancel_fork(struct task_struct *task, struct css_set *cset)
|
||||
*/
|
||||
static void cpuset_fork(struct task_struct *task)
|
||||
{
|
||||
struct cpuset *cs;
|
||||
bool same_cs;
|
||||
struct cpuset *cs, *oldcs;
|
||||
|
||||
rcu_read_lock();
|
||||
cs = task_cs(task);
|
||||
same_cs = (cs == task_cs(current));
|
||||
oldcs = task_cs(current);
|
||||
rcu_read_unlock();
|
||||
|
||||
if (same_cs) {
|
||||
if (cs == oldcs) {
|
||||
if (cs == &top_cpuset)
|
||||
return;
|
||||
|
||||
@@ -3655,10 +3805,22 @@ static void cpuset_fork(struct task_struct *task)
|
||||
|
||||
/* CLONE_INTO_CGROUP */
|
||||
mutex_lock(&cpuset_mutex);
|
||||
guarantee_online_mems(cs, &cpuset_attach_nodemask_to);
|
||||
cpuset_attach_task(cs, task);
|
||||
guarantee_online_mems(cs, &attach_ctx.nodemask_to);
|
||||
cs->old_mems_allowed = attach_ctx.nodemask_to;
|
||||
|
||||
dec_attach_in_progress_locked(cs);
|
||||
/*
|
||||
* Assume CPUs and memory nodes are updated
|
||||
* A CLONE_INTO_CGROUP operation should have taken the cgroup mutex
|
||||
* and so there shouldn't be a competing cpuset_attach() operation.
|
||||
*/
|
||||
attach_ctx.cpus_updated = attach_ctx.mems_updated = true;
|
||||
attach_ctx.task_work_queued = false;
|
||||
attach_ctx.old_cs = oldcs;
|
||||
cpuset_attach_task(cs, task);
|
||||
if (attach_ctx.task_work_queued)
|
||||
schedule_flush_migrate_mm();
|
||||
|
||||
dec_attach_in_progress_locked();
|
||||
mutex_unlock(&cpuset_mutex);
|
||||
}
|
||||
|
||||
@@ -3705,6 +3867,7 @@ int __init cpuset_init(void)
|
||||
cpumask_setall(top_cpuset.effective_xcpus);
|
||||
cpumask_setall(top_cpuset.exclusive_cpus);
|
||||
nodes_setall(top_cpuset.effective_mems);
|
||||
init_llist_node(&top_cpuset.attach_node);
|
||||
|
||||
cpuset1_init(&top_cpuset);
|
||||
|
||||
@@ -3762,23 +3925,11 @@ static void cpuset_hotplug_update_tasks(struct cpuset *cs, struct tmpmasks *tmp)
|
||||
bool remote;
|
||||
int partcmd = -1;
|
||||
struct cpuset *parent;
|
||||
retry:
|
||||
wait_event(cpuset_attach_wq, cs->attach_in_progress == 0);
|
||||
|
||||
mutex_lock(&cpuset_mutex);
|
||||
|
||||
/*
|
||||
* We have raced with task attaching. We wait until attaching
|
||||
* is finished, so we won't attach a task to an empty cpuset.
|
||||
*/
|
||||
if (cs->attach_in_progress) {
|
||||
mutex_unlock(&cpuset_mutex);
|
||||
goto retry;
|
||||
}
|
||||
|
||||
wait_attach_done_lock();
|
||||
parent = parent_cs(cs);
|
||||
compute_effective_cpumask(&new_cpus, cs, parent);
|
||||
nodes_and(new_mems, cs->mems_allowed, parent->effective_mems);
|
||||
compute_effective_nodemask(&new_mems, cs, parent);
|
||||
|
||||
if (!tmp || !cs->partition_root_state)
|
||||
goto update_tasks;
|
||||
@@ -3794,7 +3945,7 @@ static void cpuset_hotplug_update_tasks(struct cpuset *cs, struct tmpmasks *tmp)
|
||||
if (remote && (cpumask_empty(subpartitions_cpus) ||
|
||||
(cpumask_empty(&new_cpus) &&
|
||||
partition_is_populated(cs, NULL)))) {
|
||||
cs->prs_err = PERR_HOTPLUG;
|
||||
WRITE_ONCE(cs->prs_err, PERR_HOTPLUG);
|
||||
remote_partition_disable(cs, tmp);
|
||||
compute_effective_cpumask(&new_cpus, cs, parent);
|
||||
remote = false;
|
||||
@@ -4337,14 +4488,10 @@ void cpuset_nodes_allowed(struct cgroup *cgroup, nodemask_t *mask)
|
||||
* cpuset_spread_node() - On which node to begin search for a page
|
||||
* @rotor: round robin rotor
|
||||
*
|
||||
* If a task is marked PF_SPREAD_PAGE or PF_SPREAD_SLAB (as for
|
||||
* tasks in a cpuset with is_spread_page or is_spread_slab set),
|
||||
* and if the memory allocation used cpuset_mem_spread_node()
|
||||
* to determine on which node to start looking, as it will for
|
||||
* certain page cache or slab cache pages such as used for file
|
||||
* system buffers and inode caches, then instead of starting on the
|
||||
* local node to look for a free page, rather spread the starting
|
||||
* node around the tasks mems_allowed nodes.
|
||||
* If a task is marked PFA_SPREAD_PAGE and a page cache allocation uses
|
||||
* cpuset_mem_spread_node() to determine where to start looking, spread the
|
||||
* starting node around the task's mems_allowed nodes instead of starting on
|
||||
* the local node.
|
||||
*
|
||||
* We don't have to worry about the returned node being offline
|
||||
* because "it can't happen", and even if it did, it would be ok.
|
||||
|
||||
@@ -15,11 +15,6 @@ import time
|
||||
import json
|
||||
import math
|
||||
|
||||
import drgn
|
||||
from drgn import container_of
|
||||
from drgn.helpers.linux.list import list_for_each_entry,list_empty
|
||||
from drgn.helpers.linux.radixtree import radix_tree_for_each,radix_tree_lookup
|
||||
|
||||
import argparse
|
||||
parser = argparse.ArgumentParser(description=desc,
|
||||
formatter_class=argparse.RawTextHelpFormatter)
|
||||
@@ -34,6 +29,11 @@ parser.add_argument('--json', action='store_true',
|
||||
help='Output in json')
|
||||
args = parser.parse_args()
|
||||
|
||||
import drgn
|
||||
from drgn import container_of
|
||||
from drgn.helpers.linux.list import list_for_each_entry,list_empty
|
||||
from drgn.helpers.linux.radixtree import radix_tree_for_each,radix_tree_lookup
|
||||
|
||||
def err(s):
|
||||
print(s, file=sys.stderr, flush=True)
|
||||
sys.exit(1)
|
||||
|
||||
@@ -8,6 +8,7 @@
|
||||
|
||||
#define MB(x) (x << 20)
|
||||
|
||||
#define NSEC_PER_USEC 1000L
|
||||
#define USEC_PER_SEC 1000000L
|
||||
#define NSEC_PER_SEC 1000000000L
|
||||
|
||||
|
||||
@@ -426,7 +426,6 @@ static int test_cgcore_no_internal_process_constraint_on_threads(const char *roo
|
||||
ret = KSFT_PASS;
|
||||
|
||||
cleanup:
|
||||
cg_enter_current(root);
|
||||
cg_enter_current(root);
|
||||
if (child)
|
||||
cg_destroy(child);
|
||||
@@ -795,10 +794,9 @@ static int lesser_ns_open_thread_fn(void *arg)
|
||||
static int test_cgcore_lesser_ns_open(const char *root)
|
||||
{
|
||||
static char stack[65536];
|
||||
const uid_t test_euid = 65534; /* usually nobody, any !root is fine */
|
||||
int ret = KSFT_FAIL;
|
||||
char *cg_test_a = NULL, *cg_test_b = NULL;
|
||||
char *cg_test_a_procs = NULL, *cg_test_b_procs = NULL;
|
||||
char *cg_test_b_procs = NULL;
|
||||
int cg_test_b_procs_fd = -1;
|
||||
struct lesser_ns_open_thread_arg targ = { .fd = -1 };
|
||||
pid_t pid;
|
||||
@@ -813,10 +811,9 @@ static int test_cgcore_lesser_ns_open(const char *root)
|
||||
if (!cg_test_a || !cg_test_b)
|
||||
goto cleanup;
|
||||
|
||||
cg_test_a_procs = cg_name(cg_test_a, "cgroup.procs");
|
||||
cg_test_b_procs = cg_name(cg_test_b, "cgroup.procs");
|
||||
|
||||
if (!cg_test_a_procs || !cg_test_b_procs)
|
||||
if (!cg_test_b_procs)
|
||||
goto cleanup;
|
||||
|
||||
if (cg_create(cg_test_a) || cg_create(cg_test_b))
|
||||
@@ -825,10 +822,6 @@ static int test_cgcore_lesser_ns_open(const char *root)
|
||||
if (cg_enter_current(cg_test_b))
|
||||
goto cleanup;
|
||||
|
||||
if (chown(cg_test_a_procs, test_euid, -1) ||
|
||||
chown(cg_test_b_procs, test_euid, -1))
|
||||
goto cleanup;
|
||||
|
||||
targ.path = cg_test_b_procs;
|
||||
pid = clone(lesser_ns_open_thread_fn, stack + sizeof(stack),
|
||||
CLONE_NEWCGROUP | CLONE_FILES | CLONE_VM | SIGCHLD,
|
||||
@@ -863,7 +856,6 @@ static int test_cgcore_lesser_ns_open(const char *root)
|
||||
if (cg_test_a)
|
||||
cg_destroy(cg_test_a);
|
||||
free(cg_test_b_procs);
|
||||
free(cg_test_a_procs);
|
||||
free(cg_test_b);
|
||||
free(cg_test_a);
|
||||
return ret;
|
||||
|
||||
@@ -291,6 +291,8 @@ static int test_cpucg_nice(const char *root)
|
||||
|
||||
user_usec = cg_read_key_long(cpucg, "cpu.stat", "user_usec");
|
||||
nice_usec = cg_read_key_long(cpucg, "cpu.stat", "nice_usec");
|
||||
if (user_usec <= 0)
|
||||
goto cleanup;
|
||||
if (!values_close_report(nice_usec, expected_nice_usec, 1))
|
||||
goto cleanup;
|
||||
|
||||
@@ -639,6 +641,31 @@ test_cpucg_nested_weight_underprovisioned(const char *root)
|
||||
return run_cpucg_nested_weight_test(root, false);
|
||||
}
|
||||
|
||||
/*
|
||||
* Best effort attempt to get the kernel's HZ value from the config.
|
||||
* Return the HZ value if found otherwise return 1000 (the default) to
|
||||
* indicate failure.
|
||||
*/
|
||||
static long
|
||||
get_config_hz(void)
|
||||
{
|
||||
long hz = 1000;
|
||||
FILE *f;
|
||||
char cmd[256] = "zcat /proc/config.gz 2>/dev/null | grep '^CONFIG_HZ='";
|
||||
|
||||
f = popen(cmd, "r");
|
||||
|
||||
if (!f)
|
||||
return hz;
|
||||
|
||||
if (fscanf(f, "CONFIG_HZ=%ld", &hz) == EOF)
|
||||
goto out;
|
||||
|
||||
out:
|
||||
pclose(f);
|
||||
return hz;
|
||||
}
|
||||
|
||||
/*
|
||||
* This test creates a cgroup with some maximum value within a period, and
|
||||
* verifies that a process in the cgroup is not overscheduled.
|
||||
@@ -646,15 +673,18 @@ test_cpucg_nested_weight_underprovisioned(const char *root)
|
||||
static int test_cpucg_max(const char *root)
|
||||
{
|
||||
int ret = KSFT_FAIL;
|
||||
long hz = get_config_hz();
|
||||
long quota_usec = 1000;
|
||||
long default_period_usec = 100000; /* cpu.max's default period */
|
||||
long duration_seconds = 1;
|
||||
|
||||
long duration_usec = duration_seconds * USEC_PER_SEC;
|
||||
long duration_usec;
|
||||
long usage_usec, n_periods, remainder_usec, expected_usage_usec;
|
||||
char *cpucg;
|
||||
char quota_buf[32];
|
||||
|
||||
duration_usec = duration_seconds * USEC_PER_SEC * 1000 / hz;
|
||||
|
||||
snprintf(quota_buf, sizeof(quota_buf), "%ld", quota_usec);
|
||||
|
||||
cpucg = cg_name(root, "cpucg_test");
|
||||
@@ -670,8 +700,8 @@ static int test_cpucg_max(const char *root)
|
||||
struct cpu_hog_func_param param = {
|
||||
.nprocs = 1,
|
||||
.ts = {
|
||||
.tv_sec = duration_seconds,
|
||||
.tv_nsec = 0,
|
||||
.tv_sec = duration_usec / USEC_PER_SEC,
|
||||
.tv_nsec = duration_usec % USEC_PER_SEC * NSEC_PER_USEC,
|
||||
},
|
||||
.clock_type = CPU_HOG_CLOCK_WALL,
|
||||
};
|
||||
@@ -710,15 +740,18 @@ static int test_cpucg_max(const char *root)
|
||||
static int test_cpucg_max_nested(const char *root)
|
||||
{
|
||||
int ret = KSFT_FAIL;
|
||||
long hz = get_config_hz();
|
||||
long quota_usec = 1000;
|
||||
long default_period_usec = 100000; /* cpu.max's default period */
|
||||
long duration_seconds = 1;
|
||||
|
||||
long duration_usec = duration_seconds * USEC_PER_SEC;
|
||||
long duration_usec;
|
||||
long usage_usec, n_periods, remainder_usec, expected_usage_usec;
|
||||
char *parent, *child;
|
||||
char quota_buf[32];
|
||||
|
||||
duration_usec = duration_seconds * USEC_PER_SEC * 1000 / hz;
|
||||
|
||||
snprintf(quota_buf, sizeof(quota_buf), "%ld", quota_usec);
|
||||
|
||||
parent = cg_name(root, "cpucg_parent");
|
||||
@@ -741,8 +774,8 @@ static int test_cpucg_max_nested(const char *root)
|
||||
struct cpu_hog_func_param param = {
|
||||
.nprocs = 1,
|
||||
.ts = {
|
||||
.tv_sec = duration_seconds,
|
||||
.tv_nsec = 0,
|
||||
.tv_sec = duration_usec / USEC_PER_SEC,
|
||||
.tv_nsec = duration_usec % USEC_PER_SEC * NSEC_PER_USEC,
|
||||
},
|
||||
.clock_type = CPU_HOG_CLOCK_WALL,
|
||||
};
|
||||
|
||||
@@ -1,7 +1,13 @@
|
||||
// SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
#define _GNU_SOURCE
|
||||
#include <assert.h>
|
||||
#include <linux/limits.h>
|
||||
#include <pthread.h>
|
||||
#include <sched.h>
|
||||
#include <signal.h>
|
||||
#include <sys/syscall.h>
|
||||
#include <unistd.h>
|
||||
|
||||
#include "kselftest.h"
|
||||
#include "cgroup_util.h"
|
||||
@@ -232,6 +238,246 @@ static int test_cpuset_perms_subtree(const char *root)
|
||||
return ret;
|
||||
}
|
||||
|
||||
static int get_cpu_affinity(cpu_set_t *mask)
|
||||
{
|
||||
CPU_ZERO(mask);
|
||||
return sched_getaffinity(0, sizeof(*mask), mask);
|
||||
}
|
||||
|
||||
static int cpu_set_equal(cpu_set_t *dst, unsigned long mask)
|
||||
{
|
||||
cpu_set_t expected;
|
||||
|
||||
CPU_ZERO(&expected);
|
||||
assert(sizeof(mask) < CPU_SETSIZE);
|
||||
|
||||
for (int cpu = 0; cpu < sizeof(mask) * 8; ++cpu)
|
||||
if ((1UL << cpu) & mask)
|
||||
CPU_SET(cpu, &expected);
|
||||
|
||||
return CPU_EQUAL(&expected, dst);
|
||||
}
|
||||
|
||||
enum test_phase {
|
||||
AFFINITY_SETUP,
|
||||
AFFINITY_CONTROLLER_DISABLED,
|
||||
AFFINITY_COMPLETE,
|
||||
AFFINITY_ERROR
|
||||
};
|
||||
|
||||
struct thread_args {
|
||||
const char *cgroup;
|
||||
cpu_set_t *affinity_before;
|
||||
cpu_set_t *affinity_after;
|
||||
int affinity_before_ready;
|
||||
};
|
||||
|
||||
static pthread_mutex_t test_mutex = PTHREAD_MUTEX_INITIALIZER;
|
||||
static pthread_cond_t test_cond = PTHREAD_COND_INITIALIZER;
|
||||
static enum test_phase test_phase;
|
||||
|
||||
static void *affinity_thread_fn(void *arg)
|
||||
{
|
||||
struct thread_args *args = (struct thread_args *)arg;
|
||||
|
||||
if (cg_enter_current_thread(args->cgroup))
|
||||
goto fail;
|
||||
|
||||
if (get_cpu_affinity(args->affinity_before) != 0)
|
||||
goto fail;
|
||||
|
||||
pthread_mutex_lock(&test_mutex);
|
||||
args->affinity_before_ready = 1;
|
||||
pthread_cond_broadcast(&test_cond);
|
||||
|
||||
while (test_phase < AFFINITY_CONTROLLER_DISABLED)
|
||||
pthread_cond_wait(&test_cond, &test_mutex);
|
||||
pthread_mutex_unlock(&test_mutex);
|
||||
|
||||
if (get_cpu_affinity(args->affinity_after) != 0)
|
||||
goto fail;
|
||||
|
||||
|
||||
return NULL;
|
||||
|
||||
fail:
|
||||
pthread_mutex_lock(&test_mutex);
|
||||
test_phase = AFFINITY_ERROR;
|
||||
pthread_cond_broadcast(&test_cond);
|
||||
pthread_mutex_unlock(&test_mutex);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
/*
|
||||
* Test that disabling cpuset controller properly updates thread affinity.
|
||||
*
|
||||
* This test exposes a bug in cpuset_attach() where threads in child cgroups
|
||||
* don't get their affinity updated when the cpuset controller is disabled.
|
||||
*
|
||||
* Setup:
|
||||
* - Create parent cgroup with cpuset.cpus=0-1
|
||||
* - Create child A with cpuset.cpus=0-1
|
||||
* - Create child B with cpuset.cpus=1
|
||||
* - Place multithreaded process: group leader + thread_a in A, thread_b in B
|
||||
* - Disable cpuset controller on parent
|
||||
*
|
||||
* Expected: thread_b's affinity should expand from {1} to {0-1}
|
||||
* Buggy: thread_b's affinity remains {1}
|
||||
*/
|
||||
static int test_cpuset_affinity_on_controller_disable(const char *root)
|
||||
{
|
||||
char *parent = NULL, *child_a = NULL, *child_b = NULL;
|
||||
pthread_t thread_a, thread_b;
|
||||
int thread_a_created = 0, thread_b_created = 0;
|
||||
cpu_set_t affinity_a_before, affinity_a_after;
|
||||
cpu_set_t affinity_b_before, affinity_b_after;
|
||||
int ret = KSFT_FAIL;
|
||||
|
||||
parent = cg_name(root, "cpuset_affinity_test");
|
||||
if (!parent)
|
||||
goto cleanup;
|
||||
if (cg_create(parent))
|
||||
goto cleanup;
|
||||
if (cg_write(parent, "cgroup.type", "threaded"))
|
||||
goto cleanup;
|
||||
|
||||
child_a = cg_name(parent, "A");
|
||||
if (!child_a)
|
||||
goto cleanup;
|
||||
if (cg_create(child_a))
|
||||
goto cleanup;
|
||||
if (cg_write(child_a, "cgroup.type", "threaded"))
|
||||
goto cleanup;
|
||||
|
||||
child_b = cg_name(parent, "B");
|
||||
if (!child_b)
|
||||
goto cleanup;
|
||||
if (cg_create(child_b))
|
||||
goto cleanup;
|
||||
if (cg_write(child_b, "cgroup.type", "threaded"))
|
||||
goto cleanup;
|
||||
|
||||
/* Now enable cpuset controller in parent */
|
||||
if (cg_write(parent, "cgroup.subtree_control", "+cpuset"))
|
||||
goto skip;
|
||||
|
||||
/*
|
||||
* Set CPU affinity constraints
|
||||
* Skip the test if the setting of "cpuset.cpus" fails as the test
|
||||
* system may not have CPU 1.
|
||||
*/
|
||||
if (cg_write(parent, "cpuset.cpus", "0-1"))
|
||||
goto skip;
|
||||
if (cg_write(child_a, "cpuset.cpus", "0-1"))
|
||||
goto skip;
|
||||
if (cg_write(child_b, "cpuset.cpus", "1"))
|
||||
goto skip;
|
||||
|
||||
/* Move group leader (main thread) to child A */
|
||||
if (cg_enter_current(child_a))
|
||||
goto cleanup;
|
||||
|
||||
/* Create threads - they will move themselves to their respective cgroups */
|
||||
test_phase = AFFINITY_SETUP;
|
||||
|
||||
struct thread_args args_a = {
|
||||
.cgroup = child_a,
|
||||
.affinity_before = &affinity_a_before,
|
||||
.affinity_after = &affinity_a_after,
|
||||
.affinity_before_ready = 0,
|
||||
};
|
||||
if (pthread_create(&thread_a, NULL, affinity_thread_fn, &args_a))
|
||||
goto cleanup;
|
||||
thread_a_created = 1;
|
||||
|
||||
struct thread_args args_b = {
|
||||
.cgroup = child_b,
|
||||
.affinity_before = &affinity_b_before,
|
||||
.affinity_after = &affinity_b_after,
|
||||
.affinity_before_ready = 0,
|
||||
};
|
||||
if (pthread_create(&thread_b, NULL, affinity_thread_fn, &args_b))
|
||||
goto cleanup_threads;
|
||||
thread_b_created = 1;
|
||||
|
||||
pthread_mutex_lock(&test_mutex);
|
||||
while ((test_phase < AFFINITY_ERROR) &&
|
||||
(args_a.affinity_before_ready + args_b.affinity_before_ready < 2))
|
||||
pthread_cond_wait(&test_cond, &test_mutex);
|
||||
|
||||
/* If a thread failed during setup, bail out */
|
||||
if (test_phase == AFFINITY_ERROR) {
|
||||
pthread_mutex_unlock(&test_mutex);
|
||||
goto cleanup_threads;
|
||||
}
|
||||
pthread_mutex_unlock(&test_mutex);
|
||||
|
||||
if (!cpu_set_equal(&affinity_a_before, 0x3)) {
|
||||
ksft_print_msg("FAIL: thread_a initial affinity incorrect\n");
|
||||
goto cleanup_threads;
|
||||
}
|
||||
|
||||
if (!cpu_set_equal(&affinity_b_before, 0x2)) {
|
||||
ksft_print_msg("FAIL: thread_b initial affinity incorrect\n");
|
||||
goto cleanup_threads;
|
||||
}
|
||||
|
||||
/* Disable cpuset controller - this should trigger affinity update */
|
||||
if (cg_write(parent, "cgroup.subtree_control", "-cpuset"))
|
||||
goto cleanup_threads;
|
||||
|
||||
/* Signal threads to save their final affinity and exit */
|
||||
pthread_mutex_lock(&test_mutex);
|
||||
test_phase = AFFINITY_CONTROLLER_DISABLED;
|
||||
pthread_cond_broadcast(&test_cond);
|
||||
pthread_mutex_unlock(&test_mutex);
|
||||
|
||||
pthread_join(thread_a, NULL);
|
||||
pthread_join(thread_b, NULL);
|
||||
|
||||
/* Verify thread affinities AFTER disabling controller */
|
||||
if (!cpu_set_equal(&affinity_a_after, 0x3)) {
|
||||
ksft_print_msg("FAIL: thread_a final affinity incorrect\n");
|
||||
goto cleanup;
|
||||
}
|
||||
|
||||
if (!cpu_set_equal(&affinity_b_after, 0x3)) {
|
||||
ksft_print_msg("FAIL: thread_b affinity did not expand to {0-1}\n");
|
||||
goto cleanup;
|
||||
}
|
||||
|
||||
ret = KSFT_PASS;
|
||||
goto cleanup;
|
||||
|
||||
skip:
|
||||
ret = KSFT_SKIP;
|
||||
goto cleanup;
|
||||
|
||||
cleanup_threads:
|
||||
pthread_mutex_lock(&test_mutex);
|
||||
test_phase = AFFINITY_COMPLETE;
|
||||
pthread_cond_broadcast(&test_cond);
|
||||
pthread_mutex_unlock(&test_mutex);
|
||||
|
||||
if (thread_a_created)
|
||||
pthread_join(thread_a, NULL);
|
||||
if (thread_b_created)
|
||||
pthread_join(thread_b, NULL);
|
||||
|
||||
cleanup:
|
||||
/* Move back to root before cleanup */
|
||||
cg_enter_current(root);
|
||||
|
||||
cg_destroy(child_b);
|
||||
free(child_b);
|
||||
cg_destroy(child_a);
|
||||
free(child_a);
|
||||
cg_destroy(parent);
|
||||
free(parent);
|
||||
|
||||
return ret;
|
||||
}
|
||||
|
||||
|
||||
#define T(x) { x, #x }
|
||||
struct cpuset_test {
|
||||
@@ -241,6 +487,7 @@ struct cpuset_test {
|
||||
T(test_cpuset_perms_object_allow),
|
||||
T(test_cpuset_perms_object_deny),
|
||||
T(test_cpuset_perms_subtree),
|
||||
T(test_cpuset_affinity_on_controller_disable),
|
||||
};
|
||||
#undef T
|
||||
|
||||
|
||||
@@ -20,7 +20,7 @@ skip_test() {
|
||||
WAIT_INOTIFY=$(cd $(dirname $0); pwd)/wait_inotify
|
||||
|
||||
# Find cgroup v2 mount point
|
||||
CGROUP2=$(mount -t cgroup2 | head -1 | awk -e '{print $3}')
|
||||
CGROUP2=$(mount -t cgroup2 | head -1 | awk '{print $3}')
|
||||
[[ -n "$CGROUP2" ]] || skip_test "Cgroup v2 mount point not found!"
|
||||
SUBPARTS_CPUS=$CGROUP2/.__DEBUG__.cpuset.cpus.subpartitions
|
||||
CPULIST=$(cat $CGROUP2/cpuset.cpus.effective)
|
||||
@@ -495,13 +495,26 @@ REMOTE_TEST_MATRIX=(
|
||||
# Narrowing cpuset.cpus to previously sibling-excluded CPUs should
|
||||
# not return CPUs that were never actually owned.
|
||||
" C1-4:P1 . C1-2:P1 C1-3:P2 . . \
|
||||
. . . C3 . . p1:4|c11:1-2|c12:3 \
|
||||
. . . C3 . . p1:4|c11:1-2|c12:3 \
|
||||
p1:P1|c11:P1|c12:P2 3"
|
||||
# Expanding cpuset.cpus to include a previously sibling-excluded CPU
|
||||
# after the sibling has become a member should correctly request it.
|
||||
" C1-4:P1 . C1-2:P1 C1-3:P2 . . \
|
||||
. . P0 C2-3 . . p1:1,4|c11:1|c12:2-3 \
|
||||
. . P0 C2-3 . . p1:1,4|c11:1|c12:2-3 \
|
||||
p1:P1|c11:P0|c12:P2 2-3"
|
||||
# Changing a sibling partition's cpuset.cpus to overlap with another
|
||||
# sibling partition should invalidate itself and return only actually
|
||||
# allocated CPUs (effective_xcpus) to the parent.
|
||||
" C1-4:P1 . C1-2:P1 C2-4:P2 . . \
|
||||
. . . C1-2 . . p1:3-4|c11:1-2|c12:3-4 \
|
||||
p1:P1|c11:P1|c12:P-2"
|
||||
# Cpusets with empty cpuset.cpus should inherit parent's effective_cpus
|
||||
" C1-4:P1 C5-6 C1-2 . C5 . \
|
||||
. P1 P1 . . . p1:3-4|p2:5-6|c11:1-2|c12:3-4|c21:5|c22:5-6 \
|
||||
p1:P1|p2:P1|c11:P1"
|
||||
" C1-4:P1 C5-6 C1-2 . C5 . \
|
||||
. P1 P1 . O5=0 . p1:3-4|p2:6|c11:1-2|c12:3-4|c21:6|c22:6 \
|
||||
p1:P1|p2:P1|c11:P1"
|
||||
)
|
||||
|
||||
#
|
||||
@@ -513,6 +526,7 @@ write_cpu_online()
|
||||
CPU=${1%=*}
|
||||
VAL=${1#*=}
|
||||
CPUFILE=//sys/devices/system/cpu/cpu${CPU}/online
|
||||
echo $VAL > $CPUFILE || return 1
|
||||
if [[ $VAL -eq 0 ]]
|
||||
then
|
||||
OFFLINE_CPUS="$OFFLINE_CPUS $CPU"
|
||||
@@ -522,7 +536,6 @@ write_cpu_online()
|
||||
sort | uniq -u)
|
||||
}
|
||||
fi
|
||||
echo $VAL > $CPUFILE
|
||||
pause 0.05
|
||||
}
|
||||
|
||||
@@ -590,7 +603,8 @@ set_ctrl_state()
|
||||
eval $COMM $REDIRECT
|
||||
;;
|
||||
O*) VAL=${CMD#?}
|
||||
write_cpu_online $VAL
|
||||
COMM="write_cpu_online $VAL"
|
||||
eval $COMM $REDIRECT
|
||||
;;
|
||||
T*) COMM="echo 0 > $TFILE"
|
||||
eval $COMM $REDIRECT
|
||||
|
||||
@@ -14,7 +14,7 @@ skip_test() {
|
||||
[[ $(id -u) -eq 0 ]] || skip_test "Test must be run as root!"
|
||||
|
||||
# Find cpuset v1 mount point
|
||||
CPUSET=$(mount -t cgroup | grep cpuset | head -1 | awk -e '{print $3}')
|
||||
CPUSET=$(mount -t cgroup | grep cpuset | head -1 | awk '{print $3}')
|
||||
[[ -n "$CPUSET" ]] || skip_test "cpuset v1 mount point not found!"
|
||||
|
||||
#
|
||||
|
||||
@@ -199,7 +199,10 @@ static int test_hugetlb_memcg(char *root)
|
||||
int main(int argc, char **argv)
|
||||
{
|
||||
char root[PATH_MAX];
|
||||
int ret = EXIT_SUCCESS, has_memory_hugetlb_acc;
|
||||
int has_memory_hugetlb_acc;
|
||||
|
||||
ksft_print_header();
|
||||
ksft_set_plan(1);
|
||||
|
||||
has_memory_hugetlb_acc = proc_mount_contains("memory_hugetlb_accounting");
|
||||
if (has_memory_hugetlb_acc < 0)
|
||||
@@ -211,7 +214,7 @@ int main(int argc, char **argv)
|
||||
if (get_hugepage_size() != 2048) {
|
||||
ksft_print_msg("test_hugetlb_memcg requires 2MB hugepages\n");
|
||||
ksft_test_result_skip("test_hugetlb_memcg\n");
|
||||
return ret;
|
||||
ksft_finished();
|
||||
}
|
||||
|
||||
if (cg_find_unified_root(root, sizeof(root), NULL))
|
||||
@@ -233,10 +236,9 @@ int main(int argc, char **argv)
|
||||
ksft_test_result_skip("test_hugetlb_memcg\n");
|
||||
break;
|
||||
default:
|
||||
ret = EXIT_FAILURE;
|
||||
ksft_test_result_fail("test_hugetlb_memcg\n");
|
||||
break;
|
||||
}
|
||||
|
||||
return ret;
|
||||
ksft_finished();
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user