linux

mirror of https://github.com/torvalds/linux.git synced 2024-11-02 02:01:29 +00:00

History

Vincent Guittot 6b94780e45 sched/core: Use load_avg for selecting idlest group find_idlest_group() only compares the runnable_load_avg when looking for the least loaded group. But on fork intensive use case like hackbench where tasks blocked quickly after the fork, this can lead to selecting the same CPU instead of other CPUs, which have similar runnable load but a lower load_avg. When the runnable_load_avg of 2 CPUs are close, we now take into account the amount of blocked load as a 2nd selection factor. There is now 3 zones for the runnable_load of the rq: - [0 .. (runnable_load - imbalance)]: Select the new rq which has significantly less runnable_load - [(runnable_load - imbalance) .. (runnable_load + imbalance)]: The runnable loads are close so we use load_avg to chose between the 2 rq - [(runnable_load + imbalance) .. ULONG_MAX]: Keep the current rq which has significantly less runnable_load The scale factor that is currently used for comparing runnable_load, doesn't work well with small value. As an example, the use of a scaling factor fails as soon as this_runnable_load == 0 because we always select local rq even if min_runnable_load is only 1, which doesn't really make sense because they are just the same. So instead of scaling factor, we use an absolute margin for runnable_load to detect CPUs with similar runnable_load and we keep using scaling factor for blocked load. For use case like hackbench, this enable the scheduler to select different CPUs during the fork sequence and to spread tasks across the system. Tests have been done on a Hikey board (ARM based octo cores) for several kernel. The result below gives min, max, avg and stdev values of 18 runs with each configuration. The patches depend on the "no missing update_rq_clock()" work. hackbench -P -g 1 `ea86cb4b76` `7dc603c902` v4.8 v4.8+patches min 0.049 0.050 0.051 0,048 avg 0.057 0.057(0%) 0.057(0%) 0,055(+5%) max 0.066 0.068 0.070 0,063 stdev +/-9% +/-9% +/-8% +/-9% More performance numbers here: https://lkml.kernel.org/r/20161203214707.GI20785@codeblueprint.co.uk Tested-by: Matt Fleming <matt@codeblueprint.co.uk> Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org> Reviewed-by: Matt Fleming <matt@codeblueprint.co.uk> Cc: Linus Torvalds <torvalds@linux-foundation.org> Cc: Morten.Rasmussen@arm.com Cc: Peter Zijlstra <peterz@infradead.org> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: dietmar.eggemann@arm.com Cc: kernellwp@gmail.com Cc: umgwanakikbuti@gmail.com Cc: yuyang.du@intel.comc Link: http://lkml.kernel.org/r/1481216215-24651-3-git-send-email-vincent.guittot@linaro.org Signed-off-by: Ingo Molnar <mingo@kernel.org>		2016-12-11 13:10:57 +01:00
..
auto_group.c	sched/autogroup: Fix 64-bit kernel nice level adjustment	2016-11-24 05:45:02 +01:00
auto_group.h	sched, timer: Convert usages of ACCESS_ONCE() in the scheduler to READ_ONCE()/WRITE_ONCE()	2015-05-08 12:11:32 +02:00
clock.c	sched/clock: Make local_clock()/cpu_clock() inline	2016-04-13 12:25:22 +02:00
completion.c	sched/completion: Serialize completion_done() with complete()	2015-02-18 14:27:40 +01:00
core.c	sched: Extend scheduler's asym packing	2016-11-24 14:09:46 +01:00
cpuacct.c	sched/cpuacct: Avoid %lld seq_printf warning	2016-11-16 10:29:03 +01:00
cpuacct.h	sched/cpuacct: Simplify the cpuacct code	2016-03-21 11:00:28 +01:00
cpudeadline.c	sched/deadline: Split cpudl_set() into cpudl_set() and cpudl_clear()	2016-09-05 13:29:43 +02:00
cpudeadline.h	sched/deadline: Split cpudl_set() into cpudl_set() and cpudl_clear()	2016-09-05 13:29:43 +02:00
cpufreq_schedutil.c	cpufreq: schedutil: Add iowait boosting	2016-09-13 23:36:01 +02:00
cpufreq.c	cpufreq / sched: Pass flags to cpufreq_update_util()	2016-08-16 22:14:55 +02:00
cpupri.c	sched/core: Use tsk_cpus_allowed() instead of accessing ->cpus_allowed	2016-05-12 09:55:35 +02:00
cpupri.h	sched/cpupri: Remove unnecessary definitions in cpupri.h	2014-11-16 10:58:59 +01:00
cputime.c	sched/cputime: Simplify task_cputime()	2016-11-15 09:51:05 +01:00
deadline.c	sched/dl: Fix comment in pick_next_task_dl()	2016-11-23 10:23:21 +01:00
debug.c	Merge branch 'for-4.9' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/cgroup	2016-10-14 12:18:50 -07:00
fair.c	sched/core: Use load_avg for selecting idlest group	2016-12-11 13:10:57 +01:00
features.h	sched/fair: Convert arch_scale_cpu_capacity() from weak function to #define	2015-09-13 09:52:55 +02:00
idle_task.c	sched/core: Rewrite and improve select_idle_siblings()	2016-09-30 11:03:09 +02:00
idle.c	nmi_backtrace: generate one-line reports for idle cpus	2016-10-07 18:46:30 -07:00
loadavg.c	sched/core: Correct off by one bug in load migration calculation	2016-07-13 14:58:20 +02:00
Makefile	cpufreq: schedutil: New governor based on scheduler utilization data	2016-04-02 01:09:12 +02:00
rt.c	cpufreq / sched: Pass runqueue pointer to cpufreq_update_util()	2016-08-16 22:16:03 +02:00
sched.h	sched: Extend scheduler's asym packing	2016-11-24 14:09:46 +01:00
stats.c	sched: use %*pb[l] to print bitmaps including cpumasks and nodemasks	2015-02-13 21:21:37 -08:00
stats.h	sched/debug: Rename 'schedstat_val()' -> 'schedstat_val_or_zero()'	2016-09-05 13:29:46 +02:00
stop_task.c	locking/lockdep, sched/core: Implement a better lock pinning scheme	2016-05-05 09:23:59 +02:00
swait.c	wait.[ch]: Introduce the simple waitqueue (swait) implementation	2016-02-25 11:27:16 +01:00
wait.c	mm: remove per-zone hashtable of bitlock waitqueues	2016-10-27 09:27:57 -07:00