mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* Re: [PATCH] minor sched-E1 tweaks and questions
       [not found] <Pine.LNX.4.33.0201091126520.2276-100000@localhost.localdomain>
@ 2002-01-10  3:48 ` Rusty Russell
  2002-01-10  6:06   ` Davide Libenzi
  2002-01-10 13:07   ` Ingo Molnar
  0 siblings, 2 replies; 3+ messages in thread
From: Rusty Russell @ 2002-01-10  3:48 UTC (permalink / raw)
  To: mingo; +Cc: linux-kernel

In message <Pine.LNX.4.33.0201091126520.2276-100000@localhost.localdomain> you 
write:
> > Q. How can this happen in expire_task():
> > 	if (p->array != rq->active) {
> > 		p->need_resched = 1;
> > 		return;
> > 	}
> 
> if a task gets delayed by some really heavy IRQ load and the timer
> interrupt hits the task twice.

Hmm... still don't see it.  update_process_times() surely doesn't
re-enter?  And another CPU cannot load_balance() p (== current) away
from us.

Another question:

	if (likely(prev != next)) {
		rq->nr_switches++;
		rq->curr = next;
		next->cpu = prev->cpu;
		context_switch(prev, next);
		/*
		 * The runqueue pointer might be from another CPU
		 * if the new task was last running on a different
		 * CPU - thus re-load it.
		 */
		barrier();
		rq = this_rq();
	}
	spin_unlock_irq(&rq->lock);

I do not understand this comment.  How can rq (ie. smp_processor_id())
change?  Nothing sleeps here, and if it DID change, the
spin_unlock_irq() would be wrong...

> i've taken your fixes, please double-check the next patch whether all of
> them are correctly applied.

Um, missed one:

 static struct runqueue {
-	int cpu;
 	spinlock_t lock;
-	unsigned long nr_running, nr_switches, last_rt_event;
+	unsigned long nr_running, nr_switches;
 	task_t *curr, *idle;
 	prio_array_t *active, *expired, arrays[2];
-	char __pad [SMP_CACHE_BYTES];
+	int prev_nr_running[NR_CPUS];
 } runqueues [NR_CPUS] __cacheline_aligned;

You want each entry in the array to be aligned, not the whole array!
You need to define the struct runqueue to be the cacheline aligned
(using ____cacheline_aligned since it's a type), THEN put the
__cacheline_aligned after the array declaration so it gets put in the
aligned section:

 struct runqueue {
 	spinlock_t lock;
	unsigned long nr_running, nr_switches;
 	task_t *curr, *idle;
 	prio_array_t *active, *expired, arrays[2];
	int prev_nr_running[NR_CPUS];
 } ____cacheline_aligned;

 static struct runqueue runqueues [NR_CPUS] __cacheline_aligned;

This is why my __per_cpu patch was invented 8)
Cheers,
Rusty.
--
  Anyone who quotes me in their sig is an idiot. -- Rusty Russell.

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] minor sched-E1 tweaks and questions
  2002-01-10  3:48 ` [PATCH] minor sched-E1 tweaks and questions Rusty Russell
@ 2002-01-10  6:06   ` Davide Libenzi
  2002-01-10 13:07   ` Ingo Molnar
  1 sibling, 0 replies; 3+ messages in thread
From: Davide Libenzi @ 2002-01-10  6:06 UTC (permalink / raw)
  To: Rusty Russell; +Cc: Ingo Molnar, lkml

On Thu, 10 Jan 2002, Rusty Russell wrote:

> Another question:
>
> 	if (likely(prev != next)) {
> 		rq->nr_switches++;
> 		rq->curr = next;
> 		next->cpu = prev->cpu;
> 		context_switch(prev, next);
> 		/*
> 		 * The runqueue pointer might be from another CPU
> 		 * if the new task was last running on a different
> 		 * CPU - thus re-load it.
> 		 */
> 		barrier();
> 		rq = this_rq();
> 	}
> 	spin_unlock_irq(&rq->lock);
>
> I do not understand this comment.  How can rq (ie. smp_processor_id())
> change?  Nothing sleeps here, and if it DID change, the
> spin_unlock_irq() would be wrong...

If you switch you'll on the stack the rq of the previous cpu
spin_unlock_irq(&rq->lock) is fine if you do not switch and if you switch
you need to reload rq





- Davide



^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] minor sched-E1 tweaks and questions
  2002-01-10  3:48 ` [PATCH] minor sched-E1 tweaks and questions Rusty Russell
  2002-01-10  6:06   ` Davide Libenzi
@ 2002-01-10 13:07   ` Ingo Molnar
  1 sibling, 0 replies; 3+ messages in thread
From: Ingo Molnar @ 2002-01-10 13:07 UTC (permalink / raw)
  To: Rusty Russell; +Cc: linux-kernel


On Thu, 10 Jan 2002, Rusty Russell wrote:

> > > Q. How can this happen in expire_task():
> > > 	if (p->array != rq->active) {
> > > 		p->need_resched = 1;
> > > 		return;
> > > 	}
> >
> > if a task gets delayed by some really heavy IRQ load and the timer
> > interrupt hits the task twice.
>
> Hmm... still don't see it.  update_process_times() surely doesn't
> re-enter?  And another CPU cannot load_balance() p (== current) away
> from us.

the issue is when a task has not rescheduled yet even though the previous
timer tick has has told it to do so. Then we'll have the p->array !=
rq->active condition. If some other event (like a heavy SCSI irq or
something else) delays the task from getting into the scheduler for more
than a jiffy, we can get an expire_task() call again - and hit the
condition.

this situation is especially likely to happen with HZ=1024 or higher.

> Another question:
>
> 		context_switch(prev, next);
> 		/*
> 		 * The runqueue pointer might be from another CPU
> 		 * if the new task was last running on a different
> 		 * CPU - thus re-load it.
> 		 */
> 		barrier();
> 		rq = this_rq();
> 	}
> 	spin_unlock_irq(&rq->lock);
>
> I do not understand this comment.  How can rq (ie. smp_processor_id())
> change?  Nothing sleeps here, and if it DID change, the
> spin_unlock_irq() would be wrong...

the 'rq' variable is not 'constant' across context-switch if you look at
what happens on the CPU - we switch away from a task into some other task.
That other task might have a much older 'rq' variable on its kernel stack,
which might be invalid, if the (now executing) task was load-balanced over
to this CPU.

>  } runqueues [NR_CPUS] __cacheline_aligned;
>
> You want each entry in the array to be aligned, not the whole array!

hm, right you are. Will be in my next patch.

	Ingo


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2002-01-10 11:11 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
     [not found] <Pine.LNX.4.33.0201091126520.2276-100000@localhost.localdomain>
2002-01-10  3:48 ` [PATCH] minor sched-E1 tweaks and questions Rusty Russell
2002-01-10  6:06   ` Davide Libenzi
2002-01-10 13:07   ` Ingo Molnar

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®