From: Andrey Grodzovsky <Andrey.Grodzovsky@amd.com>
To: christian.koenig@amd.com, "Eric W. Biederman" <ebiederm@xmission.com>
Cc: David.Panariti@amd.com, Oleg Nesterov <oleg@redhat.com>,
amd-gfx@lists.freedesktop.org, linux-kernel@vger.kernel.org,
Alexander.Deucher@amd.com, akpm@linux-foundation.org
Subject: Re: [PATCH 2/3] drm/scheduler: Don't call wait_event_killable for signaled process.
Date: Mon, 30 Apr 2018 10:32:51 -0400 [thread overview]
Message-ID: <bceb1a1b-c453-782d-5a7d-40fa2f22c813@amd.com> (raw)
In-Reply-To: <c3c9787d-b279-8169-43d1-74eeb666ffbd@gmail.com>
On 04/30/2018 08:08 AM, Christian König wrote:
> Hi Eric,
>
> sorry for the late response, was on vacation last week.
>
> Am 26.04.2018 um 02:01 schrieb Eric W. Biederman:
>> Andrey Grodzovsky <Andrey.Grodzovsky@amd.com> writes:
>>
>>> On 04/25/2018 01:17 PM, Oleg Nesterov wrote:
>>>> On 04/25, Andrey Grodzovsky wrote:
>>>>> here (drm_sched_entity_fini) is also a bad idea, but we still want
>>>>> to be
>>>>> able to exit immediately
>>>>> and not wait for GPU jobs completion when the reason for reaching
>>>>> this code
>>>>> is because of KILL
>>>>> signal to the user process who opened the device file.
>>>> Can you hook f_op->flush method?
>
> THANKS! That sounds like a really good idea to me and we haven't
> investigated into that direction yet.
>
>>> But this one is called for each task releasing a reference to the
>>> the file, so
>>> not sure I see how this solves the problem.
>> The big question is why do you need to wait during the final closing a
>> file?
>
> As always it's because of historical reasons. Initially user space
> pushed commands directly to a hardware queue and when a processes
> finished we didn't need to wait for anything.
>
> Then the GPU scheduler was introduced which delayed pushing the jobs
> to the hardware queue to a later point in time.
>
> This wait was then added to maintain backward compability and not
> break userspace (but see below).
>
>> The wait can be terminated so the wait does not appear to be simply a
>> matter of correctness.
>
> Well when the process is killed we don't care about correctness any
> more, we just want to get rid of it as quickly as possible (OOM
> situation etc...).
>
> But it is perfectly possible that a process submits some render
> commands and then calls exit() or terminates because of a SIGTERM,
> SIGINT etc.. In this case we need to wait here to make sure that all
> rendering is pushed to the hardware because the scheduler might need
> resources/settings from the file descriptor.
>
> For example if you just remove that wait you could close firefox and
> get garbage on the screen for a millisecond because the remaining
> rendering commands where not executed.
>
> So what we essentially need is to distinct between a SIGKILL (which
> means stop processing as soon as possible) and any other reason
> because then we don't want to annoy the user with garbage on the
> screen (even if it's just for a few milliseconds).
>
> Constructive ideas how to handle this would be very welcome, cause I
> completely agree that what we have at the moment by checking PF_SIGNAL
> is just a very very hacky workaround.
What about changing PF_SIGNALED to PF_EXITING in
drm_sched_entity_do_release
- if ((current->flags & PF_SIGNALED) && current->exit_code == SIGKILL)
+ if ((current->flags & PF_EXITING) && current->exit_code == SIGKILL)
From looking into do_exit and it's callers , current->exit_code will
get assign the signal which was delivered to the task. If SIGINT was
sent then it's SIGINT, if SIGKILL then SIGKILL.
Andrey
>
> Thanks,
> Christian.
>
>>
>> Eric
>> _______________________________________________
>> amd-gfx mailing list
>> amd-gfx@lists.freedesktop.org
>> https://lists.freedesktop.org/mailman/listinfo/amd-gfx
>
next prev parent reply other threads:[~2018-04-30 14:33 UTC|newest]
Thread overview: 63+ messages / expand[flat|nested] mbox.gz Atom feed top
2018-04-24 15:30 Avoid uninterruptible sleep during process exit Andrey Grodzovsky
2018-04-24 15:30 ` [PATCH 1/3] signals: Allow generation of SIGKILL to exiting task Andrey Grodzovsky
2018-04-24 16:10 ` Eric W. Biederman
2018-04-24 16:42 ` Eric W. Biederman
2018-04-24 16:51 ` Andrey Grodzovsky
2018-04-24 17:29 ` Eric W. Biederman
2018-04-25 13:13 ` Oleg Nesterov
2018-04-24 15:30 ` [PATCH 2/3] drm/scheduler: Don't call wait_event_killable for signaled process Andrey Grodzovsky
2018-04-24 15:46 ` Michel Dänzer
2018-04-24 15:51 ` Andrey Grodzovsky
2018-04-24 15:52 ` Andrey Grodzovsky
2018-04-24 19:44 ` Daniel Vetter
2018-04-24 21:00 ` Eric W. Biederman
2018-04-24 21:02 ` Andrey Grodzovsky
2018-04-24 21:21 ` Eric W. Biederman
2018-04-24 21:37 ` Andrey Grodzovsky
2018-04-24 22:11 ` Eric W. Biederman
2018-04-25 7:14 ` Daniel Vetter
2018-04-25 13:08 ` Andrey Grodzovsky
2018-04-25 15:29 ` Eric W. Biederman
[not found] ` <311660b9-9e46-b960-3088-06e16ac3838d@amd.com>
2018-04-25 16:31 ` Eric W. Biederman
2018-04-24 21:40 ` Daniel Vetter
2018-04-25 13:22 ` Oleg Nesterov
2018-04-25 13:36 ` Daniel Vetter
2018-04-25 14:18 ` Oleg Nesterov
2018-04-25 13:43 ` Andrey Grodzovsky
2018-04-24 16:23 ` Eric W. Biederman
2018-04-24 16:43 ` Andrey Grodzovsky
2018-04-24 17:12 ` Eric W. Biederman
2018-04-25 13:55 ` Oleg Nesterov
2018-04-25 14:21 ` Andrey Grodzovsky
2018-04-25 17:17 ` Oleg Nesterov
2018-04-25 18:40 ` Andrey Grodzovsky
2018-04-26 0:01 ` Eric W. Biederman
2018-04-26 12:34 ` Andrey Grodzovsky
2018-04-26 12:52 ` Andrey Grodzovsky
2018-04-26 15:57 ` Eric W. Biederman
2018-04-26 20:43 ` Andrey Grodzovsky
2018-04-30 12:08 ` Christian König
2018-04-30 14:32 ` Andrey Grodzovsky [this message]
2018-04-30 15:25 ` Christian König
2018-04-30 16:00 ` Oleg Nesterov
2018-04-30 16:10 ` Andrey Grodzovsky
2018-04-30 18:29 ` Christian König
2018-04-30 19:28 ` Andrey Grodzovsky
2018-05-02 11:48 ` Christian König
2018-05-01 14:35 ` Oleg Nesterov
2018-05-23 15:08 ` Andrey Grodzovsky
2018-04-30 15:29 ` Oleg Nesterov
2018-04-30 16:25 ` Eric W. Biederman
2018-04-30 17:18 ` Andrey Grodzovsky
2018-04-25 13:05 ` Oleg Nesterov
2018-04-24 15:30 ` [PATCH 3/3] drm/amdgpu: Switch to interrupted wait to recover from ring hang Andrey Grodzovsky
2018-04-24 15:52 ` Panariti, David
2018-04-24 15:58 ` Andrey Grodzovsky
2018-04-24 16:20 ` Panariti, David
2018-04-24 16:30 ` Eric W. Biederman
2018-04-25 17:17 ` Andrey Grodzovsky
2018-04-25 20:55 ` Eric W. Biederman
2018-04-26 12:28 ` Andrey Grodzovsky
2018-04-24 16:14 ` Eric W. Biederman
2018-04-24 16:38 ` Andrey Grodzovsky
2018-04-30 11:34 ` Christian König
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=bceb1a1b-c453-782d-5a7d-40fa2f22c813@amd.com \
--to=andrey.grodzovsky@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=David.Panariti@amd.com \
--cc=akpm@linux-foundation.org \
--cc=amd-gfx@lists.freedesktop.org \
--cc=christian.koenig@amd.com \
--cc=ebiederm@xmission.com \
--cc=linux-kernel@vger.kernel.org \
--cc=oleg@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®