From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id B4E53EE49AF for ; Tue, 22 Aug 2023 20:31:50 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S229998AbjHVUbu (ORCPT ); Tue, 22 Aug 2023 16:31:50 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:56504 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S229667AbjHVUbt (ORCPT ); Tue, 22 Aug 2023 16:31:49 -0400 Received: from madras.collabora.co.uk (madras.collabora.co.uk [46.235.227.172]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id D76D6CC; Tue, 22 Aug 2023 13:31:46 -0700 (PDT) Received: from nicolas-tpx395.localdomain (unknown [IPv6:2606:6d00:15:bae9::7a9]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (4096 bits) server-digest SHA256) (No client certificate requested) (Authenticated sender: nicolas) by madras.collabora.co.uk (Postfix) with ESMTPSA id 36822660720C; Tue, 22 Aug 2023 21:31:44 +0100 (BST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=collabora.com; s=mail; t=1692736305; bh=eFMQNZGba+uOQx+8IPFzb+FmChWNc51CVLN7MDsgj2s=; h=Subject:From:To:Cc:Date:In-Reply-To:References:From; b=oV8jlNxBSYhMNHRzZ8i6SE3aIKhvNdUmeiRHo+blnu2JeM+x5IOtEYCrmcit7YTXs ljo0intT/hrpvmFNcw2YZP6+czl703Ta9K4j3Uq2J3fUdZeGxC1DWMmmmHYVF0ueG6 GsKcMLcfmAoYXFWnzb3MfdwsDKs7jUpEx77HOM1DU/TNYXYggdx5yS8tx+Km+2s+E/ Ml8lHuXz6PrMa0GOZv7zizqGMZFxTZjAeW6jvkzmNjarutMBW3EfaWQ+SPT6cyXPH7 b7Q0ATedmp3sjUZF6JldwIoBBaTmGKxY4kCJEAOwEv2rmXoQR5E6GC2C1Khz9cvUYa YYBuIIzSX/mdA== Message-ID: Subject: Re: Stateless Encoding uAPI Discussion and Proposal From: Nicolas Dufresne To: Hsia-Jun Li , Paul Kocialkowski Cc: linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, Hans Verkuil , Sakari Ailus , Andrzej Pietrasiewicz , Michael Tretter , Jernej =?UTF-8?Q?=C5=A0krabec?= , Chen-Yu Tsai , Samuel Holland , Thomas Petazzoni Date: Tue, 22 Aug 2023 16:31:34 -0400 In-Reply-To: <39270c5e-24ab-8ff6-d925-7718b1fef3c4@synaptics.com> References: <720c476189552596cbd61dd74d6fa12818718036.camel@collabora.com> <39270c5e-24ab-8ff6-d925-7718b1fef3c4@synaptics.com> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable User-Agent: Evolution 3.48.4 (3.48.4-1.fc38) MIME-Version: 1.0 Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi, >=20 [...] > > In cable streaming notably, the RC job is to monitor the about of bits = over a > > period of time (the window). This window is defined by the streaming ha= rdware > > buffering capabilities. Best at this point is to start reading through = HRD > > specifications, and open source rate control implementation (notably x2= 64). > >=20 > > I think overall, we can live with adding hints were needed, and if the = gop > > information is appropriate hint, then we can just reuse the existing co= ntrol. > >=20 > Why we still care about GOP here. Hardware have no idea about GOP at=20 > all. Although in codec likes HEVC, IDR and intra pictures's nalu header= =20 > is different, there is not different in the hardware coding=20 > configration. NALU header is generated by the userspace usually. >=20 > While future encoding would regard the current encoded picture as an IDR= =20 > is completed decided by the userspace. The discussion was around having basic RC algorithm in the kernel driver, possibly making use of hardware specific features without actually exposing= it all to userspace. So assuming we do that: Paul's concern is that for best result, an RC algorithm could use knowledge= of keyframe placement to preserve bucket space (possibly using the last keyfra= me size as a hint). Exposing the GOP structure in some form allow "prediction"= , so the adaption can lookahead future budget without introducing latency. There= is an alternative, which is to require ahead of time queuing of encode request= s. But this does introduce latency since the way it works in V4L2 today, we ne= ed the picture to be filled by the time we request an encode. Though, if we drop the GOP structure and favour this approach, the latency = could be regain later by introducing fence base streaming. The technique would be= for a video source (like a capture driver) to pass dmabuf that aren't filled ye= t, but have a companion fence. This would allow queuing requests ahead of time= , and all we need is enough pre-allocation to accommodate the desired look ahead.= Only issue is that perhaps this violates the fundamental of "short term" deliver= y of fences. But fences can also fail I think, in case the capture was stopped. We can certainly move forward with this as a future solution, or just don't implement future aware RC algorithm in term to avoid the huge task this inv= olves (and possibly patents?) [...] > >=20 > > Of course, the subject is much more relevant when there is encoders wit= h more > > then 1 reference. But you are correct, what the commands do, is allow t= o change, > > add or remove any reference from the list (random modification), as lon= g as they > > fit in the codec contraints (like the DPB size notably). This is the on= ly way > > one can implement temporal SVC reference pattern, robust reference tree= s or RTP > > RPSI. Note that long term reference also exists, and are less complex t= hen these > > commands. > >=20 >=20 > If we the userspace could manage the lifetime of reconstruction=20 > buffers(assignment, reference), we don't need a command here. Sorry if I created confusion, the comments was something specific to H.264 coding. Its a compressed form for the reference lists. This information is = coded in the slice header and enabled through adaptive_ref_pic_marking_mode_flag It was suggested so far to leave h264 slice headers writing to the driver. = This is motivated by H264 slice header not being byte aligned in size, so the slice_data() is hard to combine. Also, some hardware actually produce the slice_header. This needs actual hardware interface analyses, cause an H.264 slice header is worth nothing if it cannot instruct the decoder how to main= tain the desired reference state. I think this aspect should probably not be generalized to all CODECs, since= the packing semantic can largely differ. When the codec header is indeed byte aligned, it can easily be seperate and combined by application, improve the application flexibility, reducing the kernel API complexity. >=20 > It is just a problem of how to design another request API control=20 > structure to select which buffers would be used for list0, list1. > > I this raises a big question, and I never checked how this worked with = let's say > > VA. Shall we let the driver resolve the changes into commands (VP8 have > > something similar, while VP9 and AV1 are refresh flags, which are just = trivial > > to compute). I believe I'll have to investigate this further. > >=20 > > > >=20 > > [...] regards, Nicolas