Repository navigation
Chromium encoder support - #242
Merged
Merged
Conversation
zlatinski
force-pushed
the
chromium-encoder-support
branch
6 times, most recently
from
September 18, 2026 06:15
4eccd01 to
03b690c
Compare
…n ctest
The encoder library becomes something a host can configure, compile and test on
its own, outside any embedding checkout, on a CI that actually runs the tests.
A STANDALONE CMAKE TREE compiles every translation unit the embedded build
compiles, so a change that breaks the library is visible here rather than in
somebody else's checkout. SPIRV-Tools is linked whenever glslang is a static
archive rather than only under MSVC, which is the configuration that otherwise
fails to link. The compute-filter gate is expressed against the shader-compiler
backend instead of naming one compiler, so the option survives a change of
backend. The FFmpeg-dependent decoder demo builds without FFmpeg.
ONE PORTABILITY SEAM, AND IT IS LINUX. VkVideoEncoderOsAdapterLinux and
vulkan_video_encoder_os_event_linux put the event, thread and handle primitives
the encoder needs behind a single interface, so the platform differences live in
two translation units instead of being spread through the encoder. Both are
Linux-only by construction, and the file names now say so rather than implying a
portability that does not exist. An unsupported target is refused at configure
time, with the two files a port has to supply named, rather than as a preprocessor
error partway through a build. VkThreadPool and VkThreadSafeQueue gain the
liveness the encoder relies on.
DIAGNOSTICS BEHIND A LATCH. VkEncoderStdioLatch gives the library one pair of
streams and one latch. An embedder that silences the library cannot have a later
session unsilence it, which is what makes silencing a property of the process
rather than of each call site, and it is why the later features in this series
report through the API rather than through stdout.
CI THAT CAN FAIL. The registered tests gated nothing, because no workflow ran
ctest at all. enable_testing() lands at the repository root -- without that single
line CMake writes no CTestTestfile.cmake and every registered test registers and
never runs -- a CI script runs ctest with a device-free and gpu label split, and
that script refuses to report success on a label typo or on an empty selection,
because a suite that selects nothing must not pass. The two workflows are repaired
so they can fire.
CORRECTNESS. The encode DPB image barrier names the read and write accesses the
encode itself performs in its dstAccessMask. VkEncoderRenderFrame is deleted; the
class was present and had no caller.
BITSTREAM OUTPUTS ARE NOT SOURCE. The encoder's own default output names are
ignored, and only those, so a fixture a suite deliberately commits is still
visible to the repository.
CREDITS. The encode DPB read and write access mask came from Jose Lopes.
BUILD_ENCODER_COMPUTE_FILTER IS DOCUMENTED AS THE NO-OP IT IS. Removing the
duplicate add_definitions() made this the only place
VK_VIDEO_SAMPLES_COMPUTE_FILTER_SUPPORTED is DEFINED, but no source file in the
tree READS it -- grep returns the CMake file and Chromium's BUILD.gn and nothing
else -- so OFF removes a -D that no #ifdef consults and the filter is still
compiled and still used. The real control is the runtime flag
EncoderConfig::enablePreprocessComputeFilter, which no CLI or ext-API exposes. The
comment said otherwise, and cited a VK_ERROR_DEVICE_LOST on an unadvertised
input format as an A/B baseline to reproduce rather than as the defect it was;
it now says what the option actually does. The option is kept rather than deleted because wiring the macro up would let
an embedder with no GLSL compiler drop the filter's code entirely.
WHAT THE COMMENTS NO LONGER SAY. The narration this series had accumulated in
its own files -- that the branch filter "used to be unsatisfiable", that before
scripts/run_ctest_ci.sh existed CI ran no ctest so ~15 add_test() registrations
gated nothing, that earlier revisions of the stdio-latch comment described dumps
since deleted, that the compute-filter guard's audience is "no longer anyone" --
is recorded here instead. A file describes what the tree does now; the commit
carries how it got there. Concretely, and once, so it is not lost:
* Neither workflow had ever run: the branch filter named a branch this
repository does not have, so no push and no pull request ever matched --
including test.yml's `Run CTest suite` step, which invoked the script
correctly and was simply never reached. Both workflows now trigger on the
default branch and on pull requests targeting it, which gates a topic
branch through its pull request without enumerating branch names.
* There was no `ctest` anywhere under .github/workflows/, so every add_test()
in the tree registered and gated nothing. Adding a bare `ctest` was not
enough either: on a GPU-less runner the suite was 2 passed / 4 skipped /
8 failed, which is why the label split and run_ctest_ci.sh exist.
* The stdio-latch comment claimed the latch covered a CreateCodecConfig argv
dump and a positional-argument dump in ParseArguments; both dumps had
already been deleted from the tree.
* BUILD_ENCODER_COMPUTE_FILTER's comment claimed that removing a duplicate
add_definitions() made -DBUILD_ENCODER_COMPUTE_FILTER=OFF effective again,
and cited a VK_ERROR_DEVICE_LOST on an unadvertised input format as an A/B
baseline to reproduce. Neither held: no #ifdef reads the macro, and the
device loss was the defect.
* FFmpegDemuxer.cpp was listed twice in the decoder demo's source list --
unconditionally, and again under if(FFMPEG_AVAILABLE) -- so the guard could
never take effect. The file compiled with FFMPEG_AVAILABLE OFF and failed on
<libavformat/avformat.h> whenever the headers were absent. Only the
conditional append remains.
* The compute-filter guard was left undefined, which silently dropped the
filter from vk-video-enc-test as well as from embedders -- not what the
guard exists for. It defaults ON.
Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com>
Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>
Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
.settings/, .vscode/, .idea/ and *.code-workspace are machine-local editor state, not source. Nothing in the tree reads them and they differ per developer, so an accidental `git add -A` is all it takes to commit one. Ignore them, so that stops being possible. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…losing the device Without the preprocess compute filter the input image IS the encode source, so ANY requested format the driver does not advertise for the profile cannot be encoded -- there is nothing left to convert it. The unsupported set is a property of the driver and the profile, not of a particular layout or subsampling; the filter is what makes the rest reachable. The condition is reachable by default rather than only on an unusual request: EncoderConfig::input.vkFormat defaults to a format no driver advertises as an encode source. With the filter off, the format-negotiation fallback substituted the driver's first advertised format while frames kept arriving in the requested one, and the mismatch surfaced as VK_ERROR_DEVICE_LOST and a 0-byte bitstream -- a hung queue several hundred lines from its cause, and a whole configuration that could not be run. InitEncoder now refuses that combination with VK_ERROR_FORMAT_NOT_SUPPORTED and prints the formats the driver does advertise, so the caller learns both that the request was refused and what to ask for instead. Not supporting a format is a legitimate configuration; losing the device over it is not. The guard is conditioned on the preprocess filter being disabled, which nothing in this tree does, so no existing path changes. Running without the filter becomes a supported configuration, and the BUILD_ENCODER_COMPUTE_FILTER comment says so. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Two rules the view factory was not applying, both of which a caller can now trip because the encoder accepts images it did not allocate. A view may only be created over an image whose usage includes at least one view-compatible bit (VUID-VkImageViewCreateInfo-image-04441). TRANSFER_SRC and TRANSFER_DST are not among them, so a transfer-only image -- a staging copy handed in by a caller, consumed entirely through its raw VkImage by vkCmdCopyImage and barriers -- must get no view rather than an invalid one. A per-plane view reinterprets the image as a different format (R8, R8G8), which is only legal if the image was created VK_IMAGE_CREATE_MUTABLE_FORMAT. That always holds for images this library allocates; for one imported from a dma-buf, where the EXPORTER chose the create flags, it usually does not. Request plane views only when the flag is present and keep the combined view, which is all the encode path reads. Foundation for the encoder-ext work: every later topic that imports or exports an image depends on these two guards being in the view factory rather than at each call site. EXISTENCE IS A PROPERTY OF THE IMAGE, NOT OF ITS VIEW, and the decoder's frame buffer is corrected here rather than separately, because the guard above is what makes the distinction reachable. The decoder allocates its linear output image with VK_IMAGE_USAGE_TRANSFER_DST_BIT alone -- vkCmdCopyImage and image barriers take the raw VkImage and it needs nothing else -- so under the rule above it correctly gets no view. NvPerFrameDecodeResources::ImageExist() tested the VIEW, so that image reports itself non-existent for good. Nothing between there and the copy says so. GetImageSetNewLayout returns VK_SUCCESS with validImage false under an assert(), the frame buffer's overload calls CreateImage, re-checks, still gets false and returns VK_SUCCESS anyway, and GetCurrentImageResourceByIndex returns the slot index, which its caller reads as success. Each of those reports failure only through an assert that NDEBUG removes, so in a Release build the zeroed PictureResourceInfo reaches vkCmdCopyImage as a NULL VkImage with layout UNDEFINED and the driver faults on it. Track the image resource alongside the view, key ImageExist() on the image, source the picture resource's image and format from it, and bind an image view only when one exists. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Two things the device context owes an embedded host. QUEUE FAMILIES. When the caller supplies the VkDevice, the families it actually created queues on are not discoverable from VkPhysicalDevice properties, and vkGetDeviceQueue on a family the device was not created with is undefined. Take the caller's encode and compute family indices and bind those instead of the probe's picks; UINT32_MAX keeps the probed family, which is the library-owned-device path. STDIO. A library linked into a browser's GPU process must not write to stdout or stderr. The gated-logging latch behind VkEncOut()/VkEncErr() is header-only, but its one process-wide definition has to live somewhere single: VkCodecUtils is the bottom-most target every encoder translation unit already links, so it lives here. Keeping it out of the header is an ODR requirement, not a style choice. Depends on nothing above it; the encoder core and the ext implementation both build on this. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
STORAGE on a multi-planar format is a per-plane-view usage, not an image usage. Neither VK_FORMAT_G8_B8_R8_3PLANE_420_UNORM nor VK_FORMAT_G8_B8R8_2PLANE_420_UNORM advertises VK_FORMAT_FEATURE_2_STORAGE_IMAGE_BIT in either tiling, while the single-plane formats their planes are viewed as -- R8, R8G8 -- advertise it in both. So naming VK_IMAGE_USAGE_STORAGE_BIT on the YCbCr format without VK_IMAGE_CREATE_EXTENDED_USAGE_BIT is validated against the image format and refused, even though every view that will actually be created is legal. Set EXTENDED_USAGE so the check applies to the plane views that carry the storage access, and keep the pool's format bookkeeping in step. Sits on the view rules from the previous topic: it is the allocation half of the same question those answer at view-creation time. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
BuildGlslShader() returns VK_NULL_HANDLE when a compile fails, and both the vertex and fragment paths passed the result straight on. An unguarded null reaches VkPipelineShaderStageCreateInfo::module, which is where it becomes the driver's problem rather than a reportable error. Check both, and return VK_ERROR_INVALID_SHADER_NV. The fragment check sits BEFORE the cache swap, which is the part worth stating. The rebuild is skipped when the generated source is unchanged (m_fssCache.str() != imageFss.str()), so caching the source that just failed to compile would make it the current fragment shader and that test would skip the rebuild from then on -- one bad compile made permanent for the life of the object. Returning first leaves the cache holding the last source that did compile, so a later call retries. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The generated GLSL and the descriptors it is bound with had drifted apart in several places, each of which is a validation error or a wrong pixel. Configure()'s result was being dropped -- the next statement overwrote it, so a failed configure carried on as success. Chroma was fetched at a hardcoded srcPos/2, which halves both axes. That is right for 4:2:0 and wrong for 4:2:2 (chroma is full height) and for 4:4:4 (no subsampling at all), so chroma came from the wrong rows. The ratios now come from the same YcbcrVkFormatInfo the plane arithmetic uses, so they cannot drift from the descriptors generated beside them. Descriptor selection is a preference rather than a command: with autoSelect the layout picks push descriptors when VK_KHR_push_descriptor is present, instead of an unconditional PUSH_DESCRIPTOR_BIT that discarded the descriptor-buffer arm outright. Needs the view rules from the earlier topic, because which view a binding gets is what decides how its shader must declare it. THE YCBCR2RGBA GOLDEN IS RE-RECORDED HERE, with the change that invalidates it, so the decode suite does not go red in between. The output store above moves from ivec3(pos, pushConstants.dstLayer) to a plain ivec2. The binding it writes is the combined view, and VkImageResourceView types that view VK_IMAGE_VIEW_TYPE_2D whenever layerCount is 1, which every caller of this filter pins it to; storing to a layer of a non-array view is undefined, and the driver returns plausible-looking data rather than failing, so the recorded golden captured that result instead of a failure. The new output is the better one, measured rather than assumed. Ten of thirty frames differ and the rest are byte-identical, and on every one of those ten the new output is closer to the same clip decoded without the filter -- frame 1 psnr_y 9.96 dB against 3.32 dB, and 5 to 6 dB better across frames 1 to 10. The stray layer index is what those frames were carrying. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…alise Fix for d535085 ("filter: packed 4:4:4 and 4:2:2 support, and three conversion defects"), which added m_inputPackedYcbcr and m_outputPackedYcbcr to the class and appended them to the end of the constructor's initialiser list rather than at their declared position. A member is initialised in the order it is DECLARED, never in the order the initialiser list writes it. Six members disagreed: the two packed-YCbCr pointers and the two block ratios initialise before m_ycbcrPrimariesConstants, and all four before the bitfields the list placed them after. So the list described an order the compiler did not follow, and a reader checking what runs first had to reconstruct it from the declarations instead. Nothing is misinitialised today, because every initialiser here reads a constructor parameter and none reads another member. That is the whole margin: the first initialiser that reaches for a sibling would read it before it was set, and the list would look correct while it happened. The list is reordered, not the declarations, so the class layout is untouched and no member changes its position or its value. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
A host that drives the encoder frame by frame needs per-frame facts back -- what was captured for a frame, keyed by the id it submitted rather than by an internal slot number, which is the only handle the caller has. The probe owns that mapping and its lifetime: a frame's capture stays retrievable until the caller collects it, independent of when the encoder recycles the image the frame was carried in. It lands before the encoder core because VkVideoEncoder.h includes it and holds a FrameCapture by value; the core cannot compile without it. Wired into the encoder library's source list in the same commit so the build and the code arrive together. TWO DESIGN CHOICES THE HEADER STATES AS RULES, recorded here as the measurements that forced them. The arm parameter is a scoped enum rather than a bool because the sense is inverted from the obvious spelling: `bool directlyEncodable` true means "do not arm", CaptureSite's true-ish value means "DO arm". A bool would have let every existing call site keep compiling with its meaning reversed -- arming what should be refused and refusing what should be armed -- which is the same class of silent breakage as the pNext clobber that made this probe inert to begin with. IsProbeableFormat() is public because the ARM decision has to ask it. Consulted only from RecordCapture, on the first frame, a P010 registration echoed ARMED at import and was downgraded to NOT_APPLICABLE a frame later. THE MECHANISM IS DOCUMENTED alongside it, under vk_video_encoder/docs: the arm/capture/score/echo sequence and the two command-buffer sites it rides, the dead-plane predicate and the Q8 constant the ext layer static_asserts against, the row stride and what a byte-wise walk costs instead, and why a verdict is latched per registration rather than per frame. Markdown, a sequence and state diagram, and the rendered page. It gives armedRegistrationCount a section of its own, because that field is the one a reader gets wrong: probed=0 damaged=0 is what a session reports when every buffer was clean AND what it reports when nothing was ever measured, and only the armed count separates them. The document also records the reachability decision rather than leaving it to be inferred. The probe is internal: its structures are declared in the descriptor API's internal header, which is on no consumer's include path, its consumers are this library's own white-box tests, and the public encoder interface does not expose it and is not intended to. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The parts of the encode path that a host exercises and the demo does not, separated from the session object itself because none of them touch the interface the ext layer binds to -- they compile and link against the existing one. Config gains the per-codec derivations a caller can reach directly rather than through the demo's single fixed path. The H.264 DPB and the GOP structure each carried state that was only ever read back in the order the demo produced it; both now survive a caller that reconfigures mid-stream, submits frames out of order, or stops early. PSNR belongs to this topic by subject but not by dependency: it reads fields that the next commit adds to the session, so it cannot compile until that lands and it travels with it instead. Splitting on what builds rather than on what reads well is deliberate -- a commit that does not compile is not a reviewable unit. Kept ahead of the session change so that the reviewable, self-contained half of the core work is not buried inside the atomic commit that follows. FINALIZECONFIG() TAKES AN OPTIONAL DEVICE CAPABILITY ARGUMENT, because some of what it decides is a property of the hardware and not of the command line. The function runs on two paths: argv parsing, which happens before any device exists, and an embedding host, which has already probed one. A pointer that may be null carries that difference without splitting the tail in two -- null keeps the command-line answer, which is what the demo gets, and a caller that has probed reaches a device-correct answer through the same function. A second FinalizeAgainstDevice() would be a call every caller has to remember. One decision uses it. Intra refresh is a device feature (VkPhysicalDeviceVideoEncodeIntraRefreshFeaturesKHR::videoEncodeIntraRefresh), so a configuration asking for it on a device that lacks it is refused here, where the request is still attributable, rather than inside session creation. The minQp default is deliberately NOT a second one. Every consumer of that field reads it only under minQpSet, which stays clear precisely when the default applies, so an unset minQp reaches the driver as "no clamp" rather than as 20 -- checking the number against the device would decide nothing. A caller that wants a QP clamp sets one, and the codec configs check THAT against the device window and refuse it rather than narrowing a request silently. DeviceCapabilities carries only what it is asked, and defaults its members, so a field the caller does not assign reads as absent rather than as whatever the stack held. --GOPFRAMECOUNT IS PARSED AT THE WIDTH IT IS STORED. VkVideoGopStructure keeps the count in a uint32_t, and its constructor now agrees with the member and the setter. Parsing into a narrower local truncated silently, because parseUint casts without a range check: --gopFrameCount 300 reached the encoder as 44, and 256 as 0, which the codec configs read as ZERO_GOP_FRAME_COUNT and replaced with the device's preferred count. Neither is refused and neither is reported, so the stream is encoded to a GOP the caller never asked for. Zero remains legal: it is that sentinel, and requesting the device's preference is a request like any other. codecBlockAlignment is deliberately NOT among them. It holds the H.264 macroblock size for every codec, which is wrong for H.265 CTBs and AV1 superblocks -- but nothing in the tree reads it: one write here, one declaration, one constructor initialiser and no readers. Giving it a device-derived value would add correctness to a field no consumer checks, so it keeps its value and gains a comment saying that a reader must derive the real granularity from VkVideoCapabilitiesKHR::pictureAccessGranularity instead. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
parseUint and parseInt reported success for three classes of input whose
value the encoder then ran with, changed:
* NARROWING. static_cast<T> of a value too large for T keeps the low bits.
Into a uint8_t, 256 became 0 and 300 became 44. --temporalLayerCount 256
configured one layer's worth of nothing, and 0 is a meaningful sentinel
in several of these fields, so a truncation did not even have to land on
an obviously wrong number to be accepted.
* A NEGATIVE INTO AN UNSIGNED FIELD. strtoull takes a leading '-' and
wraps, so --gopFrameCount -5 was honoured as 4294967291: a GOP that never
closes, from a typo, silently.
-1 is different and is kept. It is the Video Codec SDK spelling for all
bits set -- an infinite GOP is UINT32_MAX, as the json_config README and
the idrPeriod derivation in EncoderConfigH264 both record -- so it is
accepted explicitly and yields the maximum value of the field's type,
exactly what the wrap produced. Command lines already using it are
unaffected. Every other negative is refused.
* OVERFLOW. Beyond the accumulator's range strtoull and strtoll saturate
and set ERANGE, which neither helper read.
All three are refused now, as an invalid parameter, at the point the argument
is parsed. The narrowing check -- the value must compare equal after a round
trip through T -- is what makes parsing straight into a uint8_t field safe
again, so the narrow locals need no widening: out-of-range input is rejected
rather than folded into range.
Nothing that was accepted and meant what it said changes. A caller who typed
a number the encoder could represent gets the same encode; a caller who typed
one it could not now gets told, instead of an encode configured from the
remainder.
--idrPeriod is read unsigned, which is how SetIdrPeriod stores it. Read as a
signed value it accepted every negative, and the conversion at the setter
turned each one into a period so large that no periodic IDR is emitted -- so
a mistyped -5 asked for an infinite IDR period and was given one. -1 keeps
its meaning there, on the same convention: all bits set, for a field of
whatever width. minQp and maxQp stay signed, because they are signed fields
whose -1 is their own "unset".
The verbose prints that reported these fields with %d are corrected to %u, with the promotion made explicit where the field is a uint8_t -- which is
why an infinite GOP logged itself as "-1" rather than as its value.
Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The API a host embeds this library through, its implementation, and the encoder session both are expressed in terms of. WHY THESE ARE ONE COMMIT. They were split three ways first and none of the three built. The ext API renames methods that the existing vulkan_video_encoder_ext.cpp declares `override`, so the header cannot land before the implementation. The session changes SetExternalInputFrame's signature, which that same file calls, so the core cannot land before it either. And the implementation is written against both, so it cannot land first. Separating them needs compatibility shims that exist in no other commit and would be deleted by the next one -- history that never described a state anyone had. They are one change; this records them as one. WHAT IT IS. The interface is shaped by what an embedding host cannot do. It cannot own a VkDevice, so the library adopts the caller's instance, physical device and device, and binds the queue families the caller actually created queues on. It has no visibility of internal slots, so every per-frame fact is keyed by the frame id the caller submitted. It cannot tolerate writes to stdio or to files, so bitstreams come back in memory and logging goes through the latch. It cannot stand up a session just to ask a question, so capabilities are answerable before one exists, and a codec whose encode extension the device does not advertise is refused at the point of asking rather than deep inside session creation. The session carries the corrections that discipline requires: rate control keeps its codec-specific struct and per-layer state in the BeginCoding chain, which VUID-vkCmdBeginVideoCodingKHR-pBeginInfo-08254 requires to match the state a control command established, pointing at encoder-owned storage rather than at a pool-recycled array. PSNR and the demo come along because they are compiled against the session interface this changes -- the demo target builds VkVideoEncoder.cpp from its own source list, so its build wiring moves in the same commit or the link breaks. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The AV1 arm of the encoder, and the AV1 parser it is exercised against. It lands directly on the interface commit, because its capture arm plugs into the in-memory capture funnel that commit introduces. Everything else it touches -- the AV1 parser, the AV1 codec config and the AV1 encoder -- already exists below this series, so nothing after this point has to reach backwards past the interface work to edit a parser or a codec config. IN-MEMORY CAPTURE. The AV1 encode arm captures its bitstream in memory and prepends a temporal delimiter OBU to each temporal unit. The driver emits none, so what a caller pulled out was a fragment that only a file writer would have made whole; with the prepend, every captured chunk is independently parseable. File-based output is unchanged. THE DIMENSIONS A LATER FRAME INHERITS. AV1 frame_size_with_refs() lets the current frame take its size from a reference, and SetupFrameSizeWithRefs reads decodeWidth, decodeHeight and decodeSuperResWidth back out of the referenced picture. Nothing on the AV1 path ever wrote them -- the only writers were on the VP9 path -- so on a pure AV1 stream all three were read while indeterminate. They are published in BeginPicture, which end_of_picture reaches only after the whole frame header is parsed, so all three values are final. Deliberately NOT in UpdateFramePointers, which looks like the natural home but also runs on the show_existing_frame path with a previously decoded picture, where writing would replace a correct stored dimension with an unrelated one -- turning a read of indeterminate memory into corruption of good data. And outside the allocation guard rather than inside it, so a retained picture does not keep a stale size. The narrow branch is where this becomes visible; it is not the wide exposure. The Vulkan reference codedExtent is built from decodeWidth and decodeHeight for every AV1 reference slot on every frame that has references, with no frame-size-override gate. vkPicBuffBase's constructor now names its VkPicIf base, so the base subobject is value-initialised rather than default-initialised. With the parser writing all three fields before they are read this is belt and braces; what it buys is that a MISSING write reads as zero rather than as recycled pool storage, which is the difference between a defect that reproduces and one that appears only under memory pressure. ORDER_HINT_BITS IS 8, NOT 7. The reference-order-hint writer casts to uint8_t while the DPB writer masks by this value, so at 7 the two disagree and a ref_order_hint of 128 or more can be emitted for a field that cannot hold it. 8 is also what a consumer masking order hints with 0xFF expects: at 7 the order hint wraps at frame 128 and such a consumer rejects every frame after it. The mask at the encoder's writer is taken from the advertised width rather than from the cast, so the two cannot drift apart again. QUANTIZER DEFAULTS. An unset AV1 quantizer index is left to the AV1 encoder config's own driver-preferred triple instead of inheriting an H.26x QP. The two scales are not the same scale: an ordinary H.26x QP read as an AV1 qindex is near lossless, so the encoder silently produces enormous frames while every counter reports success. The code no longer pretends the scales are shared. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
…port The filter suite ran without judging a frame: it dispatched, reported success and asserted nothing about the pixels. It now validates its output, and --list is registered as a device-free test so the case table is proven to construct on every push, while the two dispatching arms are labelled gpu and exit 77 so a GPU-less host reports SKIPPED rather than FAILED. Every other init failure still exits 1, so a real regression cannot hide in the skip arm. Adds linux_dmabuf_import, the counterpart of win32_opaque_import: it proves an image exported by one VkDevice can be imported AND USED by a different VkDevice on the same GPU. That is the property the browser's zero-copy path depends on and nothing in this tree previously tested. Depends on the shared runtime topics it exercises -- the view rules, the image pool's storage-usage handling and the filter's descriptors. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Twelve suites over the ext surface, each one written against a property a host depends on and none of which the demo's single open-push-close path reaches. They are split by what they need, because a suite that cannot tell "no GPU here" from "this is broken" teaches people to ignore it. The device-free ones -- the chained-descriptor walk, the input-format taxonomy, the config binder's disposition of enablePreprocessFilter -- are pure functions of their arguments and belong in the gating set that runs on any host. The real-device ones declare SKIP_RETURN_CODE 77 and report SKIPPED, not FAILED, on a machine with no encode-capable GPU. Several are regression tests for defects that are now fixed and would otherwise be invisible: the acquire fd's ownership on every refusal path, the release fence's export, and DrainPendingFrames() permanently disabling async assembly so that every frame submitted afterwards was encoded and then never became acquirable. The registration in the two CMakeLists files lands here rather than in a build topic of its own: add_subdirectory of a directory that does not exist yet fails at configure time, so the tests and the lines that run them cannot be separated. qfot_foreign_release_repro.c IS NOT BUILT, AND THAT IS DELIBERATE. It is a driver-defect reproducer rather than a test: it asserts nothing about this library, and a suite that failed when the driver was wrong would report this library as broken. No CMakeLists references it, so it compiles only when someone builds it by hand against the driver under investigation. The cost of that is real and worth stating: nothing keeps it compiling as the surfaces it calls move, so it is evidence with a shelf life, not a gate. THE FOUR PROFILE CASES ARE DISPATCHED. encoder-ext-filter defined CaseProfileNumbersReachTheCodecConfigUnchanged, CaseProfileDefaultIsDerivedPerCodec, CaseProfileNumbersAreReadAgainstTheCodec and CaseProfileMustAdmitTheInputDepth and called none of them, so nine checks were compiled, never run, and reported as part of a passing suite. A case that is written but not dispatched is worse than one that does not exist: the file reads as though the claim is covered. They pass -- the suite runs 159 checks, 0 failures -- so what was missing was the four call sites and nothing else. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…e it Defects in the encoder itself that became reachable, or became visible, only once a host drove the encoder in process and read its status instead of reading a file afterwards. A FRAME THAT DID NOT ENCODE FAILS ITS CALLER. A frame the encoder could not record or submit now fails, and an encode that produced no bitstream is reported as a failure rather than as a success with an empty payload. A command that asked for no bitstream is not judged on one, so the two cases stay distinguishable. THE B-FRAME CLAMP IS EFFECTIVE. The consecutive-B-frame clamp was computed and then not applied, so a GOP structure could ask for more B frames than the encoder was configured to carry. The assembly queue is sized to the burst the GOP can present rather than to a fixed depth, so a legal structure cannot overrun it. ONE DPB SLOT, BOUND ONCE. The H.264 and H.265 encode commands bind each DPB slot at most once when a reference picture holds a position in both reference lists. Binding it twice is a validation error, and the second binding is not what the codec means. Both arms report their validation output, and the arm that codes B frames runs under H.265 as well as H.264. H.264 BASELINE MUST NOT EMIT CABAC. Baseline does not carry CABAC, so a Baseline session that enables it produces a stream naming a profile it does not conform to. The entropy mode is decided from the profile, and the suite fails if someone later widens the rule to fire on requests and assumes it subsumes the clamp. THE DEBUG SWITCHES ARE READ ONCE. The encoder read its environment switches once per frame. They are read once, so the cost is off the frame path and a switch cannot change meaning mid-session. THE SESSION CREATE RESULT IS CHECKED IN RELEASE BUILDS. It was checked only under assert, so a release build carried a null session past the point of failure and died later inside the driver. GOP SEQUENCING IS TESTED DIRECTLY. The GOP machine decides encode order, B-frame counts and B-frame positions before any picture exists, and nothing else in this tree asserted those tables -- the sibling suites set a GOP length and then encode, which cannot tell a correct sequence from a plausible one. It is carried as its own subdirectory because it compiles the GOP source directly and links nothing else. A round-trip quality test is registered beside it: it encodes, decodes out of process and compares per-frame PSNR against the source, which is the only round-trip measurement here, because the encoder's own PSNR option compares the input against the RECONSTRUCTED picture, inside the encoder, and so cannot see a stream that decodes wrongly. It exits with the skip code when the encoder binary or either tool is missing, which is the skip its label requires. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
What a caller declares about the colour of its input, what the library signals in the bitstream, and the HDR10 static metadata it carries there. THE COLOUR MODEL IS DECLARED, NOT INFERRED. An input's colour model is a property of the surface, not of its Vulkan format: the packed 4:4:4 layouts are carried on RGBA-typed enumerants, so inferring the model from the format reads a Y-Cb-Cr surface as RGB. The config declares it, and what was a boolean asking whether the input is RGBA becomes a colour-space enum that can name the cases a boolean cannot. The registration gate asks its class question under the model the descriptor declares and falls back to the session's model only where the descriptor declares none, so a descriptor is judged as the surface it says it is. A colour model the format cannot carry is refused at registration, and both the refusal and the session binder name the colour model rather than the format. The model a session is declared in is immutable across Reconfigure, compared as it resolves against the input format. REFUSE OR COERCE, NEVER DIVERGE. A colour description the encoder cannot signal is either refused at the boundary or coerced to a value it can signal, and the coerced value is what the bitstream carries. What it must never do is accept the caller's description, signal something else, and leave the two disagreeing with nothing to say which is true. The signalled colour matrix is the matrix the preprocess filter applies, and the range is carried through to the filter. THE TRANSFER FUNCTION IS DECLARED. The library decides the preprocess conversion; the transfer function is the caller's declaration and is signalled as such, not derived from the conversion the library happened to choose. THE INPUT'S OWN COLOUR IS DECLARED, SEPARATELY FROM THE STREAM'S. The only input-side colour field was the transfer function, so the RGBA-to-Y'CbCr filter derived its conversion matrix from the primaries the BITSTREAM advertises -- an output field standing in for an input property. That is sound only while no primaries conversion exists, and nothing said so. A chained input-colour structure completes the declaration: the caller states what its buffer carries, the library states what it did, and neither states the other's business. An absent chain is undeclared on every axis at once and reproduces the previous behaviour exactly, so a caller that says nothing is unaffected. A declared input colour that disagrees with the bitstream's on an axis this library cannot convert is refused with the reason, which is the rule the transfer axis already applied. The comparison reads the caller's own values rather than the bound config's, because the binder rewrites an unsupplied output field to Unspecified and reading that fill back would refuse a config nobody contradicted. WHAT THE DECLARATION MEANS IS STATED PER LANE AND PER CODEC, because two things needed saying that no document said. The range axis is the one that had a field and no reader: a caller that chained a full-range declaration got a narrow-range conversion, a zero range flag and success, which is worse than offering nothing. The header stated a correct contract and the code failed it, so the code moves. On the Y'CbCr lane no range scaling happens between the caller's samples and the coded ones, and the copy filter takes its output range from its input range, so the input's range IS the stream's: the declaration is applied, and it raises the video signal type beside the flag, because an unsignalled range is normatively inferred as studio swing and reported as unknown at a decoder's API boundary. On the RGB lane the filter PRODUCES the output range from that same flag and samples its source over the full range, so a declared limited input describes a conversion that does not happen and is refused with the reason -- the rule the primaries, matrix and transfer axes already apply. Exactly one contradiction is refused, because the flag has no undeclared state and reading its false as a positive claim would refuse every caller that left it alone. AV1 has no absent state for range at all: the range is unconditional syntax, so the library commits to a value on every AV1 stream whether one was declared or not. The structure is additive -- no export, no virtual, no change to the config's own layout -- and it carries its own size and chain-prefix assertions, because those assertions are the only compatibility enforcement this interface has. THE INPUT'S IDENTITY IS DERIVED TWICE, AND THE TWO MUST AGREE. The binder derives chroma subsampling, bit depth and plane count from the caller's format; the input-parameter verification reconstructs a format from those same three. Two derivations of one quantity in opposite directions, with nothing saying they had to agree -- and whatever the reverse one produced was what the session was configured with and what the compute filter was built from. The reverse derivation cannot simply be deleted, which is why this is an assertion rather than a removal: the packed-alias arm leaves the format unwritten on purpose so the reconstruction supplies it, and that is the only route by which either packed 4:4:4 layout is nameable. One comparison covers both lanes -- a carry-through on the RGBA lane, where the reverse derivation spells no RGB layout, and a round trip on the Y'CbCr lanes -- and the proposition is the same either way. A disagreement is refused and not repaired: overwriting one side with the other picks a winner between two derivations without knowing which was wrong, and the wrong choice is a session configured for a picture the caller is not sending. HDR10 STATIC METADATA. The mastering display and content light level are carried as an H.265 SEI message and as AV1 metadata OBUs, from one payload builder shared by both arms. The metadata struct chains onto the config, so a caller that does not set it is unaffected, and a caller that chains it against a library which does not know the structure type is refused loudly by the unknown-type gate rather than silently ignored. THE ENCODE-SOURCE REQUEST IS A REQUEST ONLY ON THE DIRECT LANE. On the filter lane the declared input format describes the FILTER's input -- a three-plane layout, a depth-converted source, an RGBA surface -- and no device advertises those as encode sources, because they are the formats that exist to be converted. Matching against them fails by construction, so every healthy filtered session announced a chroma mismatch it did not have. The gate is the preprocess-filter flag rather than the colour model: the colour model caught only the RGB half of the filter lane and let its Y'CbCr half fall through. The case that survives is stated as what it is -- a layout substitution on a lane with no conversion, where the encoder reads the caller's bytes under a layout the caller did not declare -- rather than as a chroma claim the comparison cannot support. Whether that case should be a refusal rather than a diagnostic is a behavioural question and is left open. WHAT THE LIBRARY ROUTES IS A PROPERTY OF THE DECLARED PAIR. A device format-feature bit is not an answer about this library: it accepts formats the device does not advertise as encode inputs, because the preprocess filter converts them, and it does not accept every format the device does advertise. The advertised list is derived from the library's own taxonomy and is narrower than the accepted set, and the subsampling rule is stated only for the Y-Cb-Cr inputs it holds for. THE TAXONOMY IS TRIMMED TO ITS CONTRACT. Semi-planar 4:4:4 is directly encodable, and the direct set the header names is the set the classifier calls direct. Twelve-bit semi-planar input is encodable through the filter and not directly: the encode-source formats a device reports are the 8- and 10-bit semi-planar set, so a 12-bit registration routed to the direct path asked the encoder for a source format it never advertises. Its plane layout and subsampling already match; what has to be converted is the bit depth, which is the compute tier's work. That exposed a second assumption -- the plane count was answered from the class, which can describe neither RGBA nor 12-bit semi-planar -- so the plane count is read from the format's own layout description, which is what drives the input geometry and the staging pool's shape. All of it is here rather than with the capability work, because the classifier these rules live in is the one this commit gives a colour-model parameter: the routing decision and the model that parameterises it are the same function. THE MODIFIER ANSWER IS PROFILE-BLIND, AND THAT IS A LIMIT ON WHAT IT MEANS. It comes from a format-properties query with no video-profile term at all, unlike the input-format surfaces, which are keyed on codec and profile. A modifier reported as carrying the encode-input feature is not thereby usable for the profile a session will negotiate: it says the format supports the feature under that modifier, not that this codec and profile accept that format. The two answers are conjoined, and the header says so. THE CODEC ARMS' GEOMETRY TERMS ARE CORRECTED WHERE THEY WERE UNREACHABLE. Three derivations that are defined over the configured geometry did not read it. The H.265 CPB VCL factor assigned a chroma subsampling flag into a variable named for a chroma format index and tested it against an index value, so no real input could satisfy the 4:4:4 arm and every stream took the 4:2:0 factor. ITU-T H.265 Table A.8 gives 4:4:4 twice that, so the level came out at a higher tier than the profile needs and the default buffer smaller. The comment on the line stated the intent the code broke; the code now matches it, and the conversion is the one every other site already uses. The table's depth term is a second defect on the same function; it is recorded and not repaired. The AV1 sequence header's subsampling was hardcoded to 4:2:0 under a comment calling that the only chroma format this encoder admits, while the profile derivation selects a 4:4:4 profile that requires both subsampling terms to be zero. The derivation is the codec-correct side, so the comment was the stale one, and the header is written from the configured chroma. The narration in the H.264 arm is sorted rather than deleted. One block describes a configuration that cannot occur, because the field it branches on has no setter on any surface and one constructor default; it is kept because it states the intent, and deleting it would delete the record that the intent is unreachable. A second note records that the arm reads the input side of the geometry where it should read the encode side -- inert while one writer makes the two equal, and a defect the moment anything makes them differ. THE DISPOSITION TABLE DROPS THE ENUMERANT WITH NO ROW. One disposition named a class the config surface does not contain -- a value refused at initialisation because honouring it is impossible in this build -- and carried a single definition and no row anywhere in the table. Every refusal this surface makes is either a declaration that cannot be honoured or a versioning gate, and both are properties of the config rather than of the build, so the enumerant is deleted rather than given a member to justify it: inventing a member to justify an enumerant is how a table stops describing the thing it tables. Each remaining row names the full extent of what its field binds, so a row that reads as one derived value cannot hide three. THE COVERAGE JUDGES IN Y-Cb-Cr. The encode suite compares in the colour space the encoder wrote, and compares the HDR payload numerically rather than checking that a message of the right type is merely present. The filter suite's input check pins the order of its two single-plane arms, so the arm a failure names is the arm that ran. Where a comment cites the code that produces a value it describes, it names the symbol rather than a line range, because a line range in a comment eventually points at neither the write site nor the declaration. Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com> Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
What the library tells a caller it can accept, and the classification behind that answer. TWO ENTRY POINTS, NOT FOUR. Capability enumeration is two entry points rather than four near-duplicates, and both go through one implementation. The enumeration states what it probes, what it answers and what a context caches, and the context notes say what a caller can share, index and rely on. The advertised formats and the video std syntax flags leave the capability structure for two-call enumerators on the context, so the structure carries scalars only and the format list is computed rather than stored. Each enumerator leaves a defined count behind on every return but a null count pointer -- the number written on success and zero on any error -- so one wrapper can read it without knowing which of them it called. EACH ADVERTISED FORMAT NAMES ITS OWN ROUTE. An entry states whether it is the format the encoder reads, as an optimality answer, and the documentation of that answer names no conversion mechanism: directness and encode format are different questions, and the route an image takes is an implementation choice rather than part of the interface -- so the route enumeration moves to the internal header, beside the one declaration that has ever had a member of that type. The optimality answer names the filter that makes a suboptimal entry encodable, which is the half of the rule that says why the entry is on the list. A build without the preprocess filter advertises no converted format, instead of naming formats its own session binder then refuses. Driver-workaround verdicts and input-routing observables are internal, and are named nowhere on the interface; the status echo's chain carries library-internal observation structures, so a caller built against the public header alone has nothing to chain there. WHAT ROUTES WHERE IS DERIVED, NOT LISTED. The packed 4:4:4 layouts route wherever the declared colour model names them, and the routable set is derived from the format tables the compute filter itself reads rather than listed beside them where the two can drift. THE PROBE ANSWERS BEFORE A SESSION EXISTS. The capability probe is keyed on bit depth as well as profile, and states its chroma subsampling per profile instead of hardcoding it at the call. An explicitly named profile is checked against the input's chroma subsampling as well as its bit depth, so the pair a caller may feed the library is answerable before a session exists. THE ADVERTISED LIST DESCRIBES THIS DEVICE. Two surfaces answered the same question differently. The enumerator built its list from a capability snapshot probed at one subsampling and two depths; the point query answered from the caller's own bound configuration. The two disagreed on the encode format -- the field that decides the chroma subsampling and bit depth of the bitstream -- for identical arguments on the same device in the same process. Both now resolve through one function. The enumerator binds each candidate through the real binder to obtain the profile, subsampling and depth the session would use, issues a live device format query at those values, and admits the candidate only if the encoder-input format it would be handed is in that list. The snapshot survives, unchanged, as the source for the capability scalars, and no longer carries a format list for nobody to read. What that makes visible is capability the library already exercised: semi-planar 4:4:4 at both depths reaches the advertised lists, and ten-bit 4:2:0 reaches the lists it was accepted at and advertised at neither. Nothing is advertised that the point query would refuse, because the two compute the same answer. The ordering clause is restated: with each candidate queried at its own derived profile there is no single device list left to order by. The list builder keeps its ordering, uniqueness, capacity and compute-filter rules and takes an admission callback in place of a device list, so the production caller supplies the live resolver and a device-free test supplies a synthetic one. An enumeration silences the binder's per-candidate refusal diagnostics for its own duration and restores the caller's setting after: offering every routable format to a profile most of them cannot reach is what the sweep is for, so a refusal there is the answer rather than an error. THE BIND SET IS THE STANDARD'S OWN LIMITS TABLE. The profile derivation selects H.264 High 4:4:4 Predictive, H.265 Range Extensions and AV1 High from 4:4:4 input, and the library emits those streams -- yet naming the same numbers explicitly was refused as unbindable. The refusal machinery for them already existed, with a limits table stating the standard's own rule for exactly those numbers and no way to reach it. The bind set now covers every number that table states, the standard's limits refuse what the standard forbids, and the device refuses the rest -- which a caller can discover before a session exists. The header's constant block names what the binder binds rather than a subset of it, so the two cannot drift apart again. The assertions that recorded the old behaviour are inverted with it, and the one that named three numbers as unbindable is replaced rather than inverted, because two of the three are bound now and the third would have gone on passing while its stated reason had become false. ADVERTISED IS ACCEPTED. The list had become device-truthful; acceptance had not. Initialisation settled the library half of the question -- does this library route this pair, can the named profile carry it -- and asked the device nothing, so a format the device does not encode was bound, given a session, and refused by the driver at the capability query. That refusal names neither the format nor its chroma subsampling, and the library relayed it as a bare result code. The 4:2:2 and twelve-bit families are exactly that shape: routed by the taxonomy, bound by the binder, on no advertised list, and refused late by a message a caller cannot act on. So the accepted set was wider than the advertised set and a caller had to consult two surfaces to learn what it could hand in. Initialisation now gates acceptance on the same device resolve the two query surfaces answer from. The device half of that resolve is one function: the point query calls it for the pair a caller names, the enumerator calls it once per routable candidate, and initialisation calls it for the pair the session declared. Three surfaces, one statement of the rule, so a query that promises what initialisation refuses cannot be written without changing that function. The refusal carries what the driver's does not -- the format, its chroma subsampling, its bit depth, its plane count, the profile the configuration derived, and which of four reasons applies -- and it is an enumerated verdict rather than a result code for one reason: a query owes its caller a yes or a no, an initialisation boundary owes it a reason. The result code a caller sees is the point query's own, because one verdict with two spellings would put a caller back to asking which surface it was talking to. The one axis on which absence from the list is still not a refusal is the colour model, which a list of formats cannot express, because the packed 4:4:4 layouts ride RGBA enumerants. That exception is stated as the single exception it is, with the point query named as the surface that closes it. AND A SUBOPTIMAL ENTRY IS TAKEN ONLY ON THE REGISTERED LANE. The conversion runs against a registered resource, so a suboptimal format is handed in through registration followed by a registered submit; the unregistered submit refuses it, because that path has no registration to attach a conversion to. An optimal entry is accepted on either lane. That is a qualification of the membership rule rather than an exception to it -- the format is still one the session may declare; what is restricted is which submit carries it. BOTH LIVE ACCESSORS ARE NAMED AS TWO. The context header said the modifier query was the one accessor that reads the driver rather than the snapshot. Moving the input-format enumeration onto the same live resolver is what made that false, and it is the same change that makes it necessary, so a rationale outlived the mechanism it described. Both live accessors are named, each with the reason it has to be live, and the immutability claim the paragraph exists for is stated precisely: neither writes context state. The enumerator's two-call idiom had also borrowed the sibling's snapshot rule -- that the count is a property of an immutable snapshot, so the counting call and the fetching call cannot disagree. The conclusion holds and that reason does not, because this list is recomputed on every call, so the rule is restated on its own footing: the resolve is a deterministic function of context, device, codec and profile, and a context's devices do not change for its lifetime, which is the premise the snapshot itself rests on. The conclusion is kept, so it is pinned -- counted twice and fetched twice per publishable key, with all four answers required to agree entry for entry, ordering included, and the counts shown to differ across keys first, because agreement asserts nothing coming from an observable that reads the same for everything. The field table's disposition for the bitstream range request is completed in the same pass: it is still bound to the VUI flag and it is now also validated, and a table whose stated purpose is to say what happens to each field has to say both. THE BINARY THAT CREATES SESSIONS IS GATED ON VALIDATION. The input-format query tests declined a validation gate on a stated ground: nothing in the file created a session, an image or a command buffer, so there was no spec-violation surface to observe. That was a property of the file and it stopped holding when the initialisation gate arrived, which drives one initialisation per advertised entry and so runs session creation, device format selection, the filter build and the image-pool allocation. All rows are gated rather than only the one that creates sessions: the gate reads the layer output alongside the exit code, so it adds a way to fail and removes none, and a row with no spec-violation surface gates at zero and stays there. The objection the old note raised -- that declaring a gate advertises coverage a test does not have -- is answered by recording a run with no layer as unmeasured rather than as clean, instead of by declining to measure. Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com> THE 4:2:2 EXPECTATIONS ARE DERIVED FROM THE DEVICE, not hardcoded, because the two suites this commit adds would otherwise be born failing. EncoderExtInputFormatPointQuery and EncoderExtInputFormatCrossRoute both asserted that the device REFUSES 4:2:2 input. That is not a library rule but a per-device capability: Blackwell and later encode 4:2:2, earlier parts do not. On a part that has it, the point query fails with "got VK_SUCCESS, wanted VK_ERROR_FORMAT_NOT_SUPPORTED" and the cross-route control with "NV16 (4:2:2) is not advertised: it was" -- a hardware CAPABILITY reported as a library DEFECT, red on the first machine that runs it. Ask the device instead. The point-query table derives the expected result from VkEncEnumerateInputFormats and, where 4:2:2 is supported, takes the expected encodeFormat and optimality from the advertised entry, so the point query is still held to agreeing with the list rather than to a constant. The cross-route refusal control keeps its teeth. 12-bit is refused by every part this library targets at every key, so P012 stays an absolute that a merely permissive enumerator still fails. NV16 stays absolute for every NAMED profile, all of which are 4:2:0 only. Only DEFAULT, which derives the profile from the input's own subsampling, becomes a derived check: the list must contain NV16 exactly when the point query accepts it. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Every codec arm's profile, level and sequence-header syntax is defined by the standards over the values the BITSTREAM carries, and each derivation is read from the encode side rather than from the input side. WHICH SIDE IS RIGHT IS NOT A PREFERENCE. The configuration keeps the input's chroma subsampling and bit depth apart from the encode side's, and the two sets exist precisely so an encode value can differ from the input value: a chroma resampler or a device-driven depth downgrade is what would make them. Today one writer sets the encode side from the input side and nothing else touches either, so a read of the wrong one costs nothing -- and the three codec arms read different ones. H.264 read the input on both axes, so on the first change that made the two differ it would have picked a 4:4:4 profile from a 4:4:4 input and written a 4:2:0 chroma format from the encode value, inside one sequence parameter set. H.265 read the encode value for chroma and the input for depth, so the two halves of one derivation sat on opposite sides of the boundary. AV1 read the encode value for chroma and the input for depth, and its sequence header wrote the bit depth from the input while writing the subsampling from the encode side -- two members of one structure from two sides, with the profile defined against the one it read from the input. Each arm is moved for that reason and not for uniformity, and the note that recorded the hazard is deleted with it, because a defect note that outlives its defect is the next reader's wrong turn. Two more readers of that boundary move with them. The pre-session format resolver and the initialisation gate both build a video profile to ask the device about, and the profile a session actually creates is built from the encode fields -- so asking about the input's geometry would have predicted a different session than the one being guarded. The gate's refusal separates the two in its wording as well: it names the input format and its plane count as the input, and the chroma subsampling, bit depth and profile as the stream that input derives. THE ENCODE BIT DEPTH EXISTS BEFORE THE LEVEL IS SELECTED. The H.265 CPB VCL factor reads the encode bit depths for its depth term. Those were derived at session creation, which runs after level and tier selection and after the binder, so the depth term read zero at one of the function's two call sites and the real depth at the other: one function, two answers, inside one configuration. A 10-bit 4:4:4 stream selected its level and tier with the 8-bit factor and then sized its default coded picture buffer with the 10-bit one. ITU-T H.265 Table A.8's depth term is what separates Main 4:4:4 10 from Main 4:4:4 and Main 12 from Main 10, and it contributed nothing. The fix is the ordering and not the arithmetic. The encode-side bit depths are derived where the encode-side chroma subsampling already was, before any codec arm reads either, so every reader of the pair sees the same value; the zero-means-unset guards move with them, because an explicit encode depth is a request and not a default. THIS MOVES SELECTED LEVELS AND TIERS: a stream above main tier's old bitrate ceiling and below its correct one stops declaring high tier. Annex A selects the lowest tier and level whose limits the stream satisfies, so main tier is the correct answer there and high tier was a stricter claim on the decoder than the bitstream needed. Nothing about the coded pictures changes. The probe field's own comment stated the defect as a fact about where the value is read, and is corrected with it. THE GUARD DISCRIMINATES ON WHICH SIDE AN ARM READ, NOT ONLY ON THE DERIVATION. The case that accompanies the derivation could not fail on the thing it was written for: with one writer setting the encode side from the input side, the two are equal on every state the case can reach, and an assertion over equal values is satisfied identically whichever side an arm reads. Reverting all three arms to the input side leaves it green with an unchanged check count. So the two sides are made to differ without a mutation: the binder gains an internal way to state an explicit encode depth -- internal header, internal function, no public surface and no new field on any public struct -- and writes it before the derivation runs, where the zero-means-unset guard above reads it. The encode side is then the request and the input side is still the caller's format's, so the profile an arm derives says which one it read. A second pass over the routable sweep states the far end of the depth range from each input, asserts its own fixture first -- that the request landed on the encode side and moved neither the input depth nor the subsampling -- and counts, per codec arm, how many rows actually had the two sides implying different profiles, requiring that count to be non-zero. Without that count the pass could go green having compared every row against itself, which is the exact defect being repaired. WHAT IS STILL NOT GUARDED, SAID PLAINLY. The subsampling axis. Its writer is unconditional, so no request can make the encode and input subsampling differ and no test can distinguish an arm that reads one from an arm that reads the other. The depth axis discriminates for all three arms; the subsampling axis is asserted for derivation only. THE ORACLE STOPS ASSERTING AN OUT-OF-SPEC PROFILE. Its H.264 arm returned High 10 for anything above eight bits, so twelve-bit 4:2:0 rows expected 110 and twelve-bit 4:2:2 rows expected 122, neither of which ITU-T H.264 Table A-1 admits above ten bits. The derivation defect is out of scope; writing its answer into an expectation as the correct one is not. Those rows are excluded, counted, and printed with the in-spec answer beside the defect, so the exclusion is visible on every run instead of being a silently missing assertion. Every expectation is checked for spec-admissibility before it is used, from a restatement of the three standards' limits written for that purpose rather than from the library's own table, because a check that shares the library's table would agree with whatever that table says. WHAT IS PROJECTED, AND WHY BOTH SIDES ARE. The binder's probe carries the encode-side geometry beside the input's, and the profile each arm derived, even though one writer makes the two sides equal in every configuration a caller can build: a test that can see only one side cannot say which side an arm read, which is how three arms came to read different ones. CaseCodecArmsDeriveTheProfileFromTheEncodeGeometry is dispatched. Written and not called, it compiled into a suite that reported PASS without ever running it -- the third time in this series that a case has been authored and left out of the dispatch list, and the reason the suite's own check count is the thing worth watching rather than its verdict. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The public structures are pinned at build time, so a change to one of them is a build break in the reviewer's tree rather than a field silently read at the wrong offset in somebody else's. WHAT A SIZE PIN CANNOT SEE. Every public struct already carries its own size, and a size assertion catches a member appended or removed at the end. It cannot see the interior: a member removed from the middle and one added at the end leave the size unchanged, and a member inserted into a run that alignment padding absorbs leaves the size unchanged too. So each structure additionally pins the OFFSET of the member that closes each of its padding-absorbed runs. A removal or reorder among the members before that one -- which the size alone cannot show -- is then a build break rather than a silent shift of everything after it. The rationale states the absorption rule at every member width, which is the set of widths the pins are taken at. Every chainable struct is asserted to carry the structure-type and chain-pointer prefix. Every public run that a four-byte insertion would slide carries an assertion on the member that closes it, and where a structure is one run of equal-width members it is pinned member by member, because there is no closing member that can stand for the rest. The pins are described as the review gate they are, and they state the struct-extension rule they enforce. THE VERSION CONSTANT IS NOT A COMPATIBILITY COUNTER. It identifies a release. The compatibility guarantees rest on these assertions, and on the header and the implementation being vendored and rolled as one unit. So the constant is 1, and the per-number history of what each earlier value meant is removed rather than carried forward as a record no consumer can act on -- including the entry recording that one number had named two different vtables, which no artifact can be recalled to fix. THE FIELD TABLE CARRIES EXTENTS. The config field table names the extent of each field it names, and a check lays the rows across the structure so that a field with no row fails rather than passing unnoticed. THE INTERFACE RECORDS WHERE ITS OWN RULE WAS NOT HELD TO. The route to a maximum quantizer clamp of zero is named as the chained structure it would actually be, and the header records that its own extension rule has not been held to there, rather than implying it has. Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com> Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
Reconfigure answered success to everything and applied a subset. It now applies what it can carry, refuses the rest by name, and records what it applied. WHAT IT CHANGES AND WHAT IT REFUSES. The header names the fields Reconfigure changes and the fields it refuses, in the header itself rather than by reference to somewhere else, and names the members Reconfigure reads at all. The encoding fields it cannot carry are refused with a status naming them, instead of being accepted and dropped. THE CONSTANT-QP TRIPLE LANDS ON THE RIGHT FRAME. Reconfigure applies the constant-QP defaults, and the triple lands on the frame the Reconfigure preceded, like the six fields around it, rather than one frame later. An out-of-range constant quantizer is refused at the boundary instead of wrapping into a plausible value eight bits down. THE RECORD IS OF WHAT WAS APPLIED. The session's record of its own configuration is what a later Reconfigure compares against, and it was updated from the RAW config, which is not what the session ends up running on. A maximum bitrate of zero means track the average and is coerced on the way in, so a session reconfigured to a rate with a zero maximum ran capped while the record said zero. A frame-rate numerator of zero leaves the frame rate alone entirely, so both halves kept the values already in force while the record stored the zero. A zero denominator beside a non-zero numerator becomes one, and the record stored zero, which is not a frame rate at all. Each now records the applied value. This is a truthfulness fix and not a behaviour change, and the reason is worth stating because it is what makes the claim checkable: the set of members written back is disjoint from the set the immutables and the levers compare. Nothing compared reads a bitrate, a frame rate or a quantizer, so recording the applied value instead of the raw one cannot move a single refusal decision either way. THE QUANTIZER CLAMPS ARE CARRIED. The minimum and maximum QP were refused on the grounds that they reach the driver only through the codec-specific rate-control structs, which are filled once at codec init from a call site entangled with session-parameter creation. The first half is true and the second is not: the rate-control fill is a pure function of config state, and the codec rate-control command copies its output onto EVERY rate-control command rather than only the first. So the fill can be re-invoked on the encoder thread and the result rides the very command this call already causes. Three properties of that fill are handled rather than assumed. It raises the use-minimum and use-maximum flags and never lowers them, so re-invoking it in place could set a clamp and never clear one; the codec layer struct is therefore reset to its constructed state first, which is the state the fill saw at codec init and which nothing else writes. And it rewrites the live layer bitrate and frame rate from the config, which still holds what the session was built with, so those members are snapshotted and restored around it -- without that, a clamp change would silently revert a bitrate or frame rate an earlier Reconfigure had applied. A RECOGNISED STRUCTURE TYPE IS STILL A CHAIN. The chained input-colour structure is accepted at initialisation and every axis it carries is immutable for the life of the session, so a caller reusing one config for both entry points must clear the chain for this one -- and learning that by having the declaration silently ignored is exactly the accepted-and-dropped class this call refuses. Reconfigure refuses any chained structure, recognised or not, and the refusal is asserted with a recognised one rather than only with a junk pointer, which cannot tell "any chain" from "an unrecognised chain". THE H.26x ENCODER CLASSES MARK THEIR OVERRIDES CONSISTENTLY. Carrying the quantizer clamps adds two overriding methods to each codec encoder. A class in which some overriding members are marked and others are not is inconsistent, and a consumer that compiles this library with that inconsistency treated as an error cannot build it, so every overriding declaration in both classes is marked. Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com> Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
The descriptor-based encoder API sat in the public include directory, so every declaration it carried -- the session internals, the structures the implementation passes among its own layers -- was surface a host could reach and therefore surface that could not change without breaking one. Split it: what a host needs stays, and the rest moves into vulkan_video_encoder_ext_internal.h. The remaining header then leaves the include path entirely for internal/, which is not on any consumer's include path, so the descriptor API is reachable from inside this library and its own white-box tests and nowhere else. Those tests name the internal directory explicitly. That is the point: a test that reaches past the public surface should have to say so in its build. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Three test targets listed ${VK_DISPATCH_TABLE_SOURCE} in their source lists
without ${VK_DISPATCH_TABLE_HEADER}: common/libs/tests, and the drm_format_mod
and win32_opaque_import suites beneath it.
Listing the generated SOURCE orders the generation of that ONE generated
object. It orders nothing else in the target. An ordinary source in the same
target that includes the generated header -- VkCodecUtils/VkImageResource.cpp,
which vk_filter_test compiles directly -- has no declared dependency on the
generator at all, so a parallel build is free to compile it first and fail with
VkCodecUtils/HelpersDispatchTable.h: No such file or directory
Listing the generated HEADER is what orders the whole target, because a header
in a target's source list becomes a prerequisite of every object in it. Six
targets in the tree already do this; these three are brought in line with them.
The ordering is what this fixes, not an observed failure rate: seven clean
parallel builds of the unfixed tree (three at -j16, four at -j64) all
succeeded here, so the window is narrow enough on this machine that the
generator wins the race every time. It is won by luck rather than by anything
the build declares, and the two suites that do not compile VkCodecUtils
sources today are one added source away from depending on that luck too.
Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The library's embeddable surface is a Vulkan-shaped descriptor API: flat structs stamped with a structure type, extended by pNext chains, driven by free functions. The rest of this library is virtual, abstract, reference- counted C++, and so is every host that embeds it. This adds the interface in the idiom of both. WHAT IT IS. One root object reached through a single exported symbol. Capabilities and configurations come from it; a session comes from a configuration; and from a session come the roles a caller uses -- submission, bitstream retrieval, registration, completion, device identity, diagnostics. A role the implementation does not provide is a null reference rather than an error returned from a call the caller had no way to know would fail, so the capability query and the feature check are one operation done once at setup. WHY IT IS SHAPED THIS WAY. Extension moves from data layout to vtable extension. A pNext chain extends a struct the CALLER allocates, so its layout is pinned the moment any caller compiles; inheritance extends an interface the LIBRARY allocates, so nothing about layout is observable and a later version adds behaviour by deriving. One QueryInterface on the base replaces the descriptor registry -- and, because this library and Chromium both build without RTTI, it doubles as the typed downcast dynamic_cast cannot provide here. TWO RULES THE LIBRARY TAKES BACK. ResolveInputPath answers whether a configuration reaches the encoder directly, by copy, or through the compute filter, so the copy-engine-unless-the-samples-must-change rule has one implementation instead of one per host. HDR metadata is a separate interface offered only by codecs that can carry it, so an H.264 caller learns that from a null query at configuration time instead of a refusal inside session creation. VkEncClassifyFormat answers, without a device, the four things a host needs before it allocates: whether a format can be taken, by which route, what the compute filter needs of the image, and what the session will run at. None of it depends on a device, and a host picks its buffer format long before an encoder exists. WHAT THE VALUE TYPES CARRY is what a host actually reads back, which is more than a first pass suggested: a per-frame outcome, because a delivered frame can be a deadline drop carrying no bitstream; the device an image was allocated on, because all-zero means "not stated" to the cross-device check and omitting it disables that check rather than weakening it; explicit handle ownership, because the default transfers and a caller that assumes otherwise double-closes; an acquire and release fence pair, because that is what orders a producer against the encode; the ownership echo as an out-parameter, because it is true on the refusal path where no id is returned; and the device identity, because two instances give different handles for the same hardware and a host that adopted a device must be able to confirm the encoder took the one it meant. RequireImportExtensions refuses a device lacking what the import paths need, naming the missing extension, rather than accepting it and failing at the first import with a message about a handle. Result is a code and a pointer to a string literal -- sixteen bytes, trivially copyable, no allocation on a path that runs once per frame. The header compiles at C++17 while the implementation stays C++20, and names nothing Xlib defines as a macro. The implementation is a translation onto the descriptor API and holds no encoder logic. What it adds is that the rules above are stated once here rather than in each host. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Four suites on the interface, plus the support every one needs and nothing about any one of them. The interface suite drives what the interface promises: a role resolves and an absent one is a null reference; a role keeps its session alive because it shares the session's reference count; the H.264-carries-no-HDR rule is a null query at configuration time rather than a refusal at session creation; frames go in and a bitstream comes out by both the inline and the registered path; a session reports the device it runs on, by identity rather than by handle; and a reconfiguration applies the rate while refusing the resolution by name. Completion notifications coalesce, so it checks the callback fires and never more often than frames complete rather than asserting one per frame and pinning a promise the library does not make. The drain suite separates the two calls that end things. A drain that quietly ends the completion surface leaves a session that still accepts frames and retires none: no error, no crash, a caller waiting forever. Only a second batch submitted after a drain catches it. The release-fence suite checks the property that makes a fence a fence: it is handed over BEFORE it signals. The floor is half the run rather than one, because a fence exported after the fact comes back signalled on nearly every frame and "> 0" passes as soon as one frame races ahead. It runs at 1920x1080 because at a small enough frame the encode finishes inside the submission call and the check would report the frame size instead. A submission refused with NotReady is parked and re-sent, which is what that code means. The floor guard compiles the public header alone at exactly -std=c++17 with the window-system defines set. Those defines are as much the point as the standard: Xlib's Success, None and Status are object-like macros, so they bite only a translation unit that has included Xlib -- which is every browser and compositor, and none of this tree's own targets. A check that could not be made is counted as skipped, not passed, so a run on a machine that can measure nothing reports that rather than a clean sheet. Report::Summarise refuses to pass a run that measured nothing. Returning success on zero checks makes a suite whose assertions were all skipped -- or all deleted -- indistinguishable from one that ran and held. Zero checks with a skip recorded now reports SKIP, which the three tests declare through SKIP_RETURN_CODE 77; zero checks with no skip recorded is a failure, because a harness that ran nothing and did not say why has not measured the library. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
An OWN-mode context is cached for the process lifetime so that repeated create/destroy cycles reuse one VkInstance rather than issuing a second vkCreateInstance, which a sandboxed process may no longer be permitted to do. The cache held the only reference and nothing ever released it, so vkDestroyInstance never ran. Under a driver that audits allocations at exit that is reported, accurately, as a leak: ???NVDBG_MALLOC: final-report LEAKED 22589 BYTES in 22 leaks GLNV: Assert#1 !"MEMORY LEAK" All 22 are instance- and physical-device-scoped -- vkinstance.cpp, vkphysicaldevice.cpp, nvVkVideoPhysicalDevice.cpp -- which is the instance that was never destroyed and nothing else. Device-scoped resources were already being freed. The rule the cache enforces is "never re-created", not "never destroyed", and those are separable. VkEncRetireOwnContexts() latches retirement and then releases the cache: the instance is destroyed, and any later OWN-mode create is refused rather than allowed to stand a second one up. That is a stronger guarantee than before, where a post-teardown create was merely impossible because teardown never happened. Release stays a call and not a destructor deliberately. It runs vkDestroyInstance, so it has to happen while the process can still reach the driver -- at a point the caller picks, not at static destruction, where the ordering against the loader is undefined. Reached publicly as IPlatformLifetime::Retire() through QueryInterface, so the exported surface is unchanged. A caller that simply runs to process exit need not call it at all. Measured on the VM, same binary either way: without Retire(), 22589 bytes in 22 leaks and the assert; with it, neither. ctest 60/60. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
A session created from IEncoderPlatform is created ON the platform's context, and a context already carries the instance and physical device it was built against. The config nevertheless copied both out of PlatformCreateInfo, so a platform given a caller's VkInstance produced a config that named them a second time. The encoder refuses that pairing rather than guess which one wins: [EncoderExt] config.externalInstance / externalPhysicalDevice set on a session created with CreateVulkanVideoEncoderExtOnContext. The context already supplies both; remove them from the config. CreateSession therefore failed for every adopted-instance platform, which is every session a browser creates. A platform that creates its own instance was unaffected -- the handles it copied were null -- so the standalone tools and the whole ctest suite passed throughout. externalDevice and the queue-family indices still pass through. A context never creates a VkDevice, so a caller supplying one is adding something the context does not have, which is the opposite of restating it. Co-Authored-By: Vishesh Garg <visgarg@nvidia.com> Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
zlatinski
force-pushed
the
chromium-encoder-support
branch
3 times, most recently
from
September 18, 2026 22:59
18bef1f to
215c874
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Chromium encode support — change request summary
The problem this solves
The samples library is built to be driven by its own main(): it creates the Vulkan instance and device, parses argv, reads and writes files, prints to stdio, and exits. A browser can do none of that.
Chromium's GPU process already owns Vulkan, is sandboxed, has no filesystem to write to, and needs the encoder as a linked-in component with a stable surface. These 28 commits make the library embeddable
without changing what it does.
The public header goes from 58 lines to 1,126: from one free function to 13 role interfaces (IEncoderPlatform, IEncoderSession, IFrameSubmitter, IBitstreamSource, ICompletionSignal, IResourceRegistry,
IDiagnostics, …) reached through Query<>, behind exactly two exported symbols — VkEncCreatePlatform and VkEncClassifyFormat.
Roles carry per-interface revision ids ("vk.video.enc.IEncoderConfig/1"), so capability can be added later without breaking the ABI. The old descriptor-based vulkan_video_encoder_ext.h moves out of
include/ into internal/ — the interface is now a translation over it, and nothing outside the library depends on sType/pNext chains.
728541e abstract reference-counted interface · 76e6a1d descriptor API becomes an implementation detail · 3d55b51 the embeddable interface and the session it reshapes
A caller supplies its VkInstance and VkPhysicalDevice; the library creates only the VkDevice. An OWN-mode platform caches its instance for the process lifetime so repeated create/destroy cycles never
issue a second vkCreateInstance — which a sandboxed process may no longer be permitted to do. IPlatformLifetime::Retire() releases it at a moment the caller picks, by reference release rather than forced
teardown, so retiring under a live session leaves that session usable.
4fcfca9 adopt a caller's Vulkan objects · 5161a1f release the process-wide instance · 4a43c5a stop naming the instance twice on an adopted platform
No file output, no argv parsing on the embedded path, and stdio behind a single switch — a GPU process cannot write to disk, and stray stdio writes are at best lost and at worst trip sandbox diagnostics.
Encoded bitstreams are returned in memory, keyed by the caller's frame id, because a browser tracks frames by its own identifiers.
cabb2d2 content probe keyed by the caller's frame id · 4fcfca9 (stdio switch) · fe44dc8 (OS primitives behind one Linux seam)
Capability enumeration before initialisation, so a caller can choose a format instead of discovering a refusal at session creation. An input-format taxonomy covering what browsers hand over — packed
4:4:4 and 4:2:2, 10/12-bit, RGBA through the compute filter — plus colour model, transfer function and HDR10 static metadata. An unsupported input is now refused with a reason rather than taken to the
driver and losing the device.
1b2695b capability advertisement and the input-format taxonomy · e56d9a4 colour model, transfer function, HDR10 · a38f6cd refuse an input format rather than losing the device · 7e8e138 codec parameters
derived from bitstream geometry
Compute-filter descriptors matched to the shaders actually generated; storage requested the way a multi-planar format can grant it; image views following what the image supports; a shader module that
failed to compile refused rather than used; config, DPB and GOP corrections; AV1 capture, parser fixes and quantizer defaults; a command-line number the parser would silently change now rejected.
cdab403 3a91369 18a1d67 793095f 060a869 b4bcf6a 5211881 61530c6
Reconfigure applies what it can carry and refuses the rest rather than accepting-and-ignoring. Struct layouts are pinned by static assertion, so a field added to an extension struct fails the build
instead of silently shifting a caller's ABI.
9f01d15 Reconfigure · 025a854 layout pins as the struct-extension review gate
Standalone build, OS primitives behind one Linux seam, generated dispatch table made a prerequisite of every object, and 25 new test suites (47 files) — covering the interface a host actually drives,
cross-device dma-buf import, and the filter's output judged rather than merely produced. Total suite is 60 ctest tests.
fe44dc8 6070163 1d77ac0 6b208b3 f56f65d 086da40
What reviewers should know
One caveat worth stating plainly: the two largest commits (3d55b51, ~17.6k insertions, and the interface test suite, ~18.6k) are "introduce a subsystem" commits. They don't split into
independently-reviewable pieces without inventing commits that don't build — but if you'd rather review the colour/HDR commit (e56d9a4, 26 files) split into its three concerns, that one genuinely does
separate.