Skip to content

Chromium encoder support - #242

Merged
zlatinski merged 28 commits into
mainfrom
chromium-encoder-support
Oct 4, 2026
Merged

zlatinski merged 28 commits into
mainfrom
chromium-encoder-support

Conversation

@zlatinski

@zlatinski zlatinski commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Chromium encode support — change request summary

The problem this solves

The samples library is built to be driven by its own main(): it creates the Vulkan instance and device, parses argv, reads and writes files, prints to stdio, and exits. A browser can do none of that.
Chromium's GPU process already owns Vulkan, is sandboxed, has no filesystem to write to, and needs the encoder as a linked-in component with a stable surface. These 28 commits make the library embeddable
without changing what it does.

  1. A public, embeddable interface — the headline change

The public header goes from 58 lines to 1,126: from one free function to 13 role interfaces (IEncoderPlatform, IEncoderSession, IFrameSubmitter, IBitstreamSource, ICompletionSignal, IResourceRegistry,
IDiagnostics, …) reached through Query<>, behind exactly two exported symbols — VkEncCreatePlatform and VkEncClassifyFormat.

Roles carry per-interface revision ids ("vk.video.enc.IEncoderConfig/1"), so capability can be added later without breaking the ABI. The old descriptor-based vulkan_video_encoder_ext.h moves out of
include/ into internal/ — the interface is now a translation over it, and nothing outside the library depends on sType/pNext chains.

728541e abstract reference-counted interface · 76e6a1d descriptor API becomes an implementation detail · 3d55b51 the embeddable interface and the session it reshapes

  1. Running on someone else's Vulkan

A caller supplies its VkInstance and VkPhysicalDevice; the library creates only the VkDevice. An OWN-mode platform caches its instance for the process lifetime so repeated create/destroy cycles never
issue a second vkCreateInstance — which a sandboxed process may no longer be permitted to do. IPlatformLifetime::Retire() releases it at a moment the caller picks, by reference release rather than forced
teardown, so retiring under a live session leaves that session usable.

4fcfca9 adopt a caller's Vulkan objects · 5161a1f release the process-wide instance · 4a43c5a stop naming the instance twice on an adopted platform

  1. Operating inside a sandbox

No file output, no argv parsing on the embedded path, and stdio behind a single switch — a GPU process cannot write to disk, and stray stdio writes are at best lost and at worst trip sandbox diagnostics.
Encoded bitstreams are returned in memory, keyed by the caller's frame id, because a browser tracks frames by its own identifiers.

cabb2d2 content probe keyed by the caller's frame id · 4fcfca9 (stdio switch) · fe44dc8 (OS primitives behind one Linux seam)

  1. Taking the frames a browser actually produces

Capability enumeration before initialisation, so a caller can choose a format instead of discovering a refusal at session creation. An input-format taxonomy covering what browsers hand over — packed
4:4:4 and 4:2:2, 10/12-bit, RGBA through the compute filter — plus colour model, transfer function and HDR10 static metadata. An unsupported input is now refused with a reason rather than taken to the
driver and losing the device.

1b2695b capability advertisement and the input-format taxonomy · e56d9a4 colour model, transfer function, HDR10 · a38f6cd refuse an input format rather than losing the device · 7e8e138 codec parameters
derived from bitstream geometry

  1. Correctness fixes the encode path needed once a host drove it

Compute-filter descriptors matched to the shaders actually generated; storage requested the way a multi-planar format can grant it; image views following what the image supports; a shader module that
failed to compile refused rather than used; config, DPB and GOP corrections; AV1 capture, parser fixes and quantizer defaults; a command-line number the parser would silently change now rejected.

cdab403 3a91369 18a1d67 793095f 060a869 b4bcf6a 5211881 61530c6

  1. Reconfiguration and review gates
    Reconfigure applies what it can carry and refuses the rest rather than accepting-and-ignoring. Struct layouts are pinned by static assertion, so a field added to an extension struct fails the build
    instead of silently shifting a caller's ABI.

9f01d15 Reconfigure · 025a854 layout pins as the struct-extension review gate

  1. Build and test

Standalone build, OS primitives behind one Linux seam, generated dispatch table made a prerequisite of every object, and 25 new test suites (47 files) — covering the interface a host actually drives,
cross-device dma-buf import, and the filter's output judged rather than merely produced. Total suite is 60 ctest tests.

fe44dc8 6070163 1d77ac0 6b208b3 f56f65d 086da40

What reviewers should know

  • The decoder is barely touched — 4 files, +108/−9, in AV1 parsing and the frame buffer. This is an encode change.
  • Two files removed: the old public vulkan_video_encoder_ext.h (now internal/) and VkEncoderRenderFrame.* (a render path an embedded encoder has no use for).
  • The bulk is concentrated: vulkan_video_encoder_ext.cpp (+12,501) and the two largest test mains (+5,413, +3,352) are over a third of the diff.
  • Upstream absorbed 24 of our earlier commits already — this branch is rebased on top of them, so those are not re-proposed here.

One caveat worth stating plainly: the two largest commits (3d55b51, ~17.6k insertions, and the interface test suite, ~18.6k) are "introduce a subsystem" commits. They don't split into
independently-reviewable pieces without inventing commits that don't build — but if you'd rather review the colour/HDR commit (e56d9a4, 26 files) split into its three concerns, that one genuinely does
separate.

@zlatinski
zlatinski force-pushed the chromium-encoder-support branch 6 times, most recently from 4eccd01 to 03b690c Compare September 18, 2026 06:15
vish-garg and others added 24 commits September 19, 2026 01:46
…n ctest

The encoder library becomes something a host can configure, compile and test on
its own, outside any embedding checkout, on a CI that actually runs the tests.

A STANDALONE CMAKE TREE compiles every translation unit the embedded build
compiles, so a change that breaks the library is visible here rather than in
somebody else's checkout. SPIRV-Tools is linked whenever glslang is a static
archive rather than only under MSVC, which is the configuration that otherwise
fails to link. The compute-filter gate is expressed against the shader-compiler
backend instead of naming one compiler, so the option survives a change of
backend. The FFmpeg-dependent decoder demo builds without FFmpeg.

ONE PORTABILITY SEAM, AND IT IS LINUX. VkVideoEncoderOsAdapterLinux and
vulkan_video_encoder_os_event_linux put the event, thread and handle primitives
the encoder needs behind a single interface, so the platform differences live in
two translation units instead of being spread through the encoder. Both are
Linux-only by construction, and the file names now say so rather than implying a
portability that does not exist. An unsupported target is refused at configure
time, with the two files a port has to supply named, rather than as a preprocessor
error partway through a build. VkThreadPool and VkThreadSafeQueue gain the
liveness the encoder relies on.

DIAGNOSTICS BEHIND A LATCH. VkEncoderStdioLatch gives the library one pair of
streams and one latch. An embedder that silences the library cannot have a later
session unsilence it, which is what makes silencing a property of the process
rather than of each call site, and it is why the later features in this series
report through the API rather than through stdout.

CI THAT CAN FAIL. The registered tests gated nothing, because no workflow ran
ctest at all. enable_testing() lands at the repository root -- without that single
line CMake writes no CTestTestfile.cmake and every registered test registers and
never runs -- a CI script runs ctest with a device-free and gpu label split, and
that script refuses to report success on a label typo or on an empty selection,
because a suite that selects nothing must not pass. The two workflows are repaired
so they can fire.

CORRECTNESS. The encode DPB image barrier names the read and write accesses the
encode itself performs in its dstAccessMask. VkEncoderRenderFrame is deleted; the
class was present and had no caller.

BITSTREAM OUTPUTS ARE NOT SOURCE. The encoder's own default output names are
ignored, and only those, so a fixture a suite deliberately commits is still
visible to the repository.

CREDITS. The encode DPB read and write access mask came from Jose Lopes.

BUILD_ENCODER_COMPUTE_FILTER IS DOCUMENTED AS THE NO-OP IT IS. Removing the
duplicate add_definitions() made this the only place
VK_VIDEO_SAMPLES_COMPUTE_FILTER_SUPPORTED is DEFINED, but no source file in the
tree READS it -- grep returns the CMake file and Chromium's BUILD.gn and nothing
else -- so OFF removes a -D that no #ifdef consults and the filter is still
compiled and still used. The real control is the runtime flag
EncoderConfig::enablePreprocessComputeFilter, which no CLI or ext-API exposes. The
comment said otherwise, and cited a VK_ERROR_DEVICE_LOST on an unadvertised
input format as an A/B baseline to reproduce rather than as the defect it was;
it now says what the option actually does. The option is kept rather than deleted because wiring the macro up would let
an embedder with no GLSL compiler drop the filter's code entirely.

WHAT THE COMMENTS NO LONGER SAY. The narration this series had accumulated in
its own files -- that the branch filter "used to be unsatisfiable", that before
scripts/run_ctest_ci.sh existed CI ran no ctest so ~15 add_test() registrations
gated nothing, that earlier revisions of the stdio-latch comment described dumps
since deleted, that the compute-filter guard's audience is "no longer anyone" --
is recorded here instead. A file describes what the tree does now; the commit
carries how it got there. Concretely, and once, so it is not lost:

  * Neither workflow had ever run: the branch filter named a branch this
    repository does not have, so no push and no pull request ever matched --
    including test.yml's `Run CTest suite` step, which invoked the script
    correctly and was simply never reached. Both workflows now trigger on the
    default branch and on pull requests targeting it, which gates a topic
    branch through its pull request without enumerating branch names.
  * There was no `ctest` anywhere under .github/workflows/, so every add_test()
    in the tree registered and gated nothing. Adding a bare `ctest` was not
    enough either: on a GPU-less runner the suite was 2 passed / 4 skipped /
    8 failed, which is why the label split and run_ctest_ci.sh exist.
  * The stdio-latch comment claimed the latch covered a CreateCodecConfig argv
    dump and a positional-argument dump in ParseArguments; both dumps had
    already been deleted from the tree.
  * BUILD_ENCODER_COMPUTE_FILTER's comment claimed that removing a duplicate
    add_definitions() made -DBUILD_ENCODER_COMPUTE_FILTER=OFF effective again,
    and cited a VK_ERROR_DEVICE_LOST on an unadvertised input format as an A/B
    baseline to reproduce. Neither held: no #ifdef reads the macro, and the
    device loss was the defect.
  * FFmpegDemuxer.cpp was listed twice in the decoder demo's source list --
    unconditionally, and again under if(FFMPEG_AVAILABLE) -- so the guard could
    never take effect. The file compiled with FFMPEG_AVAILABLE OFF and failed on
    <libavformat/avformat.h> whenever the headers were absent. Only the
    conditional append remains.
  * The compute-filter guard was left undefined, which silently dropped the
    filter from vk-video-enc-test as well as from embedders -- not what the
    guard exists for. It defaults ON.

Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com>

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
.settings/, .vscode/, .idea/ and *.code-workspace are machine-local editor
state, not source. Nothing in the tree reads them and they differ per
developer, so an accidental `git add -A` is all it takes to commit one.

Ignore them, so that stops being possible.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…losing the device

Without the preprocess compute filter the input image IS the encode source, so
ANY requested format the driver does not advertise for the profile cannot be
encoded -- there is nothing left to convert it. The unsupported set is a
property of the driver and the profile, not of a particular layout or
subsampling; the filter is what makes the rest reachable.

The condition is reachable by default rather than only on an unusual request:
EncoderConfig::input.vkFormat defaults to a format no driver advertises as an
encode source. With the filter off, the format-negotiation fallback substituted
the driver's first advertised format while frames kept arriving in the requested
one, and the mismatch surfaced as VK_ERROR_DEVICE_LOST and a 0-byte bitstream --
a hung queue several hundred lines from its cause, and a whole configuration that
could not be run.

InitEncoder now refuses that combination with VK_ERROR_FORMAT_NOT_SUPPORTED and
prints the formats the driver does advertise, so the caller learns both that the
request was refused and what to ask for instead. Not supporting a format is a
legitimate configuration; losing the device over it is not.

The guard is conditioned on the preprocess filter being disabled, which nothing
in this tree does, so no existing path changes. Running without the filter
becomes a supported configuration, and the BUILD_ENCODER_COMPUTE_FILTER comment
says so.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Two rules the view factory was not applying, both of which a caller can
now trip because the encoder accepts images it did not allocate.

A view may only be created over an image whose usage includes at least one
view-compatible bit (VUID-VkImageViewCreateInfo-image-04441). TRANSFER_SRC
and TRANSFER_DST are not among them, so a transfer-only image -- a staging
copy handed in by a caller, consumed entirely through its raw VkImage by
vkCmdCopyImage and barriers -- must get no view rather than an invalid one.

A per-plane view reinterprets the image as a different format (R8, R8G8),
which is only legal if the image was created VK_IMAGE_CREATE_MUTABLE_FORMAT.
That always holds for images this library allocates; for one imported from a
dma-buf, where the EXPORTER chose the create flags, it usually does not.
Request plane views only when the flag is present and keep the combined
view, which is all the encode path reads.

Foundation for the encoder-ext work: every later topic that imports or
exports an image depends on these two guards being in the view factory
rather than at each call site.

EXISTENCE IS A PROPERTY OF THE IMAGE, NOT OF ITS VIEW, and the decoder's frame
buffer is corrected here rather than separately, because the guard above is what
makes the distinction reachable.

The decoder allocates its linear output image with VK_IMAGE_USAGE_TRANSFER_DST_BIT
alone -- vkCmdCopyImage and image barriers take the raw VkImage and it needs
nothing else -- so under the rule above it correctly gets no view.
NvPerFrameDecodeResources::ImageExist() tested the VIEW, so that image reports
itself non-existent for good.

Nothing between there and the copy says so. GetImageSetNewLayout returns
VK_SUCCESS with validImage false under an assert(), the frame buffer's overload
calls CreateImage, re-checks, still gets false and returns VK_SUCCESS anyway, and
GetCurrentImageResourceByIndex returns the slot index, which its caller reads as
success. Each of those reports failure only through an assert that NDEBUG removes,
so in a Release build the zeroed PictureResourceInfo reaches vkCmdCopyImage as a
NULL VkImage with layout UNDEFINED and the driver faults on it.

Track the image resource alongside the view, key ImageExist() on the image,
source the picture resource's image and format from it, and bind an image view
only when one exists.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Two things the device context owes an embedded host.

QUEUE FAMILIES. When the caller supplies the VkDevice, the families it
actually created queues on are not discoverable from VkPhysicalDevice
properties, and vkGetDeviceQueue on a family the device was not created
with is undefined. Take the caller's encode and compute family indices and
bind those instead of the probe's picks; UINT32_MAX keeps the probed
family, which is the library-owned-device path.

STDIO. A library linked into a browser's GPU process must not write to
stdout or stderr. The gated-logging latch behind VkEncOut()/VkEncErr() is
header-only, but its one process-wide definition has to live somewhere
single: VkCodecUtils is the bottom-most target every encoder translation
unit already links, so it lives here. Keeping it out of the header is an
ODR requirement, not a style choice.

Depends on nothing above it; the encoder core and the ext implementation
both build on this.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
STORAGE on a multi-planar format is a per-plane-view usage, not an image
usage. Neither VK_FORMAT_G8_B8_R8_3PLANE_420_UNORM nor
VK_FORMAT_G8_B8R8_2PLANE_420_UNORM advertises
VK_FORMAT_FEATURE_2_STORAGE_IMAGE_BIT in either tiling, while the
single-plane formats their planes are viewed as -- R8, R8G8 -- advertise it
in both. So naming VK_IMAGE_USAGE_STORAGE_BIT on the YCbCr format without
VK_IMAGE_CREATE_EXTENDED_USAGE_BIT is validated against the image format
and refused, even though every view that will actually be created is legal.

Set EXTENDED_USAGE so the check applies to the plane views that carry the
storage access, and keep the pool's format bookkeeping in step.

Sits on the view rules from the previous topic: it is the allocation half
of the same question those answer at view-creation time.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
BuildGlslShader() returns VK_NULL_HANDLE when a compile fails, and both the
vertex and fragment paths passed the result straight on. An unguarded null
reaches VkPipelineShaderStageCreateInfo::module, which is where it becomes the
driver's problem rather than a reportable error.

Check both, and return VK_ERROR_INVALID_SHADER_NV.

The fragment check sits BEFORE the cache swap, which is the part worth stating.
The rebuild is skipped when the generated source is unchanged
(m_fssCache.str() != imageFss.str()), so caching the source that just failed to
compile would make it the current fragment shader and that test would skip the
rebuild from then on -- one bad compile made permanent for the life of the
object. Returning first leaves the cache holding the last source that did
compile, so a later call retries.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The generated GLSL and the descriptors it is bound with had drifted apart
in several places, each of which is a validation error or a wrong pixel.

Configure()'s result was being dropped -- the next statement overwrote it,
so a failed configure carried on as success.

Chroma was fetched at a hardcoded srcPos/2, which halves both axes. That is
right for 4:2:0 and wrong for 4:2:2 (chroma is full height) and for 4:4:4
(no subsampling at all), so chroma came from the wrong rows. The ratios now
come from the same YcbcrVkFormatInfo the plane arithmetic uses, so they
cannot drift from the descriptors generated beside them.

Descriptor selection is a preference rather than a command: with autoSelect
the layout picks push descriptors when VK_KHR_push_descriptor is present,
instead of an unconditional PUSH_DESCRIPTOR_BIT that discarded the
descriptor-buffer arm outright.

Needs the view rules from the earlier topic, because which view a binding
gets is what decides how its shader must declare it.

THE YCBCR2RGBA GOLDEN IS RE-RECORDED HERE, with the change that invalidates it,
so the decode suite does not go red in between.

The output store above moves from ivec3(pos, pushConstants.dstLayer) to a plain
ivec2. The binding it writes is the combined view, and VkImageResourceView types
that view VK_IMAGE_VIEW_TYPE_2D whenever layerCount is 1, which every caller of
this filter pins it to; storing to a layer of a non-array view is undefined, and
the driver returns plausible-looking data rather than failing, so the recorded
golden captured that result instead of a failure.

The new output is the better one, measured rather than assumed. Ten of thirty
frames differ and the rest are byte-identical, and on every one of those ten the
new output is closer to the same clip decoded without the filter -- frame 1
psnr_y 9.96 dB against 3.32 dB, and 5 to 6 dB better across frames 1 to 10. The
stray layer index is what those frames were carrying.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…alise

Fix for d535085 ("filter: packed 4:4:4 and 4:2:2 support, and three
conversion defects"), which added m_inputPackedYcbcr and m_outputPackedYcbcr
to the class and appended them to the end of the constructor's initialiser
list rather than at their declared position.

A member is initialised in the order it is DECLARED, never in the order the
initialiser list writes it. Six members disagreed: the two packed-YCbCr
pointers and the two block ratios initialise before m_ycbcrPrimariesConstants,
and all four before the bitfields the list placed them after. So the list
described an order the compiler did not follow, and a reader checking what
runs first had to reconstruct it from the declarations instead.

Nothing is misinitialised today, because every initialiser here reads a
constructor parameter and none reads another member. That is the whole
margin: the first initialiser that reaches for a sibling would read it
before it was set, and the list would look correct while it happened.

The list is reordered, not the declarations, so the class layout is
untouched and no member changes its position or its value.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
A host that drives the encoder frame by frame needs per-frame facts back --
what was captured for a frame, keyed by the id it submitted rather than by
an internal slot number, which is the only handle the caller has.

The probe owns that mapping and its lifetime: a frame's capture stays
retrievable until the caller collects it, independent of when the encoder
recycles the image the frame was carried in.

It lands before the encoder core because VkVideoEncoder.h includes it and
holds a FrameCapture by value; the core cannot compile without it. Wired
into the encoder library's source list in the same commit so the build and
the code arrive together.

TWO DESIGN CHOICES THE HEADER STATES AS RULES, recorded here as the measurements
that forced them.

The arm parameter is a scoped enum rather than a bool because the sense is
inverted from the obvious spelling: `bool directlyEncodable` true means "do not
arm", CaptureSite's true-ish value means "DO arm". A bool would have let every
existing call site keep compiling with its meaning reversed -- arming what should
be refused and refusing what should be armed -- which is the same class of silent
breakage as the pNext clobber that made this probe inert to begin with.

IsProbeableFormat() is public because the ARM decision has to ask it. Consulted
only from RecordCapture, on the first frame, a P010 registration echoed ARMED at
import and was downgraded to NOT_APPLICABLE a frame later.

THE MECHANISM IS DOCUMENTED alongside it, under vk_video_encoder/docs: the
arm/capture/score/echo sequence and the two command-buffer sites it rides,
the dead-plane predicate and the Q8 constant the ext layer static_asserts
against, the row stride and what a byte-wise walk costs instead, and why a
verdict is latched per registration rather than per frame. Markdown, a
sequence and state diagram, and the rendered page.

It gives armedRegistrationCount a section of its own, because that field is
the one a reader gets wrong: probed=0 damaged=0 is what a session reports
when every buffer was clean AND what it reports when nothing was ever
measured, and only the armed count separates them.

The document also records the reachability decision rather than leaving it
to be inferred. The probe is internal: its structures are declared in the
descriptor API's internal header, which is on no consumer's include path,
its consumers are this library's own white-box tests, and the public encoder
interface does not expose it and is not intended to.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The parts of the encode path that a host exercises and the demo does not,
separated from the session object itself because none of them touch the
interface the ext layer binds to -- they compile and link against the
existing one.

Config gains the per-codec derivations a caller can reach directly rather
than through the demo's single fixed path. The H.264 DPB and the GOP
structure each carried state that was only ever read back in the order the
demo produced it; both now survive a caller that reconfigures mid-stream,
submits frames out of order, or stops early.

PSNR belongs to this topic by subject but not by dependency: it reads
fields that the next commit adds to the session, so it cannot compile until
that lands and it travels with it instead. Splitting on what builds rather
than on what reads well is deliberate -- a commit that does not compile is
not a reviewable unit.

Kept ahead of the session change so that the reviewable, self-contained
half of the core work is not buried inside the atomic commit that follows.

FINALIZECONFIG() TAKES AN OPTIONAL DEVICE CAPABILITY ARGUMENT, because some of
what it decides is a property of the hardware and not of the command line.

The function runs on two paths: argv parsing, which happens before any device
exists, and an embedding host, which has already probed one. A pointer that may
be null carries that difference without splitting the tail in two -- null keeps
the command-line answer, which is what the demo gets, and a caller that has
probed reaches a device-correct answer through the same function. A second
FinalizeAgainstDevice() would be a call every caller has to remember.

One decision uses it. Intra refresh is a device feature
(VkPhysicalDeviceVideoEncodeIntraRefreshFeaturesKHR::videoEncodeIntraRefresh), so
a configuration asking for it on a device that lacks it is refused here, where the
request is still attributable, rather than inside session creation.

The minQp default is deliberately NOT a second one. Every consumer of that field
reads it only under minQpSet, which stays clear precisely when the default
applies, so an unset minQp reaches the driver as "no clamp" rather than as 20 --
checking the number against the device would decide nothing. A caller that wants
a QP clamp sets one, and the codec configs check THAT against the device window
and refuse it rather than narrowing a request silently. DeviceCapabilities
carries only what it is asked, and defaults its members, so a field the caller
does not assign reads as absent rather than as whatever the stack held.

--GOPFRAMECOUNT IS PARSED AT THE WIDTH IT IS STORED. VkVideoGopStructure keeps
the count in a uint32_t, and its constructor now agrees with the member and the
setter. Parsing into a narrower local truncated silently, because parseUint casts
without a range check: --gopFrameCount 300 reached the encoder as 44, and 256 as
0, which the codec configs read as ZERO_GOP_FRAME_COUNT and replaced with the
device's preferred count. Neither is refused and neither is reported, so the
stream is encoded to a GOP the caller never asked for.

Zero remains legal: it is that sentinel, and requesting the device's preference
is a request like any other.

codecBlockAlignment is deliberately NOT among them. It holds the H.264 macroblock
size for every codec, which is wrong for H.265 CTBs and AV1 superblocks -- but
nothing in the tree reads it: one write here, one declaration, one constructor
initialiser and no readers. Giving it a device-derived value would add correctness
to a field no consumer checks, so it keeps its value and gains a comment saying
that a reader must derive the real granularity from
VkVideoCapabilitiesKHR::pictureAccessGranularity instead.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
parseUint and parseInt reported success for three classes of input whose
value the encoder then ran with, changed:

  * NARROWING. static_cast<T> of a value too large for T keeps the low bits.
    Into a uint8_t, 256 became 0 and 300 became 44. --temporalLayerCount 256
    configured one layer's worth of nothing, and 0 is a meaningful sentinel
    in several of these fields, so a truncation did not even have to land on
    an obviously wrong number to be accepted.

  * A NEGATIVE INTO AN UNSIGNED FIELD. strtoull takes a leading '-' and
    wraps, so --gopFrameCount -5 was honoured as 4294967291: a GOP that never
    closes, from a typo, silently.

    -1 is different and is kept. It is the Video Codec SDK spelling for all
    bits set -- an infinite GOP is UINT32_MAX, as the json_config README and
    the idrPeriod derivation in EncoderConfigH264 both record -- so it is
    accepted explicitly and yields the maximum value of the field's type,
    exactly what the wrap produced. Command lines already using it are
    unaffected. Every other negative is refused.

  * OVERFLOW. Beyond the accumulator's range strtoull and strtoll saturate
    and set ERANGE, which neither helper read.

All three are refused now, as an invalid parameter, at the point the argument
is parsed. The narrowing check -- the value must compare equal after a round
trip through T -- is what makes parsing straight into a uint8_t field safe
again, so the narrow locals need no widening: out-of-range input is rejected
rather than folded into range.

Nothing that was accepted and meant what it said changes. A caller who typed
a number the encoder could represent gets the same encode; a caller who typed
one it could not now gets told, instead of an encode configured from the
remainder.

--idrPeriod is read unsigned, which is how SetIdrPeriod stores it. Read as a
signed value it accepted every negative, and the conversion at the setter
turned each one into a period so large that no periodic IDR is emitted -- so
a mistyped -5 asked for an infinite IDR period and was given one. -1 keeps
its meaning there, on the same convention: all bits set, for a field of
whatever width. minQp and maxQp stay signed, because they are signed fields
whose -1 is their own "unset".

The verbose prints that reported these fields with %d are corrected to %u, with the promotion made explicit where the field is a uint8_t -- which is
why an infinite GOP logged itself as "-1" rather than as its value.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The API a host embeds this library through, its implementation, and the
encoder session both are expressed in terms of.

WHY THESE ARE ONE COMMIT. They were split three ways first and none of the
three built. The ext API renames methods that the existing
vulkan_video_encoder_ext.cpp declares `override`, so the header cannot land
before the implementation. The session changes SetExternalInputFrame's
signature, which that same file calls, so the core cannot land before it
either. And the implementation is written against both, so it cannot land
first. Separating them needs compatibility shims that exist in no other
commit and would be deleted by the next one -- history that never described
a state anyone had. They are one change; this records them as one.

WHAT IT IS. The interface is shaped by what an embedding host cannot do. It
cannot own a VkDevice, so the library adopts the caller's instance,
physical device and device, and binds the queue families the caller
actually created queues on. It has no visibility of internal slots, so
every per-frame fact is keyed by the frame id the caller submitted. It
cannot tolerate writes to stdio or to files, so bitstreams come back in
memory and logging goes through the latch. It cannot stand up a session
just to ask a question, so capabilities are answerable before one exists,
and a codec whose encode extension the device does not advertise is refused
at the point of asking rather than deep inside session creation.

The session carries the corrections that discipline requires: rate control
keeps its codec-specific struct and per-layer state in the BeginCoding
chain, which VUID-vkCmdBeginVideoCodingKHR-pBeginInfo-08254 requires to
match the state a control command established, pointing at encoder-owned
storage rather than at a pool-recycled array.

PSNR and the demo come along because they are compiled against the session
interface this changes -- the demo target builds VkVideoEncoder.cpp from
its own source list, so its build wiring moves in the same commit or the
link breaks.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The AV1 arm of the encoder, and the AV1 parser it is exercised against. It lands
directly on the interface commit, because its capture arm plugs into the
in-memory capture funnel that commit introduces. Everything else it touches --
the AV1 parser, the AV1 codec config and the AV1 encoder -- already exists below
this series, so nothing after this point has to reach backwards past the
interface work to edit a parser or a codec config.

IN-MEMORY CAPTURE. The AV1 encode arm captures its bitstream in memory and
prepends a temporal delimiter OBU to each temporal unit. The driver emits none,
so what a caller pulled out was a fragment that only a file writer would have made
whole; with the prepend, every captured chunk is independently parseable.
File-based output is unchanged.

THE DIMENSIONS A LATER FRAME INHERITS. AV1 frame_size_with_refs() lets the
current frame take its size from a reference, and SetupFrameSizeWithRefs reads
decodeWidth, decodeHeight and decodeSuperResWidth back out of the referenced
picture. Nothing on the AV1 path ever wrote them -- the only writers were on the
VP9 path -- so on a pure AV1 stream all three were read while indeterminate.

They are published in BeginPicture, which end_of_picture reaches only after the
whole frame header is parsed, so all three values are final. Deliberately NOT in
UpdateFramePointers, which looks like the natural home but also runs on the
show_existing_frame path with a previously decoded picture, where writing would
replace a correct stored dimension with an unrelated one -- turning a read of
indeterminate memory into corruption of good data. And outside the allocation
guard rather than inside it, so a retained picture does not keep a stale size.

The narrow branch is where this becomes visible; it is not the wide exposure. The
Vulkan reference codedExtent is built from decodeWidth and decodeHeight for every
AV1 reference slot on every frame that has references, with no frame-size-override
gate.

vkPicBuffBase's constructor now names its VkPicIf base, so the base subobject is
value-initialised rather than default-initialised. With the parser writing all
three fields before they are read this is belt and braces; what it buys is that a
MISSING write reads as zero rather than as recycled pool storage, which is the
difference between a defect that reproduces and one that appears only under
memory pressure.

ORDER_HINT_BITS IS 8, NOT 7. The reference-order-hint writer casts to uint8_t
while the DPB writer masks by this value, so at 7 the two disagree and a
ref_order_hint of 128 or more can be emitted for a field that cannot hold it.
8 is also what a consumer masking order hints with 0xFF expects: at 7 the
order hint wraps at frame 128 and such a consumer rejects every frame after
it. The mask at the encoder's writer is taken from the advertised width
rather than from the cast, so the two cannot drift apart again.

QUANTIZER DEFAULTS. An unset AV1 quantizer index is left to the AV1 encoder
config's own driver-preferred triple instead of inheriting an H.26x QP. The two
scales are not the same scale: an ordinary H.26x QP read as an AV1 qindex is near
lossless, so the encoder silently produces enormous frames while every counter
reports success. The code no longer pretends the scales are shared.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
…port

The filter suite ran without judging a frame: it dispatched, reported
success and asserted nothing about the pixels. It now validates its output,
and --list is registered as a device-free test so the case table is proven
to construct on every push, while the two dispatching arms are labelled gpu
and exit 77 so a GPU-less host reports SKIPPED rather than FAILED. Every
other init failure still exits 1, so a real regression cannot hide in the
skip arm.

Adds linux_dmabuf_import, the counterpart of win32_opaque_import: it proves
an image exported by one VkDevice can be imported AND USED by a different
VkDevice on the same GPU. That is the property the browser's zero-copy path
depends on and nothing in this tree previously tested.

Depends on the shared runtime topics it exercises -- the view rules, the
image pool's storage-usage handling and the filter's descriptors.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Twelve suites over the ext surface, each one written against a property a
host depends on and none of which the demo's single open-push-close path
reaches.

They are split by what they need, because a suite that cannot tell "no GPU
here" from "this is broken" teaches people to ignore it. The device-free
ones -- the chained-descriptor walk, the input-format taxonomy, the config
binder's disposition of enablePreprocessFilter -- are pure functions of
their arguments and belong in the gating set that runs on any host. The
real-device ones declare SKIP_RETURN_CODE 77 and report SKIPPED, not
FAILED, on a machine with no encode-capable GPU.

Several are regression tests for defects that are now fixed and would
otherwise be invisible: the acquire fd's ownership on every refusal path,
the release fence's export, and DrainPendingFrames() permanently disabling
async assembly so that every frame submitted afterwards was encoded and
then never became acquirable.

The registration in the two CMakeLists files lands here rather than in a
build topic of its own: add_subdirectory of a directory that does not exist
yet fails at configure time, so the tests and the lines that run them
cannot be separated.

qfot_foreign_release_repro.c IS NOT BUILT, AND THAT IS DELIBERATE. It is a
driver-defect reproducer rather than a test: it asserts nothing about this
library, and a suite that failed when the driver was wrong would report this
library as broken. No CMakeLists references it, so it compiles only when
someone builds it by hand against the driver under investigation.

The cost of that is real and worth stating: nothing keeps it compiling as the
surfaces it calls move, so it is evidence with a shelf life, not a gate.

THE FOUR PROFILE CASES ARE DISPATCHED. encoder-ext-filter defined
CaseProfileNumbersReachTheCodecConfigUnchanged, CaseProfileDefaultIsDerivedPerCodec,
CaseProfileNumbersAreReadAgainstTheCodec and CaseProfileMustAdmitTheInputDepth and
called none of them, so nine checks were compiled, never run, and reported as
part of a passing suite. A case that is written but not dispatched is worse
than one that does not exist: the file reads as though the claim is covered.
They pass -- the suite runs 159 checks, 0 failures -- so what was missing was
the four call sites and nothing else.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…e it

Defects in the encoder itself that became reachable, or became visible, only once a
host drove the encoder in process and read its status instead of reading a file
afterwards.

A FRAME THAT DID NOT ENCODE FAILS ITS CALLER. A frame the encoder could not record
or submit now fails, and an encode that produced no bitstream is reported as a
failure rather than as a success with an empty payload. A command that asked for no
bitstream is not judged on one, so the two cases stay distinguishable.

THE B-FRAME CLAMP IS EFFECTIVE. The consecutive-B-frame clamp was computed and then
not applied, so a GOP structure could ask for more B frames than the encoder was
configured to carry. The assembly queue is sized to the burst the GOP can present
rather than to a fixed depth, so a legal structure cannot overrun it.

ONE DPB SLOT, BOUND ONCE. The H.264 and H.265 encode commands bind each DPB slot at
most once when a reference picture holds a position in both reference lists.
Binding it twice is a validation error, and the second binding is not what the
codec means. Both arms report their validation output, and the arm that codes B
frames runs under H.265 as well as H.264.

H.264 BASELINE MUST NOT EMIT CABAC. Baseline does not carry CABAC, so a Baseline
session that enables it produces a stream naming a profile it does not conform to.
The entropy mode is decided from the profile, and the suite fails if someone later
widens the rule to fire on requests and assumes it subsumes the clamp.

THE DEBUG SWITCHES ARE READ ONCE. The encoder read its environment switches once
per frame. They are read once, so the cost is off the frame path and a switch
cannot change meaning mid-session.

THE SESSION CREATE RESULT IS CHECKED IN RELEASE BUILDS. It was checked only under
assert, so a release build carried a null session past the point of failure and
died later inside the driver.

GOP SEQUENCING IS TESTED DIRECTLY. The GOP machine decides encode order, B-frame
counts and B-frame positions before any picture exists, and nothing else in this
tree asserted those tables -- the sibling suites set a GOP length and then encode,
which cannot tell a correct sequence from a plausible one. It is carried as its own
subdirectory because it compiles the GOP source directly and links nothing else. A
round-trip quality test is registered beside it: it encodes, decodes out of process
and compares per-frame PSNR against the source, which is the only round-trip
measurement here, because the encoder's own PSNR option compares the input against
the RECONSTRUCTED picture, inside the encoder, and so cannot see a stream that
decodes wrongly. It exits with the skip code when the encoder binary or either tool
is missing, which is the skip its label requires.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
What a caller declares about the colour of its input, what the library signals in
the bitstream, and the HDR10 static metadata it carries there.

THE COLOUR MODEL IS DECLARED, NOT INFERRED. An input's colour model is a property
of the surface, not of its Vulkan format: the packed 4:4:4 layouts are carried on
RGBA-typed enumerants, so inferring the model from the format reads a Y-Cb-Cr
surface as RGB. The config declares it, and what was a boolean asking whether the
input is RGBA becomes a colour-space enum that can name the cases a boolean cannot.
The registration gate asks its class question under the model the descriptor
declares and falls back to the session's model only where the descriptor declares
none, so a descriptor is judged as the surface it says it is. A colour model the
format cannot carry is refused at registration, and both the refusal and the
session binder name the colour model rather than the format. The model a session is
declared in is immutable across Reconfigure, compared as it resolves against the
input format.

REFUSE OR COERCE, NEVER DIVERGE. A colour description the encoder cannot signal is
either refused at the boundary or coerced to a value it can signal, and the coerced
value is what the bitstream carries. What it must never do is accept the caller's
description, signal something else, and leave the two disagreeing with nothing to
say which is true. The signalled colour matrix is the matrix the preprocess filter
applies, and the range is carried through to the filter.

THE TRANSFER FUNCTION IS DECLARED. The library decides the preprocess conversion;
the transfer function is the caller's declaration and is signalled as such, not
derived from the conversion the library happened to choose.

THE INPUT'S OWN COLOUR IS DECLARED, SEPARATELY FROM THE STREAM'S. The only
input-side colour field was the transfer function, so the RGBA-to-Y'CbCr filter
derived its conversion matrix from the primaries the BITSTREAM advertises -- an
output field standing in for an input property. That is sound only while no
primaries conversion exists, and nothing said so. A chained input-colour
structure completes the declaration: the caller states what its buffer carries,
the library states what it did, and neither states the other's business. An
absent chain is undeclared on every axis at once and reproduces the previous
behaviour exactly, so a caller that says nothing is unaffected. A declared input
colour that disagrees with the bitstream's on an axis this library cannot convert
is refused with the reason, which is the rule the transfer axis already applied.
The comparison reads the caller's own values rather than the bound config's,
because the binder rewrites an unsupplied output field to Unspecified and reading
that fill back would refuse a config nobody contradicted.

WHAT THE DECLARATION MEANS IS STATED PER LANE AND PER CODEC, because two things
needed saying that no document said. The range axis is the one that had a field
and no reader: a caller that chained a full-range declaration got a narrow-range
conversion, a zero range flag and success, which is worse than offering nothing.
The header stated a correct contract and the code failed it, so the code moves.
On the Y'CbCr lane no range scaling happens between the caller's samples and the
coded ones, and the copy filter takes its output range from its input range, so
the input's range IS the stream's: the declaration is applied, and it raises the
video signal type beside the flag, because an unsignalled range is normatively
inferred as studio swing and reported as unknown at a decoder's API boundary. On
the RGB lane the filter PRODUCES the output range from that same flag and samples
its source over the full range, so a declared limited input describes a
conversion that does not happen and is refused with the reason -- the rule the
primaries, matrix and transfer axes already apply. Exactly one contradiction is
refused, because the flag has no undeclared state and reading its false as a
positive claim would refuse every caller that left it alone. AV1 has no absent
state for range at all: the range is unconditional syntax, so the library commits
to a value on every AV1 stream whether one was declared or not.

The structure is additive -- no export, no virtual, no change to the config's own
layout -- and it carries its own size and chain-prefix assertions, because those
assertions are the only compatibility enforcement this interface has.

THE INPUT'S IDENTITY IS DERIVED TWICE, AND THE TWO MUST AGREE. The binder derives
chroma subsampling, bit depth and plane count from the caller's format; the
input-parameter verification reconstructs a format from those same three. Two
derivations of one quantity in opposite directions, with nothing saying they had
to agree -- and whatever the reverse one produced was what the session was
configured with and what the compute filter was built from. The reverse
derivation cannot simply be deleted, which is why this is an assertion rather
than a removal: the packed-alias arm leaves the format unwritten on purpose so
the reconstruction supplies it, and that is the only route by which either packed
4:4:4 layout is nameable. One comparison covers both lanes -- a carry-through on
the RGBA lane, where the reverse derivation spells no RGB layout, and a round
trip on the Y'CbCr lanes -- and the proposition is the same either way. A
disagreement is refused and not repaired: overwriting one side with the other
picks a winner between two derivations without knowing which was wrong, and the
wrong choice is a session configured for a picture the caller is not sending.

HDR10 STATIC METADATA. The mastering display and content light level are carried as
an H.265 SEI message and as AV1 metadata OBUs, from one payload builder shared by
both arms. The metadata struct chains onto the config, so a caller that does not
set it is unaffected, and a caller that chains it against a library which does not
know the structure type is refused loudly by the unknown-type gate rather than
silently ignored.

THE ENCODE-SOURCE REQUEST IS A REQUEST ONLY ON THE DIRECT LANE. On the filter
lane the declared input format describes the FILTER's input -- a three-plane
layout, a depth-converted source, an RGBA surface -- and no device advertises
those as encode sources, because they are the formats that exist to be converted.
Matching against them fails by construction, so every healthy filtered session
announced a chroma mismatch it did not have. The gate is the preprocess-filter
flag rather than the colour model: the colour model caught only the RGB half of
the filter lane and let its Y'CbCr half fall through. The case that survives is
stated as what it is -- a layout substitution on a lane with no conversion, where
the encoder reads the caller's bytes under a layout the caller did not declare --
rather than as a chroma claim the comparison cannot support. Whether that case
should be a refusal rather than a diagnostic is a behavioural question and is
left open.

WHAT THE LIBRARY ROUTES IS A PROPERTY OF THE DECLARED PAIR. A device format-feature
bit is not an answer about this library: it accepts formats the device does not
advertise as encode inputs, because the preprocess filter converts them, and it does
not accept every format the device does advertise. The advertised list is derived from
the library's own taxonomy and is narrower than the accepted set, and the subsampling
rule is stated only for the Y-Cb-Cr inputs it holds for.

THE TAXONOMY IS TRIMMED TO ITS CONTRACT. Semi-planar 4:4:4 is directly encodable, and
the direct set the header names is the set the classifier calls direct. Twelve-bit
semi-planar input is encodable through the filter and not directly: the encode-source
formats a device reports are the 8- and 10-bit semi-planar set, so a 12-bit
registration routed to the direct path asked the encoder for a source format it never
advertises. Its plane layout and subsampling already match; what has to be converted is
the bit depth, which is the compute tier's work. That exposed a second assumption --
the plane count was answered from the class, which can describe neither RGBA nor 12-bit
semi-planar -- so the plane count is read from the format's own layout description,
which is what drives the input geometry and the staging pool's shape. All of it is here
rather than with the capability work, because the classifier these rules live in is the
one this commit gives a colour-model parameter: the routing decision and the model that
parameterises it are the same function.

THE MODIFIER ANSWER IS PROFILE-BLIND, AND THAT IS A LIMIT ON WHAT IT MEANS. It
comes from a format-properties query with no video-profile term at all, unlike
the input-format surfaces, which are keyed on codec and profile. A modifier
reported as carrying the encode-input feature is not thereby usable for the
profile a session will negotiate: it says the format supports the feature under
that modifier, not that this codec and profile accept that format. The two
answers are conjoined, and the header says so.

THE CODEC ARMS' GEOMETRY TERMS ARE CORRECTED WHERE THEY WERE UNREACHABLE. Three
derivations that are defined over the configured geometry did not read it.

The H.265 CPB VCL factor assigned a chroma subsampling flag into a variable named
for a chroma format index and tested it against an index value, so no real input
could satisfy the 4:4:4 arm and every stream took the 4:2:0 factor. ITU-T H.265
Table A.8 gives 4:4:4 twice that, so the level came out at a higher tier than the
profile needs and the default buffer smaller. The comment on the line stated the
intent the code broke; the code now matches it, and the conversion is the one
every other site already uses. The table's depth term is a second defect on the
same function; it is recorded and not repaired.

The AV1 sequence header's subsampling was hardcoded to 4:2:0 under a comment
calling that the only chroma format this encoder admits, while the profile
derivation selects a 4:4:4 profile that requires both subsampling terms to be
zero. The derivation is the codec-correct side, so the comment was the stale one,
and the header is written from the configured chroma.

The narration in the H.264 arm is sorted rather than deleted. One block describes
a configuration that cannot occur, because the field it branches on has no setter
on any surface and one constructor default; it is kept because it states the
intent, and deleting it would delete the record that the intent is unreachable.
A second note records that the arm reads the input side of the geometry where it
should read the encode side -- inert while one writer makes the two equal, and a
defect the moment anything makes them differ.

THE DISPOSITION TABLE DROPS THE ENUMERANT WITH NO ROW. One disposition named a
class the config surface does not contain -- a value refused at initialisation
because honouring it is impossible in this build -- and carried a single
definition and no row anywhere in the table. Every refusal this surface makes is
either a declaration that cannot be honoured or a versioning gate, and both are
properties of the config rather than of the build, so the enumerant is deleted
rather than given a member to justify it: inventing a member to justify an
enumerant is how a table stops describing the thing it tables. Each remaining row
names the full extent of what its field binds, so a row that reads as one derived
value cannot hide three.

THE COVERAGE JUDGES IN Y-Cb-Cr. The encode suite compares in the colour space the
encoder wrote, and compares the HDR payload numerically rather than checking that a
message of the right type is merely present. The filter suite's input check pins the
order of its two single-plane arms, so the arm a failure names is the arm that ran. Where a comment cites the code that
produces a value it describes, it names the symbol rather than a line range,
because a line range in a comment eventually points at neither the write site nor
the declaration.

Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com>
Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
What the library tells a caller it can accept, and the classification behind that
answer.

TWO ENTRY POINTS, NOT FOUR. Capability enumeration is two entry points rather than
four near-duplicates, and both go through one implementation. The enumeration
states what it probes, what it answers and what a context caches, and the context
notes say what a caller can share, index and rely on. The advertised formats and
the video std syntax flags leave the capability structure for two-call
enumerators on the context, so the structure carries scalars only and the format
list is computed rather than stored. Each enumerator leaves a defined count
behind on every return but a null count pointer -- the number written on success
and zero on any error -- so one wrapper can read it without knowing which of them
it called.

EACH ADVERTISED FORMAT NAMES ITS OWN ROUTE. An entry states whether it is the
format the encoder reads, as an optimality answer, and the documentation of that
answer names no conversion mechanism: directness and encode format are different
questions, and the route an image takes is an implementation choice rather than
part of the interface -- so the route enumeration moves to the internal header,
beside the one declaration that has ever had a member of that type. The optimality
answer names the filter that makes a suboptimal entry encodable, which is the half
of the rule that says why the entry is on the list. A build without the preprocess
filter advertises no converted format, instead of naming formats its own session
binder then refuses. Driver-workaround verdicts and input-routing observables are
internal, and are named nowhere on the interface; the status echo's chain carries
library-internal observation structures, so a caller built against the public
header alone has nothing to chain there.

WHAT ROUTES WHERE IS DERIVED, NOT LISTED. The packed 4:4:4 layouts route wherever
the declared colour model names them, and the routable set is derived from the format
tables the compute filter itself reads rather than listed beside them where the two
can drift.

THE PROBE ANSWERS BEFORE A SESSION EXISTS. The capability probe is keyed on bit
depth as well as profile, and states its chroma subsampling per profile instead of
hardcoding it at the call. An explicitly named profile is checked against the
input's chroma subsampling as well as its bit depth, so the pair a caller may feed
the library is answerable before a session exists.

THE ADVERTISED LIST DESCRIBES THIS DEVICE. Two surfaces answered the same
question differently. The enumerator built its list from a capability snapshot
probed at one subsampling and two depths; the point query answered from the
caller's own bound configuration. The two disagreed on the encode format -- the
field that decides the chroma subsampling and bit depth of the bitstream -- for
identical arguments on the same device in the same process. Both now resolve
through one function. The enumerator binds each candidate through the real binder
to obtain the profile, subsampling and depth the session would use, issues a live
device format query at those values, and admits the candidate only if the
encoder-input format it would be handed is in that list. The snapshot survives,
unchanged, as the source for the capability scalars, and no longer carries a
format list for nobody to read.

What that makes visible is capability the library already exercised: semi-planar
4:4:4 at both depths reaches the advertised lists, and ten-bit 4:2:0 reaches the
lists it was accepted at and advertised at neither. Nothing is advertised that
the point query would refuse, because the two compute the same answer. The
ordering clause is restated: with each candidate queried at its own derived
profile there is no single device list left to order by.

The list builder keeps its ordering, uniqueness, capacity and compute-filter
rules and takes an admission callback in place of a device list, so the
production caller supplies the live resolver and a device-free test supplies a
synthetic one. An enumeration silences the binder's per-candidate refusal
diagnostics for its own duration and restores the caller's setting after:
offering every routable format to a profile most of them cannot reach is what the
sweep is for, so a refusal there is the answer rather than an error.

THE BIND SET IS THE STANDARD'S OWN LIMITS TABLE. The profile derivation selects
H.264 High 4:4:4 Predictive, H.265 Range Extensions and AV1 High from 4:4:4
input, and the library emits those streams -- yet naming the same numbers
explicitly was refused as unbindable. The refusal machinery for them already
existed, with a limits table stating the standard's own rule for exactly those
numbers and no way to reach it. The bind set now covers every number that table
states, the standard's limits refuse what the standard forbids, and the device
refuses the rest -- which a caller can discover before a session exists. The
header's constant block names what the binder binds rather than a subset of it,
so the two cannot drift apart again. The assertions that recorded the old
behaviour are inverted with it, and the one that named three numbers as
unbindable is replaced rather than inverted, because two of the three are bound
now and the third would have gone on passing while its stated reason had become
false.

ADVERTISED IS ACCEPTED. The list had become device-truthful; acceptance had not.
Initialisation settled the library half of the question -- does this library
route this pair, can the named profile carry it -- and asked the device nothing,
so a format the device does not encode was bound, given a session, and refused by
the driver at the capability query. That refusal names neither the format nor its
chroma subsampling, and the library relayed it as a bare result code. The 4:2:2
and twelve-bit families are exactly that shape: routed by the taxonomy, bound by
the binder, on no advertised list, and refused late by a message a caller cannot
act on. So the accepted set was wider than the advertised set and a caller had to
consult two surfaces to learn what it could hand in.

Initialisation now gates acceptance on the same device resolve the two query
surfaces answer from. The device half of that resolve is one function: the point
query calls it for the pair a caller names, the enumerator calls it once per
routable candidate, and initialisation calls it for the pair the session
declared. Three surfaces, one statement of the rule, so a query that promises
what initialisation refuses cannot be written without changing that function. The
refusal carries what the driver's does not -- the format, its chroma subsampling,
its bit depth, its plane count, the profile the configuration derived, and which
of four reasons applies -- and it is an enumerated verdict rather than a result
code for one reason: a query owes its caller a yes or a no, an initialisation
boundary owes it a reason. The result code a caller sees is the point query's
own, because one verdict with two spellings would put a caller back to asking
which surface it was talking to.

The one axis on which absence from the list is still not a refusal is the colour
model, which a list of formats cannot express, because the packed 4:4:4 layouts
ride RGBA enumerants. That exception is stated as the single exception it is,
with the point query named as the surface that closes it.

AND A SUBOPTIMAL ENTRY IS TAKEN ONLY ON THE REGISTERED LANE. The conversion runs
against a registered resource, so a suboptimal format is handed in through
registration followed by a registered submit; the unregistered submit refuses it,
because that path has no registration to attach a conversion to. An optimal entry
is accepted on either lane. That is a qualification of the membership rule rather
than an exception to it -- the format is still one the session may declare; what
is restricted is which submit carries it.

BOTH LIVE ACCESSORS ARE NAMED AS TWO. The context header said the modifier query
was the one accessor that reads the driver rather than the snapshot. Moving the
input-format enumeration onto the same live resolver is what made that false, and
it is the same change that makes it necessary, so a rationale outlived the
mechanism it described. Both live accessors are named, each with the reason it
has to be live, and the immutability claim the paragraph exists for is stated
precisely: neither writes context state. The enumerator's two-call idiom had also
borrowed the sibling's snapshot rule -- that the count is a property of an
immutable snapshot, so the counting call and the fetching call cannot disagree.
The conclusion holds and that reason does not, because this list is recomputed on
every call, so the rule is restated on its own footing: the resolve is a
deterministic function of context, device, codec and profile, and a context's
devices do not change for its lifetime, which is the premise the snapshot itself
rests on. The conclusion is kept, so it is pinned -- counted twice and fetched
twice per publishable key, with all four answers required to agree entry for
entry, ordering included, and the counts shown to differ across keys first,
because agreement asserts nothing coming from an observable that reads the same
for everything. The field table's disposition for the bitstream range request is
completed in the same pass: it is still bound to the VUI flag and it is now also
validated, and a table whose stated purpose is to say what happens to each field
has to say both.

THE BINARY THAT CREATES SESSIONS IS GATED ON VALIDATION. The input-format query
tests declined a validation gate on a stated ground: nothing in the file created
a session, an image or a command buffer, so there was no spec-violation surface
to observe. That was a property of the file and it stopped holding when the
initialisation gate arrived, which drives one initialisation per advertised entry
and so runs session creation, device format selection, the filter build and the
image-pool allocation. All rows are gated rather than only the one that creates
sessions: the gate reads the layer output alongside the exit code, so it adds a
way to fail and removes none, and a row with no spec-violation surface gates at
zero and stays there. The objection the old note raised -- that declaring a gate
advertises coverage a test does not have -- is answered by recording a run with
no layer as unmeasured rather than as clean, instead of by declining to measure.

Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com>

THE 4:2:2 EXPECTATIONS ARE DERIVED FROM THE DEVICE, not hardcoded, because the
two suites this commit adds would otherwise be born failing.

EncoderExtInputFormatPointQuery and EncoderExtInputFormatCrossRoute both asserted
that the device REFUSES 4:2:2 input. That is not a library rule but a per-device
capability: Blackwell and later encode 4:2:2, earlier parts do not. On a part that
has it, the point query fails with "got VK_SUCCESS, wanted
VK_ERROR_FORMAT_NOT_SUPPORTED" and the cross-route control with "NV16 (4:2:2) is
not advertised: it was" -- a hardware CAPABILITY reported as a library DEFECT,
red on the first machine that runs it.

Ask the device instead. The point-query table derives the expected result from
VkEncEnumerateInputFormats and, where 4:2:2 is supported, takes the expected
encodeFormat and optimality from the advertised entry, so the point query is
still held to agreeing with the list rather than to a constant.

The cross-route refusal control keeps its teeth. 12-bit is refused by every part
this library targets at every key, so P012 stays an absolute that a merely
permissive enumerator still fails. NV16 stays absolute for every NAMED profile,
all of which are 4:2:0 only. Only DEFAULT, which derives the profile from the
input's own subsampling, becomes a derived check: the list must contain NV16
exactly when the point query accepts it.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Every codec arm's profile, level and sequence-header syntax is defined by the
standards over the values the BITSTREAM carries, and each derivation is read from
the encode side rather than from the input side.

WHICH SIDE IS RIGHT IS NOT A PREFERENCE. The configuration keeps the input's
chroma subsampling and bit depth apart from the encode side's, and the two sets
exist precisely so an encode value can differ from the input value: a chroma
resampler or a device-driven depth downgrade is what would make them. Today one
writer sets the encode side from the input side and nothing else touches either,
so a read of the wrong one costs nothing -- and the three codec arms read
different ones. H.264 read the input on both axes, so on the first change that
made the two differ it would have picked a 4:4:4 profile from a 4:4:4 input and
written a 4:2:0 chroma format from the encode value, inside one sequence
parameter set. H.265 read the encode value for chroma and the input for depth, so
the two halves of one derivation sat on opposite sides of the boundary. AV1 read
the encode value for chroma and the input for depth, and its sequence header
wrote the bit depth from the input while writing the subsampling from the encode
side -- two members of one structure from two sides, with the profile defined
against the one it read from the input. Each arm is moved for that reason and not
for uniformity, and the note that recorded the hazard is deleted with it, because
a defect note that outlives its defect is the next reader's wrong turn.

Two more readers of that boundary move with them. The pre-session format resolver
and the initialisation gate both build a video profile to ask the device about,
and the profile a session actually creates is built from the encode fields -- so
asking about the input's geometry would have predicted a different session than
the one being guarded. The gate's refusal separates the two in its wording as
well: it names the input format and its plane count as the input, and the chroma
subsampling, bit depth and profile as the stream that input derives.

THE ENCODE BIT DEPTH EXISTS BEFORE THE LEVEL IS SELECTED. The H.265 CPB VCL
factor reads the encode bit depths for its depth term. Those were derived at
session creation, which runs after level and tier selection and after the binder,
so the depth term read zero at one of the function's two call sites and the real
depth at the other: one function, two answers, inside one configuration. A 10-bit
4:4:4 stream selected its level and tier with the 8-bit factor and then sized its
default coded picture buffer with the 10-bit one. ITU-T H.265 Table A.8's depth
term is what separates Main 4:4:4 10 from Main 4:4:4 and Main 12 from Main 10,
and it contributed nothing.

The fix is the ordering and not the arithmetic. The encode-side bit depths are
derived where the encode-side chroma subsampling already was, before any codec
arm reads either, so every reader of the pair sees the same value; the
zero-means-unset guards move with them, because an explicit encode depth is a
request and not a default. THIS MOVES SELECTED LEVELS AND TIERS: a stream above
main tier's old bitrate ceiling and below its correct one stops declaring high
tier. Annex A selects the lowest tier and level whose limits the stream
satisfies, so main tier is the correct answer there and high tier was a stricter
claim on the decoder than the bitstream needed. Nothing about the coded pictures
changes. The probe field's own comment stated the defect as a fact about where
the value is read, and is corrected with it.

THE GUARD DISCRIMINATES ON WHICH SIDE AN ARM READ, NOT ONLY ON THE DERIVATION.
The case that accompanies the derivation could not fail on the thing it was
written for: with one writer setting the encode side from the input side, the two
are equal on every state the case can reach, and an assertion over equal values
is satisfied identically whichever side an arm reads. Reverting all three arms to
the input side leaves it green with an unchanged check count. So the two sides
are made to differ without a mutation: the binder gains an internal way to state
an explicit encode depth -- internal header, internal function, no public surface
and no new field on any public struct -- and writes it before the derivation runs,
where the zero-means-unset guard above reads it. The encode side is then the
request and the input side is still the caller's format's, so the profile an arm
derives says which one it read. A second pass over the routable sweep states the
far end of the depth range from each input, asserts its own fixture first -- that
the request landed on the encode side and moved neither the input depth nor the
subsampling -- and counts, per codec arm, how many rows actually had the two
sides implying different profiles, requiring that count to be non-zero. Without
that count the pass could go green having compared every row against itself,
which is the exact defect being repaired.

WHAT IS STILL NOT GUARDED, SAID PLAINLY. The subsampling axis. Its writer is
unconditional, so no request can make the encode and input subsampling differ and
no test can distinguish an arm that reads one from an arm that reads the other.
The depth axis discriminates for all three arms; the subsampling axis is asserted
for derivation only.

THE ORACLE STOPS ASSERTING AN OUT-OF-SPEC PROFILE. Its H.264 arm returned High 10
for anything above eight bits, so twelve-bit 4:2:0 rows expected 110 and
twelve-bit 4:2:2 rows expected 122, neither of which ITU-T H.264 Table A-1 admits
above ten bits. The derivation defect is out of scope; writing its answer into an
expectation as the correct one is not. Those rows are excluded, counted, and
printed with the in-spec answer beside the defect, so the exclusion is visible on
every run instead of being a silently missing assertion. Every expectation is
checked for spec-admissibility before it is used, from a restatement of the three
standards' limits written for that purpose rather than from the library's own
table, because a check that shares the library's table would agree with whatever
that table says.

WHAT IS PROJECTED, AND WHY BOTH SIDES ARE. The binder's probe carries the
encode-side geometry beside the input's, and the profile each arm derived, even
though one writer makes the two sides equal in every configuration a caller can
build: a test that can see only one side cannot say which side an arm read, which
is how three arms came to read different ones.

CaseCodecArmsDeriveTheProfileFromTheEncodeGeometry is dispatched. Written and
not called, it compiled into a suite that reported PASS without ever running
it -- the third time in this series that a case has been authored and left out
of the dispatch list, and the reason the suite's own check count is the thing
worth watching rather than its verdict.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The public structures are pinned at build time, so a change to one of them is a
build break in the reviewer's tree rather than a field silently read at the wrong
offset in somebody else's.

WHAT A SIZE PIN CANNOT SEE. Every public struct already carries its own size, and a
size assertion catches a member appended or removed at the end. It cannot see the
interior: a member removed from the middle and one added at the end leave the size
unchanged, and a member inserted into a run that alignment padding absorbs leaves
the size unchanged too. So each structure additionally pins the OFFSET of the
member that closes each of its padding-absorbed runs. A removal or reorder among
the members before that one -- which the size alone cannot show -- is then a build
break rather than a silent shift of everything after it.

The rationale states the absorption rule at every member width, which is the set of
widths the pins are taken at. Every chainable struct is asserted to carry the
structure-type and chain-pointer prefix. Every public run that a four-byte
insertion would slide carries an assertion on the member that closes it, and where
a structure is one run of equal-width members it is pinned member by member,
because there is no closing member that can stand for the rest. The pins are
described as the review gate they are, and they state the struct-extension rule
they enforce.

THE VERSION CONSTANT IS NOT A COMPATIBILITY COUNTER. It identifies a release. The
compatibility guarantees rest on these assertions, and on the header and the
implementation being vendored and rolled as one unit. So the constant is 1, and the
per-number history of what each earlier value meant is removed rather than carried
forward as a record no consumer can act on -- including the entry recording that
one number had named two different vtables, which no artifact can be recalled to
fix.

THE FIELD TABLE CARRIES EXTENTS. The config field table names the extent of each
field it names, and a check lays the rows across the structure so that a field with
no row fails rather than passing unnoticed.

THE INTERFACE RECORDS WHERE ITS OWN RULE WAS NOT HELD TO. The route to a maximum
quantizer clamp of zero is named as the chained structure it would actually be, and
the header records that its own extension rule has not been held to there, rather
than implying it has.

Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com>
Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
Reconfigure answered success to everything and applied a subset. It now applies
what it can carry, refuses the rest by name, and records what it applied.

WHAT IT CHANGES AND WHAT IT REFUSES. The header names the fields Reconfigure
changes and the fields it refuses, in the header itself rather than by reference to
somewhere else, and names the members Reconfigure reads at all. The encoding fields
it cannot carry are refused with a status naming them, instead of being accepted
and dropped.

THE CONSTANT-QP TRIPLE LANDS ON THE RIGHT FRAME. Reconfigure applies the
constant-QP defaults, and the triple lands on the frame the Reconfigure preceded,
like the six fields around it, rather than one frame later. An out-of-range
constant quantizer is refused at the boundary instead of wrapping into a plausible
value eight bits down.

THE RECORD IS OF WHAT WAS APPLIED. The session's record of its own configuration is
what a later Reconfigure compares against, and it was updated from the RAW config,
which is not what the session ends up running on. A maximum bitrate of zero means
track the average and is coerced on the way in, so a session reconfigured to a rate
with a zero maximum ran capped while the record said zero. A frame-rate numerator
of zero leaves the frame rate alone entirely, so both halves kept the values already
in force while the record stored the zero. A zero denominator beside a non-zero
numerator becomes one, and the record stored zero, which is not a frame rate at
all. Each now records the applied value.

This is a truthfulness fix and not a behaviour change, and the reason is worth
stating because it is what makes the claim checkable: the set of members written
back is disjoint from the set the immutables and the levers compare. Nothing
compared reads a bitrate, a frame rate or a quantizer, so recording the applied
value instead of the raw one cannot move a single refusal decision either way.

THE QUANTIZER CLAMPS ARE CARRIED. The minimum and maximum QP were refused on the
grounds that they reach the driver only through the codec-specific rate-control
structs, which are filled once at codec init from a call site entangled with
session-parameter creation. The first half is true and the second is not: the
rate-control fill is a pure function of config state, and the codec rate-control
command copies its output onto EVERY rate-control command rather than only the
first. So the fill can be re-invoked on the encoder thread and the result rides the
very command this call already causes.

Three properties of that fill are handled rather than assumed. It raises the
use-minimum and use-maximum flags and never lowers them, so re-invoking it in place
could set a clamp and never clear one; the codec layer struct is therefore reset to
its constructed state first, which is the state the fill saw at codec init and which
nothing else writes. And it rewrites the live layer bitrate and frame rate from the
config, which still holds what the session was built with, so those members are
snapshotted and restored around it -- without that, a clamp change would silently
revert a bitrate or frame rate an earlier Reconfigure had applied.

A RECOGNISED STRUCTURE TYPE IS STILL A CHAIN. The chained input-colour structure
is accepted at initialisation and every axis it carries is immutable for the life
of the session, so a caller reusing one config for both entry points must clear
the chain for this one -- and learning that by having the declaration silently
ignored is exactly the accepted-and-dropped class this call refuses. Reconfigure
refuses any chained structure, recognised or not, and the refusal is asserted
with a recognised one rather than only with a junk pointer, which cannot tell
"any chain" from "an unrecognised chain".

THE H.26x ENCODER CLASSES MARK THEIR OVERRIDES CONSISTENTLY. Carrying the
quantizer clamps adds two overriding methods to each codec encoder. A class in
which some overriding members are marked and others are not is inconsistent, and
a consumer that compiles this library with that inconsistency treated as an error
cannot build it, so every overriding declaration in both classes is marked.

Co-Authored-By: Tony Zlatinski <tzlatinski@nvidia.com>
Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Vishesh Garg <visgarg@nvidia.com>
The descriptor-based encoder API sat in the public include directory, so
every declaration it carried -- the session internals, the structures the
implementation passes among its own layers -- was surface a host could reach
and therefore surface that could not change without breaking one.

Split it: what a host needs stays, and the rest moves into
vulkan_video_encoder_ext_internal.h. The remaining header then leaves the
include path entirely for internal/, which is not on any consumer's include
path, so the descriptor API is reachable from inside this library and its own
white-box tests and nowhere else.

Those tests name the internal directory explicitly. That is the point: a test
that reaches past the public surface should have to say so in its build.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Three test targets listed ${VK_DISPATCH_TABLE_SOURCE} in their source lists
without ${VK_DISPATCH_TABLE_HEADER}: common/libs/tests, and the drm_format_mod
and win32_opaque_import suites beneath it.

Listing the generated SOURCE orders the generation of that ONE generated
object. It orders nothing else in the target. An ordinary source in the same
target that includes the generated header -- VkCodecUtils/VkImageResource.cpp,
which vk_filter_test compiles directly -- has no declared dependency on the
generator at all, so a parallel build is free to compile it first and fail with

    VkCodecUtils/HelpersDispatchTable.h: No such file or directory

Listing the generated HEADER is what orders the whole target, because a header
in a target's source list becomes a prerequisite of every object in it. Six
targets in the tree already do this; these three are brought in line with them.

The ordering is what this fixes, not an observed failure rate: seven clean
parallel builds of the unfixed tree (three at -j16, four at -j64) all
succeeded here, so the window is narrow enough on this machine that the
generator wins the race every time. It is won by luck rather than by anything
the build declares, and the two suites that do not compile VkCodecUtils
sources today are one added source away from depending on that luck too.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
zlatinski and others added 4 commits September 19, 2026 01:46
The library's embeddable surface is a Vulkan-shaped descriptor API: flat
structs stamped with a structure type, extended by pNext chains, driven by
free functions. The rest of this library is virtual, abstract, reference-
counted C++, and so is every host that embeds it. This adds the interface
in the idiom of both.

WHAT IT IS. One root object reached through a single exported symbol.
Capabilities and configurations come from it; a session comes from a
configuration; and from a session come the roles a caller uses --
submission, bitstream retrieval, registration, completion, device identity,
diagnostics. A role the implementation does not provide is a null reference
rather than an error returned from a call the caller had no way to know
would fail, so the capability query and the feature check are one operation
done once at setup.

WHY IT IS SHAPED THIS WAY. Extension moves from data layout to vtable
extension. A pNext chain extends a struct the CALLER allocates, so its
layout is pinned the moment any caller compiles; inheritance extends an
interface the LIBRARY allocates, so nothing about layout is observable and
a later version adds behaviour by deriving. One QueryInterface on the base
replaces the descriptor registry -- and, because this library and Chromium
both build without RTTI, it doubles as the typed downcast dynamic_cast
cannot provide here.

TWO RULES THE LIBRARY TAKES BACK. ResolveInputPath answers whether a
configuration reaches the encoder directly, by copy, or through the compute
filter, so the copy-engine-unless-the-samples-must-change rule has one
implementation instead of one per host. HDR metadata is a separate
interface offered only by codecs that can carry it, so an H.264 caller
learns that from a null query at configuration time instead of a refusal
inside session creation.

VkEncClassifyFormat answers, without a device, the four things a host needs
before it allocates: whether a format can be taken, by which route, what
the compute filter needs of the image, and what the session will run at.
None of it depends on a device, and a host picks its buffer format long
before an encoder exists.

WHAT THE VALUE TYPES CARRY is what a host actually reads back, which is
more than a first pass suggested: a per-frame outcome, because a delivered
frame can be a deadline drop carrying no bitstream; the device an image was
allocated on, because all-zero means "not stated" to the cross-device check
and omitting it disables that check rather than weakening it; explicit
handle ownership, because the default transfers and a caller that assumes
otherwise double-closes; an acquire and release fence pair, because that is
what orders a producer against the encode; the ownership echo as an
out-parameter, because it is true on the refusal path where no id is
returned; and the device identity, because two instances give different
handles for the same hardware and a host that adopted a device must be able
to confirm the encoder took the one it meant.

RequireImportExtensions refuses a device lacking what the import paths
need, naming the missing extension, rather than accepting it and failing at
the first import with a message about a handle.

Result is a code and a pointer to a string literal -- sixteen bytes,
trivially copyable, no allocation on a path that runs once per frame. The
header compiles at C++17 while the implementation stays C++20, and names
nothing Xlib defines as a macro.

The implementation is a translation onto the descriptor API and holds no
encoder logic. What it adds is that the rules above are stated once here
rather than in each host.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Four suites on the interface, plus the support every one needs and nothing
about any one of them.

The interface suite drives what the interface promises: a role resolves and
an absent one is a null reference; a role keeps its session alive because
it shares the session's reference count; the H.264-carries-no-HDR rule is a
null query at configuration time rather than a refusal at session creation;
frames go in and a bitstream comes out by both the inline and the
registered path; a session reports the device it runs on, by identity
rather than by handle; and a reconfiguration applies the rate while
refusing the resolution by name. Completion notifications coalesce, so it
checks the callback fires and never more often than frames complete rather
than asserting one per frame and pinning a promise the library does not
make.

The drain suite separates the two calls that end things. A drain that
quietly ends the completion surface leaves a session that still accepts
frames and retires none: no error, no crash, a caller waiting forever. Only
a second batch submitted after a drain catches it.

The release-fence suite checks the property that makes a fence a fence: it
is handed over BEFORE it signals. The floor is half the run rather than
one, because a fence exported after the fact comes back signalled on nearly
every frame and "> 0" passes as soon as one frame races ahead. It runs at
1920x1080 because at a small enough frame the encode finishes inside the
submission call and the check would report the frame size instead. A
submission refused with NotReady is parked and re-sent, which is what that
code means.

The floor guard compiles the public header alone at exactly -std=c++17 with
the window-system defines set. Those defines are as much the point as the
standard: Xlib's Success, None and Status are object-like macros, so they
bite only a translation unit that has included Xlib -- which is every
browser and compositor, and none of this tree's own targets.

A check that could not be made is counted as skipped, not passed, so a run
on a machine that can measure nothing reports that rather than a clean
sheet.

Report::Summarise refuses to pass a run that measured nothing. Returning
success on zero checks makes a suite whose assertions were all skipped --
or all deleted -- indistinguishable from one that ran and held. Zero checks
with a skip recorded now reports SKIP, which the three tests declare through
SKIP_RETURN_CODE 77; zero checks with no skip recorded is a failure, because
a harness that ran nothing and did not say why has not measured the library.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>

Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
An OWN-mode context is cached for the process lifetime so that repeated
create/destroy cycles reuse one VkInstance rather than issuing a second
vkCreateInstance, which a sandboxed process may no longer be permitted to do.
The cache held the only reference and nothing ever released it, so
vkDestroyInstance never ran. Under a driver that audits allocations at exit
that is reported, accurately, as a leak:

  ???NVDBG_MALLOC: final-report LEAKED 22589 BYTES in 22 leaks
  GLNV: Assert#1 !"MEMORY LEAK"

All 22 are instance- and physical-device-scoped -- vkinstance.cpp,
vkphysicaldevice.cpp, nvVkVideoPhysicalDevice.cpp -- which is the instance
that was never destroyed and nothing else. Device-scoped resources were
already being freed.

The rule the cache enforces is "never re-created", not "never destroyed", and
those are separable. VkEncRetireOwnContexts() latches retirement and then
releases the cache: the instance is destroyed, and any later OWN-mode create
is refused rather than allowed to stand a second one up. That is a stronger
guarantee than before, where a post-teardown create was merely impossible
because teardown never happened.

Release stays a call and not a destructor deliberately. It runs
vkDestroyInstance, so it has to happen while the process can still reach the
driver -- at a point the caller picks, not at static destruction, where the
ordering against the loader is undefined.

Reached publicly as IPlatformLifetime::Retire() through QueryInterface, so the
exported surface is unchanged. A caller that simply runs to process exit need
not call it at all.

Measured on the VM, same binary either way: without Retire(), 22589 bytes in
22 leaks and the assert; with it, neither. ctest 60/60.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
A session created from IEncoderPlatform is created ON the platform's context,
and a context already carries the instance and physical device it was built
against. The config nevertheless copied both out of PlatformCreateInfo, so a
platform given a caller's VkInstance produced a config that named them a
second time. The encoder refuses that pairing rather than guess which one
wins:

  [EncoderExt] config.externalInstance / externalPhysicalDevice set on a
  session created with CreateVulkanVideoEncoderExtOnContext. The context
  already supplies both; remove them from the config.

CreateSession therefore failed for every adopted-instance platform, which is
every session a browser creates. A platform that creates its own instance was
unaffected -- the handles it copied were null -- so the standalone tools and
the whole ctest suite passed throughout.

externalDevice and the queue-family indices still pass through. A context
never creates a VkDevice, so a caller supplying one is adding something the
context does not have, which is the opposite of restating it.

Co-Authored-By: Vishesh Garg <visgarg@nvidia.com>
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
@zlatinski
zlatinski force-pushed the chromium-encoder-support branch 3 times, most recently from 18bef1f to 215c874 Compare September 18, 2026 22:59
@zlatinski
zlatinski merged commit ab9cf18 into main Oct 4, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants