Repository navigation
General validation errors and test fixes - #239
Merged
Merged
Conversation
zlatinski
force-pushed
the
general-validation-errors-and-test-fixes
branch
from
September 10, 2026 18:09
7838c2e to
fdfe89c
Compare
... add 12-bit coverage, and make a correct rejection provable.
Carries the shared harness changes (config schema, result scoring, reporting)
as well as the decode suite, because that is where they originate; the encode
commit that follows builds on them.
THE MD5 CHECK HAD BEEN SKIPPED FOR EVERY CELL.
The decode harness recorded expected_output_md5 for 48 of its 52 samples and
verified none of them. vk-video-dec-test defaults to a Y4M container and,
when the -o path does not end in .y4m, writes to "<path>.y4m" instead of the
requested path. The harness names its output decoded_<name>.yuv, so the file
never existed where the check looked:
if (should_verify_md5 and output_file and output_file.exists() and ...)
output_file.exists() was False, the block was skipped, and every run was
scored on the decoder's exit code alone. The suite reported 100% while
verifying nothing about the decoded pixels. Raw output is now requested
explicitly unless the sample pins a format in extra_args. The goldens are
md5s of RAW YUV, confirmed rather than assumed: h264_4k_main's stored
716a1a1999bd67ed129b07c749351859 reproduces with --yuv and not with the Y4M
default (58d62074d022c01ef50fb8b8d6abe8a3).
before: 52 tests, 52 "passed", 0 failed (MD5 never checked)
after: 52 tests, 40 passed, 10 FAILED (all 10 MD5 mismatches; every
input file present, so none were missing-content failures)
Of those 10, 7 were stale goldens and 3 were real. Each of the 7 was decided
against an INDEPENDENT ffmpeg reference, never against the golden -- two
conforming decoders must agree bit-for-bit -- and all 7 are byte-identical to
ffmpeg over the full clip, so the decoder was right and the recorded value
was wrong. Byte-identity was re-confirmed at mint time in the same run that
produced each md5, so the new value cannot be "whatever the decoder emitted
that day". Re-minting on a green run alone would have buried the other 3.
av1_basic_10bit av1_cdef_10bit av1_forward_key_frame_10bit
av1_loop_filter_10bit av1_lossless_10bit av1_orderhint_10bit
vp9_320x240_10bits
The remaining 3 are Argon vectors and are TEST issues, not decode defects --
established by measurement, and the skip entries now say so and are scoped to
all drivers rather than radv:
av1_argon_test787 6 frames across 5 geometries (39x223, 59x82, 105x89
x2, 192x189, 35x32). The recorded "OBU frame header
parsing issue" described an apparent W39 H90 vs
W105 H89 disagreement; there is no single resolution,
we report the first frame's and ffmpeg's probe the
modal one. A single flat raw-YUV md5 cannot express
this, so no correct decoder could satisfy it.
av1_argon_test9354_2 14 frames across 8 geometries. The "resolution change
issue" is accurate but describes the STREAM, which
changes resolution by design -- that is the
conformance feature under test.
av1_argon_test1019 multi-resolution as well, and no usable reference:
dav1d fails on the stream ("Unknown Metadata OBU type
0") and writes 0 bytes, so ffmpeg cannot arbitrate.
Both need per-frame validation to be testable at all.
12-BIT COVERAGE.
Four cells -- HEVC Main 12 4:2:0, Main 4:2:2 12, Main 4:4:4 12, and VP9
Profile 2 12-bit 4:2:0 -- each verified byte-identical to an independent
ffmpeg reference over the first 16 frames. Provenance is recorded in each
description and reproducible: all derive from one public HDR10 master via the
memory-compression repo's scripts/gen_12bit_decode_clips.sh, which truncates
at 72s, keeps 256 COUNTED frames (not trusted from the container header) and
transcodes to each chroma/codec, so a difference between the clips is a
format difference and never a content difference.
NEGATIVE CELLS.
Some profiles the hardware simply cannot do. The right behaviour is a clean
rejection, but the suite had no way to say so: the only options were to leave
the cell red forever or to skip it, and both make "correctly rejected" and
"never tested" the same empty result. Nothing would go red the day a driver
started accepting one of these and emitting garbage -- and no positive cell
can catch that, because there is no positive cell for a profile the hardware
does not implement.
New schema on both suites: expected_result: "unsupported", plus an optional
expected_vk_result naming the exact VkResult the rejection must carry. The
cell passes only on exit EX_UNAVAILABLE (69) WITH that result in the output.
It fails if the run succeeded, if the rejection cited a different result (a
missing file or an unrelated capability would otherwise score as "correctly
unsupported"), or if the app crashed. An unrecognised VkResult name is an
error rather than a check that quietly does nothing.
Two decode cells, both verified on RTX 5080 (GB203):
vp9_profile3_12_422_4k_unsupported 422 12-bit; NVDEC has no VP9 4:2:2
h264_high444_unsupported NVDEC has no H.264 4:4:4
The second is the interesting one: NVENC DOES encode 4:4:4, and that
asymmetry makes "the decoder should accept it too" an easy and wrong
conclusion.
Positive control, because a negative cell that cannot fail proves nothing.
Two deliberately-wrong cells were run on hardware: a supported profile
claimed unsupported failed with "completed successfully", and a
genuinely-rejected profile asserting the wrong VkResult failed with "was
never reported". Both paths, plus the crash and generic-error paths, are
pinned by 12 unit tests in tests/unit_tests/test_expected_rejection.py.
Inputs come from scripts/gen_negative_test_content.sh, built from the public
YUVs the suite already downloads and gated afterwards -- a "4:4:4" stream
that is quietly 4:2:0 would make the rejection an accident.
These cells assert NVIDIA hardware behaviour; on an implementation that
really supports one of these profiles the cell fails, which is the correct
signal to look. tests/README.md documents that, and says to resolve it with
a driver-scoped skip-list entry rather than by deleting the cell.
Decode suite on RTX 5080: 58 tests, 53 passed, 0 failed, 4 skipped. It was
52/52 green while verifying nothing.
CONTENT THAT IS GENERATED, NOT DOWNLOADED.
The 12-bit clips and the negative-cell inputs have no source_url: they are
built by a script rather than fetched from the sample bucket, so a fresh
checkout does not have them. That was handled badly on two counts.
download_sample_assets() filtered URL-less samples out of the fetch list and
then reported success for the ones that remained, so the caller was told every
resource was present and the tests failed one by one on a missing file. It now
names the files it cannot fetch, says they are generated locally, and does not
claim to have downloaded them.
A cell whose content is absent AND has no URL to fetch it from is now SKIPPED
rather than failed, with a reason pointing at the generator script named in the
sample's description. Absence means "this machine does not have the content",
not "the decoder is broken", and failing it would paint every checkout without
the generated content red. A cell that HAS a source_url and is still missing
remains an error -- that is a genuine download failure.
Negative cells respect that too: scoring a skipped cell as "expected a
rejection, got skipped" would invent a failure from a test that never ran.
Measured both ways on RTX 5080, by hiding the generated content and restoring
it: without it, 80 tests / 68 passed / 0 failed / 11 skipped; with it, 80 tests
/ 75 passed / 0 failed / 4 skipped.
GENERATE THE CHEAP CONTENT IN CI, SKIP THE REST.
The negative-cell inputs are derived from public YUVs the suite already
downloads, so CI now builds them: apt gains the ffmpeg CLI (it had only
libavformat-dev, the build libraries) and a step runs
scripts/gen_negative_test_content.sh straight into tests/resources/. Measured
at under a second plus a 13 MB fetch, and it turns three SKIPs into three real
cells. Their source_checksum is dropped in favour of the generator's own gates
-- exact byte count, an ffprobe profile/pix_fmt probe, and a peak sample above
the 10-bit range -- because a re-encoded artifact's bytes depend on the local
x264 build, so a hash would fail on a version difference while saying nothing
about whether the format is right.
The 12-bit clips stay skipped, and the reason is a measurement rather than
effort: their master is 1.07 GB from a third-party CDN, and the transcodes are
4K x265 and libvpx-vp9 over 256 frames. Hosting the five generated clips
(60 MB) on the sample bucket is the way to make those cells run anywhere.
A NEGATIVE CELL NEEDS A DEVICE BEFORE IT CAN ASSERT ANYTHING.
Generating that content exposed a real defect in the scoring. A runner with no
Vulkan ICD -- which is what the CI job is -- exits EX_UNAVAILABLE for every
cell, so a negative cell would have been marked NOT_SUPPORTED and then failed
for not reporting the specific VkResult. That is a verdict about the
environment, not the driver.
The scoring now requires evidence that a physical device was selected before it
insists on the exact VkResult; without that marker the cell reports
NOT_SUPPORTED like everything else in that environment. A device that WAS
selected and rejected for a different reason still fails, which is the case
worth catching. Pinned by unit tests for both, plus one for the skipped case.
The scoring moved to tests/libs/video_test_expected_result.py: it is a
self-contained decision about one TestResult, and the framework base was over
the module length limit.
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
... and assert the rejections NVENC must give.
Validation-by-decode was all-or-nothing -- a global --no-validate-with-decoder
-- which is wrong for profiles that are legal to ENCODE but cannot be DECODED
by the same hardware. h264_high444_profile is exactly that: NVENC encodes
H.264 High 4:4:4 Predictive, NVDEC has no H.264 4:4:4 at any bit depth, so the
decode capability query correctly returns
VK_ERROR_VIDEO_PROFILE_FORMAT_NOT_SUPPORTED_KHR and the round-trip can never
succeed. Confirmed by decoding the encoder's own output directly, and
confirmed NOT to be a regression from the demuxer change earlier in this
series (yuv444p both before and after).
The only options were to fail that cell forever or to skip it, and skipping
throws away the ENCODE coverage, which works and is worth having. Per-sample
validate_with_decoder (default true) keeps the encode half tested and drops
only the decode-back step, with the reason recorded in the sample's own
description rather than in a skip list somewhere else.
NEGATIVE CELLS, using the expected_result schema added in the preceding
commit. Three profiles NVENC cannot do, all verified on RTX 5080 (GB203):
h265_12bit_unsupported the driver advertises 12-bit under
__VK_VIDEO_DECODE_USAGE_MASK only; NVENC has no
12-bit encode on any shipping chip. Decode at
12-bit IS supported and has four positive cells,
so this keeps the encode half of that asymmetry
measured -- 12-bit encode quietly starting to
"work", or silently downshifting to 10-bit,
becomes a failure rather than a silence.
av1_high_profile NVENC encodes AV1 Main (profile 0) only; High
requires 4:4:4.
av1_professional_profile ditto; Professional requires 4:2:2 or 12-bit.
The two AV1 cells already existed and scored a plain N/S. That is a soft
outcome: an N/S and a silently-accepted profile are both non-failures, so
those cells would have turned green, not red, on a regression.
The 12-bit raw input comes from scripts/gen_negative_test_content.sh and is
gated on content, not on its filename -- samples must exceed the 10-bit range
(peak 3716) or it is a 10-bit file with a 12-bit name and the encoder would be
rejected for the wrong reason.
Encode suite on RTX 5080: 22 tests, 22 passed, 0 failed, 0 not-supported
(was 18 passed / 1 failed / 2 not-supported).
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
A raw golden hashes the decoded surface as it comes out of the dump path,
so its byte layout follows the decode output image. The same pixels can
therefore hash differently on another architecture -- observed on
RTX 3080 Ti vs RTX 5080 for HEVC Main 12 4:2:0, where every sample matched
an independent ffmpeg decode but the raw hash did not. Chasing that costs
real time, and the usual fix (re-mint the golden on the new part) silently
gives up cross-architecture coverage.
Two mechanisms:
expected_output_y4m_md5 -- Y4M is planar and self-describing, so the
hash covers pixels rather than surface layout. Preferred when present.
resolve_expected_md5() -- a golden may be a plain string or a map keyed
by GPU-name fragment ("RTX 50", "RTX 30", "default"). Longest matching
fragment wins, matching is case-insensitive, and "default" is the
fallback. For values that legitimately differ, not for papering over
ones that should not.
42 cells gain a Y4M golden, each minted from a run that was verified two
ways first: deterministic across repeat runs, and sample-exact against an
independent ffmpeg decode. A golden minted from an unverified run just
freezes whatever the decoder did that afternoon.
hevc_main12_420_4k gets no Y4M golden. Its runs went through the compute
filter -- "--enablePostProcessFilter 0" selected filter type 0 instead of
disabling it -- and the filter corrupts this clip from frame 8 onward at a
row boundary that moves between runs. Anything minted from those runs is
untrustworthy, so the cell keeps only its raw golden and its description
now records what must be re-verified before a portable one can be added.
The previous description blamed the HEVC DPB for that corruption, which
was wrong; it is a filter defect, and saying so in the cell stops the next
reader from re-investigating the decoder.
tests/README.md gains a Goldens section covering when to use which, and
how to mint one without freezing a bad run.
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
--enablePostProcessFilter takes a filter TYPE, not a boolean. The decoder's
default is -1, which disables the post-process pass; 0 is a long-standing
legacy value that selects the first filter.
The harness passed "0" on every decode cell that did not explicitly request
a filter, intending "off". It therefore ran decode PLUS a compute shader on
those cells for as long as the option has been wired up. That is why a
filter defect presented as a decoder defect: the decode-only cells were
never decode-only.
Fixed by omitting the option so the decoder's own default applies. It
cannot be spelled explicitly -- the argument parser rejects "-1" with
"we don't allow values starting with `-` by default".
Measured on the Linux VM (GA102, driver 620.18), HEVC Main 12 4:2:0 4K,
16 frames:
filter off (omitted) MD5 c100ec5f... identical across 3 runs, and
16/16 frames sample-exact against an independent
ffmpeg decode
filter type 0 3 different MD5s across 3 runs, same byte size
So the decoder is correct and the compute post-process path is
nondeterministic on this clip. The filter defect is real and now has
numbers behind it; it is not addressed here.
hevc_main12_420_4k keeps its Y4M golden. It was minted on Blackwell and
reproduces bit-exactly on GA102 with the filter genuinely off, so it is
portable as intended -- it had only ever looked wrong because the harness
was running it through the filter. Its description no longer claims a DPB
defect, which was never the cause.
Full decode suite on the VM after this change: 58 tests, 52 passed,
0 failed, 0 crashed, 2 not supported.
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The compute filter's submit waited its decode-to-filter semaphore with
const VkPipelineStageFlags2KHR waitDecoderStageMasks =
VK_PIPELINE_STAGE_2_VIDEO_DECODE_BIT_KHR;
but that submit goes to the COMPUTE queue. A semaphore wait stage mask may
only name stages the target queue family supports
(VUID-VkSemaphoreSubmitInfo-stageMask-03929), and no compute family exposes
the video-decode stage. On the test GPU (RTX 3080 Ti):
queueProperties[0]: GRAPHICS | COMPUTE | TRANSFER | SPARSE
queueProperties[2]: COMPUTE | TRANSFER | SPARSE
queueProperties[3]: TRANSFER | SPARSE | VIDEO_DECODE
whichever of 0 or 2 GetComputeQueueFamilyIdx() returns, the video-decode
stage is not in it.
Wait at ALL_COMMANDS instead. The gated work is the dispatch, but the
command buffer also begins with a layout transition, which is a
transfer-class operation, so naming ALL_COMMANDS covers the whole buffer
and does not have to be revisited as the recording grows.
Note this is a latent correctness issue, not the cause of the compute-filter
corruption fixed in 0583a311: swapping this mask to ALL_COMMANDS was tested
in isolation against a control and did not change that corruption at all
(still first bad frame 8, still a different md5 every run). It is also not
reported by the Khronos validation layers, which do not appear to implement
this check -- the queue family properties above are the evidence.
Verified: 2 runs byte-identical at c100ec5f... with 0/16 frames differing
from an independent ffmpeg decode, i.e. no behavioural change.
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The YCbCr compute filter's generated GLSL declares its per-plane images as
image2DArray and addresses them with ivec3(pos, layer):
imageStore(outputImageY, ivec3(lumaPos + ivec2(x, y), pushConstants.dstLayer), …)
imageLoad (inputImageY, ivec3(pos, pushConstants.srcLayer))
but the per-plane views were created as VK_IMAGE_VIEW_TYPE_2D whenever the
view covered a single layer. A view's type must match the Dim and Arrayed
operands of the shader's OpTypeImage
(VUID-vkCmdDispatch-viewType-07752), so those accesses were undefined even
though the layer index is always 0 in practice -- the per-slot view already
selects the layer.
Create the per-plane views, which exist specifically to be bound as storage
images to the filter, as VK_IMAGE_VIEW_TYPE_2D_ARRAY. A one-layer array view
is legal and matches what the shader declares. The combined view used for
display is untouched and stays 2D.
Before: 20 occurrences of VUID-vkCmdDispatch-viewType-07752 per run, naming
inputImageY, inputImageCbCr, outputImageY and outputImageCbCr.
After: zero.
No behavioural change -- 2 runs byte-identical at c100ec5f... with 0/16
frames differing from an independent ffmpeg decode, and the YCBCR2RGBA
filter (type 2) still produces output.
This was NOT the cause of the compute-filter corruption fixed in 0583a311:
it fires identically on a 4:4:4 clip that decodes exactly over 32 frames.
Fixing it removes undefined behaviour that was masking nothing here but
would eventually bite on a multi-layer output.
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
UpdateImageDescriptorSets() applied the caller's single imageLayout to every descriptor it wrote. For a YCbCr input that layout is the sampled-image layout, so the per-plane descriptors -- which are VK_DESCRIPTOR_TYPE_STORAGE_IMAGE -- were pushed as VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL. A storage image is only defined in VK_IMAGE_LAYOUT_GENERAL (VUID-VkDescriptorImageInfo-imageView-06711), so the descriptor never described the access the shader actually performs. Decide the layout per descriptor instead: a STORAGE_IMAGE always gets GENERAL, everything else keeps the caller's layout. Both RecordCommandBuffer() overloads now also declare the input side GENERAL unconditionally rather than picking SHADER_READ_ONLY_OPTIMAL whenever a YCbCr sampler happens to exist. GENERAL is valid for both sampled and storage reads, the filter never transitions its input to SHADER_READ_ONLY_OPTIMAL itself, and no caller hands it an input in that layout -- the decoder hands over a VIDEO_DECODE_DPB_KHR image it moves to GENERAL, and the encoder hands over a GENERAL/TRANSFER_SRC linear staging image. Decode suite 58 tests / 52 passed / 0 failed and encoder suite 22 / 16 / 0 unchanged; decode output byte-identical. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…atch The compute post-process filter reads the decoder's DPB layer as a storage image, which is only defined in VK_IMAGE_LAYOUT_GENERAL. Nothing ever moved it there: the decoder transitions a DPB layer exactly once, from VK_IMAGE_LAYOUT_UNDEFINED to VK_IMAGE_LAYOUT_VIDEO_DECODE_DPB_KHR on its first use, and it stays in VIDEO_DECODE_DPB_KHR for the rest of the stream. So the image's actual layout never matched the layout its descriptor declared. Emit the transition in the filter's command buffer before the dispatch, and transition back to VIDEO_DECODE_DPB_KHR after it, so the decoder's own layout tracking -- which still believes the layer is in VIDEO_DECODE_DPB_KHR -- stays true and later frames can keep referencing the layer without a transition. Pairs with the preceding commit, which makes the storage-image descriptors declare GENERAL; together they make the declared and actual layouts agree. Decode suite 58 tests / 52 passed / 0 failed and encoder suite 22 / 16 / 0 unchanged; decode output byte-identical. Costs about 1% of frame time, inside the run-to-run spread. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…4:2:2 profiles STD_VIDEO_H264_PROFILE_IDC_HIGH_10 and _HIGH_422 were added to StdVideoH264ProfileIdc in Vulkan-Headers 1.4.360. Referencing them unguarded makes these two files compile only where the headers in the include path are at least that new -- true on Linux, which builds against the copy FetchContent pulls into _deps, and false against an installed Windows SDK of 1.4.341, where the enum stops at HIGH_444_PREDICTIVE. The Windows renderer build fails with C2065 on both identifiers. Define them when absent. The values are not ours to choose: 110 and 122 are the profile_idc numbers fixed by the H.264 specification, and are exactly what 1.4.360 assigns. Where the header does define them the shim is inert. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…to a fault InitEncoder() clamped encodeWidth/encodeHeight to videoCapabilities.maxCodedExtent but left encodeMaxWidth/encodeMaxHeight at the unclamped request, and those are what the video session and DPB are sized from. For example, asking H.264 for 7680x4320 on a part whose maximum is 4096x4096 therefore built a session larger than the device admits, and the encoder faulted with an access violation. Two changes. Fail with VK_ERROR_FORMAT_NOT_SUPPORTED when the requested extent exceeds the profile's maximum: silently encoding a cropped picture is worse than an error, because the caller compares the result against a reference at the size it asked for and sees a quality failure with no cause. And clamp the session maximums alongside the extent so the remaining paths cannot build an over-sized session at all. The min-clamp is left as it was -- padding up to minCodedExtent is benign and does not change the picture the caller gets back. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
The per-codec video-encode extensions are requested as optional, so a device without one is still selected and the miss is only warned about. Initialisation then continues and the encoder faults dereferencing state the driver never created. After the physical device is chosen, check that it advertises the extension the requested codec needs and return VK_ERROR_EXTENSION_NOT_PRESENT when it does not. H.264 and H.265 are covered by the same switch, so the next part that lacks one of those fails the same way rather than faulting. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…h bits 10-bit HEVC encodes came out as ~3.7 KB of flat content where 8-bit produced ~1.9 MB from the same scene. The bitstream was structurally valid -- 62 frames, yuv420p10le -- but near-constant, so PSNR failed and test_yuv_baseline.py reported two failures. DetectInputMsbShift() treated "all low bits clear" as the test for left-aligned input. That does not hold: in P010 the bottom (16 - bpp) bits are explicitly UNDEFINED, and a renderer dump of a 10-bit surface carries noise there, spanning the full 16-bit range. Such data failed BOTH the left- and right-aligned tests and fell through to the "ambiguous" branch, which returns the documented default shift -- multiplying every sample by 64 and saturating the frame to white. Only one of the two signals is reliable. "All high bits clear" proves the data is right-aligned and must be shifted. Its converse does not follow, so treat any sample setting a bit above the bit depth as proof the data is NOT right-aligned and return 0: shifting could only overflow. The genuinely degenerate case (all bits clear, i.e. all-zero data) keeps the documented default. Verified with 10-bit 1280x720 h265 3,764 -> 1,657,714 bytes and test_yuv_baseline.py 6/2 -> 8/0, with the 8-bit control byte-identical at 1,947,686. The same file measured 3.09 dB read as yuv420p10le versus 38.53 dB read as yuv420p16le, independently confirming the data is left-aligned. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
… chain vkCmdBeginVideoCodingKHR reported VUID-vkCmdBeginVideoCodingKHR-pBeginInfo-08254 on nearly every frame -- 14 times in an 8-frame H.264 encode: The video encode rate control information specified when beginning the video coding scope does not match the currently configured rate control state: VkVideoEncodeH264RateControlInfoKHR is not in the pNext chain but the current device state for its gopFrameCount member is set (30). The spec requires the rate-control chain passed to BeginCoding to MATCH the state configured on the session. That state is established by CmdControlVideoCodingKHR with the codec-specific struct chained on -- it is the head of pControlCmdChain -- but the cached copy used for every later frame's BeginCoding set pNext = nullptr, dropping exactly that struct. Cache it alongside the base struct and re-link the two. The cache has to be codec-agnostic because this base class does not know its codec, hence the union over the H264/H265/AV1 rate-control structs; both halves are encoder-owned storage, so the chain stays valid for every frame that reuses the cached state. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
Both are reported by any client that creates a device without video queues --
the filter test app is exactly that ("Num Decode Queues: 0, Num Encode Queues: 0").
Neither is specific to the filter or to the test suite; they were simply never
seen because nothing enabled the validation layers on this path.
VUID-VkDeviceQueueCreateInfo-pQueuePriorities-parameter: the priorities vector
was sized by max(numDecodeQueues, numEncodeQueues), but the graphics, present,
compute and transfer entries below each ask for queueCount = 1 regardless of how
many VIDEO queues were requested. With no video queues the vector is empty, and
.data() on an empty vector is nullptr -- a null pQueuePriorities against a
queueCount of 1. Size it to at least one entry.
VUID-VkDeviceCreateInfo-pNext-pNext: VkPhysicalDeviceVideoEncodeIntraRefreshFeaturesKHR
was chained unconditionally, but VK_KHR_video_encode_intra_refresh is requested
only by the encoder path, so every other client chained a feature struct into a
device that never enabled its extension. Chain it only when the extension is
enabled; the queried value is never read, so skipping it costs nothing. Note the
message misleads -- the extension postdates the validation layer's headers, so it
prints as "unknown VkStructureType (1000552004)", which reads like a
header-version complaint rather than the real error.
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…ate() overloads
VkImageResourceView::Create has two overloads that both build per-plane views and
disagreed on the type: the 7-arg one forces VK_IMAGE_VIEW_TYPE_2D_ARRAY, the
5-arg one left it following layerCount. Per-plane storage views are bound to the
YCbCr compute filter, whose generated GLSL declares them image2DArray and
addresses them with ivec3(pos, layer), and a view's type must match the
Dim/Arrayed operands of the shader's OpTypeImage
(VUID-vkCmdDispatch-viewType-07752).
Because the two overloads disagreed, the same VUID was reported with OPPOSITE
"is/should-be" directions depending on which path built the resource, and no
shader-side setting could satisfy both. Make them agree. The combined view
deliberately keeps following layerCount: presentation and the decoder consume it
as a sampled 2D image.
This also carries the matching shader-side change, and the two cannot be
separated. ShaderGeneratePlaneDescriptors' imageArray argument must follow the
OUTPUT IMAGE TYPE, because that is what decides which view the descriptor binds:
multi-planar -> per-plane view, now always 2D_ARRAY -> image2DArray, ivec3
packed -> SINGLE-plane, so no per-plane view exists; it binds the
combined view, which follows layerCount -> image2D, ivec2
Hence (m_outputPackedYcbcr == nullptr) rather than a bare true. Passing true
unconditionally compiles no shader at all for the packed 4:4:4 formats --
"imageStore: no matching overloaded function found" -- taking AYUV and Y410 down
while vk_filter_test stays 63/63 and both encode matrices stay green;
verify_packed_formats.py is the only suite that catches it. Chasing it from the
store side instead (forcing the packed store to ivec3) makes the shader compile
and then fails at dispatch with the VUID pointing the other way.
Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…ion errors Validation was reachable only through --verbose, behind a log firehose, so in practice nobody enabled it -- which is how a per-dispatch VUID (vkCmdDispatch-viewType-07752) survived a week in the filter with the suite reporting 63/63 the whole time. Two halves. VulkanDeviceContext now COUNTS ERROR-severity messages that survive the suppressed-id list, and vk_filter_test fails the run on a nonzero count: a suite that only prints VUIDs is not a gate. And --validate enables the layers without --verbose, so the gate is usable in CI. The two device-creation VUIDs this would otherwise trip on are fixed in their own commit earlier in this series, so this one adds only the gate. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…ng 9 FLAG_ENABLE_Y_SUBSAMPLING makes the generated shader declare and statically use binding 9 (subsampledImageY), but the test left execDesc.numOutputs at 1, so the dispatch ran with that descriptor never written -- VUID-vkCmdDispatch-None-08114. The case still reported PASS, because writes to an unbound descriptor are simply discarded. Map the subsampled image explicitly into output slot 1 rather than by position, since TC091 appends it to a dual-output config while the filter wants it in slot 1. It also declared that output as NV12 with a "placeholder - actually R8" comment. That is not harmless: a multi-planar format gets a per-plane view, which is now always VK_IMAGE_VIEW_TYPE_2D_ARRAY, while the shader declares subsampledImageY as a plain image2D -- VUID-vkCmdDispatch-viewType-07752. Add a real R8 test format and use it, which is also what the AQ path allocates for this target. Found by the --validate gate. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
…s them VkVideoEncodeIntraRefreshCapabilitiesKHR was chained into the video-capabilities query unconditionally. On a device that does not expose VK_KHR_video_encode_intra_refresh that is VUID-VkVideoCapabilitiesKHR-pNext-pNext. Unchained, the struct stays zero-initialised, which is the correct "no intra-refresh support" answer for the callers that read it. This is the same class as the device-creation chain fixed in VulkanDeviceContext, but not the same case: there the app had not ENABLED the extension, so the fix removed a real error. Here it is supported, so this is defensive only. Signed-off-by: Tony Zlatinski <tzlatinski@nvidia.com>
zlatinski
force-pushed
the
general-validation-errors-and-test-fixes
branch
from
September 10, 2026 20:21
fdfe89c to
685ec98
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.