v170
v170:
HIGHLIGHTS:
- Core: added
ALLOW_UPDATE_AFTER_SETsupport (a must have VKUPDATE_AFTER_BINDfeature) - Core: dynamic vertex stride (moved from
VertexStreamDesctoCmdSetVertexBuffers) - Core: added ability to create multiple different structured buffer views for a single buffer
- Core: added
CmdZeroBuffer, working with a resource (not a descriptor), implemented using copies (likevkCmdFillBuffer) - Core: fixed multisampling
- Core:
PlaneBitsimprovements - VK: storage
formatdecoration clarifications - Imgui: added extension
NRIImguifor comfortable ImGui rendering (CMake optionNRI_ENABLE_IMGUI_EXTENSION, no ImGui dependency) - Streamer: simplified and improved (with breaking changes)
- RayTracing: opacity micromap support (DXR 1.2 + VK EXT, not NVAPI)
- RayTracing: numerous improvements (with breaking changes)
- numerous improvements and bug fixes, including thread safety (samples work on NVIDIA, AMD, INTEL and APPLE)
BREAKING CHANGES:
- Core:
NRICompatibility.hlslirenamed toNRI.hlsl - Core:
CmdClearStorageBufferandCmdClearStorageTexturemerged intoCmdClearStorage - Core:
CmdSetVertexBuffersarguments grouped intoVertexBufferDescfor comfortable use (added stride, which was removed fromVertexStreamDesc) - Core: changed order of
buffer/textureandmemoryin[Buffer/Texture]MemoryBindingDescto avoid confusion whatoffsetis for - Core:
boolsinPipelineLayoutDescare grouped intoPipelineLayoutBits flags - Core:
DeviceDescfields grouped into structs:aaaBbb=>aaa.bbb(mostly) - Core:
DeviceDesc::isAaaSupported=>features.aaa - Core:
DeviceDesc::isShaderAaaSupported=>shaderFeatures.aaa - Streamer: breaking changes (easy to follow)
- Wrappers:
nriCreateDeviceFromVkDevicerenamed tonriCreateDeviceFromVKDeviceto match the style - SwapChain:
AcquireNextSwapChainTexturerenamed toAcquireNextTexture(can returnOUT_OF_DATEin VK) - SwapChain: breaking changes to respect explicit synchronization between SwapChain textures and rendering (explained usage, see samples)
- RayTracing: breaking changes (easy to follow)
- DeviceCreation:
shaderExtRegister=>d3dShaderExtRegister
DETAILS:
- Core: added
Result::INVALID_AGILITY_SDK, explainedResultvalues - Core: added comments
Threadsafe: yes/no/?, clarified implicit expectations - Core: added comments for buffer and texture usage bits
- Core: added comments clarifying
R, W, RWaccess - Core: added specialized
AccessBits::SCRATCH_BUFFERflag (which can be used withStageBits::ACCELERATION_STRUCTURE) - Core: added
planeBitstoTextureRegionDesc - Core: added
NRI_FORMAT(format)macro toNRI.hlsl, explained usage - Core: added
CmdZeroBuffer(copy based, matchesvkCmdFillBuffer) - Core: added
structureStridetoBufferViewDescto allow passing individual structure strides per buffer view (instead of using the predefined one at a buffer creation time) - Core: added C (in addition to C++) detection to
NRI.hlsl - Core: added
ALLOW_UPDATE_AFTER_SETflag forPipelineLayout,DescriptorPoolandDescriptorRange(no breaking changes, descriptors/data already assumedstatic while set at executewithout this addition) - Core: added
UpdateAfterSetlimits toDeviceDesc(can be significantly greater than non-UpdateAfterSet limits) - Core: added
tiers.resourceBindingtoDeviceDesc - Core: added
SHADER_BINDING_TABLEtoAccessBitssince it's a separate access flag in VK - Core: added
features.pipelineStatisticstoDeviceDescsince it can be unsupported on some platforms - Core: added
availabilitymacro to each extension (unlocks some programmatic magic) - Core: clarified
AccelerationStructurecompatible vertex formats in the main header - Core: clarified
Map/Unmapusage (for modern GAPIs these functions do not invoke any GAPI calls) - Core: clarified
CmdClearStorage - Core: relaxed
AbortExecution, which is not called in some cases (seeResultcomments) - Core: regrouped
AccessBitsandLayout(no functional changes) - Core: grouped unreadable
wallof fields inDeviceDescintostructs(lots of 1:1 breaking changes) - Core: fixed incorrect
videoMemorySizedetection forotheradapters (notDESCRETEand notINTEGRATED) - Core: improved
AdapterDescsorting algorithm - Core: matched logic for
nriEnumerateAdaptersacross GAPIs (simplification) - Core: made
ShadingRateCombiner::KEEPthe 1st in enum to match VK default behavior if not provided - Core: explained usage of
structureStride - Core: reduced
DeviceDescsize - Core:
major & minor versionsmerged into justversion - Imgui: added new extension
- Helper: fixed
UploadDataalignment issues after removing hardcoded embedded alignment - Helper: improved
UploadDataperformance - Helper: filtered out redundant barriers in
DataUpload - Helper: fixed
FormatProps::isSignedforSNORMformats - Streamer: fixed alignment issues- Streamer: fixed multi-threaded access
- Streamer: added
dynamicBufferSizetoStreamerDescto allow static allocation of the dynamic buffer (and disable re-allocation) - Streamer: utilized more flexible
ResourceAllocatorextension - Streamer: removed
AddUpdateRequest-CopyStreamerUpdateRequestscombo with the requirement to keep CPU memory in-between, nowStream[Buffer/Texture]Dataimmediately streams data and returnsbuffer-offsetpair for immediate use - Upscaler: fixed multi-threaded access (actually it's a lie since most of underlying SDKs are single-threaded)
- Upscaler: optimized descriptor sets updates for NIS
- Wrappers: minor improvements for D3D, added acceleration structure flags
- SwapChain: removed incorrectly used internal sempahores, added explicit ones (explained usage)
- RayTracing: added opacity micromap support (DXR 1.2 or VK KHR extension, not NVAPI)
- RayTracing: added
enableD3D12RayTracingValidationtoDeviceCreationDesc - RayTracing: added optional
optimizedSizeforAccelerationStructureandMicromapneeded for copy and compaction - RayTracing: extended query type to query
currentandcompactedsizes ofAccelerationStructuresandMicromaps - RayTracing: lots of
under the hoodimprovements - DeviceCreation: added
d3dZeroBufferSizeneeded forCmdZeroBufferimplementation (may be improved for D3D12) - ResourceAllocator: added basic memory leaks reporting in DEBUG mode
- D3D11/D3D12/VK: fixed multi-threaded access to
DescriptorPoolandCommandAllocator - D3D11/D3D12/VK: reduced size of
SwapChain - D3D11/D3D12/VK: reduced size of
DescriptorSetto 32-40 bytes and eliminated dynamic memory allocations - D3D11/D3D12/VK: all
DescriptorSetget allocated once duringDescriptorPoolcreation - D3D12: re-enabled message
D3D12_MESSAGE_ID_COMMAND_LIST_STATIC_DESCRIPTOR_RESOURCE_DIMENSION_MISMATCHfor AgilitySDK (since modern validation does understand AS used outside of RayGen shaders) - D3D12: added
planeBitshandling in enhanced barriers - D3D12: added auto handling of MSAA- (4Mb) and non-MSAA (64 Kb) memory heap alignment (MSAA textures transparently go into a special
MemoryType) - D3D12: added
tight alignmentsupport (https://devblogs.microsoft.com/directx/agility-sdk-1-716-0-preview-tight-alignment/) - D3D12: added missing locks for command signatures creation
- D3D12: relaxed
GetSubresourceIndexcalculations (same results, no asserts) - D3D12: fixed
NRI_AGILITY_SDK_DIRusage - D3D12: fixed
COLOR_ATTACHMENTdescriptor creation for an MSAA texture - D3D12: hooked up message callback mechanism (if supported)
- D3D12: fixed incorrect
rayTracingShaderGroupIdentifierSizeandrayTracingShaderTableMaxStride - D3D12: fixed
WriteShaderGroupIdentifiers - D3D12: implemented persistent mapping to reach parity with VK
- D3D12: improved
D3D12_HEAP_TYPE_CUSTOMsupport - D3D12: properly queried ray tracing
position fetchfeature - D3D12: fixed query type for devices not supporting mesh shaders
- D3D12: hooked up
structureStridefromBufferViewDesc - D3D12: implemented
DEVICE_UPLOADtoHOST_UPLOADfallback forResourceAllocator - D3D12: reduced size of
DescriptorD3D12 - D3D12: fixed
StageBits::ACCELERATION_STRUCTUREflag by adding missingD3D12_BARRIER_SYNC_EMIT_RAYTRACING_ACCELERATION_STRUCTURE_POSTBUILD_INFOflag - D3D12: unregistered callbacks in Device destructor (could lead to crashes)
- D3D12: fixed not working
WriteAccelerationStructuresSizesandCopyQueries(for acceleration structure sizes) - D3D12: improved code around
QueryPoolinternal buffer used for acceleration structure sizes - D3D12: added optional support for optimized clear values (PR #142)
- D3D12/VK: added forgotten acceleration structure flags here and there
- D3D12/VK: respected allowed 0 values in
TextureDescpassed toAllocateTexture - D3D12/VK: reported
message IDin message callbacks - D3D12/D3D11: disabled misused
ForcedSampleCount(no analog in VK) - D3D12/D3D11: properly queried
shader clockfeature - VK: fixed
uploadBufferTextureSliceAlignmentcalculation - VK: added missing
VkRayTracingPipelineInterfaceCreateInfoKHRfor ray tracing pipeline creation - VK: silently clamp
SwapChainDesc::textureNumto the supported range, returned byGetPhysicalDeviceSurfaceCapabilities2KHR - VK: improved surface formats sorting in
SwapChain - VK: fixed
VK_USE_PLATFORM_Xmacro usage - VK: fixed wrong
BottomLevelGeometryDescconversion - VK: fixed inability to use
NULLforVertexInputDesc - VK: increased device creation robustness if some queue families are not supported
- VK: getting rid of some deprecated names
- VK: added proper VK memory flags for host visible memory for VMA
- VK: properly muted some informational messages caused during
vkCreateInstance - VK: removed
storage read/write without formaterror suppression - VK: clarified
NRI_FORMATusage (unknownis allowed, except some mobile devices) - VK: added
storageReadWithoutFormatandstorageWriteWithoutFormatshader features toDeviceDesc - D3D11: allowed disobeying of the spec to allow to create multiple different structured views for a single buffer by utilizing RAW buffers under the hood (D3D11_MESSAGE_ID_DEVICE_SHADERRESOURCEVIEW_BUFFER_TYPE_MISMATCH error message is silently suppressed, since the solution works). This behavior is only enabled if a buffer has been created with
structureStride = 4. Other sizes (>4) follow the spec and initialize a structured buffer as usual. In other words SB views are treated as RAW if a RAW functionality has been requested during creation - D3D11: fixed
D3D11_MAPPED_SUBRESOURCEdata usage, ifpitchis 0 on some HW (but why?) - D3D11: fixed
GetDatausage inQueryPool(must be wait on) - NONE: simplified code
- Validation: return 0-initialized object pointers on errors
- Validation: various improvements
- CMake:
NRI_ENABLE_AGILITY_SDK_SUPPORTmade visible in the parent project - CMake: Agility SDK version made
CACHE STRINGto become in CMake-GUI (like other options) - CMake: fixed path to Agility SDK
- CMake: fixed detection of
wayland-client.h - CMake:
NRI.hlsladded toIncludefolder - CMake: removed unnecessary SSE4.1 flag
- CMake: updated to v3.30
- CMake: switched to URL instead of GIT_REPOSITORY almost everywhere to avoid getting
.gitfolders - CMake: original archives get removed after deployment to respect disk space
- CMake: added X11/Wayland availability auto-detection
- Actions: greatly reduced tasks time
- updated ShaderMake removing the requirement to install latest and greatest Windows SDK (AgilitySDK is auto downloaded by NRI, DXC is auto downloaded by ShaderMake)
- added a lot of missing
NRI_CALL - lots of changes to respect
nri::usage (not needed in.hppfiles) - resolved some TODOs
- improved validation
- improved docs
- improved robustness
- updated README
- polishing