Skip to content

Releases: janblade/OptiScaler-F5-DLSSNR-Multipass

v0.1.22 - F5 - Tune for This Scene + Eye Adaptation

Choose a tag to compare

F5(Fork From a Fork From a Fork) v0.1.22 Release:

Two new exposure tools for DLSS-NR: Tune for this scene finds the model input brightness that gives NR the most detail, and Eye adaptation stops Automatic exposure from pumping. Both work on D3D12 and Vulkan. Only OptiScaler.dll changes.

New

Tune for this scene

A Tune for this scene button under the brightness slider of Automatic exposure and of Game exposure.

  • What it does: it tries the slider across its useful range on the picture on screen and measures, at each step, how much detail NR adds and how much it flickers. It also checks how much of the picture the model sees as crushed black or clipped white. Then it offers the best step: Apply sets the slider, Keep leaves it as it was.
  • How to use it:
    • Hold the camera still while it runs (4 to 8 seconds). The picture gets brighter and darker on purpose.
    • A pause screen or a photo mode works best. In RDR2 the idle camera sways, so use photo mode.
    • If the camera moves, it stops and says so. Nothing changes until you press Apply.
  • The result is an offset on the exposure, so it keeps following the scene afterwards. Tuning once on a typical shot is usually enough.
  • Several model passes: it runs and measures only the first pass. The first pass is the only one that sees the game's picture, so the result holds for any pass count. Your other passes skip the few seconds of the run and resume after it.
  • Follow the game's exposure no longer needs turning off: Tune measures against Automatic's own exposure. Follow re-learns during the run, so the tuned value holds once Follow takes over again.
  • Details under the result shows the score for every step.

Eye adaptation (Automatic exposure)

  • What it does: Automatic exposure now adapts to a brightness change over about a second, like an eye. Before, it jumped every frame. A camera zoom, or a brief shot of a dark crowd or a bright floor, no longer pumps the picture's brightness and tone. In NBA 2K27 that pumping showed as the picture turning yellow after a made basket.
  • Camera cuts: a cut the game announces is followed at once.
  • New slider: Eye adaptation, under Automatic exposure, from 0 to 5 s. The default is 1 s: about two thirds of a change is followed after that long. Off behaves exactly like v0.1.21.
  • Follow the game's exposure: eye adaptation doesn't apply while Automatic follows the game's exposure, because the game's own exposure already adapts.

Fixed

  • Vulkan crash after a device change: after the graphics device was lost and recreated (driver reset, some alt-tab cases), Vulkan could reuse Automatic exposure's image from the old device and crash. It is now rebuilt on the new device.
  • Vulkan per-frame buffers: Vulkan could rewrite per-frame shader settings while a frame still in flight was reading them. It now has room for the extra work these features add.

Upgrading

  • Replace OptiScaler.dll, or whatever name your setup uses for it, such as dxgi.dll or winmm.dll. Keep your OptiScaler.ini, nvngx.dll_dlssnr.dll and nvngx_dlssnr.dll.
  • Automatic exposure behaves differently by default: it now adapts over 1 s instead of changing every frame. To get the old behaviour, set Eye adaptation to off, or add AutoExposureAdaptSeconds = 0 under [DlssNr] in the ini.

Tested

  • Card: RTX 5070 Ti.
  • NBA 2K27 (D3D12, 3 passes): Tune on Game exposure and Automatic, including repeat runs from different starting values.
  • Red Dead Redemption 2 (Vulkan, 2 passes): Tune with Follow the game's exposure on, and eye adaptation.

Unchanged

Everything else is the same as in v0.1.21, including the default NVIDIA path (nvngx.dll_dlssnr.dll at the package root) and the vendor-neutral backend in Optional\ (AMD / Intel).

Known issues

  • #58: Vulkan with the DLSS-on-D3D12 backend can run DLSS-NR twice per frame.
  • #59: with the overlay open, a click can count twice in games that read the mouse through both window messages and DirectInput.
  • #60: overlay colours may be off when OptiScaler forces HDR output.

v0.1.21 - F5 - Plain FP16 Reuse Bottleneck Fix

Choose a tag to compare

F5(Fork From a Fork From a Fork) v0.1.21 Release:

A hotfix for Reuse bottleneck on the plain FP16 kernels. Only OptiScaler.dll changes.

Fixed

  • Flicker with Reuse bottleneck on the plain FP16 kernels, for good this time. Some modified DLSS-NR DLLs run the model on plain FP16 kernels instead of NVIDIA's FP8 ones. The v0.1.20 fix removed one cause, but users still saw flicker with Reuse on, especially at low frame rates and with 2 or more passes.
    • Cause: with those kernels, the memory the ViT result is kept in for the next frame is shared with another part of the network, which overwrites it every frame. On a reused frame the rest of the model read that other data instead of the ViT result.
    • Fix: on a reused frame OptiScaler now still runs the last, very small kernel of the ViT run. It rebuilds that result from a copy that does survive between frames. Reuse still skips the rest of the run.
    • Cost: about 0.05 ms per reused frame on an RTX 5070 Ti; Reuse still saves about 1.3 ms per pass there.
    • NVIDIA's own FP8 kernels are not affected and work exactly as in v0.1.20.

Upgrading

  • Replace OptiScaler.dll (or whatever name your setup uses for it, such as dxgi.dll or winmm.dll). Nothing else changes: keep your OptiScaler.ini, nvngx.dll_dlssnr.dll and nvngx_dlssnr.dll.
  • If you turned Reuse bottleneck: plain FP16 kernels off because of flicker, you can turn it back on.

Unchanged

  • Everything else is the same as in v0.1.20, including the default NVIDIA path (nvngx.dll_dlssnr.dll at the package root) and the vendor-neutral backend in Optional\ (AMD / Intel).
  • Tested on an RTX 5070 Ti in NBA 2K27 (D3D12), 2 passes, at 20 and about 50 fps: no flicker with Reuse on, with both the plain FP16 and the FP8 kernels.

Known issues

  • #58: Vulkan with the DLSS-on-D3D12 backend can run DLSS-NR twice per frame.
  • #59: with the overlay open, a click can count twice in games that read the mouse through both window messages and DirectInput.
  • #60: overlay colours may be off when OptiScaler forces HDR output.

v0.1.20 - F5 - Bottleneck Reuse Fix + Upstream Overlay/Input Updates

Choose a tag to compare

F5(Fork From a Fork From a Fork) v0.1.20 Release:

This release fixes flicker with Reuse bottleneck on the plain FP16 kernels, and gives Reuse one setting per kernel set. The pass presets no longer change Reuse. From upstream OptiScaler it brings the latest overlay and HDR menu rendering, input system fixes, and a faster DX12 resource tracker for OptiFG with Hudfix. The update check now uses the new repository name and no longer floods the log.

Fixed

  • Flicker with Reuse bottleneck on the plain FP16 kernels. Some modified DLSS-NR DLLs run the model on plain FP16 kernels instead of NVIDIA's FP8 ones. With Reuse on, dark, lamp-lit distant objects flickered. On a reused frame the ViT run is skipped, and the kernels on either side of the gap could overlap on the GPU. OptiScaler now puts a GPU barrier where the skipped run would have been. On NVIDIA's own FP8 kernels this costs nothing measurable: 2.630 ms per reused frame with the barrier and 2.630 ms without, on an RTX 5070 Ti.
  • Update check.
    • It now uses the new repository name, janblade/OptiScaler-F5-DLSSNR-Multipass.
    • When the network isn't reachable yet while the game starts, it tries once more after 10 seconds.
    • A failed check used to be written to OptiScaler.log on every frame the menu was drawn, thousands of lines per session. It's now logged once.

New

  • One Reuse bottleneck checkbox per kernel set.
    • Reuse bottleneck: FP8 kernels (VitEvery) is used with NVIDIA's DLL and FP8-based builds.
    • Reuse bottleneck: plain FP16 kernels (new VitEveryPlain) is used with the plain FP16 kernels.
    • OptiScaler reads which kernels the model is running and applies only the matching checkbox. Both are on by default.
    • A new Kernel set in use line in the menu shows which set is running.

Changed

  • The 1/2/3-pass preset buttons no longer change Reuse bottleneck. In v0.1.19 they switched it on. They now leave both checkboxes as you set them.

From upstream OptiScaler

  • Overlay and HDR menu: better HDR output handling, menu rendering fixes under HDR, and cleaner resource ownership for the DX12 overlay. The Vulkan-with-D3D12 bridge gets upstream's fixes and improvements. The DLSS-NR passes on that bridge are unchanged.
  • Input system: fixes from upstream, including work to prevent input deadlocks. HID mouse hooks are now off. This fork's overlay input over DirectInput (Assetto Corsa + CSP) is kept: clicks in the menu, gamepads and wheels blocked while the menu is open, and Alt+F4 always works.
  • DX12 resource tracker, faster: cdozdil's performance series for the tracker behind Hudfix. It adds early rejects, fewer locks, per-command-list binding tracking, a fast descriptor copy path and heap caching. It only runs with OptiFG and Hudfix (FG input "Upscaler"), so that's where it can help.
    • New option HUDFixPersistentBindings (default on): the new binding tracking system for hudless resources. Set it to false for the old Hudfix behaviour.
    • The relaxed hudless resolution check is now up to 5% per axis instead of 32 pixels.
    • Upstream removed the "disable Hudfix" quirk for The Last of Us and Rise of the Tomb Raider. If the HUD misbehaves there with OptiFG, please report it.

Upgrading

  • VitEveryPlain is new. Missing from your ini means on. Your existing VitEvery setting still applies to the FP8 kernels.
  • If you turned Reuse off, a preset button no longer turns it back on.

Known issues

  • #58: Vulkan with the DLSS-on-D3D12 backend can run DLSS-NR twice per frame.
  • #59: with the overlay open, a click can count twice in games that read the mouse through both window messages and DirectInput.
  • #60: overlay colours may be off when OptiScaler forces HDR output.

Unchanged

  • The default NVIDIA path (nvngx.dll_dlssnr.dll at the package root) and the vendor-neutral backend in Optional\ (AMD / Intel) are the same as in v0.1.19.
  • Tested on an RTX 5070 Ti in NBA 2K27 (D3D12):
    • DLSS-NR with FP8 and plain FP16 kernels, with and without Reuse;
    • OptiFG with Hudfix;
    • DLSS-G frame generation input;
    • the overlay menu and input.

v0.1.19 - F5 - Automatic Exposure by Default + Bottleneck Reuse On

Choose a tag to compare

@janblade janblade released this 26 Sep 06:57
af7e5fa

F5(Fork From a Fork From a Fork) v0.1.19 Release:

Automatic exposure is now the default, at one simple brightness for every game. Reusing the ViT bottleneck is on by default. The 1/2/3-pass presets have new values, upstream OptiScaler's latest game quirks are included, and the vendor-neutral backend for AMD and Intel offers new fp16 kernels.

"Automatic" below is HDR input → Exposure source: Automatic exposure (WhitePointSource=3).

New

  • Automatic exposure is the default exposure source. Before, a fresh install used Game exposure (WhitePointSource=1), which gave the model a picture far too dark in NBA 2K27. Automatic measures the HDR frame itself and needs nothing from the game.

  • One default brightness for every game: +1.5 EV. v0.1.18 tried to tell unexposed games (RDR2) apart from pre-exposed games (NBA 2K27, Cyberpunk 2077) and picked +4.3 or +2.3 EV. That guess was unreliable: a dim RDR2 scene reads exactly like a pre-exposed frame. Now every game starts at +1.5 EV on Model input brightness, and the slider sets your own value.

  • Reuse bottleneck is on by default (VitEvery=2). Each pass reuses the network's ViT bottleneck on every other frame, which saves about 0.35 ms per pass on average (about 0.7 ms per pass on the frames that skip it, RTX 5070 Ti, 2048x1152 model). The 1/2/3-pass preset buttons used to switch reuse off without saying so; they now keep it on.

  • With several passes, all passes reuse on the same frame. They compute the bottleneck together on one frame and all reuse it on the next, so the picture never depends on data older than one frame. (The first upload of this pre-release let the passes take turns instead; that flickered on camera motion at 3 passes, because a pass then built its bottleneck on a pass that was reusing, which made the data two frames old, like VitEvery=3. Fixed in the current zip.)

  • New 1/2/3-pass preset values (style, intensity, local structure, local tone, skin structure; auto skin mask on in every pass):

    Preset Style Intensity Structure Tone Skin
    1 pass Natural 1.80 1.80 1.80 -1.00
    2 pass Natural 1.00 1.00 1.00 -1.00
    3 pass Natural 0.75 1.50 0.46 1.00

Changed

  • Following the game's exposure is on by default only for RDR2 (rdr2.exe, playrdr2.exe), the game measured to hand over its frame before applying its own exposure. For other games it stays off, because in a pre-exposed game following would apply the game's exposure a second time. If another game behaves like RDR2, tick Follow the game's exposure under Automatic, or set AutoExposureFollowGame=true.
  • Re-calibrate button under Automatic: learns the calibration against the game's exposure again (about 2 s), for example when it was learned during a cutscene or loading screen. Do it in an ordinary daylight scene, not snow, night or indoors: the brightness learned there is kept for the whole game (in RDR2, a calibration in snow left the rest of the game about 0.4 EV darker).
  • The follow status in the menu reads Off / Not available yet / Learning / Calibration.

Fixed

  • A small, very bright spot (the sun, a stadium light) could lift Automatic's black cutoff. The cutoff was taken from the plain average of the frame, which a few extremely bright tiles can raise a long way. It is now taken from an average that leaves those tiles out. In RDR2, Automatic's reading moved in step with the game's own exposure across a 9x brightness range.
  • D3D12: the follow calibration kept learning while Follow was off. Switching Follow on later then started from a stale calibration. It now learns only while Follow is on, as on Vulkan, and starts learning when you switch it on.

Vendor-neutral backend (AMD / Intel, Optional\)

  • New dot2add kernels offered next to the packed-fp16 ones. dot2add multiplies two fp16 pairs and adds them into an fp32 total, which RDNA GPUs run as a single native instruction. At startup the backend times every kernel it may use (plain fp32, packed fp16, dot2add, DirectML) per layer and keeps the fastest, so a kernel that is slower on your GPU is simply not used. nr_port.ini now ships with attn16=3 and lin16=3 (offer everything). attn16=1 / lin16=1 restore v0.1.18's kernel set exactly.
  • Speed report: prof=1 in nr_port.ini writes the GPU time per kernel to nr_port.log every ~5 s, and the log's gpu: line now names the adapter and which kernels were picked. Off by default. If you report AMD or Intel speed, please turn it on and attach the log.
  • Tested on an RTX 5070 Ti only (NBA 2K27, 1664x944 model): 29.16 ms per NR frame with everything offered, against 28.98 ms with v0.1.18's set and 32.48 ms plain fp32. NVIDIA has no fast dot2add, so the timing never picks it there. Not yet measured on AMD or Intel, where it is meant to help.

Game compatibility (from upstream OptiScaler)

  • New game quirks: Granblue Fantasy Relink (fakenvapi off, fixes broken rendering), Trails in the Sky 2nd Chapter (FSR2 DX11 inputs), Sword and Fairy 7 (fixes a crash when the upscaler starts), FBC: Firebreak and CONTROL Resonant (no DXGI spoofing; the Streamline spoof unlocks DLSS).
  • Updated quirks: Dragon's Dogma 2, Rise of the Tomb Raider and No Man's Sky get upstream's latest fixes.
  • Not tested here: none of these games is installed on the test machine.

Upgrading

  • If your ini says WhitePointSource=auto, you move to Automatic. OptiScaler saves a setting equal to its default as auto, so anyone who was on Game exposure (the old default) has auto and now gets Automatic. To keep Game exposure, set WhitePointSource=1. An explicit 0 or 2 is kept.
  • A brightness you saved earlier (AutoExposureTrim) still wins over the +1.5 EV default. Set it to auto, or press Reset next to the slider, to use the default.
  • If your ini says VitEvery=1, reuse stays off. Set it to auto or 2 to turn it on.
  • A Follow the game's exposure choice you made yourself is kept.

Unchanged

  • The default NVIDIA path (nvngx.dll_dlssnr.dll at the package root) is the same as in v0.1.18. AMD and Intel users: use the files in Optional\, since the NVIDIA path needs the NVIDIA NGX core.
  • Tested on an RTX 5070 Ti: the exposure changes in RDR2 (Vulkan), the bottleneck reuse in NBA 2K27 and The Blood of Dawnwalker (D3D12, 3 passes).

v0.1.18 - F5 - Automatic Exposure Follows the Game

Choose a tag to compare

@janblade janblade released this 25 Sep 05:06
1102b59

F5(Fork From a Fork From a Fork) v0.1.18 Release:

Automatic exposure now picks a sensible brightness for each game on its own, follows the game's own exposure in games like RDR2, and no longer over-brightens letterboxed cutscenes. Also: the debug views work on RDR2-type games, and two opt-in diagnostics.

"Automatic" below is HDR input → Exposure source: Automatic exposure (WhitePointSource=3). The multipass presets select it.

New

  • Automatic exposure picks its default brightness per game. Some games hand over their frame before applying their own exposure (RDR2); most apply it first (NBA 2K27, Cyberpunk 2077, The Witcher 3). OptiScaler now tells the two apart in the first seconds of play and sets Model input brightness to match:

    • games that hand over the frame unexposed (e.g. RDR2): +4.3 EV
    • games that expose it first: +2.3 EV. +4.3 EV there put yellow highlights in NBA 2K27's player shadows.

    The menu shows it under the slider as "Default for this game: +4.3 EV (unexposed frame detected)" or "+2.3 EV (pre-exposed frame detected)". Moving the slider overrides it; Reset goes back to the detected default. Previously every game started at the same 0 EV, which left the model a dark picture in RDR2 (median 0.29).

  • The multipass presets now use Automatic exposure at that per-game default. They used Game exposure at a fixed setting, which gave NBA 2K27's model a dark picture.

  • Automatic follows the game's own exposure (RDR2-type games, D3D12 and Vulkan). On games that hand over their frame unexposed, Automatic spends the first ~2 seconds learning how its own metering relates to the exposure the game supplies, then follows the game's exposure from then on. Brightness now moves exactly with the game: eye adaptation, cutscene cuts, menus and fades. The brightness slider keeps its meaning. On by default; turn it off with Follow the game's exposure under Automatic, or AutoExposureFollowGame=false. Games that expose their frame themselves are not affected.

    • On Vulkan it follows a few frames behind the game. In RDR2 on Vulkan the transitions showed no visible flash, and it learned the same calibration as on D3D12 (+3.62 EV).
  • Two diagnostics, both off by default ([DlssNr] in OptiScaler.ini, or Compare in the menu):

    • FrameStats=true ("Log frame brightness stats"): every ~2 seconds, logs the frame NR is given (format, brightness percentiles), the game's and Automatic's exposure, the white point in use, and the measured brightness of the picture the model is shown.
    • KernelProfile=true ("Log NR kernel profile"): logs which NVIDIA NR kernels run (fp8 or fp16 variants) and where their GPU time goes. It samples 3 of every 240 evaluations; with it off, nothing is recorded.

Fixed

  • Automatic exposure was ~3 EV too bright in letterboxed cutscenes (RDR2). The black bars counted as extremely dark scene content and pulled the metering down. Black bars and black borders are now left out, and a fade to black keeps the last good exposure instead of jumping. In RDR2, the model's picture in cutscenes now matches gameplay (median 0.65-0.86 vs 0.69-0.74; before: 0.86-0.97).
  • Debug views were near-black on RDR2-type games. Debug view 1/2/3 (model input, model output, the edit) were scaled by the Paper white slider instead of the frame's own brightness, so on games that expose their frame later they looked ~100x too dark, even though the model was given a normal picture. They now show at the game's brightness and follow the brightness slider.

Improved

  • HDR menu: the HDR input section now sits after the presets, and the exposure Trim is shown as a brightness slider in EV (+ is brighter).

Upgrading

  • A brightness you saved earlier (AutoExposureTrim in OptiScaler.ini) still wins over the per-game default. Set it to auto, or press Reset next to the slider, to use the detected default.
  • If you are not on a preset, the exposure source is unchanged (Game exposure by default). Pick Automatic exposure to get the new behaviour.

Unchanged

  • The default NVIDIA path (nvngx.dll_dlssnr.dll at the package root) and the vendor-neutral backend under Optional\ are the same as in v0.1.17.
  • Tested on an RTX 5070 Ti in RDR2 (D3D12 and Vulkan), NBA 2K27, Cyberpunk 2077 and The Witcher 3 (D3D12).

v0.1.17 - F5 - Streamline Hook Race Fix + Faster Vendor-Neutral Port

Choose a tag to compare

F5(Fork From a Fork From a Fork) v0.1.17 Release:

A reliability fix for games that load DLSS through Streamline, and a faster vendor-neutral Neural Rendering backend.

Fixed

  • "DLSS options will be disabled" at game start, intermittently (Streamline games): OptiScaler hooks each Streamline plugin (sl.dlss, sl.dlss_g, sl.reflex, ...) as the game loads it. Two of those hook installs could run at the same time, which Microsoft Detours does not allow. When that happened, the log showed Failed to hook DLSS: 10DD / Failed to unhook DLSS: 10DD, Streamline wrote a minidump and stopped loading its remaining plugins, and DLSS disappeared from the game's menu for that session. The installs now run one at a time. A stress test that reproduced hundreds of these failures per second on the old code shows none with the fix. If you still see 10DD in OptiScaler.log, please report it with the log attached.

Improved

  • Vendor-neutral Neural Rendering backend (Optional\): packed fp16 Linear kernels (lin16=1 in nr_port.ini, on by default). The backend's own matrix-multiply kernels gain a variant that does the multiplications in packed half precision, two values per instruction, while keeping the running sums in full precision. At start the backend times it against the existing kernels and DirectML for every layer, and uses it only where it is faster.
    • RTX 5070 Ti, 768×448 model resolution: 8.40 → 8.05 ms GPU time per NR frame. In NBA 2K27 on the Before-SR Potato preset, the arena went from about 76–77 fps to 80 fps.
    • Image difference against the previous kernels: 67–77 dB PSNR, at most 0.41/255 after 12 frames of history feedback. No visible change.
    • It applies only with fp16=1, on GPUs that report shader model 6.2 with native 16-bit operations. Anywhere else, the previous kernels run unchanged. Set lin16=0 to turn it off.
    • Not yet measured on AMD or Intel. Because the per-layer timing picks the fastest kernel on each GPU, it should never be slower than before; AMD reports are welcome.
    • To upgrade, copy both Optional\nvngx.dll_dlssnr.dll and Optional\nr_port.ini over your previous copies in the game folder. The first launch re-times the kernels once, then caches the result in nr_tile_cache.txt.

Unchanged

  • The default NVIDIA path (nvngx.dll_dlssnr.dll at the package root) and every OptiScaler.ini default are the same as in v0.1.16.

v0.1.16 - F5 - Multipass Quality Upgrade

Choose a tag to compare

F5(Fork From a Fork From a Fork) v0.1.16 Release:

The multipass quality control from v0.1.15 -- a new dial on how much of each extra pass's raw answer the next pass actually receives -- is this branch's biggest quality change, now out of pre-release. This build also consolidates everything else shipped on release/0.1.14 across v0.1.14, v0.1.15 and this release into one changelog.

Added

  • Multipass pass-feedback control (DlssNr.PassFeedback, default 1.0, unchanged behaviour): each extra multipass step hands the next pass its own raw output as if it were a fresh rendered frame, which drifts further from what the model was trained on with every added pass -- the "richer at 2, synthetic at 4+" falloff. This slider, under multipass "Passes", under-relaxes that hand-off instead of always passing the full raw answer through. At 1.0 every existing configuration behaves exactly as before; try 0.7 / 0.5 / 0.3 if you run 3+ passes and want to soften that falloff.
  • Run DLSS Neural Rendering without an NVIDIA GPU (opt-in, AMD/Intel): a separate, vendor-neutral build of the model runner runs the model itself in D3D12 compute + DirectML instead of going through NVIDIA's NGX core, which this feature no longer hard-requires to start. This zip ships a prebuilt copy under Optional\, along with its tuned nr_port.ini; see INSTALL-DLSSNR.md, "Run without an NVIDIA GPU".
    • First real non-NVIDIA hardware results: confirmed running correctly on an AMD RX 6700 XT (RDNA2) -- the menu correctly reports "Model backend: vendor-neutral port" and produces genuine neural-reconstructed output, not a fallback. Two real measurements so far, on the same GPU:

      • 1 Pass / Potato (50% model resolution): ~16.6 fps (59.27 ms NR pass), using weight data decoded for older-generation Nvidia DLSS-NR runtime.
      • 1 Pass / Medium (80% model resolution): 25 fps base, ~75 fps with 3x frame generation, using weight data decoded from NVIDIA's current-generation (RTX 50) DLSS-NR runtime. Despite the higher resolution, this ran noticeably faster than the Potato result -- the current-generation model appears to be a meaningfully lighter network, not just re-trained weights on the same architecture. (The port only ever reads weight tensors out of a real Nvidia DLSS-NR DLL; it never executes that DLL's own code, running everything through its own D3D12/DirectML compute instead -- which is what makes AMD support possible at all.)

      playable with FG, the Medium result is a real, meaningful step up, and RDNA2's lack of dedicated matrix/AI acceleration hardware is still the likely main bottleneck either way. This needs its own AMD-specific perf pass, not a judgment that the backend doesn't work.

    • RDNA4 (RX 9070 XT) potential -- not tested, a guesstimate only: RDNA4 adds real dedicated matrix/AI accelerator hardware and FP8 support that RDNA2 simply doesn't have, plus ~3.7x more raw FP32 throughput from CU count and clocks alone. If DirectML maps this workload's GEMMs onto those AI accelerators effectively, a 5-10x improvement over these 6700 XT results is architecturally plausible; if it doesn't, expect something closer to the ~3.7x raw-compute scaling alone. Whether AMD's DirectML driver actually exploits RDNA4's matrix units for this workload is genuinely unknown -- this is architecture-based reasoning, not a measurement, and nobody has run this on a 9070 XT yet.

    • Still untested: Intel Arc.

  • The Neural Rendering status line now names which model backend is running (NVIDIA NGX vs. vendor-neutral port).
  • Model-reduced-resolution display: the menu now rounds the reduced size actually submitted to the model to a multiple of 16 (matching what the model receives) and shows it.

Fixed

  • Streamline dual-runtime crash (slInit() failed with error code 0x18. The DLSS options will be disabled.): fixed a collision between OptiScaler's own private Streamline runtime (used for DLSS Frame Generation) and the game's own Streamline init, which could throw an exception inside NVIDIA's Streamline SDK and make the game disable its DLSS menu entirely. The game's plugin loads are now kept isolated from OptiScaler's own already-active ones.
  • DLSS-NR Highlight guard: in Composed mode, the guard now bounds brightening only -- darkening is no longer floored, matching the control's own name and help text. Replace mode's own, separate guard is unaffected and still bounds both directions. This removes a floor that was previously added for a measured Nioh 3 regression (a dark-scene, raised-paper-white flicker); the tradeoff is judged worth it, but if you hit whole-frame flicker or a contrast/saturation collapse in a dark, high-paper-white scene, this is the first place to look.
  • Shutdown/teardown safety: fixed several ExitProcess-time crash/deadlock risks in the NR and NGX shutdown paths.
  • RTX 40 MFG unlock: fixed a real unsynchronized read/write race in MfgUnlock::LastStatus().
  • Kingdom Come: Deliverance II: native HDR10 output now stays correct when frame generation is on.
  • Vulkan menu: the DLSS-NR menu now shows the real Vulkan game-exposure reading instead of a stale/incorrect value (display-only).
  • An internal undefined-behaviour fix in a path-comparison helper newly exercised by the Streamline fix above (no user-facing behaviour change on its own).

Credits

Several of this release's fixes are adapted from wilsjo2's fork -- the shutdown/teardown safety work, the KCD2 HDR quirk patch, the MFG LastStatus race fix, the Streamline dual-runtime isolation fix, and (from @mattjaas) the Highlight guard change. Full attribution, including third-party licenses, is in docs/CREDITS.md.

Notes

  • Same release/0.1.14 branch as v0.1.14 and v0.1.15 -- no new branch cut.
  • Two things are still worth knowing even though this is no longer marked pre-release: the vendor-neutral backend is now confirmed running on real AMD hardware across two settings tiers, but not yet at a playable framerate on RDNA2 (see above -- Intel still entirely untested), and the Highlight guard change's Nioh 3 tradeoff hasn't been broadly field-tested beyond this branch's own testing.

v0.1.15 - F5 - Multipass Tuning + Model Preview

Choose a tag to compare

F5(Fork From a Fork From a Fork) v0.1.15 Release:

A menu display fix, a model-usage display feature, and a new multipass tuning control -- all built on the same release/0.1.14 branch as the previous release.

Fixed

  • Vulkan menu: the DLSS-NR menu now shows the real Vulkan game-exposure reading instead of a stale/incorrect value (display-only, does not affect quality or output).

Added

  • Model-reduced-resolution display: the menu now rounds the reduced size actually submitted to the model to a multiple of 16 (matching what the model receives) and shows it, so what you see reflects what the model sees.
  • Multipass pass-feedback control (DlssNr.PassFeedback, default 1.0, unchanged behaviour): a new slider under multipass "Passes" that under-relaxes how much of an extra pass's raw output the next pass receives, instead of always handing it the full restored answer. At 1.0 every existing configuration behaves exactly as before -- this is purely additive and opt-in. Lower values are an experimental knob for Passes >= 3 if the current "richer at 2, synthetic at 4+" falloff bothers you; try 0.7 / 0.5 / 0.3 and judge for yourself, there is no verified "better" value yet.

Notes

  • Same release/0.1.14 branch as the previous release -- no new branch cut, no unrelated changes pulled in.
  • Vendor-neutral (AMD/Intel) port backend and its tuned nr_port.ini are still shipped under Optional\, same as the previous release; see INSTALL-DLSSNR.md, "Run without an NVIDIA GPU".

v0.1.14 - F5 - Experimental - Vendor-Neutral - DLSSNR

Choose a tag to compare

F5(Fork From a Fork From a Fork) v0.1.14 Release:

Neural Rendering shutdown safety, an HDR fix for Kingdom Come: Deliverance II's DLSSG output, a real data-race fix in the RTX 40 MFG unlock, and an opt-in vendor-neutral model backend for GPUs without NVIDIA NGX support.

Fixed

  • Shutdown/teardown safety: fixed several ExitProcess-time crash/deadlock risks in the NR and NGX shutdown paths — a leaked-not-destructed state singleton (avoids static-destruction-order races with other DLLs' own detach callbacks), an early bail on DLL_PROCESS_DETACH when lpReserved != nullptr (no DLL unloading/logging/GPU cleanup is safe then), and a reentrancy guard consolidating the NGX D3D12 shutdown paths.
  • RTX 40 MFG unlock: fixed a real unsynchronized read/write race in MfgUnlock::LastStatus() — a caller on the menu/overlay thread could read a std::string field mid-write from the render/loader-hook threads with no synchronization at all.
  • Kingdom Come: Deliverance II: native HDR10 output now stays correct when frame generation is on (a game-specific quirk patch, gated by executable name, applied at most once).

Added NR Support for Non-Nvidia (Experimental)

  • Vendor-neutral NR model backend (opt-in, AMD/Intel): Neural Rendering no longer hard-requires the NVIDIA NGX core to start. A separate, vendor-neutral build of the model runner runs the model itself in D3D12 compute + DirectML instead of going through NGX. This repository's own source does not include that build — it comes from a separate toolchain — but this zip ships a prebuilt copy under Optional\, along with its tuned nr_port.ini (fp16/DirectML acceleration, vit_every=2); see INSTALL-DLSSNR.md, "Run without an NVIDIA GPU" for how to use it. Not yet validated on real AMD or Intel hardware — only exercised with the build forced in on an NVIDIA machine (27,600+ frames, no NGX core touched, no errors). Reports from non-NVIDIA hardware welcome before this is considered fully supported.
  • The Neural Rendering status line now names which model backend is running (NVIDIA NGX vs. vendor-neutral port).

Credits

Two of this release's fixes are adapted from wilsjo2's fork:

  • The shutdown/teardown safety work is adapted from commit dac290ed.
  • The KCD2 HDR quirk patch is adapted near-verbatim from commit 34dfe6d9.
  • The MFG LastStatus race fix follows the same lock-and-return-by-value approach as commit bf91eebd; the write-side half of the fix was found and written independently during review.

Full attribution, including third-party licenses, is in docs/CREDITS.md.

Notes

This is a pre-release: the vendor-neutral backend is untested on real AMD/Intel hardware, and none of this build's changes have been exercised in-game beyond the automated/manual checks noted above and in each change's own testing notes.

v0.1.13 - F5 - Optional Bottleneck Reuse and Update Notice Fix

Choose a tag to compare

@janblade janblade released this 21 Sep 09:00
126269b

F5(Fork From a Fork From a Fork) v0.1.13 Release:

New: optional bottleneck reuse (off by default)

  • A new Reuse bottleneck every other frame checkbox in the DLSS-NR menu, next to Model precision. It runs the model's coarsest stage only every other frame and reuses the last result in between, saving roughly a tenth of the model's GPU time. It applies immediately, no restart.
  • Ini: [DlssNr] VitEvery = 1 (off, default) or 2 (on).
  • It works with NVIDIA's own model only. Scene cuts always recompute, and fast camera motion can look slightly softer.
  • Aimed at RTX 20/30, Tested in RTX3060. Tester feedback: Smooth at 80% model reso x2FG with Reshade even.

Fixed: update notice on the latest version

  • Builds since v0.1.9 announced an update to themselves because the internal version number was never bumped. It now matches the release (0.1.13), and the packaging script refuses a release whose tag and version differ. Builds of v0.1.9 to v0.1.12 will keep showing the notice until you update to this one.

Everything else is v0.1.12

  • NVIDIA residual presets, the built-in RTX 40 MFG unlock (off unless you turn it on; still not confirmed on RTX 40 hardware), and the bundled DLSS Frame Generation runtime and RTX 20/30 MFG files (NVIDIA and third-party files with their licence texts next to them under OptiScaler\streamline\ and OptiScaler\dlssg_sm86\).
  • Single-player games only, and the DLL is unsigned, so Windows SmartScreen or antivirus may warn.