Releases: perseval-BLR/NeuralScreen-AMD
Release list
v0.3.24 - the flicker fix is the default
v0.3.24
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
The flicker is fixed, and it is now the default
The four corners of the 2x2 are in, and they close it. Same card (RX 7900 XTX, #3), same maximised window, work scale 0.65 in every run but the sweep:
| run | upscale runs on | motion vectors | flicker | frames with >=1024 | staging re-creations |
|---|---|---|---|---|---|
| 1 | shared module | off | yes | 75 / 75 | 0 |
| 2 | private copy | off | yes | 70 / 73 | 0 |
| 3 | private copy | on | no | 0 / 77 | 0 |
| 4 | shared module | on | no | 0 / 77 | 90 |
| 5 | private copy | on, scale sweep | no | 0 / 224 | 86 |
Two things become visible at once, and they are the whole story:
- The vectors are the fix. With them bound, not one frame in 77 - or in the 224 of the sweep - carries corruption, and the flicker is gone on screen. Without them it is there in both other runs.
- The private copy is what keeps the fix cheap. The vectors alone (run 4) work and take the runtime into its staging re-create loop: 90 re-creations for the 77 frames. On the private copy, run 3 does the same 77 frames with zero re-creations.
So both are on by default now, and they are tied together: the vectors are bound only while the private copy is actually in force. That is in the build, not in the instructions - a machine where the copy fails to load gets the old picture rather than the loop.
Run 5 also answers a question left from last time: the sweep passed through 0.65, 0.75, 0.80 and 0.95 before reaching 1.00, and none of them produced a corrupted frame. Until now only 0.65 had been tested with the vectors on, so the fix is not a property of one work scale.
What is new
The two arms ship switched on. NS_AMD_UPSCALE_PRIVATE and NS_AMD_UPSCALE_MV default to on; set either to exactly 0 to reproduce the old picture.
The flicker probes are now an A/B against the fix. NeuralScreen-probe.vbs is the fixed configuration, so the two launchers that used to ask for what everyone now gets were retired. The pair that replaces them switches the fix off, one variable at a time:
NeuralScreen-probe-mv-off.vbs- vectors off (NS_AMD_UPSCALE_MV=0). The alternating output comes back; this is the defect.NeuralScreen-probe-shared.vbs- private copy off (NS_AMD_UPSCALE_PRIVATE=0). The vectors are refused on that path by the build, so this is the old picture, not a broken one.
The log says which arm ran, not which was asked for - and it now names the copy as the reason when the vectors are absent, because on a machine where the copy failed to load that is the whole explanation.
If you can help
The reporter on #3 has already sent the runs this build was written from - there is nothing further to do there unless the picture on your own card disagrees.
Everyone else who has seen the flicker: run NeuralScreen-probe.vbs (double-click, no variable to set) in the same window and at the same work scale you saw it before, until the black frames would have shown or 20-30 seconds of no flicker. Then send the diagnostic package and native\dlssnr_on_amd.log. If your card still flickers, the two A/B launchers above are what separates the two causes on your machine.
What did not change
The neural pass, the capture path, the menu and every other setting are v0.3.23's. Frame Generation is still refused on a Radeon, and the recording still writes even dimensions.
The other open reports are a different fault and are not addressed here: the fault inside D3D12Core.dll on #4 and #5 happens before any frame is processed, and #2 is the hybrid Radeon + NVIDIA detection path.
v0.3.23 - the arm the three probe runs could not isolate gets a launcher
v0.3.23
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
What the runs showed
#3 (RX 7900 XTX). The third run was clean: no flicker, and not one frame above 1024 in the upscale's output - 0 of 68 samples, against 34 of 69 on the baseline and 59 of 68 on the private copy alone. The reporter checked that the run was real (upscale on in 68 of 68 samples, 67 frames processed, 14 network jobs), so those numbers come from a run that worked, not from one that did nothing.
Two things in it are still open, and he said so himself:
- the presented surface sits at 0.916 x the native anchor, where the clean frames of the other runs sat at 1.016 and 1.306 - so the result may be arriving attenuated, and that is not yet confirmed either way;
- the run changed two things at once, and "motion vectors without the private copy" has never been run in a build where it could answer anything.
#6 (RX 9070 XT). A second reporter on a different card generation, and his log shows the same defect from a different angle: with the upscale in the path, 4 of 12 sampled frames present a black surface (mean 0.0000 against a live native anchor of 0.3434); with it out of the path, none do. In the 1:1 segment, where the upscale does not exist, it is clean. Same class as #3, on another GPU.
What is new
A launcher for the arm the three runs could not isolate: NeuralScreen-probe-mv-shared.vbs.
It is the per-frame probe with motion vectors bound on the shared upscaler module - the same file the runtime's hooks are in. The other two launchers cover the private copy without vectors and the private copy with them; this one is the fourth corner of that 2x2, and it is the one #3 asked for. The private copy is deliberately not set, so the runtime can see this dispatch.
Worth saying plainly: this is the combination that in v0.3.21 drove the runtime into a staging re-create loop (80 re-creations for 4 network jobs). v0.3.22 did not fix that loop - it moved the upscale onto a private copy so the runtime never saw it. So this run may hit the loop again. If it does, the picture still answers the question it was built for (does the flicker stop), and the worker's log says which happened.
Nothing changes for the default path: the launcher is a test, the arms are off unless a launcher turns them on, and the defaults are v0.3.22's.
If you can help
Same window, work scale below 1.00, and the same order as before, so the runs stay comparable:
NeuralScreen-probe.vbsNeuralScreen-probe-private.vbsNeuralScreen-probe-mv.vbsNeuralScreen-probe-mv-shared.vbs(new)
Run each until it flickers (or 20-30 seconds if it does not), close it, then send the diagnostic package and native\dlssnr_on_amd.log. The runtime's log is appended to, so sending it once at the end is fine if you note which run was which.
What did not change
Defaults, the neural pass, the capture path and the menu are v0.3.22's. The flicker is still open. The vectors now have a clean run behind them and a way to isolate them; the attenuation question in #3 is unanswered, and that is the next thing to look at.
v0.3.22 - the upscale can run where the runtime cannot see it
v0.3.22
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
What the v0.3.21 runs showed
#3 (RX 7900 XTX). The composition line answered its question: the white frames are ordinary finite numbers (84-100% of the picture above 1024), not NaN or infinity, and the network's own surface never carries any. So they are made inside the FSR upscale step.
Two more things came out of that log:
- The frame after a history reset is always the first bad one. The reset frame itself is clean, every time. The fault lives in the upscaler's frame-to-frame history.
- With motion vectors bound, even the warm-up frames were clean - frames where the neural pass had not run yet, and which on the default path alternate clean / 83.8% wrong. That points at the missing vectors.
But that second run was not clean evidence, and the reporter said so first: with vectors, the neural runtime - which follows the dispatch that has motion vectors - stopped ignoring the upscale and re-created its staging 80 times for 4 network jobs. The run measured that loop as well as the vectors.
#1 (RX 9070 XT). Both runs stopped before the first frame, at the FSR setup, six launches in a row. The upscaler DLL, the driver and the code at that point are the same as in v0.3.20, so this looks like the state of the machine rather than the release - a restart A/B is asked for in the thread. It cannot be ruled out from the logs, which is why this build logs that step.
What is new
The upscale can run where the runtime cannot see it. NS_AMD_UPSCALE_PRIVATE=1 runs the upscale step on a second copy of the same FidelityFX DLL, loaded from a private folder, that the runtime never hooked. Same file, same name, only the folder differs. The log says whether the copy loaded; if it did, the runtime's own log stops saying it is ignoring a dispatch without motion vectors.
Two launchers use it, and they differ in exactly one thing:
NeuralScreen-probe-private.vbs- the probe, with the upscale on the private copy. The control.NeuralScreen-probe-mv.vbs- the same, plus motion vectors for the upscale. This launcher changed: in v0.3.21 it bound the vectors on the shared DLL; now it uses the private copy, so the runtime stays out of the test.
A stop before the first frame now names its place. Each FSR setup call is logged before and after (FSR setup: ... calling / returned), and if one has not returned after 4 seconds, a second line says which one is still open. It only reads and logs.
Both arms are tests, not fixes, and are off unless a launcher turns them on. Defaults are unchanged.
If you can help
Same window, work scale below 1.00 (0.65 on a maximised window), one run each, in this order:
NeuralScreen-probe.vbsNeuralScreen-probe-private.vbsNeuralScreen-probe-mv.vbs
Run each until it flickers (or for 20-30 seconds if it does not), close it, then send the diagnostic package and native\dlssnr_on_amd.log after each run - the runtime's log is appended to, so it is fine to send it once at the end if you note which run was which. All three are slow on purpose.
If you are on #1 and the program stops before the first frame: restart Windows first, then run the plain NeuralScreen.vbs once. If it still stops, the log will now say where.
What did not change
Defaults, the neural pass, the capture path and the menu are v0.3.21's. The flicker is still open, narrowed to the upscaler's history and, most likely, its missing motion vectors. These three runs decide the next step.
v0.3.21 - the probe says what the upscale's output is made of
v0.3.21
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
What your v0.3.20 runs showed
Two probe runs at work scale 0.65, one on an RX 7900 XTX (#3) and one on an RX 9070 XT (#1), both with the upscale in the path. The FSR upscale's input is flat on every frame. Its output alternates, and the two states of each pair are complementary: the share of the picture that is wrong on one frame is the share that is right on the next.
| machine, work -> display | frame 1 | frame 2 |
|---|---|---|
| 9070 XT, 1664x936 -> 2560x1440 | all correct (0.047) | all wrong (reads 0.0000) |
| 9070 XT, 1792x1006 -> 2560x1440 | ~75% wrong (0.011) | ~25% wrong (0.033) |
| 7900 XTX, 2496x1356 -> 3840x2160 | 99.94% wrong (63,962) | 0.04% wrong (23.4) |
So the upscale rewrites every pixel on every frame, and the rewrite goes wrong on alternate frames. It is not a frame that fails to update. And it fails the same way on two card generations whose FidelityFX DLL takes different paths, which points at what we hand the upscale, not at one implementation of it. The runtime log agrees it is not the runtime: it follows our first dispatch and leaves the upscale alone.
Two instruments for the next run
The probe says what the upscale's output is made of. The mean could not: it reads a NaN as 0 and an infinity as 1, so a surface full of NaN reads exactly 0.0000, the same as a black one, and "64,000" could not say how many pixels sat there. The per-frame probe now prints, for the upscale's output and for the network's own surface beside it as a control, the share of colour values that are NaN, infinite, exactly zero, and above 1024. The mean is unchanged, so these logs compare with every earlier one.
NeuralScreen-probe-mv.vbs runs one change on top of the probe. The upscale step has never been given motion vectors, and they are the one input of that call the FidelityFX API does not mark optional. This launcher hands it a motion-vector surface of its own, with the vector scale at zero, so the only thing that changes is whether one is bound. The log names the arm once. It is a test, not a fix, and it is off unless you use this launcher.
One honest risk: the runtime picks the dispatch it processes by "has motion vectors". With this arm the upscale has them too, so the runtime may follow it instead. Its own log says which (staging ready: colour <display size> ... motion <work size> means it took the upscale). If that happens, it is a result, not a broken run - send it anyway.
If you can help
Same window, same work scale below 1.00 (0.65 on a maximised window is where it shows):
- Double-click
NeuralScreen-probe.vbs, run until it flickers, close it. - Double-click
NeuralScreen-probe-mv.vbs, same thing. - Send the diagnostic package and
native\dlssnr_on_amd.logfrom each run.
Both launchers run slowly (a full readback every frame). They are for one diagnosis each, not for playing.
Also in this build (from the tree since v0.3.20)
- The per-frame probe line carries
reset=andupscale=, and the session summary counts both, so a history reset can no longer be mistaken for an alternation. NS_AMD_UPSCALE_EXPOSURE=1is a second, separate arm: it gives the upscale the same 1x1 exposure the network dispatch gets. Off by default, named in the log.- The panel's keep-on-top check decides by process (the shell, our own windows, the captured app) and ignores hidden windows, so using the taskbar no longer re-raises the picture over a state that was already correct.
What did not change
Defaults, the neural pass, the capture path and the menu are v0.3.20's. The flicker is still open. It is narrowed to one call and its inputs, and these two runs are how the next step gets chosen.
v0.3.20 - the per-frame probe is a one-click launcher
v0.3.20
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
The per-frame probe is now a double-click
The probe that answers "which surface is at fault" was already in v0.3.19, and it was already out of reach of the people it was built for.
It reads the surfaces every 300th frame. A readback is a full GPU-to-CPU sync, so that cadence is deliberate - but a diagnostic run is short: one reporter's three launches ran 51, 73 and 216 frames, so the modulo never fired once and the numbers that would say which surface alternates are absent from his log. The instrument existed; its cadence put it out of reach of the reports it was made for.
Turning it to every frame takes an environment variable, and that is where it stopped. Asked to set it, a reporter answered:
Sorry that I have no experience with coding, and I don't know how to set the NS_AMD_PROBE_EACH=1 environment variable. PLS tell me how I can do that in detail.
He is not short of willingness - he had already attached three diagnostic packages. The instruction was wrong for the audience. This project already had the answer for exactly this situation: NeuralScreen-diag.vbs exists so that nobody ever touches NS_PHASE by hand. There is now NeuralScreen-probe.vbs beside it, and that is the whole instruction: double-click it instead of the usual launcher.
Same pre-flight checks as the other two launchers (the interpreter, the runtime, the worker), the same tray behaviour, the same log. The picture runs slowly while it is up, because that is what a per-frame readback costs. It is a diagnosis for one run, not a way to play - the README says so in both languages.
tests/test_probe_launcher.py keeps it honest: it starts each of the three launchers for real, reads what the child process inherited, and fails if the variable is set in the wrong scope - the one mistake that reads perfectly in the source and does nothing at runtime.
Frame Generation is refused on a Radeon instead of crashing the worker
Carried from v0.3.19, unchanged - repeating it here because these comments are where people arrive.
A reporter on a 9070 XT turned the FG switch on and the worker died 21 ms later:
[fg] UI: on, 2x
[crash] ACCESS_VIOLATION (0xC0000005) at nvngx.dll + 0xCA05,
tried to read address 0x0, amd active=1 frames=1293
nvngx.dll is our own worker, so that fault was our null dereference, not the runtime declining politely. Frame Generation is an NGX feature end to end; a Radeon has no NGX, and the path was entered anyway.
The refusal is decided in one place, FgRequested - the single gate both FG paths ask - and it returns false on a Radeon before it consults NS_FRAMEGEN or the menu switch. The guard reads the adapter actually chosen, not "is the AMD pass live": a pass that failed is exactly when someone starts trying the other switches. The menu now draws the row as unavailable with the reason under it, and the hotkey refuses with the same text.
What did not change
The neural pass itself, the FSR chain, the capture path and the menu are v0.3.19's. This release is the launcher and the two READMEs.
The flicker is still open. It is localised, not fixed: the FSR upscale's output surface is the single remaining candidate, from three independent measurements, and NeuralScreen-probe.vbs is how the fifth measurement gets taken. If your picture flickers, run it once and send the package.
v0.3.19 - Frame Generation is refused on a Radeon instead of crashing the worker
v0.3.19
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
Frame Generation is refused on a Radeon, instead of crashing the worker
A reporter on a 9070 XT turned the FG switch on and the worker died 21 ms later:
04:02:01.745 [fg] UI: on, 2x
04:02:01.766 [crash] ACCESS_VIOLATION (0xC0000005) at nvngx.dll + 0xCA05,
tried to read address 0x0, amd active=1 frames=1293
[main] the worker CRASHED (exit code 3221225477)
nvngx.dll is our own worker, so that fault is our dereference of a null pointer - not the runtime declining politely. Frame generation is an NGX feature end to end (NVSDK_NGX_D3D12_*, served by nvngx_dlssg.dll); a Radeon has no NGX, and the path was entered anyway. The switch was offered, unchecked, on the one card where it can only take the session down - overlay, recording and pass with it.
The refusal is decided in one place, FgRequested, the single gate every FG path passes through (PresentFrame and the HDR presenter both ask it). It returns false on a Radeon before it consults NS_FRAMEGEN or the menu switch: a remembered preference cannot outrank a card with no runtime. The guard reads g_radeon_present - the adapter actually picked - not "is the AMD pass live": a pass that failed is exactly when someone starts looking at the other switches.
The log says so when the switch goes on, and the menu draws the row as unavailable with the reason under it instead of offering a live switch. The hotkey refuses the same way.
The probe now reads the one link it never read
Every other surface in the frame had a number before this one did: the captured frame, the conversion's output, the network's surface, and the surface the present copies from. The FSR upscale's output - up_out, between the network and the composite - had none, and that is exactly where a reporter's own bisect pointed.
The bisect: Work Scale to 1:1, no rebuild. The alternation stopped. Measured from his recording rather than taken on his word:
| work buffer | output | frames alternating | |
|---|---|---|---|
| before | 1102x756 | 1698x1164 | 231 of 231 |
| at 1:1 | 1816x1221 | 1816x1221 | 2 of 329 |
That slider moves two things at once, and it is worth saying so plainly: at 1:1 dispatch B is skipped, and the network also runs at a larger size (the work buffer is the source frame). A second reporter's run separates them by sliding in three steps - 0.65, 0.80, 1:1 - and the network's own residual is the same at all three (0.042 / 0.034 / 0.037, measured per frame), while the presented surface swings from 3.5x its anchor to 0.93x. The network is not what moves. What is left between the network and the screen is this one surface.
NS_AMD_PROBE_EACH=1 on a frame that flickers now answers it: the upscale's mean is printed on its own line, per frame, beside the network's and the presented surface's. A value that alternates there puts the fault in the upscale pass; a steady value with an alternating presented surface puts it in the composite. At 1:1 the line is absent, because the surface is.
The probe still runs every 300th frame by default - a readback is a full GPU-to-CPU sync, and this is a diagnosis run, not for playing.
The menu is baked where the user saw it, also when the mode is chosen from the menu
_window_layer is the origin the saved frame's panel is shifted by, and it was written in exactly two places: when the captured window moves, skipped while the menu is open because the user may be dragging the panel; and when the menu closes.
Choosing a window mode from an open menu ran neither. The panel was laid out against the screen and then blitted as if the frame still sat at the previous window's origin - the reporter saw it in his recording as "transparent / incorrectly composited". Measured on the pre-fix code: the baked panel lands 349 px from where he saw it, exactly the window's origin x.
switch_window now writes the geometry it just measured, before the rebuild that consumes it, and clearing it is its own call for the way back to the desktop.
Nothing else changed
The pass itself, the exposure curve, the runtime selection and the open questions are all as they were in v0.3.18.
Suite: all green, on an RTX 5070 Ti. Nothing here was run on a Radeon. Every new check was written to fail on the code as it was: the upscale probe on five separate properties (state, existence guard, spec size, and both log lines), the baked panel on the static write and on the end-to-end sequence, and the FG refusal on ten - the guard's presence, the flag it reads, the include order, the log line, the menu flag, the dead row, the click that must send nothing, and the NGX card that must stay untouched. The behavior is exercised, not only grepped: the menu is built, the row is clicked, and the refusal watcher is run against a real log line.
v0.3.18 - the probe cannot report a zero it never measured
v0.3.18
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
The instrument can no longer report a zero it never measured
Both of these were named when the surface probe was built, and both survived into every build since. They are the reason a report could look complete and still be unreadable.
The probe now poisons its readback buffer before every copy. The buffer is allocated once and reused for each measured surface, so a copy that never executed left it as it was allocated - zeros - and the probe reported a perfectly black surface it had never sampled. That is the one failure it cannot tell from a real black frame, which is the thing it exists to diagnose. The sentinel is checked on the raw bytes before any mean is computed, and a buffer still carrying it makes the probe refuse to answer and say so, instead of reporting the zero.
Its failure count now reaches the log, with the totals: the surface probe failed on N of M passes - the means above come from the passes that succeeded, not from every frame. The counter existed and was never printed, so a partial failure was invisible and the printed means silently described an unknown subset of frames. Printed only when it happened, so a healthy report grows no line.
Nothing else changed
The fixes open, the questions waiting on reporters, and the behaviour of the pass are all as they were in v0.3.17. This build is the instrument being made honest, so that the next log - the flicker run with NS_AMD_PROBE_EACH=1, or the interop A/B on the D3D12Core fault - is one we can actually read.
Suite: 165 checks, all green, on an RTX 5070 Ti. Nothing here was run on a Radeon. The new check fails on each of the four properties separately: no poison, poison after the copy, a fall-through on an untouched buffer, and an unprinted failure count.
v0.3.17 - the empty recording was an odd width, and my earlier answer was wrong
v0.3.17
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
Record: the empty file is fixed
A 924-byte .mp4 with no frames in it, reported twice. The cause was in the log all along, in two lines:
[record] video codec: hevc_amf
[record] encoding aborted: [Errno 1313558101] Unknown error occurred: 'avcodec_send_frame()'
1313558101 is 0x4E4B4E55 - the ASCII bytes "UNKN". That is AMF's AMF_UNKNOWN: a bare refusal with no cause attached, on the FIRST frame.
The cause is the frame size. yuv420p subsamples chroma 2x2, so BOTH dimensions must be even - an odd width has no chroma column to sit on. This code already knew about the height and said nothing about the width, and a borderless window measures 1059 px in windowed mode. In the reporter's own run the window was 1059x720: the height was fine and only the width was wrong.
Why it never showed up here: NVENC tolerates the odd width. AMF does not. The whole fault was invisible on an NVIDIA machine by construction.
Both dimensions are now made even before the codec is opened, and a frame that arrives one pixel larger is cropped to match - handing an encoder a frame of a size other than its own is the same refusal by another route.
One thing this fix cannot be proven by on this bench, stated rather than implied: removing the crop does not fail the live test, because NVENC accepts it. The crop is defensive - it makes the frame match the declared stream exactly, which is what AMF is entitled to require - and it is guarded by the source-level check instead. That is recorded in the test file itself.
Correction
In issue #1 I said recording needs the pixels back on the CPU every frame, and that is wrong. The log from that reporter's own run shows the button working, the encoder being chosen, recording starting, and the result pixels arriving through shared memory - 14 MB a frame. It failed at the encoder, on a size. The recording path was never missing frames; it was handed a frame the codec would not take.
Still open
Flicker on a Radeon (3.33 Hz, one presented frame per state). Not fixed, not claimed. The instrument is a run with NS_AMD_PROBE_EACH=1, which has not been done yet.
The menu in a button-taken windowed screenshot - reported again on v0.3.15, so my earlier placement fix did not settle it.
The D3D12Core fault on the RX 7600 / Windows 10 machine. v0.3.16 added the A/B arm to test it with; the answer needs that reporter's log.
Suite: 164 checks, all green, on an RTX 5070 Ti. Nothing here was run on a Radeon. The new checks fail on the odd-width version of the recorder.
v0.3.16 - Interop is an A/B arm, and the log says which one ran
v0.3.16
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
For one reporter, and for anyone whose card faults at load
Interop is no longer pinned. It was written to 1 unconditionally, so the only way to try the other value was a rebuild - which means the one variable we suspect of a whole fault class had never been tested on a reporting machine.
The suspicion, stated as a suspicion: issue #4 (RX 7600, Windows 10) dies with a read at offset 0x18C inside D3D12Core.dll, last job -1, before the engine has done anything - a device-child object whose back-reference is NULL. The same shape is open upstream in their own standalone host (#84), and they name zero-copy interop as where it happens. Our host hands resources over the same way.
What this build adds:
NS_AMD_INTEROP=0 resources are handed over by copy, not zero-copy
NS_AMD_INTEROP=1 the value every working report has run (default)
Set nothing and you get the same behaviour as before - that matters, because every report so far ran zero-copy and changing the default would make them non-comparable.
The default is deliberately NOT changed, and that is not caution for its own sake. The upstream report of the same fault says their copy path fails too (0x887A0005). So 0 is not known to be better than 1; it is the other arm of a test. If you are the reporter with the D3D12Core fault, that one variable is what we need, and the log now says which arm ran:
[amd] interop on: zero-copy shared textures (the default)
[amd] interop OFF: resources are handed over by copy, not zero-copy (NS_AMD_INTEROP=0)
One trap worth knowing: the Interop key in dlssnr_on_amd.ini is NOT the control. The runtime reads that file in its DllMain and this host writes the field after that, so the file's value is overwritten. Edit the file and nothing will change - that is why the variable exists.
Not fixed here
The D3D12Core fault itself. This gives us the arm to test it with, not the fix. If both arms fault, the fault is not interop and we will have learned that from a log rather than from another guess.
Everything else open in v0.3.15 - the flicker, Record writing an empty file, and the menu in a button-taken windowed screenshot. Unchanged.
Suite: all green, on an RTX 5070 Ti. Nothing here was run on a Radeon. The new check fails on the pinned version and on four separate ways of breaking the switch.
v0.3.15 - the probe cannot copy a surface its spec does not describe
v0.3.15
Sensitive to flashing images? See the note in v0.3.13 - it still applies.
If you are on v0.3.13, this replaces it. That build processes ZERO frames on a Radeon - it never worked, and it should not have been published. Two reporters spent an evening on it. The cause is below, and it is mine.
The regression in v0.3.13
The probe added in that build killed the device it was measuring. It reads back the captured frame to say whether a picture was lost at the capture, at our conversion or at the dispatch. It declared that frame at the CLIENT's work size (v.w/v.hgt) - 2496x1404 on an RX 7900 XTX, 1664x936 on an RX 9070 XT - while v.color.tex and v.output are created at the composited size, which is full resolution whenever the frame is upscaled: 3840x2160 and 2560x1440 on those two cards. A readback copy whose footprint does not match the resource removes the device, and a removed device fails the frame's own fence wait at once instead of timing out. That is why did not retire appeared in the same MILLISECOND as the probe's numbers, then worker exit 7 at frame 0, every launch, with the engine's own log never reaching engine init ok - the process was gone before the engine got its turn.
Fixed, and made impossible to repeat. The frame spec is now built from the size the resources are created at. The reader also verifies its spec against the resource's own description - size and format - and REFUSES to copy a mismatch, reporting it as the defect it is. A diagnostic that can remove the device is worth less than no diagnostic.
Not the crash fix. 0xC0000409 on exit stays fixed; that path is untouched. A new instrument carried a new defect and broke a build that had been working.
Also in this build, from the unreleased v0.3.14
The engine is handed this frame's exposure, not a value frozen at startup. The AMD path had a 1x1 exposure texture written once at surface creation and never touched again. The host already computed a per-frame value from the captured frame's own luminance, and here it was calculated and thrown away. That is the input the working upstream fork sends with every dispatch.
This path has its own exposure curve, because the shared one cannot come down. The host's curve runs inside 1.00 .. 1.10 - it can only ever raise exposure. On a reporter's bright outdoor frames it returns essentially the 1.0 that was hard-coded, while his complaint is the other direction ("sky and grass wash out to white"). So this path gets a curve symmetric about a centre: dark content lifts, bright content comes down. NS_AMD_PW=0 returns to the old behaviour, NS_AMD_PW_CENTER moves the centre (default 0.35, upstream's value).
NOT VERIFIED ON A RADEON. This machine has no AMD card. The exposure shape is reasoned from the fork and from a reporter's numbers, not measured here, and the centre was deliberately not fitted to one report.
Diagnostic, for the flicker that is still open
A flicker report could not be read, and now it can. The surface probe ran only every 300th frame. A reporter's package alternates between a correct frame and a blown-white one at his present period, and its three launches ran 51, 73 and 216 frames, so the probe never fired once. NS_AMD_PROBE_EACH=1 makes it run every frame, and it now also measures the surface the present actually copies from - the point a per-frame alternation would differ in. Off by default: the readback is a real GPU sync.
Still open
Flicker on a Radeon. Not fixed, and not claimed. The tool above is what will answer it; a run with NS_AMD_PROBE_EACH=1 will.
Record writes an empty file on this path.
The probe's own cadence put it out of reach of the reports it was built for. It is reachable now; treat that as untested until a log says otherwise.
Corrected
The black picture was never a second FSR dispatch - both staging sets are ours. The frame was not lost between the capture and the network's input.
Suite: all green, on an RTX 5070 Ti. Nothing here was run on a Radeon. The check added for this regression fails on the v0.3.13 source and passes on this one.