Skip to content

v1.2.6

Choose a tag to compare

@github-actions github-actions released this 11 Sep 15:50

Download

Ready to run - the filters are already inside. Unzip, run aura-engine.exe, drop a track on it. The app opens on the filter that came with it.

Bundle Filter Output rates Size
Starter 1M taps every FS multiplier (FS2-FS16) 134 MB Download
Standard 10M taps FS8 (352.8 / 384 kHz) 326 MB Download
Reference 30M taps FS8 (352.8 / 384 kHz) 966 MB Download

Unzip more than one into the same folder and the app offers every tap count you have.

Already have the filters, or want to build from source? The plain app zip is below, and the individual filter packs are on the 1.0.0 release.

Checksums (SHA-256)

bca15c1fcac25c43c1f47fdee26ba5af5156f266860c83273dca7b2145c29d40  aura-engine-v1.2.6-bundle-1M-windows-x64.zip
233887a5a018b396df12cfdf4cdfff09d45209788230f39476363888d3a43961  aura-engine-v1.2.6-bundle-10M-windows-x64.zip
ea746982b703ce730e7be98cc0fa6c3bd2708b8bd2a736b60f38b50831382868  aura-engine-v1.2.6-bundle-30M-windows-x64.zip

What changed in 1.2.6

The card was idle two thirds of the time.

A user with a new RTX 5060 Ti measured his conversions, asked whether to buy
more system memory or a card with more of its own, and sent four session logs
with the question. The logs answered something nobody had asked: neither
purchase was what was holding him back. Video memory never ran out, and the
card itself was working only about a third of the batch.

The most expensive single stage of a file turned out to be writing the output.
FLAC compresses better when the predictor order is chosen per block, so every
block is encoded with four candidate orders and the smallest kept. That search
is worth having and it stays — it was simply being done one candidate after
another, on one thread, while the rest of the machine waited. On that user's
profile it was 35% of the time spent on a file, more than either convolution
pass.

The four candidates never depended on each other, so they are now fitted at the
same time. Measured over six runs with nothing else changing:

one after another   91.92  93.99  93.28   mean 93.06 s
all four at once    70.39  72.76  70.55   mean 71.23 s

Just under a quarter off, on any machine with cores to spare — including the
ones with no discrete graphics at all. The file that comes out is the same
file: a tie between two candidates still resolves the way it always did, and a
test encodes the same audio twice and compares the bytes.

Two files at a time was a guess. That limit was chosen cautiously and never
measured. Measured now, on six files at 30M taps with Hybrid-Phase on: three at
once is about 12% better than two, and four gains nothing while making the
result less predictable. So three it is, on machines with the cores to keep
three fed. A smaller machine keeps two, because what a smaller machine wants
has not been measured. Neither number risks anything — a converter that cannot
fit waits its turn rather than taking memory the system needs.

Still to come. The peak memory one file needs is still dominated by holding
two complete copies of the output, one per phase branch, before they are
blended. On a 16 GB machine that is what stops a second file from starting, and
it is next.