-
Notifications
You must be signed in to change notification settings - Fork 2
29 August 2026
Now we have the actual value of the first subband, 3.
We may continue the decoding by asking, how many other “buckets”, or subbands, do we need to assign a value? John-K answers this by stating that at 16kHz (which the LTSDM runs at), the number of subbands that exist is 14. Now, how do we find the other values? Here lies another lesson in compression:
Encoding direct values is an expensive way to store data (uncompressed). We can code the whole frequency value for each subband, however coding “15894” is comparatively large to what a1800 does. To support the theory behind this codec, we can assume that sounds within frequencies are similar to their neighbours, they are not completely random. This accounts for smooth vocal transition sounds as opposed to alarm-bell-esque abrupt audio. This is optimized for narration, as our speech flows from one syllable to the next. A1800 encodes the differences between subbands instead of the whole value. Since we know that frequencies are similar to their neighbours, we can justify that adding “2” is much smaller to record as opposed to “1140”. This reduction in data is a part of the compression.
Imagine that one bucket is very full, while the other is barely filled with anything at all: would not that difference potentially be of a significant size, just like in the uncompressed problem earlier. We need a system that will get the differences between frequencies in the smallest amount of data stored. This is where Huffman Coding comes in…
Huffman Coding is a way to compress data based on the frequency of its values. The more a specific value exists in the dataset, the shorter the compressed code. The rarer the value in a dataset, the longer we can permit its compressed code. This results in more short codes that outweigh the occurrence of longer codes. This produces a compressed dataset that can be unwound by its unique Huffman code.
| ⟵ Older | Table of Contents | Newer ⟶ |
|---|