Understanding Digital Audio

Revision as of 12:52, 9 March 2015 by Meambobbo (talk | contribs) (File Formats)

Revision as of 12:52, 9 March 2015 by Meambobbo (talk | contribs) (File Formats)

Digital Audio is the representation of sound waves digitally.

Basic Principles

The most basic way to think of digital audio is like animation or video. By playing a number of individual "frames" quickly enough, it gives the appearance of motion. In the case of audio, the frames are called samples. Unlike video, however, digital audio is not an illusion - the digital representation of the sound is capable of storing all the information contained in and can be converted back into identical sound waves, at least theoretically.

The amount of sound information a digital audio signal can represent is determined by the bit depth and the sample rate. Bit depth is the number of bits per sample. Each sample represents the amplitude of the sound wave at that point in time. The higher the bit-depth, the more distinct values are possible for each sample, allowing more and more accuracy in recording. The number of samples per unit of time is called the sample rate, typically represented in kHz, which is the number of thousand samples per second. The sample rate determines the highest frequency that can be accurately represented. The Nyquist Frequency is 1/2 the sample rate and is the highest-pitched frequency that can be accurately represented. A sample rate of 44.1 kHZ means the highest-representable frequency is 22.05 kHZ, which is roughly the maximum frequency that humans are capable of hearing. Typical values are 16-bit/44.1 kHZ (used on typical consumer CD's) and 24-bit/96 kHZ (for commercial recording).

A device called an analog-to-digital converter (ADC) turns analog electrical waves into digital signals. Conversely, a digital-to-analog converter (DAC) turns digital audio into analog electrical signals. ADC's and DAC's typically feature a low-pass filter to remove high-frequencies that cannot be accurately represented by their maximum sample rate. This is at the Nyquist frequency. Without filtering out such frequencies during the analog-to-digital conversion, certain frequencies may be incorrectly represented as lower frequencies, which is called aliasing. Similarly, the digital-to-analog conversion will produce frequencies above the Nyquist frequency which are aliases of lower frequencies. They should be filtered out for accuracy (however, it shouldn't matter if the Nyquist frequency exceeds the range of human hearing).

While digital audio is theoretically capable of perfectly representing analog sound, in practice, devices can inaccurately measure amplitude or mis-time a sample. The quality of such devices is paramount to accurate recording and reproduction of sound. Commercial audio production uses bit depths and sample rates higher than the final mastered product to ensure any error is minimized. While there is no need for frequencies between 22.05 kHZ and 48 kHZ (as no human can hear them), the higher sample rate prevents signal degradation as digital audio streams are manipulated and mixed together.

File Formats

Digital audio can be represented by a number of different file formats, which mainly pertain to the compression scheme they use (or at least support), as well as additional features such as Digital Rights Management (DRM).

It's important to understand that file formats do not necessarily indicate the exact codec used to encode the audio they contain. They are best thought of as containers, with a defined header format. This allows various software to know how to extract the header information, where to find the audio data in the file, and how to decode it. A codec is what is actually used to encode or decode the audio. Most audio player software has its own codecs, while some can refer to external software. Most file formats are limited, however, to support only a few codecs. Thus, it is common to associate compression/encoding schemes with file formats.

Compression uses various techniques to reduce the file size of digital audio. The main concern with compression is whether it noticeably degrades the audio signal - many schemes result in a loss of quality, known as lossy compression. Such compression is based upon human perception - reducing detail in places that it is least likely be noticed. Of course, different implementations may be more noticeable than others; and many feature different settings, allowing trade-offs between the compression ratio and quality of audio.

Lossless compression, on the other hand, results in no signal degradation, and the file can be converted back and forth from compressed and uncompressed formats without any change in the sample data. Compared to lossy compression, it typically features a much smaller compression ratio, and thus results in larger file sizes.

Digital Rights Management (DRM) refers to various methods that attempt to restrict decoding media files to authorized devices/users. ITunes used DRM for a while before abandoning its use. Many web stores still use forms of DRM, although it is now less common for music.

The main formats for storing uncompressed audio are .wav (for MS Windows), .aiff (for Apple Mac), and .au (Java, Unix). The primary lossy compression formats are .mp3, .m4a, .ogg, .aac, and .wma. A popular lossless file types is .flac. Windows media audio (.wma) supports both lossy and lossless encoding; however, it is generally assumed to use the lossy codec. .m4p and .wma are the primary formats of DRM-enabled audio files. Almost all of these file types are capable of additional audio formats; however, these are their most common usages.

DAW softare typically works exclusively with uncompressed audio files. Each DAW may have encoders to create various file types with various encoding during rendering or bouncing, and you may be able to import other formats into a project; but internally the DAW will prefer to work with raw audio data. Imported audio is converted to uncompressed formats and stored as such. This prevents it from devoting resources to encoding or decoding audio unnecessarily.