Digital Audio – Quality Issues


Free Download Mp4Gain
picture

Digital Audio – Quality Issues

Digital Audio Quality

Relatively recently, the concept of “multimedia” was included in our discourse, and now the computer is increasingly used as an entertainment center. Now the computer is forced to reproduce the sound that exists in it in the form of numbers.

Digital Audio Quality issues

Just as some connoisseurs of sound argue about the advantages of “tube” sound over “transistor” sound, there is an endless debate about which is better: digital or analog sound. Let’s try to figure it out.

For our ears, sound is air vibrations with a frequency of 20 Hz to 20 kHz, and the upper limit depends on age: in children it is 22-24 kHz, and in old age the perceived frequency decreases, up to 8 -12 kHz.

The frequencies of the indicated limits are perceived as vibrations, higher, they are not perceived by a person.

However, not all the detection bandwidth is used with the same intensity, so speech is clearly perceived in the range of 500 to 3500 Hz. But for listening to music, this is not enough. Ideally, the reproduced sound should not differ from the sound field of the microphone. That is, the recording and playback equipment must not introduce distortions within the limits of human perception.

The sound we hear from the speaker is electromechanically converted to an electrical signal during recording; then there is the amplification and processing of the analog electrical signal; analog to digital conversion; digital signal processing; frequency correction; recording procedure.

After the digitized sound is stored and transmitted. During playback, digital signal processing occurs first; follows the conversion from digital to analog; analog signal processing and amplification; electromechanical conversion to sound vibrations.

All of these procedures introduce their own distortions. The process of recording and sound processing takes place, as a rule, on studio equipment, which performs much better than home audio equipment. Therefore, although there are distortions, they are significantly less than the distortions introduced by home equipment at the playback stage. With amateur sound recording, errors appear in the recording stages.

The electromechanical conversion produced by the studio microphone produces a very weak signal that needs amplification.

Even in the ideal conditions of a professional recording studio, due to acoustic noise, the dynamic range of recorded music can be narrower than that provided by 16-bit audio.

When recording from multiple microphones, the signal is necessarily processed: channel volume levels are selected, noise is filtered, etc. Furthermore, the dynamic range of the signal is reduced, which leads to a significant increase in noise. But without this procedure, it would sound unsatisfactory when playing back the recording on a home computer.

The sound path has its own distortions, which can be divided into three groups:

1. Linear distortions are caused by the amplitude-frequency characteristic of the sound path and are a change in the ratio of the amplitudes and phases of various frequency components. Frequencies that were originally missing from the signal do not appear.

2. Non-linear distortion: a change in the shape of the original signal, which leads to the appearance of frequencies that are absent in the incoming signal, but depend on it.

3. Interference: the appearance of strange frequencies in the sound path that are not associated with the useful signal. Interference appears, for example, by electromagnetic interference, penetration into the sound path of the frequency of the supply voltage, etc.

However, all these distortions occur only in analog circuits (hence speculation about the frequency response of a digital output makes specialists smile). But don’t forget about the superficial defects of CDs, DVDs, and other optical storage media that store sound, leading to data loss.

The digitization of the signal is also associated with a lot of distortion, but first let’s look at the difference between analog and digital signals.

In an analog signal, the voltage changes smoothly over time, the signal is continuous. The digital signal is discrete, its value changes instantly. Furthermore, discretion is manifested in both frequency and amplitude region. Any change in signal value is sampled, and as a result, the values ​​are rounded to the nearest whole number.


Free Download Mp4Gain
picture


Mp4Gain Main Window
picture


Mp4Gain Features
picture


Free Download Mp4Gain
picture

Audio encoding: secrets revealed

Audio encoding: secrets revealed

audio encoding

Audio settings for video capture and transmission.
As people directly related to the AV sphere, we constantly talk about audio coding and audio codecs, but what is it?

Audio Encoding

An audio codec is essentially a device or algorithm that can encode and decode a digital audio signal.

In practice, the audio waves that are transmitted over the air are continuous analog signals. The signals are converted to digital format by a device called an analog-to-digital converter (ADC), and the reverse conversion device is called a digital-to-analog converter (DAC). The codec is located between these two functions and it is it that allows you to adjust some important parameters for the successful capture, recording and transmission of an audio signal: codec algorithm, sample rate, bit depth and data transfer rate.

The three most popular audio codecs are Pulse-Code Modulation (PCM), MP3, and Advanced Audio Coding (AAC). The choice of codec determines the compression rate and the recording quality. PCM is a codec used by computers, CDs, digital phones, and sometimes SACD. The source of the PCM signal is sampled at regular intervals, and each sample is the digital magnitude of the analog signal. PCM is the simplest option for digitizing an analog signal.

With the correct parameters, this digitized signal can be completely converted back to analog without any loss. Unfortunately, this codec, which provides almost complete identity with the original audio, is not very cheap, which results in large files, and these files are not suitable for streaming. We recommend using PCM to record digital images for your sources or when doing audio post-processing.

Fortunately, we always have the option of choosing a different codec that can compress digital data (compared to PCM) based on some helpful observations on the behavior of sound waves. But in this case, you have to make a compromise: all alternative algorithms are associated with “losses”, since it is impossible to completely restore the original signal, but nevertheless the result is so good that most users will not be able to notice the difference.

MP3 is an audio encoding format that uses a digital data compression algorithm that allows you to save the audio signal in smaller files. The MP3 codec is the most used by users to record and store music files. We recommend using MP3 to stream audio content as it requires less network bandwidth.

AAC is a newer audio encoding algorithm that is the successor to MP3. AAC has become the standard for MPEG-2 and MPEG-4 formats. In fact, this is also a digital data compression codec, but with less quality loss than MP3, when encoded with the same bit rate. We recommend using this codec for online streaming.

Sampling frequency (kHz, kHz)
Sample rate (or sample rate): the frequency with which the signal is digitized, stored, processed, or converted from analog to digital. Time sampling means that the signal is represented by a number of its samples (samples) taken at regular intervals.

Measured in hertz (Hz, Hz) or kilohertz (kHz, kHz,) 1 kHz equals 1000 Hz. For example, 44100 samples per second can be labeled 44100 Hz or 44.1 kHz. The selected sample rate will determine the maximum playback frequency and, as follows from Kotelnikov’s theorem, to fully restore the original signal, the sample rate must be twice the highest frequency in the signal spectrum.

As you know, the human ear can pick up frequencies between 20 Hz and 20 kHz. Given these parameters and the values ​​shown in the table below, you can understand why 44.1 kHz was chosen as the sampling frequency for CD and is still considered a very good frequency for recording.

What formats are used to represent digital audio?

What formats are used to represent digital audio?

Audio Formats

The format is used in two different ways.

Digital Audio Formats

When using a specialized medium or recording method and special read / write devices, the concept of format includes both physical characteristics of a sound carrier: the dimensions of a cassette with a magnetic tape or disk, the tape itself, or a disc, recording method, signal parameters, encoding and error protection principles, etc. .P. When using a universal information medium of wide application, for example, a flexible computer or a hard disk, the format is understood only as a method of encoding a digital signal, the peculiarities of the arrangement of bits and words and the structure of service information; all the “low-level” part directly related to working with the media, in this case, remains under the control of the computer and its operating system.

Of the specialized digital audio formats and media, the following are the best known today:

CD (Compact Disc) is a 120mm or 90mm single sided optical laser read / write disc, containing a maximum of 74 minutes of stereo sound at 44.1 kHz sampling rate and 16 linear quantization bits. The system is offered by Sony and Philips and is called CD-DA (Compact Disc – Digital Audio). For error protection, Cross Interleaved Reed-Solomon code (CIRC) and Hamming code 8-14 modulation (Eight to Fourteen Modulation, EFM) are used. A distinction is made between stamped compact discs (CD) write-only (CD-R) and rewritable (CD-RW).
PCM decoder (PCM deck): a system for converting the digital audio signal into a pseudo-video signal compatible with popular video formats (NTSC, PAL / SECAM) and vice versa. PCM decoders are used in combination with home (VHS) or studio (S-VHS, Beta, U-Matic) VCRs, using them as read / write devices. The devices operate with 16-bit linear quantization at sample rates of 44.056 kHz (NTSC) and 44.1 kHz (PAL / SECAM) and can record a two- or four-channel digital signal. In fact, such a decoder is a modem (modulator-demodulator) for a video signal.
S-DAT (Fixed Head Digital Audio Tape – Fixed Head Digital Audio Tape) is a system similar to a conventional cassette recorder, in which recording and reading is performed by a block of thin film fixed heads in a 3.81 mm wide tape in a double-sided cassette with dimensions of 86 x 55.5 x 9.5 mm. It implements two- or four-channel 16-bit recording at 32, 44.1, and 48 kHz.
R-DAT (Rotating Head Digital Audio Tape) is a VCR-like system with cross-tilted rotating head recording. The most popular tape-based digital recording format, R-DAT systems are often referred to simply as DAT. The R-DAT uses a 73 x 54 x 10.5mm cassette, with a 3.81mm wide tape, and the cassette and tape system itself is very similar to a typical VCR. The basic belt speed is 8.15mm / s, the rotation speed of the main unit is 2000rpm. R-DAT operates with a two-channel signal (on some models, four channels) at sample rates of 44.1 and 48 kHz with 16-bit linear quantization and 32 kHz with 12-bit non-linear quantization. To guard against errors, a double Reed-Solomon code and modulation with an 8-10 code are used. Cassette capacity – 80. .240 minutes depending on speed and belt length. Domestic DAT recorders are usually equipped with a phonogram illegal copy protection system, which does not allow recording from the analog input at a frequency of 44.1 kHz, as well as direct digital copying in the presence of SCMS prohibition codes (Serial Code Managenent System). Studio tape recorders have no such restrictions.
DASH (Digital Audio Stationary Head) is a 6.3 and 12.7 mm wide magnetic tape recording system with fixed heads. Belt speed is 19.05, 38.1, 76.2 cm / sec. Implements 16-bit recording with sample rates of 44.056, 44.1 and 48 kHz from 2 to 48 channels.
ADAT (Alesis DAT) is a proprietary system for recording eight-channel audio on S-VHS videotape, developed by Alesis. It uses linear quantization of 16 bits at 48 kHz, the capacity of the cassette is up to 60 minutes per channel. ADAT tape recorders can be cascaded so that a 128-channel synchronous recording system can be assembled.

Digital Audio Coding

Digital Audio Coding

Digital Audio coding

Digital audio technologies are used to record, process, mass produce, and distribute sound, including the recording of songs, instrumental pieces, podcasts, sound effects, and other sounds.

Digital Audio Coding

Today’s online music distribution relies on digital recording and data compression. The availability of music as data files instead of physical objects has significantly reduced distribution costs. Before the advent of digital sound, the music industry distributed and sold music, selling physical copies in the form of records and cassettes.

Using online and digital audio distribution systems such as iTunes, companies sell digital audio files to consumers that the consumer receives over the Internet. An analog audio system converts the physical waveforms of sound into electrical representations of these waveforms using a transducer such as a microphone. The sounds are then stored on analog media, such as magnetic tape, or transmitted through analog media, such as a telephone line or radio. For playback, the process is reversed: an electrical audio signal is amplified and then converted back to physical waveforms through a speaker.

Analog audio retains its fundamental waveform characteristics when stored, converted, dubbed, and amplified. Analog audio signals are prone to noise and distortion due to the inherent characteristics of electronic circuits and related devices. Interference in a digital system does not result in an error, unless the interference is large enough to cause one character to be misinterpreted as another character or to be out of sequence.

Therefore, it is generally possible to have a completely error-free digital audio system in which there is no noise or distortion between converting to digital and converting to analog. The digital audio signal can be further encoded to correct any errors that may occur during storage or transmission of the signal. This technique, known as channel coding, is necessary for broadcast or recorded digital systems to maintain bit fidelity. Modulation from eight to fourteen is a channel code used on audio CDs. Conversion process

The life cycle of sound from its source through ADC, digital processing, DAC, and finally again as sound. A digital audio system begins with an ADC, which converts an analog signal into digital. The ADC operates at the specified sample rate and converts to a known bit resolution.

For example, CD audio has a sample rate of 44.1 kHz (44,100 samples per second) and a resolution of 16 bits for each stereo channel. Analog signals that have not yet been band-limited should be damped before conversion to avoid interpolation distortion caused by audio signals above the Nyquist frequency (half the sample rate).

The digital audio signal can be stored or transmitted. Digital audio can be stored on a CD, digital audio player, hard drive, USB flash drive, or any other digital storage device. The digital signal can be modified by digital signal processing, where effects can be filtered or applied. Frequency transform Sample rates, including increasing and decreasing the sample rate, can be used to match signals that have been encoded with a different sample rate to a common preprocessing sample rate. Audio compression methods such as MP3, Advanced Audio Coding, Ogg Vorbis, or FLAC are commonly used to reduce file size.

Digital audio can be transmitted through digital audio interfaces such as AES3 or MADI. Digital audio can be transmitted over a network using Audio over Ethernet, Audio over IP, or other standards and media transmission systems. For playback, digital audio must be converted back to an analog signal using a DAC. According to the Nyquist-Shannon Sampling Theorem, with some practical and theoretical limitations, the bandwidth-limited version of the original analog signal can be accurately reconstructed from the digital signal. –

Digital audio file formats wav, mp3, aiff, ogg, flac, m4a

Digital audio file formats wav, mp3, aiff, ogg, flac, m4a

digital audio formats

The last five years gave a great boost to the development of portable and stationary audio systems, and with this support for a variety of digital audio formats.

DIGITAL AUDIO FORMATS

Small pocket devices have a large internal memory and fixed audio equipment has become even smarter and more demanding. That is why, now, we can not save space on the player and download songs that weigh between 15 and 30 MB each, but at home, listen to digital music in a quality equal to the sound of an analog vinyl.

Description of popular digital audio formats
However, the most widespread audio formats still have their pros and cons, and even in an urgent matter like digital audio, a “panacea” has not yet been found. Classic digital audio formats are divided into “compressed” and “uncompressed” streams, as well as “lossless” formats, which exclude loss of sound.

Description of digital audio formats Description of digital audio formats

Wav audio format
The waveform audio file format (WAVE, WAV – “in waveform”) is a file format for storing a recording of an uncompressed digitized audio sequence. In general, this is the most common format for working in the studio and in broadcasting. allows you to get the most honest sound quality. For example, the standard audio CD format is an LPCM audio stream, with parameters: 2ch (stereo), 44-100Hz, 16bit.

Mp3 audio format
MPEG-1/2 Audio Layer 3: (MP3) is the most popular digital format for storing compressed audio. The MP3 format uses a special algorithm designed to greatly reduce the size of the original file. This format allows you to keep the audio close to the original sound, but thanks to a variety of settings, extremely small size.
Compared to the standard audio CD format, a file in MP3 format and a bit rate of 128 kbps will be approximately 1/11 the size of the original file.

FLAC audio format
FLAC (Free Lossless Audio Codec) is a popular free codec designed for lossless compression of audio data. What does that mean? Unlike lossy audio codecs such as MP3 or OGG, the FLAC audio codec does not remove any information from the audio stream. This format is ideal for audiophiles who create their own music collections and listen to music on high-quality equipment.

Ogg audio format
OGG is a format that has not gained great popularity, but is nonetheless used by a fairly large audience. The OGG format, similar to MP3, compresses audio with loss of quality, but is fundamentally different in practical conversions. This made it possible to get better quality with a smaller file size and to display this codec as absolutely independent. In addition to similar formats that convert lossy audio, OGG has the ability to adjust container properties.

Aiff audio format
The Audio Interchange File Format (AIFF) is a fairly universal audio file format developed by Apple, which is used to store audio data. Like its counterpart, the WAV format, it is uncompressed audio and is widely used in professional recordings and music production.
The .aiff and .aif files created by Apple Loops are used by GarageBand and Logic Audio music editors.

M4a audio format
Apple Losseles (also known as Apple Lossless Encoder, ALE or Apple Lossless Audio Codec, ALAC) (m4a) is another Apple development. This audio format refers to uncompressed audio, which provides lossless playback. It is a fairly specific format, which is mainly supported by products of the creator company, and in some cases, as in the iPhone system sounds, where it is possible to use exclusively the m4a format.

Sound file resolution. Audio encoding and processing

Sound file resolution. Audio encoding and processing

Digital audio

Basic concepts

udio encoding

The sampling frequency (f) determines the number of samples stored in 1 second;

1 Hz (one hertz) is one count per second,

and 8 kHz is 8000 samples per second

The encoding depth (b) is the number of bits required to encode the level of

Memory capacity for data storage 1 channel (mono)

(to store information about a sound with a duration of t seconds, encoded with a sampling rate of f Hz and a encoding depth of b bits, 1 bit of memory is required)
For 2-channel (stereo) recording, the amount of memory required to store data for one channel is multiplied by 2

I = f b t 2

Units of measurement I – bits, b – bits, f – Hertz, t – seconds Sampling frequency 44.1 kHz, 22.05 kHz, 11.025 kHz

Audio encoding
Basic theoretical provisions

Sound time sampling. In order for a computer to process sound, a continuous audio signal must be converted to a discrete digital form using time sampling. A continuous sound wave is divided into separate small time sections, for each section a certain value of sound intensity is set.

Therefore, the continuous dependence of the loudness of the sound at time A (t) is replaced by a discrete sequence of loudness levels. On the graph, this appears to replace a smooth curve with a sequence of “steps.”

Sampling frequency. A microphone connected to the sound card is used to record analog audio and convert it to digital format. The quality of the digital sound obtained depends on the number of measurements of the sound volume level per unit time, that is, sampling rate. The more measurements are made in 1 second (the higher the sampling frequency), the more accurately the “ladder” of the digital audio signal repeats the curve of the analog signal.

Audio sample rate is the number of measurements of the volume of a sound per second, measured in Hertz (Hz). Let us denote the sampling frequency with the letter f.

The audio sample rate can vary between 8000 and 48000 sound volume measurements per second. One of three frequencies is selected for encoding: 44.1 KHz, 22.05 KHz, 11.025 KHz.

Audio encoding depth. Each “step” is assigned a specific value for the sound volume level. Loudness levels can be seen as a set of possible states N, for which encoding a certain amount of information b is required, which is called the audio encoding depth.

Audio encoding depth is the amount of information required to encode the discrete volume levels of digital audio.

If the encoding depth is known, then the number of digital audio loudness levels can be calculated using the formula N = 2b. Let the audio encoding depth be 16 bit, then the number of sound volume levels is:

N = 2 b = 2 16 = 65 536.

During the encoding process, each sound volume level is assigned its own 16-bit binary code, the lowest sound level will correspond to the code 0000000000000000 and the highest – 1111111111111111.

The quality of digitized sound. The higher the sampling frequency and depth of the sound, the better the sound of the digitized sound. The lowest quality of digitized sound, corresponding to the quality of telephone communication, is obtained at a sampling rate of 8000 times per second, a sampling rate of 8 bits, and by recording an audio track (“mono” mode). The highest quality of digitized sound, corresponding to the quality of an audio CD, is achieved with a sampling rate of 48,000 times per second, a sampling rate of 16 bits and the recording of two audio tracks (stereo mode) .

Video codecs and containers.

Video codecs and containers.

Video Codec

This article is intended to refer here to those who are trying to “convert” something, without understanding what they are doing and why.

Video Codecs

To work as efficiently as possible with any object, you need to understand how it works. If the video file is for you a mysterious black box, inside which mysterious things happen, perhaps not without the help of black magic, then your effectiveness will be minimal.

So. All information on the computer is in the form of files. This, I hope, is not a surprise to anyone. Here we will start from this basic concept.

Any video file must be a container. A container is a repository of content. There are multi-structure storages – these are container formats. For example, a bento box is an example of a container. You can put sushi or tempura on it. What can you put in a video container? Well, at least image and sound, one at a time. This is a set without which there is nothing to do. What can you put to the maximum? The modern Matryoshka container allows you to put various video and audio tracks, text and graphic subtitles, fonts to display them, images and I don’t know what else.

Going back to the bento box example, note that miso cannot be poured into it; will flow in fig. Not all containers can accept all flows. There are compatibility restrictions that make life difficult.

Container examples: mpeg, avi, mkv, mp4, ogm, vob, mov, rm, divx, asf. You don’t have to look closely at the list to understand that these are standard file extensions. Of course. Because file = container.

Streams or tracks are stored inside the container. These streams have a format called a codec. And this difference must be understood with particular clarity. The container is a file format. And the codec is the stream format it contains. They are two independent things. Yes, there are some inextricably linked containers and codecs. For example, the Real Media container can only store real video and real audio streams. And vice versa, these formats cannot be stored in any other container (almost, as I have already been corrected). But they are still different concepts that should not be confused.

The codec concept usually includes the following aspects:
1) The actual data storage format.
2) Software that allows you to encode information in this format and / or decode it from it.

Examples of video codecs: divx, xvid, avc, x264, vp6, vp7, mpeg-1, mpeg-2, huffyuv.
Examples of audio codecs: mp3, ogg, ac3, aac.

While containers are generally distinguished by file extensions, codecs are distinguished by the four-character FourCC code.

The codec concept is usually associated with a kind of compression. Raw (uncompressed) streams also have their own formats, but they do not require decoding, and therefore the concept of codec is generally not applied to them.

Now let’s take a look at the most popular containers, codecs, and related issues. As a general rule, the problems we have are of two types: related to reproduction and related to editing.

MPEG is one of the oldest containers. It can store only video in mpeg-1 format and audio in mp2 format. And in a friendly way, with quite strict restrictions on the size of the image and the bitrate of the sound. Due to the age and primitiveness of the format, almost all players and publishers understand it. But for the same reasons, it became almost impossible to meet him. Nobody needs these things.

AVI is also quite old, but it is still a very useful container. It’s good because, again, all the players and all the editors get it. Almost all mpeg-based formats fit into it, as well as many that support them. The following video formats do not fit avi: avc (aka Nero AVC or Nero H.264), wmv below version 9, as well as any tinsel like actual video, which was originally designed to be incompatible with anything in the world. By sounds, supposedly anything, except Vorbis ogg.

OGM is where Vorbis ogg goes. Because the format was created on the basis of this very ogg. At the moment, he is practically ousted by the matryoshka because he can do the same, only better. It is also not compatible with any conventional software.

MKV is a nesting doll that can fit just about anything except flash video. But due to its complexity and versatility, it is still possible to do with it only things like: mount, look and dismount.

MP4 is actually modern MPEG. It only takes things that are compatible with the MPEG standard, but at the same time includes its latest updates.

Compressed audio encoding formats.

Compressed audio encoding formats.

audio encoding

MP3 (or rather, MPEG 1 Audio Level 3): no comment, compatible everywhere and by everyone, the lack of this “eternal” format is one: only two channels, which limits its use in cinema systems at home modern.
Multi-channel MP3 (5.1) MPEG 2 Audio Level 3.

audio encoding
WMA: Windows Media Audio, formally a better and more modern competitor to Microsoft’s mp3. It is not used much, although it is widely compatible with hardware.
OGG Vorbis is a best modern mp3 competitor from the open source community. Deprived of any license restrictions, it is used more and more frequently.
AAC: Advanced Audio Coding is Apple’s main audio format built into all of its iPads, iPhones, iTunes, etc. The main advantage is that it is technically more advanced than mp3, allowing sample rates of up to 96 kHz and theoretically a completely insane number of channels in one file, up to 48. It is also used in digital satellite radio. Just as mp3 is a compressed format, the quality of 96Kbps AAC is comparable to the quality of 128Kbps of mp3 (we are talking about two channels in both cases).
Dolby Digital (AC-3) is probably the most popular standard for digital audio in cinematography, due to the fact that it appeared on the market as early as 1995, it exists in two versions: DD2.0 (for high-quality stereo sound) and DD5 .1 – five full channels and one defective for a subwoofer. Players are compatible with all of them for obvious reasons, the bitrate is 640Kbps in all cases.
Dolby Digital Plus or E-AC-3 is an attempt to improve on the usual Dolby Digital, but the previous generation decoders and receivers do not support tracks in the Dolby Digital Plus format, the reasons for this are radical changes: the number of channels increased to 7.1, the bit rate – to 1, 7 Mbps This will not go through S / PDIF (when transmitting via such a cable, you will have to use downmix on DD5.1 ​​or on DTS with quality loss), but HDMI normally copes with Dolby Digital Plus as of version 1.3, you can find such tracks on Blu-Ray discs …
Dolby TrueHD – We practically have 8 tracks almost uncompressed at 96 KHz / 24 bits or 6 at 192 KHz / 24 bits, the total bit rate reaches 18 Mbit / sec, which requires decoding in the player and transmission to the receiver in the analog path, or using HDMI 1.3 or higher. For Blu-Ray, this audio coding system is optional.
DTS is a lossy digital audio coding system for cinemas, which later appeared on DVD, it is analogous to Dolby Digital 5.1, but somewhat more flexible, allowing in addition to 2.0 and 5.1 to use other schemes, such as 4.0 and 4.1, there is also a choice between two fixed bit rates of 1500 Kbps and 750 Kbps. In the first case, DTS clearly outperforms Dolby Digital in sound quality; in the second, the difference between systems is controversial.
DTS-HD is a further evolution of DTS, the number of channels has been brought to 7.1 in 96KHz / 24bit mode, the bit rate can be selected between 6Mbps and 3Mbps, it is an optional audio format for Blu-Ray. The situation with the sound transmission to the receiver is almost the same as with DolbyTrueHD.

Lossless or uncompressed compressed audio encoding formats.

LPCM is simply uncompressed audio. It is usually stereo. It should not be confused with a WAV file, it is a container and there may be something other than PCM WAV inside.
APE is a specific lossless audio compression format. Loved by audiophiles.
Flac is its competitor and analog, the differences between them are beyond the scope of this review.
Lossless audio
Lossless apple

Subtitle formats.
SRT: text format, can be attached as a separate file with the same extension. Compared to the first versions of this format, the design possibilities have been significantly increased. It can also exist within MKV.
SUB / IDX is a graphic subtitle format extracted from DVD. It can fit MKV or MP4.
s2k, ssa, ass: some more advanced text formats, ass can be placed inside MKV.
smi is a textual format based on SGML, the direct ancestor of HTML.
PGS is a graphical subtitle format, the main one for Blu-Ray, but it can also exist in ts and MKV containers.

ABOUT DIGITAL AUDIO FORMATS

ABOUT DIGITAL AUDIO FORMATS

Digital Audio Formats

Today, there are several digital audio formats that are superior in quality to compact discs and are available on both physical media and the Internet. What are advanced sound lovers listening to now? Let’s find out.

Digital Audio Formats

The capabilities and quality of the CD-DA format were initially limited by the capabilities of CD as a medium. Legend has it that the standard 74-minute compact disc capacity was chosen in order to be able to record long classical pieces without splitting into two discs. And to be absolutely precise, this figure appeared thanks to Beethoven’s Ninth Symphony: it lasts exactly 74 minutes. Another default parameter was the 44.1 kHz sample rate. This figure defines the upper limit of the reproduced frequency range. For a CD that had to reproduce frequencies up to 20 kHz, this was the lowest possible carrier frequency. As a result, the only field of maneuver was the bit depth, the level of which was 16 bits. With regard to sound recording, bit depth determines its dynamic range and resolution.

The CD cannot be copied into the memory of the computer in the usual way, since we usually copy files. To save a CD-DA, you need a special program, a program that allows you to convert data recorded on an audio disc to PCM format (WAV file). A properly organized CD-DA ripping process allows you to get a completely identical digital copy on your hard drive. Audio CDs are generally saved on a computer as a large FLAC audio file (also WAV, WV, or APE) with a CUE index card or as separate tracks.

As the best digital audio format, the CD did not last that long, just over ten years. In the mid-nineties, the first format appeared that allows for better sound quality. HDCD was an improved version of CD-DA. Their difference consisted in a special recording algorithm that made it possible to save additional data on the sampling depth in a standard CD format. With an HDCD decoder, the output signal received not 16, but 20 bits, which did not give the standard of 96, but up to 120 dB of dynamic range and a very noticeable increase in recording resolution. At the same time, devices without an HDCD decoder played discs like normal CD-DAs. Interestingly, when saving such a disk on a PC in the same way,

The next leap in terms of sound quality came at the beginning of the new millennium. Two HD audio formats were introduced to the audiophile audience at once, appearing almost simultaneously. DVD-Audio, a further development of the traditional recording method and promoted by Panasonic and Toshiba. It is capable of recording 24-bit / 192 kHz in stereo mode and 24-bit / 96 kHz in multi-channel mode.

The SACD format competed with it, which, by the way, looked much less like a normal CD, although it was called “super CD”. Super Audio CD, developed by Sony, was based on the revolutionary DSD encoding algorithm. This digitizing method assumed one-bit sampling at an ultra-high frequency of 2.8224 MHz. The encoding and decoding principles of a DSD stream are much simpler than in high-bit formats and are essentially closer to the principles of analog technology. At the same time, the SACD format retains all the advantages of the advanced digital format and has output characteristics comparable to DVD-Audio in both sound quality and number of channels.

Both DVD-Audio and SACD were designed with a high level of copy protection, but inquisitive minds have already won over both formats, so if desired, the content of both disc types can be saved to a PC as images. ISO (without changing the structure and original codec) or FLAC tracks in 24-bit / 96 kHz or 24-bit / 192 kHz. Almost simultaneously with the DVD-Audio and SACD formats, another original format for publishing high-quality music was born: DAD 24/96. DAD stands for Digital Audio Disk, but it is essentially a DVD-Video with a high-quality still image and sound that can be played on any standard DVD player or PC.

Obviously, with this approach, Blu-ray media, with its HD sound formats, recorded in high quality without compression, is quite applicable for recording music in high quality. However, at the moment there are few such publications, and a special version of the BD-Audio format has every chance of not seeing the light of day, as the sale of high-quality audio material is already very active on the Internet. Anyone who does not want to convert DVD-Audio, DAD and SACD discs to the FLAC format on their own can officially buy albums already converted in 24-bit / 96 kHz or 24-bit / 192 kHz quality.

Audio encoding and processing

Audio encoding and processing

Audio processing

Sound information. Sound is a wave that travels through air, water, or other medium with a continuously changing intensity and frequency.

Audio processing

A person perceives sound waves (air vibrations) with the help of hearing in the form of sound of different volume and pitch. The higher the intensity of the sound wave, the louder the sound, the higher the frequency of the wave, the higher the pitch of the sound

The human ear perceives sound at a frequency of 20 vibrations per second (low sound) to 20,000 vibrations per second (high sound).

A person can perceive sound in a wide range of intensities, in which the maximum intensity is 10 14 times greater than the minimum (one hundred thousand billion times). To measure the volume of sound, a special unit “decibel” (dbl) is used (Table 5.1). Decreasing or increasing the sound volume by 10 dB corresponds to a decrease or increase in sound intensity by 10 times.

Table 5.1. Sound volume
Sound Volume in decibels
Lower limit of human ear sensitivity 0
Leaf whisper ten
Conversation 60
Horn 90
Jet engine 120
Pain threshold 140
Sound time sampling. In order for a computer to process sound, a continuous audio signal must be converted to a discrete digital form using time sampling. A continuous sound wave is divided into separate small time sections, for each section a certain value of sound intensity is set.

Therefore, the continuous dependence of the loudness of the sound at time A (t) is replaced by a discrete sequence of loudness levels. On the graph, this appears to replace a smooth curve with a sequence of “steps”

Sampling frequency. A microphone connected to the sound card is used to record analog sound and convert it to digital format. The quality of the digital sound obtained depends on the number of measurements of the sound volume level per unit of time, that is, the sampling frequency. The more measurements that are made in 1 second (the higher the sampling frequency), the more accurately the “ladder” of the digital audio signal repeats the curve of the dialogue signal.

Audio sample rate is the number of audio volume measurements in one second.

The audio sample rate can vary between 8000 and 48000 sound volume measurements per second.

Audio encoding depth. Each “step” is assigned a specific value for the sound volume level. Loudness levels of sound can be viewed as a set of possible states N, for which a certain amount of information I is required, which is called audio coding depth.

Audio encoding depth is the amount of information required to encode the discrete volume levels of digital audio.

If the known encoding depth, the number of digital audio volume levels can be calculated using the formula N = 2 I. Let the audio encoding depth be 16 bit, then the number of sound volume levels is:

N = 2 I = 2 16 = 65 536.

During the encoding process, each sound volume level is assigned its own 16-bit binary code, the lowest sound level will correspond to the code 0000000000000000 and the highest – 1111111111111111.

The quality of digitized sound. The higher the sampling frequency and depth of the sound, the better the sound of the digitized sound. The lowest quality of digitized sound, corresponding to the quality of telephone communication, is obtained at a sampling rate of 8000 times per second, a sampling rate of 8 bits, and by recording an audio track (“mono” mode). The highest quality of digitized sound, corresponding to the quality of an audio CD, is achieved with a sampling rate of 48,000 times per second, a sampling rate of 16 bits and the recording of two audio tracks (stereo mode) .

It should be remembered that the higher the quality of the digital sound, the greater the volume of information in the audio file. It is possible to estimate the volume of information of a digital stereo sound file with a duration of 1 second with an average sound quality (16 bits, 24,000 measurements per second). To do this, the encoding depth must be multiplied by the number of measurements in 1 second and multiplied by 2 (stereo sound):

16 bits × 24,000 × 2 = 768,000 bits = 96,000 bytes = 93.75 KB.

Sound editors. Sound editors allow you not only to record and play sound, but also to edit it.