In what format is it better to listen to music? Part 2


Free Download Mp4Gain
picture

In what format is it better to listen to music? Part 2

Audio Formats

The reference value of the audible range for humans is 16 Hz to 20 kHz, but you cannot hear and be aware of all incoming sounds simultaneously.

audio files

Hearing is discreet and your hearing sensitivity is not linear.

Modern psychoacoustic models accurately assess human hearing and are constantly improving. In fact, despite the guarantees of music lovers, musicians and audiophiles, to the inexperienced middle ear, the initial appearance of MP3 in maximum quality has become extremely noticeable. There are exceptions, they cannot cease to exist. But they are not always easily noticed by blind listening.

Formats using psychoacoustic compression models
There are many of these formats for lossy audio compression. The most common today are the following.

OGG (Vorbis)
In general, a file with the * .ogg extension is a “container”: it can contain multiple sound recordings with their own tags and characteristics. Most of the time, the files stored in it are compressed with the Ogg Vorbis codec, although others can be used, including MP3 or FLAC.

Its main advantages include a wide range of possible parameters during encoding: the audio sampling frequency can reach 192 kHz, the bit depth is 32 bits. By default, OGG uses a variable bit rate (although this is not shown on the properties screen), which can go up to 1000 kbps.

MP3
Unlike the free OGG, MP3 was developed by the Fraunhofer Society, an association of German institutes for applied research, which is very important for modern acoustics. Among audiophiles, by the way, this is an extremely respected office, yet they don’t like to admit it. But its developments are closely watched.

Unlike OGG, it can have variable (VBR) and constant (CBR) bit rate. By the way, it was thanks to MP3 that it was discovered that not all recordings can be encoded with high quality with a variable bit rate (see the above reasons, the encoding algorithms and their results in this case may be different when encoding the same source ).


Free Download Mp4Gain
picture


Mp4Gain Main Window
picture


Mp4Gain Features
picture


Free Download Mp4Gain
picture

In what format is it better to listen to music?

In what format is it better to listen to music?

Lossy compression

Understanding digital audio formats is not easy. It is even more difficult to come to an unequivocal conclusion in which format it is better to listen to music.

Lossy Formats

If you look at the audio format comparison table on Wikipedia, your eyes will start to flutter with columns of silent numbers. Let’s try to find out what’s behind this.
In what format is it better to listen to music? Three lost whales
Share this
sixteen
Let’s make a reservation right away that the article talks ONLY about general characteristics and will not include some details. Moving forward, Lifehacker will conduct its own unbiased investigation. And today we will try to generalize the already known experience in one way or another.

There is an analog and a figure.

The analog is good, but short-lived and inconvenient. Therefore, analog media, despite high vinyl sales, will not be making a comeback.

Digital audio can be of three main types:

in a format that does not use compression;
in a format that uses lossless compression;
in a format that uses lossy compression.
At first glance, lossless formats are more promising. This is not always the case, as we will discuss in more detail in one of the following materials. Uncompressed formats make no sense other than storing the master recordings needed to create audio content. They are easier to restore. Storing and listening to home recordings is superfluous.

Of the many parameters of digital audio, the user must first be concerned with sample rate (the accuracy of digitizing an analog signal in time), bit depth (the accuracy of digitizing in amplitude – volume) , the bit rate (the amount of information contained in the file in terms of one second).

Today we will talk about lossy.

For compressed sound, the concept of the psychoacoustic model is very important – the ideas of scientists and engineers about how a person perceives sound. The ear perceives the entire spectrum of acoustic waves entering it. However, the brain processes the signals.

Lossy compression format at a glance

Lossy compression format at a glance

lossy compression

“As you know, the music we listen to consists of a set of signals, each of which has its own characteristics, including loudness.

LOSSY COMPRESSION

The human hearing aid is designed so that we do not distinguish or poorly distinguish a weak (low) signal from the background of a strong (strong) signal. This principle forms the basis of modern means of compression (compression) of audio data.

If we imagine that a signal of a certain length is divided into many parts, and each part is processed in such a way that a weaker signal, which is difficult to distinguish from the background of a strong one, falls under the knife, and one remains a signal louder, then this will be an approximate audio compression model. Consequently, the level of data compression will depend on how many parts (samples) the original file will be divided into and how many weak signals from each individual sample will be removed (what the bit rate will be, the number of bits in a sample). sample of a specified duration). This coding principle is called lossy coding or lossy coding.

Ogg Vorbis is a completely open and patent-free audio format that allows you to store and transmit audio information with high sound quality (44.1-48.0 kHz sample rate, more than 16 bits, polyphony (multi-channel audio) ) and bit rates ranging from 16 to 512 Kbps per channel. In this case, the number of processed channels can reach 255.

MP3 or MPEG-1 Layer 3 audio is by far the most popular format for storing and transmitting compressed data. This format was developed by the Fraunhofer Institut, Germany. “Http://ru.wikibooks.org/wiki/Compression_Audio_data_with_lossy

Comparative tests

Sound Forge 7.0 (Spectral Analysis / Spectrum Analysis function) was used for the analysis of the sound signal.

“Spectral analysis is a signal processing technique that can reveal the frequency content of a signal. Solving the problems of spectral analysis is possible through the use of the fast Fourier transform, which makes it possible to determine the contribution of individual components of the vibration spectrum to the overall vibration picture. “Http: //masters.donntu. edu.ua/2007/fema/belinskaya/library/a4/art4.htm

The following graphs were obtained in the form of an amplitude distribution in the frequency domain, the spectrum of the signal is presented using a Blackman-Harris / Blackman-Harris window and a maximum sampling frequency (FFT size) of 65536, this gives allows you to analyze the smallest details of the signal at frequencies around 20,000 Hz, without smoothing.

The analysis of the spectrum of the compressed signal assumes the presence of a recording of the original quality, for this we use a licensed audio CD made in the USA “Kevin Yost – Bongo Madness”, with standard characteristics 44100 Hz / 16 bit

The rich electronic sound spans the entire frequency spectrum and captures even the inaudible range (20,000 Hz to 22,000 Hz), as can be seen in the graph below. Considering that it is generally possible to notice codec compression at higher frequencies, the 10-20 kHz range will be considered.

Lossy compression: Compress audio and video

Lossy compression: Compress audio and video

Lossy cmpression

High-quality digitized audio requires a large amount of disk space. Attempts to reduce file size using standard file cabinets do not yield significant gains due to the specificity of the audio data. However, it is possible to achieve a fairly significant level of compression of the audio information using special methods based on the analysis of the data structure and subsequent compression with some loss.

Lossy Compression

The real possibility of sound processing comparable in quality to existing analog examples did not appear until the late 1980s. In 1988, the International Organization for Standardization (ISO) formed the MPEG (Moving Image Experts Group) committee. , whose main task is to develop standards for the encoding of moving images, sound and their combination. During the ten years of its existence, the committee has developed a series of norms on this subject. As a result, summarizing the extensive research in this area, several specific formats were recommended for storing data, which are excellent in quality of results and data flow.

Currently, the three most common video storage standards are MPEG-1, MPEG-2, and MPEG-4. Within the first two formats, there are also formats for storing audio information: Layer-1, Layer-2 and Layer-3. These three audio formats are defined for MPEG-1 and minor extensions are used in MPEG-2. The three formats are similar to each other, but use different levels of compromise between compression and complexity. Layer-1 is the simplest level, it does not require significant compression costs, but it also provides a negligible compression ratio. Layer-3 level: the most time consuming and provides the best compression. Recently, this format has gained immense popularity. It is often called MP3. This name is associated with the extension of the audio files stored in this format.

Founded idea, in which all audio signal loss compression methods – ignore the subtle details of the original sound, which are outside of what the human ear perceives. Here several points can be highlighted.

Noise level. Sound compression is based on a simple fact: if a person is near a loud siren, they are unlikely to hear the conversation of the people who are nearby. Also, this happens not because a person pays close attention to a loud sound, but to a greater extent because the human ear actually misses out sounds that are in the same frequency range as a louder sound. This effect is called masking, it changes with the difference in volume and frequency of the sound.

The second point is the division of the audio frequency band into subbands, each of which is further processed separately. The encoding program extracts the loudest sounds in each band and uses this information to determine an acceptable noise level for that band. The best encoding programs also take into account the influence of adjacent bands. A very loud sound in one band can affect the masking effect and nearby bands.

Another point of the codification is the use of a psychoacoustic model based on the peculiarities of the human perception of sound. Compression The use of this model is based on removing obviously inaudible frequencies with more careful preservation of sounds that are clearly distinguishable by the human ear. Unfortunately, there can be no exact mathematical formulas here. The human perception of sound is a complex process, not fully understood, so the choice of compression methods is based on analyzing listening and comparing compressed sounds differently by teams of experts. But here there are practically limitless possibilities in the field of improving psychoacoustic models. Most of the existing algorithms to encode the human voice are based on the high predictability of said signal; Universal MPEG compression algorithms have tried to apply this technique with variable success.

Another compression technique is the use of so-called joint stereo. It is known that the human hearing aid can only determine the direction of the mid frequencies, the high and low sound, so to speak, separately from the source. This means that these background frequencies can be encoded into a mono signal. In addition to all this, compression uses the difference in the complexity of the flows in the channels. For example, if there is total silence on the right channel for some time, this “reserved” place is used to improve the quality of the left channel.