FLAC versus MP3: Does it make sense to use a “lossless” audio codec?


Free Download Mp4Gain
picture

Free Lossless Audio Codec (FLAC) is an audio format that is unknown to the public, but is particularly loved by the most demanding audiophiles: unlike MP3, AAC and their partners, FLAC is lossless, which means that it compresses audio with no loss of information. The advantage is the superior quality and the certainty that a 1: 1 copy of the original can only be made from the files. The disadvantage is that the tracks “weigh” significantly more. Is it a winning engagement or not?

mp3 vs flac

Let’s go to the conclusions

If you have original, rare and / or valuable audio recordings that you want to keep indefinitely for years (even if the original media wear out), FLAC is the optimal choice.

But if you only make it a matter of quality, think twice about it: it might not be worth it.

Never convert from MP3 to FLAC – it would take up extra space for free.

best audio format
lossless

FLAC is … an audio codec

Let’s start with FLAC being an audio codec: that is, it is used to compress music or other sound sequences so that they take up less space than storing the same information directly.

To get an idea of ​​how basic this is, keep in mind that an hour of uncompressed audio (no video) takes 620MB.

FLAC is … “free”

Then there is the word “free” which should be interpreted as “free” and “free”. FLAC is distributed in open source mode (GPL license). This means that its specifications can be used by anyone without paying any commission.

In contrast, there are MP3s that must be used within software and device manufacturers by Thomson Consumer Electronics and the Fraunhofer Society.

FLAC is … lossless

The third aspect concerns the type of compression used. While MP3 and AAC reduce the weight of the file by permanently eliminating frequencies and nuances that are generally unrecognizable to the human ear, FLAC retains every last bit present in the source and then applies only a number of specific optimizations, before the file is saved result in reducing the size on the hard drive. When the file is opened, however, the process is reversed and FLAC returns the original audio perfectly.

The procedure is similar in many ways to that of compressing in zip format: when the file is unpacked, we get the perfectly preserved initial file again. The difference is that FLAC was specially developed for working with audio and significantly reduced the size of the source file.

Lossless = quality + flexibility

Audiophiles complain that the “cuts” in the MP3 codec are too heavy and that the quality is unacceptably affected. In contrast, the performance at FLAC corresponds 100% to the original “master”.

Added to this is the aspect of optimal data storage: FLAC supporters point out that a “ripped” CD in this format can later be recreated from the files themselves and that a bit-by-bit result is achieved that corresponds to the original. However, the same procedure used for MP3 extraction would produce a different, lower quality disc.

The disadvantages: size and compatibility

The disadvantage is that FLAC files in megabytes are much heavier than compressing them with MP3. Although the actual efficiency depends on the sound characteristics of the respective source, an average reduction of 40-50% can be expected: For example, an hour of audio ranges from approximately 600 MB of the uncompressed format to 300 MB in the optimal case

With MP3, compression is much more intensive – the same hour of compressed audio at 160 kbps (or very high quality anyway) is expected to be around 70 MB.

Then there is the compatibility problem: MP3 is natively compatible with any Smart TV, radio, PC, smartphone or media player that is still in circulation. FLAC, on the other hand, can only be played natively on Android, Linux and Windows 10. On the other platforms, if possible, you need to download a dedicated player or convert songs in advance.


Free Download Mp4Gain
picture


Mp4Gain Main Window
picture


Mp4Gain Features
picture


Free Download Mp4Gain
picture

Flac vs Mp3, differences

In today’s world, it is important to understand the difference between the different audio files available.

The most common and current files are practically the files in the formats MP3, Flac and WAV.

Is there really a difference between MP3, Flac and WAV files? The answer is absolutely yes.

flac vs mp3

The difference is in the audio quality that these files can play.

What is the difference between MP3, Flac and Wav files? MP3 files are of lower quality because they are more compact and smaller. Flac files are a kind of compromise. With files in Flac format, it can be of very good quality and remain true to the original, which in any case is compressed by a certain percentage. Finally, there are the WAV files that do not use compression.

Therefore, the quality is better, but the size of these files is quite large, it is not compatible with any device these days.

It may not be easy to understand which audio file to use for your work, especially if you are a beginner in this area.

However, you don’t have to be afraid of it.

flac vs mp3

Once you’ve learned all the differences between the three audio files, you can really fix the problems by always using the appropriate file.

What are the differences between the categories in detail and when should you use a specific audio file?

Here’s a complete analysis for each audio file format that really helps you understand everything you need to know about MP3, Flac, and WAV files.

What is an MP3 file?
It starts with one of the most common files in the world of information technology, namely the one called MP3.

MP3 files have been around for years, so their development is common.

But what an MP3 file really represents.

Well, a general audio file is a series of numbers obtained by sampling the analog signal.

This scan responds to some parameters, which are the frequency measured in Khz and the resolution expressed in bits.

The MP3 file represents the most compressed form of an audio file, so to speak.

Finally, you need to understand that the MP3 file can remove all unnecessary parts of the digital file from the sequence and the final sampling, taking advantage of some imperfections of the human body to give it a clear and clean melody.

On the other hand, the MP3 file significantly reduces the quality of the sounds played.

In fact, all the different nuances of a certain melody come to the bone.

An MP3 is small if you speak it from the perspective of the memory. You should think of it as a kind of concentrate that gives you a remarkable but not 100% complete end result.

In the most extreme cases, an MP3 file can reduce the original tones and nuances of music or melodies to a percentage of 90%.

However, these formats are widely appreciated and used because they are not only practical and direct, but are now compatible with all technological devices, e.g. B. MP3 players, for which we recommend that you read our guide.

This means you can take them with you at any time and any product you have can read an MP3 file.

What is a flac file?
So at this point you need to understand what a Flac file is.

Well, it should be said that the Flac file has some major differences from its MP3 counterpart.

In fact, a Flac file is much more complex than a regular file and can be reduced by 90%.

First, you need to understand that Flac is actually an acronym that stands for Free Lossless Audio Codec.

It is always a format that is somewhat compressed, but files with this name have a certain property.

In fact, the file is lossless with no loss and maintains higher fidelity than the original sound.

With a Flac file, you have a clearer quality of the audio file, so you can clearly hear some details that can be lost if you use a different audio format.

The limited storage space when using the Flac format is very small and can reach a maximum of 50%.

However, when using such files, you should be aware that their use on the storage hard drive is important.

Not surprisingly, Flac files take up a lot more space than regular MP3s. In some cases, a special reader must be downloaded to read them. Many home theaters and receivers support this compression format.

MP3 – Compression criteria

MP3 – Compression criteria

To perform such compression, the MP3 format is based on a simple concept: filter a digital piece of music and eliminate all unnecessary information, thus reducing space.

mp3 compression

The human ear is an almost perfect instrument but it also has its limits. The human ear pass band extends from 20 Hz to 20,000 Hz, but is much more sensitive to those in the midrange, 700 to 6,000 Hz, where most of the information is concentrated.
The study of auditory perception is a matter of psychoacoustics that mainly analyzes 2 factors that are later used in MP3 encoding:

Mp3 – Auditory perception

In the area of ​​sounds, only a few can be heard by the human ear. The following figure shows these areas that represent the different sound frequencies. Only those in the white area are audible from our ear.

The sounds that the ear perceives are only those of the white areas

Masking

Masking is nothing more than the superposition of weak sounds with loud sounds. It almost always happens that the sounds of different instruments overlap each other. In cases where the loudest sound completely covers the lowest, there is a so-called masking. In MP3 files, masking allows you to remove the information from the weakest sounds, which, however, because they are not perceived by the ear, are virtually irrelevant.

mp3 audio masking

MP3 – The Name

The name MP3 comes from the MPEG standard, which means Moving Picture Experts Group. This group was created specifically for the development of systems and standards used in video compression. DVD movies and satellite broadcasts (DBS) use the MPEG standard to efficiently compress video information.

MPEG compression includes a subsystem for sound compression with three different compression levels (layers) depending on the quality of the information. Layer-3 is the one used for the MP3 standard, which stands for MPEG Layer-3.

MP3 – Step by step compression

The MP3 Encoder is that program that analyzes the uncompressed digital file (for example, a Wav file) and transforms it into an MP3 file.

The audio signal is filtered and divided into 576 areas (called subbands) through a process that uses DCT (Discrete Cosine Transformation) and manages to eliminate all unnecessary frequencies. The human ear, as already said, perceives sounds only beyond a certain threshold so that all the audio below is not encoded.

At this point, the resulting signal is passed through the psychoacoustic model in which the masking thresholds of which we spoke earlier are identified. This is done using Discrete Fourier Transformation (DFT).

During the masking of the 576 subbands, the frequencies to be masked are determined and therefore can be removed.

After masking, the defined Stereo Ensemble process is applied. Below a certain frequency, the ear cannot perceive the spatial position of the sounds, so they can be recorded on a single channel (therefore, in mono format) with significant space savings.

Once the file is ready, the data is re-analyzed and compressed using Hufmann encoding which enables a data reduction (without loss of information) of approximately 20%.

At this point, after all the data has been collected, the encoder proceeds to create the bit stream that will form the final MP3 file.

Mp3, description of audio compression technique

Mp3, description of audio compression technique

Digitization

Sound is a continuous wave that propagates through air or other media, formed by pressure differences, so that it can be detected by measuring the pressure level at a point. Sound waves have the proper and studyable characteristics of waves in general, such as reflection, refraction and diffraction.

To the Being a continuous wave, a digitization process is required to represent it as a series of numbers. Currently, most of the operations performed on sound signals are digital, since both storage and
Processing and transmitting the signal in digital form offers very significant advantages over analog methods. Digital technology is more advanced and offers greater possibilities, less sensitivity to transmission noise and the ability to include error protection codes, as well as encryption. With the appropriate decoding mechanisms, moreover, they can be processed simultaneously signals of different types transmitted by the same channel. The main disadvantage of the digital signal is that it requires a much greater bandwidth than that of the analog signal, hence an exhaustive study is carried out regarding data compression, some of whose techniques will be the center of our study.

Digitalization of the audio

The digitization process consists of two phases: sampling and quantization. At sampling divides the time axis into segments
discrete: the sampling frequency will be the inverse
the time between a measurement and the
following. At this time the
quantization, which, in its simplest form,
it simply consists of measuring the value of the signal
in breadth and save it.

Nyquist’s theorem

Nyquist’s theorem ensures that the frequency required to sample a signal that has its highest components at a given frequency f is at least 2f. Therefore, being the upper range of human hearing around 20 Khz, the frequency that guarantees adequate sampling for any audible sound will be around 40 Khz.
Specifically, to obtain high quality sound, frequencies of 44’1 Khz are used,
in the case of CD, for example, and up to 48 Khz, in the case of DAT. Other typical values ​​are submultiples of the first, 22 and 11 Khz.

Depending on the nature of the application, of course, the appropriate frequencies can be much lower, such that the voice process is usually performed at a frequency between 6 and
20 Khz. or even less. Regarding quantization, it is evident that the more bits used for the division of the amplitude axis, the “finer” the partition will be and therefore the less error when attributing a specific amplitude to the sound at each moment.

For example, 8 bits offer 256 levels of quantization and 16,65536. The dynamic range of human hearing is about 100 dB. The axis division can be carried out at equal intervals or according to a specific density function, seeking more resolution in certain sections if the signal in question has more components in
certain zone of intensity, as we will see in the coding techniques.

The complete process is usually called PCM (Pulse Code Modulation) and we will refer to it hereinafter. It has been described in a very simplistic way, mainly because it is widely treated and is well known, being
another the field of study of this work. However, we will go into detail at any time that is necessary for the development of the exhibition.

Coding and Compression.

Before describing coding and compression systems, we must pause in a brief analysis of human auditory perception, to understand why a significant amount of the information provided by PCM can be discarded.

The heart of the matter, as far as we are concerned, is based on a phenomenon known as masking.

The human ear perceives a frequency range between 20 Hz. And 20 Khz.

Firstly, the sensitivity is greater in the area around 2-4 Khz., So that the sound is more difficult to hear the closer to the ends of the scale.

Second is masking, the properties of which are used extensively by the most interesting algorithms: when the component at a certain frequency of a signal has high energy, the ear cannot perceive lower energy components at close frequencies, both lower and higher.

At a certain distance from the masking frequency, the effect is reduced so much that it is negligible; the range of frequencies in which the phenomenon occurs is called the critical band.

The components that belong to the same critical band influence each other and do not affect nor are affected by those that appear outside it. The width of the critical band is different according to the
frequency in which we are located and is given by certain data that shows that it is greater with frequency.

It should be noted that these data are obtained by psychoacoustic experiments, which are carried out with experts trained in
sound perception, giving rise to psychoacoustic models with their impressions.

This we have described is the so-called simultaneous or frequency masking.

There is also the so-called asynchronous or time masking, as well as other phenomena of hearing that are not relevant in this point. For now, let’s focus on the idea that certain signal frequency components support higher noise than we would generally consider to be tolerable, and therefore require fewer bits to be encoded if the encoder is endowed.
of the right algorithms to solve masks.

Digitizing the signal using PCM is the simplest form of signal encoding, and is used by both CDs and DAT systems. Like still digitizing, it adds noise to the signal, generally undesirable. As we have seen, the fewer bits used in sampling and quantization, the greater the error in
accept discrete values ​​for the continuous signal, that is, the higher the noise.

To avoid that the noise reaches an excessive level, it is necessary to use a large number of bits, so that at 44.1 Khz. and using 16 bits to quantize the signal, one of the two channels on a CD produces more than 700 kilobits per second (kbps). As we will see,
Much of this information is unnecessary and takes up bandwidth that could be freed, at the cost of increasing the complexity of the decoder system and incurring some loss of quality.

The compromise between bandwidth, complexity and quality
it is the one that produces the different market standards and will form the essential part of our study.

Mp3: What is it really?

Mp3: What is it really?

MP3 is a data format that gets its name from an algorithm
encoding called MPEG 1 Layer 3, which, in turn, is an audio compression system that allows you to store sound with a quality similar to that of a CD and with a very high compression ratio, on the order of 1:11

In practice, this means that about 11 audio CDs can be recorded on a CD-Rom, that is, approximately 150 songs.
The encoding system that MP3 uses is a loss algorithm. That is, the original sound and the one that we obtain later are not identical.

This is because MP3 takes advantage of the deficiencies of the human ear and eliminates all the information that we are not able to perceive. A multitude of studies of acoustic perception have been carried out, discovering that there are a series of effects that can aid the coding of sound with the aim of reducing as much as possible the amount of useless or redundant information. The most important are: The limits of hearing. Our ear only works with frequencies that go between 20 Hz and 20 Khz
approximately, so the remaining frequencies are disposable.

Masking effect.

It is one that occurs when two signals of similar frequency are
overlap. So we can only perceive the one that
it has more volume and, therefore, the one with a smaller volume is
liable to be removed

Stereo redundancy.

There are redundancies between the tonal and non-tonal components of the sound on the two stereo channels, and furthermore
below a certain frequency the human ear is not capable of
perceive the directionality of the sound, so below these
frequencies it is even possible to encode a single channel together with
complementary information to restore the spatial feeling for the other channel.

To carry out this “loss of information” action, a system called Subband Coding is used, a process by which the signal is broken down into subbands through a filter bank.

These subbands are then compared to the original using a psychoacoustic model that is responsible for determining which bands can be removed and which cannot.

Depending on the quality we want to obtain, more or less will be eliminated
bands. To end the process, the resulting subbands are quantized and encoded, and the final result is compressed using a standard algorithm, thus obtaining the resulting MP3 file. The encoding process is much more complicated than the decoding process, so it takes much longer to encode an MP3 file than to play it.

This perceptual coding algorithm was developed by the company MPEG (Moving Picture Expert Group) in conjunction with the Franunhofer Institute of Technology, and has been standardized as an ISO standard.

Impossible to detect the difference between MP3 and FLAC

Impossible to detect the difference between MP3 and FLAC

New comparative study carried out by experts debunks the myth that the FLAC sound (Free Lossless Audio Codec (FLAC) or Codec free lossless audio compression, is considerably better than that obtained in MP3 or AAC files.

According to Wikipedia, FLAC is a format of the Ogg project to encode audio without loss of quality; that is, the initial file can be completely recomposed, although with the disadvantage that the resulting file takes up much more space than would be obtained by applying lossy compression.

Other formats, such as MP3 or ACC (Advanced Audio Coding), irreversibly lose part of the original information when compressing the file, in exchange for a great saving in file size.

The site Trusted Reviews has published an analysis called “Sounds Good to Me”, the conclusion of which is that there is no considerable difference in sound in FLAC and MP3, at least for the average user’s auditory perception.

In their study, Trusted Reviews made the assumption that there is a difference, and that people with developed hearing abilities could hear the difference between a 192 kbps MP3 file and a FLAC file, both obtained from the same original CD.

Only one user notices the difference between Mp3 and FLAC

The previous assumption was not confirmed in the facts, since among the seven people who participated in the analysis there was only one who detected the difference between FLAC and MP3. In many cases, MP3 at a rate of 192 kpbs had a higher score than FLAC.

This result was further strengthened when comparing 320 kbps MP3 files with FLAC files, since half answered correctly, and half were wrong. The percentage of participants who preferred MP3 was even higher.

The trial used an iBasso D3 Python USB DAC and Beyer Dynamic DT770 Pro headphones.

It should be noted that studies of this type have a certain margin of error. According to experts in comparisons such as the one carried out by Trusted Reviews, there are psychological factors, such as many users quickly forgetting their perception of what they have just heard, so the order in which the test is carried out is highly relevant. Non-expert users are also influenced by their mood and, in fact, by the music they listen to.

Audio quality in different formats (flac vs. mp3)

In this post I am going to talk about what differentiates music from mp3 and flac. First, and before you begin, go ahead with the following:

The quality of a musical hearing depends (and a lot) on the audio card and the musical equipment (amplifier, headphones / speakers) used, and on the other hand it also depends on the sensitivity of one’s ear. A newborn with perfect hearing can hear from 20 Hz to 20,000 Hz. The normal thing for a young person is to hear nothing above 18 kHz, although some people with exceptional hearing can hear 20 kHz or even more. , and a person 25 years of age or older begins to lose hearing from 15,000 – 16,000 Hz. In addition to the frequency response (quantitative aspect), the qualitative aspect is equally or more important: that the waves at each frequency are produced in the most similar way to the original source.

Having said that, we fully enter the subject at hand.

Many people think that an mp3 sounds like the quality of a CD. This is not exact. Apart from the fact that a CD sounds with the quality that those who have recorded it have given it, mp3s are formats with loss, and that means that a good part of the original information is discarded to save space. The trick is that the information that is discarded is, as a rule, information that is “hidden” among the rest of the information. To give a simple example so that the idea is understood, if a person is speaking to me at a normal volume and suddenly a helicopter passes in front of us, the sound of the helicopter will eclipse in my ears the voice of that person; the wave of his voice will continue to reach my ears but I will not perceive the sound. Another example, so that I am also understood: if we could play two very similar pianos at exactly the same time in such a way that their vibrations coincided, the mp3 would “say” that “one of the pianos is left over”. This type of operation (but, of course, at a much more subtle level, of microscopic changes) is what is done so that the initial 40 or 50 MB that a song occupies on the CD are reduced, at most, to 9 MB or less, depending on the bitrate (128, 160, 192, 256, 320 kbps) of the mp3.

But all that information that the mp3 removes at a stroke is information that, from the original source, would reach us, and it is information that would affect us emotionally (an mp3 violin can hardly give us goosebumps), although consciously most of the time we do not know how to express the difference in words. The same happens, for example, when a person is recreated in virtual reality: sooner or later we will know that this person is not real, because virtual reality technology has not yet managed to recreate the microscopic details that we are capable of capturing and that make us identify a person as real and not virtual.

Other differences between an mp3 and a wav (Microsoft’s uncompressed wave file) or a flac (Free Lossless Audio Codec, free lossless audio codec) are noticeable after spending a long time listening to music. The mp3 ends up giving you a headache, while the original sound doesn’t. And to this we must add that there are certain songs that have the musical information arranged in such a way that the mp3 algorithm is not able to “guess” what it is that you are not going to be able to listen to, and the result is that there is a notable loss quality, especially in the treble. In fact, a 128 kbps mp3 cuts all frequencies starting at around 15 kHz, and this is something that most people with normal hearing can easily perceive.

So that you can hear the REAL differences that exist between the different audio formats, I have prepared several tracks in which I have done the following:

1) I have loaded the song from the original disc.

2) I have recorded it in different formats: flac, mp3 to 128, mp3 to 160, mp3 to 192, mp3 to 256 and mp3 to 320 kbps.

3) I have then loaded all the waves into the Sound Forge Pro 10.0 program.

4) I have synchronized all the waves bit by bit. This is necessary because the mp3 introduces a certain lag of milliseconds with respect to the original.

5) I have copied each mp3 wave (lossy quality) and mixed it on the flac wave (original quality, without loss) with the reverse polarity. If both waves were identical, the result would be silence. But instead, as the mp3 has less information (the wave has fewer resolution points, so that it is understood) there is a residual noise that corresponds, neither more nor less, to what the mp3 has less than the flac added to what the mp3 has more than the flac (the mp3 not only loses information; it also introduces noise that was not in the original recording).

6) I have recorded everything on flac. Contrary to what most people think, the fact of converting an mp3 to a higher quality format does not add quality, since the additional information “cannot be invented” by the mp3, and it is still absent. An mp3 transferred to CD continues to sound like an mp3.

Important note: In order to listen to the files, your player must be able to play flac. First of all, associate the files with the .flac extension to your player so that it opens them when you click on them. If they still don’t sound, then install the necessary codec or plugin.

As a player I recommend the AIMP2 or the Foobar2000; both are free and give exceptional audio quality (they reproduce the sound as it is recorded, without any attachments of any kind). For my taste, the best of the two is the Foobar2000, because it is also more stable and lightweight. The Winamp and the Windows Media Player color the sound (or in other words, they equalize it), which can be interesting if you have low-quality audio equipment and play mp3s at low bitrate (128 kbps), but, If the equipment is hi-fi and the music is encoded in a lossless format or played directly from the original CD, then the difference between Winamp or WMP and AIMP2 or Foobar2000 is quite noticeable.

How an MP3 compresses music

We all know that MP3 was the audio format that quickly became popular and the main reason is because it took up much less space than the WAV format that has no compression and therefore was very difficult to transfer via internet from one computer to another.

And then it was when the MP3 made its appearance because it had a very good sound and yet it took between 7 and 10 times less space than the original file.

We all know that this caused people to easily exchange music files online and this changed even the way the music industry works thereafter.

But although we all know that MP3 takes up less space, it is very few people who understand that in the first place in MP3 what it does is compress the music. But it also uses some other procedures to make music take up less disk space, Today we will briefly explain how this mp3 performs this compression.

Remove inaudible sounds

One of the first things MP3 does is to analyze the music file and eliminate all those frequencies that are not audible to the human ear but nevertheless occupy a space in the original file. Then the MP3 saves a lot of space without losing quality by eliminating sound frequencies that the human ear cannot hear.

Eliminate redundancy

Another of the mechanics that is used for an mp3 saves space is to eliminate redundant sounds. And with that we understand sounds that sound very similar and basically occupy the same Soundtracks. Therefore, the ear will only perceive some. And then the MP3 eliminates those redundant sounds that will not be heard by the human ear.

Sound masking

Acoustics and audio specialists have long discovered that when the human ear perceives more than one sound simultaneously it is very likely that one of them masks the others.

The Sound perception produces that when a person perceives 2 sounds of different intensity at the same time the weakest sound, with less volume, is inaudible to the one who is listening. This, as we indicated earlier, is what is called the sound perception and the MP3 is based a lot on the sound perception to be able to eliminate sounds under this principle of sound masking.

In other words, in MP3 you decide which sound will mask others and then eliminate these others.

It should be noted that when one decides if the MP3 encodes at 128 kilo bytes per second or at 320 kbs it is modifying the amount of sounds that will be eliminated in the masking. Well, at 320 to eliminate very few sounds and as I lowered the number of kbs it will eliminate more sounds which the person can produce if he can distinguish a difference between the original audio file and the encoded file.

How is an mp3 file compressed?

How is an mp3 file compressed?

The MP3 file takes up less space but loses information from the original recording, so it is a lossy compression. The question is, what is the algorithm for scrapping those details of music? How are they removed from the recording? Don’t they really matter and we don’t perceive those losses?

MP3 and auditory masking

The algorithm for MP3 compression eliminates details of the original music based on the phenomenon of the sound masking of our sense of hearing, a psychoacoustic phenomenon so daily that surely many will not have paid attention before, and that it is necessary to know to understand the MP3 .

Imagine that we are talking to someone on the street, a car passes by and suddenly we stop hearing our interlocutor. Why have we stopped hearing the other person? If we had recorded this situation with a microphone we would see that both sounds, the voice and the car, would have been perfectly recorded …

This phenomenon occurs because there are situations in which our sense of hearing gives prominence to one sound and ignores another if both are simultaneous, what is called sound masking, and that depends on well-defined causes that can be summarized as follows.

A sound can mask another when they reach the ear simultaneously depending on their relative frequencies and volumes. As seen in the figure, at the loudest sound our ear creates a new limit of hearing or masking at that time. If another simultaneous sound is under that frequency environment, we will not perceive it.

Temporary masking

When there is a sound of sufficient power to be masking, there are moments before and after that we will not perceive other sounds, depending on how closely they are in time and their relative volume, with the behavior represented in the figure. As you can see, a sound can be masked whether it occurs immediately after the masking, or if it occurs before!

The MP3 compression algorithm

When we perform an MP3 compression, the coding algorithm divides the music into a multitude of short-lived fragments. Each of these fragments are analyzed individually in many frequency bands, to be able to detect if in any of them there is any masking sound that is masking sounds of the other bands of the fragment, and therefore are inaudible or expendable. In that case, what you will do is encode that fragment with fewer bits than the original fragment, so resolution of the more subtle details (those details that have been dispensable) will be lost and the background noise of the fragment will increase.

The amount of bit reduction for that fragment will depend on the quality sought in the encoding. If we set it to high quality, it will reduce the resolution of the fragment only just enough so that the new background noise is still masked by the masking sound that was detected in that fragment.

Therefore, and according to the masking theory, no change will be perceived after the resolution reduction: neither by the loss of the details that were already originally masked, nor by the new background noise, which will remain imperceptible by also maintaining below that masking sound detected.

After this process, the fragment could have been encoded with fewer bits, occupying less information than the original. Once this attempt at bit reduction has been repeated with all the multitude of fragments into which the original file had been divided, the song is reconstructed and a compressed file is obtained that will now take up less space.

In addition to this masking-based coding, finally an “Huffman” arithmetic coding is applied to the resulting bits, similar to that performed in a “.zip” compression. This process will not entail additional quality losses.

Sound quality in MP3 files

The sound quality of the compression depends on the size that we want the compressed song to occupy, therefore the bitrate we indicate when performing the compression. If we choose a high bitrate, the algorithm will not be forced to eliminate much information, so it will eliminate really inaudible details according to the masking curves. But if we want the file to take up less space and choose a lower bitrate, the algorithm will have to be more drastic overcoming the most imperceptible masking curves, and it will be inevitable that the loss of information will be noticed.

For example, in the most common 128 kbps MP3s a few years ago, the quality is significantly lower than the original for most people, if a direct comparison is made. On the other hand, an MP3 file with the maximum bitrate of 320 kbps hardly loses information, and is practically indistinguishable from the original in most cases.