Psychoacoustic Threshold Estimation in MP3


Free Download Mp4Gain
picture

Psychoacoustic Threshold Estimation in MP3

Psychoacoustic Threshold Estimation in MP3

Let’s talk about Psychoacoustic Threshold Estimation in MP3

Psychoacoustic threshold estimation in MP3 encoding is a crucial element for efficient compression. In my experience, this process plays a significant role in how audio is perceived by listeners after compression. It’s based on the principles of psychoacoustics, which examine how humans perceive sound. Essentially, psychoacoustic models allow MP3 encoding to remove parts of the audio that are inaudible to the human ear, making the file size smaller without compromising perceived quality. To understand it better, think of how you might ignore background noise when focusing on a conversation in a crowded room. Similarly, MP3 compression removes sounds that would not be heard by a listener under normal conditions.

In MP3 encoding, threshold estimation is done by analyzing the signal’s frequency spectrum. The human ear is more sensitive to certain frequencies and less sensitive to others. By determining which parts of the audio are inaudible based on these sensitivities, MP3 compression algorithms can selectively remove these frequencies. The result is a compressed file that maintains the most important parts of the sound while discarding unnecessary details.

The Role of Psychoacoustics in MP3 Compression

When discussing MP3 compression, psychoacoustics comes into play to ensure the best balance between sound quality and file size. It’s as though I’m packing a suitcase for a trip—choosing the essentials and leaving behind the non-essentials. In MP3 encoding, psychoacoustic models aim to identify which audio frequencies are masked by others, allowing them to be discarded without a noticeable loss in quality.

These psychoacoustic models use data about human hearing perception. For instance, our ears are more sensitive to mid-range frequencies than to low or high frequencies. When encoding an MP3, the algorithm uses this knowledge to reduce the representation of low and high frequencies, especially if they are masked by louder sounds in the mid-range. This approach reduces the file size, making it more efficient while maintaining an acceptable sound quality.

Psychoacoustic Models: Key Techniques for Estimation

Psychoacoustic models are essential for estimating thresholds in MP3 encoding. The two main models used in MP3 compression are the MPEG-1 Layer III and the more complex MPEG-2 Layer III. These models implement specific techniques to determine which parts of the audio signal can be discarded without affecting the perceived quality.

  • Critical Bands: The human ear perceives sounds in frequency groups called critical bands. Each critical band includes frequencies that are close enough together that they affect each other’s perception. When encoding, psychoacoustic models assess these bands and eliminate those that won’t affect the listener’s experience.
  • Masking Effect: This is a phenomenon where a louder sound makes it difficult to hear a quieter sound. The MP3 encoder uses this principle to discard sounds masked by others, reducing the file size.
  • Threshold of Hearing: The threshold of hearing refers to the quietest sound that the average human ear can detect. Sounds below this threshold are effectively inaudible and can be removed during encoding.

Practical Example: How Psychoacoustic Threshold Estimation Works

Imagine you’re listening to your favorite song on your smartphone. The song is compressed into an MP3 file, but somehow it still sounds amazing. What’s happening behind the scenes is the psychoacoustic threshold estimation. For example, if you’re listening to a powerful guitar solo, the MP3 algorithm may eliminate some of the higher frequencies from the background sounds like drums or cymbals that are masked by the louder guitar notes.

From my experience, it’s much like watching a movie with a powerful soundtrack. When the action is intense, the quieter background sounds fade into the background. The MP3 encoder mimics this behavior, focusing on what’s essential to the listener’s perception of the music and discarding less important details. It’s a brilliant way to optimize audio files while preserving the listening experience.

The Benefits of Psychoacoustic Threshold Estimation in MP3

The main benefit of psychoacoustic threshold estimation is the reduction in file size. The more efficient the compression, the smaller the file size, which makes it easier to store and stream audio. This is particularly crucial in a world where bandwidth is often limited, and storage space can be at a premium.

Another benefit is the preservation of sound quality. As an audio professional, I’ve found that effective psychoacoustic modeling ensures that what’s important to the listener remains intact. The algorithm removes what isn’t necessary, but it does so without compromising the overall experience. For example, it’s as if you’re cleaning up a painting by removing minor smudges that no one would notice anyway. The final image (or audio) still looks great but is lighter.

Latest Words on Psychoacoustic Threshold Estimation in MP3

Psychoacoustic threshold estimation is an essential process for MP3 compression. It ensures that audio files are as small as possible while maintaining the best possible quality. From my expertise, understanding psychoacoustics is key to understanding how modern audio compression works. These methods allow for the efficient storage of high-quality sound without sacrificing too much bandwidth or space.

At the end of the day, MP3 encoding wouldn’t be nearly as efficient or effective without psychoacoustic threshold estimation. It’s a fascinating blend of human perception and technology that allows us to enjoy high-quality audio in a convenient format. In cases where precise audio management is critical, using specialized software can further enhance the quality of the compressed file, and Mp4Gain offers a reliable option in this area.

What is psychoacoustic threshold estimation in MP3 encoding?

Psychoacoustic threshold estimation in MP3 encoding is the process of determining which parts of an audio signal are inaudible to the human ear and can be discarded to reduce file size without affecting perceived sound quality.

How does psychoacoustic modeling affect MP3 compression?

Psychoacoustic modeling reduces MP3 file sizes by removing audio frequencies that are masked by louder sounds, ensuring only the most essential elements of the sound are preserved for optimal listening quality.

What is the masking effect in psychoacoustics?

The masking effect is when louder sounds make it difficult to hear quieter ones. MP3 encoders exploit this effect to remove inaudible sounds, making the file more efficient without sacrificing quality.

Why are some frequencies removed in MP3 compression?

Some frequencies are removed in MP3 compression because they are outside the human ear’s sensitivity range or are masked by louder sounds, making them unnecessary for a high-quality listening experience.

How do critical bands influence MP3 encoding?

Critical bands are frequency ranges that the human ear perceives as a group. MP3 encoders use this information to determine which sounds in a frequency band are crucial and which can be discarded without affecting quality.

What are the benefits of psychoacoustic threshold estimation for MP3 files?

The main benefit of psychoacoustic threshold estimation is reduced file size while maintaining sound quality. This is particularly important for efficient storage and streaming of audio files.

How does psychoacoustic modeling enhance listening experience?

Psychoacoustic modeling enhances the listening experience by focusing on the most important frequencies and discarding unnecessary ones, resulting in a clear, high-quality sound that doesn’t take up much storage space.

What is the threshold of hearing in psychoacoustics?

The threshold of hearing refers to the faintest sound that can be perceived by the average human ear. Sounds below this threshold are removed during MP3 encoding because they are inaudible.

How does psychoacoustic threshold estimation improve MP3 file size efficiency?

Psychoacoustic threshold estimation improves MP3 file size efficiency by removing audio frequencies that would go unnoticed by the listener, making the file smaller without sacrificing quality.

Comments:

I’ve always been amazed by how much smaller MP3 files are compared to other formats. This article really breaks down why that is so clearly! The psychoacoustic principles are fascinating.

– AudioFan99

Really interesting read! I never realized that so much of the sound is actually removed when encoding an MP3. This helps explain why high-quality audio formats like FLAC sound so much better.

– MusicLover123

I had no idea that psychoacoustic models played such a big role in MP3 quality. I wonder how much it varies across different types of audio, like classical versus rock music.

– CuriousJoe

Great explanation! Would love to know more about how these models evolve over time and how they’ve impacted newer audio formats.

– SoundGeek2024

I’ve been looking for a deeper dive into how MP3 compression works, and this article really filled in the gaps. So cool to see the science behind it!

– TechieGuy

 


Free Download Mp4Gain
picture


Mp4Gain Main Window
picture


Mp4Gain Features
picture


Free Download Mp4Gain
picture

Implementing CBR in MP3 Compression

Implementing CBR in MP3 Compression

Implementing CBR in MP3 Compression

Implementing CBR in MP3 Compression
Implementing CBR in MP3 Compression

Let’s talk about Implementing CBR in MP3 Compression

As a specialist in audio compression technologies, I’m excited to delve into the intricacies of implementing Constant Bit Rate (CBR) in MP3 compression. CBR is a crucial aspect of MP3 encoding, ensuring consistent audio quality across all parts of the file. Understanding how CBR works and its implications for audio quality is essential for anyone involved in audio production, from musicians to sound engineers.

The Basics of CBR Encoding

Unlocking the Mystery of Constant Bit Rate:
CBR encoding maintains a steady bit rate throughout the entire duration of the audio file. Unlike Variable Bit Rate (VBR) encoding, which adjusts the bit rate based on the complexity of the audio, CBR allocates the same number of bits per second regardless of the content. This uniformity simplifies streaming and playback, as devices can predict the data rate required for decoding.

Ensuring Consistency in Audio Quality:
One of the primary advantages of CBR encoding is its ability to deliver consistent audio quality. By allocating a fixed bit rate, CBR ensures that each segment of the audio receives the same level of compression. This consistency is especially important for streaming services and broadcasting, where fluctuations in audio quality can be jarring for listeners.

Implementing CBR in MP3 Compression

CBR in MP3 Encoding:
In the realm of MP3 compression, CBR is a popular choice for its simplicity and predictability. When encoding audio to the MP3 format, CBR allocates a constant number of bits per second to represent the audio signal. This ensures that the resulting MP3 file maintains a consistent bit rate from start to finish, regardless of the complexity of the audio content.

Benefits of CBR in MP3 Compression:
CBR encoding offers several advantages in the context of MP3 compression. Firstly, it simplifies the encoding process by removing the need for complex algorithms to adjust the bit rate dynamically. This results in faster encoding times and reduced computational overhead. Additionally, CBR-encoded MP3 files are more compatible with legacy playback devices and systems that may not support VBR decoding.

Challenges and Considerations

Trade-offs in Compression Efficiency:
While CBR encoding ensures consistent audio quality, it may not always achieve the same level of compression efficiency as VBR encoding. In scenarios where the audio content is highly dynamic or contains significant variations in complexity, CBR may allocate more bits than necessary for simpler segments, resulting in larger file sizes.

Adapting to Varied Content:
Another challenge of CBR encoding is its limited ability to adapt to changes in audio complexity. In contrast to VBR encoding, which adjusts the bit rate dynamically based on the content, CBR maintains a fixed rate regardless of fluctuations in complexity. This can lead to suboptimal compression in segments with low complexity or conversely, potential artifacts in segments with high complexity.

Latest Words on Implementing CBR in MP3 Compression

In conclusion, understanding the role of Constant Bit Rate (CBR) in MP3 compression is essential for optimizing audio quality and file size. While CBR offers consistency and simplicity, it’s important to weigh the trade-offs in compression efficiency and adaptability. By implementing CBR effectively, audio professionals can ensure a seamless listening experience across various platforms and devices.

Comments:

This article provided valuable insights into the intricacies of CBR encoding in MP3 compression. As a music producer, I appreciate the clarity and depth of explanation.

– BeatMaster

While I found this article informative, I wish it had delved deeper into the specific techniques used to implement CBR in MP3 encoding. Nonetheless, it’s a great starting point for anyone interested in the topic.

– AudioEnthusiast

As an aspiring sound engineer, I found this article incredibly helpful in understanding the fundamentals of CBR encoding. The examples provided made the concepts easy to grasp.

– SoundSavvy

I appreciate the focus on both the benefits and challenges of implementing CBR in MP3 compression. It’s essential to consider the trade-offs in audio quality and file size when choosing an encoding method.

– MusicTechie

This article shed light on a topic I’ve always been curious about. Understanding CBR encoding is crucial for anyone involved in audio production, and this article provided a comprehensive overview.

– AudioExplorer

Quantization in MP3

Quantization in MP3: Balancing Compression and Quality

Quantization in MP3
Quantization in MP3
Quantization in MP3
Quantization in MP3

Let’s Talk About MP3 Quantization

Quantization in MP3
Quantization in MP3

Having spent years immersed in the realm of audio encoding, I’m here to shed light on the intricate dance between compression and quality in MP3 quantization. Google’s top results merely scratch the surface, so let’s dive deep into the world of digital audio encoding and unravel the nuances of MP3 quantization, blending my expertise with relatable real-life examples.

The Essence of MP3 Quantization

MP3 quantization, a vital aspect of audio compression, resembles a delicate balancing act. Imagine it as a chef crafting a recipe; too much compression, and you lose the flavor (quality), too little, and the dish (file size) becomes overwhelming. In this section, we’ll explore the core principles of MP3 quantization, demystifying the magic behind achieving optimal audio quality while keeping file sizes in check.

  • Bits and Bytes: Understanding the Basics
  • Quantization Levels: Fine-Tuning Audio Precision
  • Trade-offs: Balancing Quality and File Size

Bits and Bytes: Understanding the Basics

At the heart of MP3 quantization lies the concept of bits and bytes. Think of them as the canvas for a painting. The more bits we have, the finer the details and richer the colors. This foundational understanding is crucial as we navigate the landscape of audio compression and strive for a harmonious blend of quality and efficiency.

Quantization Levels: Fine-Tuning Audio Precision

Quantization levels are akin to a painter’s palette, each level representing a shade of sound. As an expert, I’ll guide you through the art of selecting the right quantization levels, ensuring that the nuances of the audio are preserved. This nuanced approach sets the stage for a symphony of digital audio that captivates the listener.

Trade-offs: Balancing Quality and File Size

In the realm of MP3 quantization, there’s a perpetual trade-off between quality and file size. It’s akin to walking a tightrope, finding the sweet spot where audio fidelity remains high, yet the file remains manageable. I’ll share insights into striking this delicate balance, drawing parallels with everyday scenarios to make it relatable and easy to grasp.

Latest Words on MP3 Quantization

As we navigate the complexities of MP3 quantization, I’ll provide fresh perspectives that go beyond the standard discourse. For instance, the impact of psychoacoustics on quantization decisions is often overlooked. Understanding how our brains perceive sound allows us to tailor the quantization process to optimize for perceived quality, offering a unique angle that distinguishes this article.

Going Beyond the Basics

While many articles skim the surface, I’ll take you on a journey into advanced territories. Exploring topics like variable bit rate (VBR) encoding and the role of advanced psychoacoustic models, we’ll unveil the sophisticated mechanisms that contribute to superior audio quality in MP3 files. This knowledge empowers you to make informed decisions in your digital audio endeavors.

Quantization Myths Unveiled

Let’s debunk common misconceptions surrounding MP3 quantization. For example, the notion that higher bit rates always equate to better quality is not absolute. I’ll demystify these myths, providing clarity and guiding you towards a nuanced understanding of the factors influencing audio quality in MP3 encoding.

Optimizing MP3 Files for Different Platforms

Not all platforms are created equal, and neither should your MP3 files be. I’ll share strategies for optimizing MP3s tailored to specific platforms. Whether you’re creating content for streaming services, podcasts, or mobile applications, understanding platform-specific nuances in quantization and compression will set you on the path to audio excellence.

Let’s Talk Real-Life Applications

Bringing it all together, I’ll delve into real-life applications of MP3 quantization. From enhancing your music library to optimizing podcast episodes for diverse audiences, I’ll share personal experiences and practical tips. Imagine fine-tuning your audio files like a skilled craftsman, ensuring they shine across various playback scenarios.

Comments:

This article opened my eyes to the intricacies of MP3 compression. More articles like this, please!

– AudioExplorer

Great breakdown! However, I’d love a deeper dive into VBR encoding techniques.

– TechAudioGeek

Finally, someone addressing the myths! Clear, concise, and enlightening.

– MythBusterListener

Can you share your thoughts on MP3 quantization for podcasters? Looking for practical advice.

– PodcasterPro

As a musician, I appreciate the analogies! Helped me grasp the technicalities effortlessly.

– MusicalSoul

This article left me craving more insights into optimizing MP3s for streaming platforms.

– StreamMaster

Thanks for the myths clarification! I’ve been misguided for so long.

– TruthSeeker

Could you explore the environmental impact of different quantization strategies? Curious to know!

– EcoListener

Kudos for making a complex topic so accessible. Looking forward to more insights!

– ClarityEnthusiast

Great article, but I wish there was more focus on mobile app optimization for music.

– MobileMusicBuff

Personal anecdotes made it so relatable. Excited to apply these principles to my projects!

– ProjectCreator

The Science of Audio Encoding: Technical Aspects

The Science of Audio Encoding: Technical Aspects

The Science of Audio Encoding
The Science of Audio Encoding
The Science of Audio Encoding
The Science of Audio Encoding

Audio encoding is the process of converting analog sound into digital data. This data can then be stored or transmitted in a variety of formats, such as WAV, MP3, or AAC.

There are two main types of audio encoding: lossless and lossy. Lossless encoding preserves all of the original sound data, resulting in high-quality audio but large file sizes. Lossy encoding removes some of the original sound data, resulting in smaller file sizes but lower sound quality.

The process of audio encoding can be divided into three main steps: sampling, quantization, and compression.

Sampling

The first step in audio encoding is sampling. In this step, the analog sound signal is converted into a series of discrete values. The number of times per second that the sound signal is sampled is called the sample rate. Higher sample rates result in more accurate representations of the original sound signal, but they also result in larger file sizes.

Quantization

The second step in audio encoding is quantization. In this step, each sample value is rounded to the nearest integer value. The number of bits used to represent each sample value is called the bit depth. Higher bit depths result in more accurate representations of the original sound signal, but they also result in larger file sizes.

Compression

The third and final step in audio encoding is compression. In this step, the digital audio data is compressed to reduce its file size. There are a number of different compression algorithms that can be used, each with its own advantages and disadvantages.

The most common compression algorithms for audio encoding are:

  • MP3: MP3 is a lossy compression algorithm that is widely used for storing and transferring audio files. MP3 files are typically much smaller than WAV files, while still providing good sound quality.
  • AAC: AAC is another lossy compression algorithm that offers better sound quality than MP3. AAC files are typically slightly larger than MP3 files, but they offer a noticeable improvement in sound quality.
  • FLAC: FLAC is a lossless compression algorithm that offers similar sound quality to WAV, but with much smaller file sizes. FLAC files are a good choice for people who want the best possible sound quality without sacrificing file size.

Final Words

Audio encoding is a complex process that involves converting analog sound into digital data. The quality of the audio that is encoded can be affected by a number of factors, including the sample rate, bit depth, and compression of the audio file.

If you are looking for the best possible sound quality, you should use a lossless audio format such as WAV or FLAC. However, if you need to store or transfer audio files over a network, you should use a lossy audio format such as MP3 or AAC.

MP3 finally goes into the public domain

MP3 finally goes into the public domain

mp3

Open Source

Mp3 Public Domain

Perhaps many did not think so, but the mp3 standard so well known to all had problems with the purity of patents. On April 23, 2017, the last patents expired and the format was finally free. Technicolor has officially stopped collecting royalties from manufacturers of software and embedded solutions.

License

Although hardware mp3 decoding is built into all other coffee machines, until recently its use in commercial projects required royalties from the developer: Fraunhofer Society. In 2005 alone, the amount paid was one hundred million euros. Most of the patents became invalid in the European Union in 2012. However, some of them continued to operate in the United States due to peculiarities of local law. What does this news bring to the community? At least now it will be possible to compile Gentoo and listen to music at the same time immediately on the base distribution. Many distributions will be able to provide support for the standard to the main repository. Now, for example, Ubuntu itself requires the installation of non-free components from a separate Ubuntu Restricted Extras meta-package to support mp3.

Bourbon vanilla vs vanillin

How does this standard, which has been the main standard in this area for 24 years, despite many more advanced free options? mp3 is in many ways similar in principle to its cousin in the photo world: JPEG. Due to the imperfection of our hearing aid and the peculiarities of psychoacoustics, it is possible to “discard” those parts of the audio spectrum that do not make a significant contribution to the musical pattern. In particular, in the illustration above, you can see how the amount of information encoded in the high-frequency region increases.

High frequencies are often sacrificed for the sake of preserving detail in the lower region – vocals, most instruments (thanks for the comment, KorDen32). Standard values ​​of cutoff frequencies for the lame encoder:

CBR 096 kbps: 14000 – 15000 Hz;
CBR 112 kbps: 15000-15600 Hz;
CBR 128 kbps: 16000 – 16500 Hz;
CBR 160 kbps: 16500-17500 Hz;
CBR 192 kbps: 18000-18700 Hz;
CBR 224 kbps: 19000-19400 Hz;
CBR 256 kbps: 19500-19700 Hz;
CBR 320 kbps: 20,000 – 21,000 Hz.

The method can be compared to the creativity of flavor chemists. You’ve probably noticed that strawberry gum is very conventionally strawberry, and there isn’t enough lemon in synthetic lemon tea. Any natural flavoring composition contains dozens and even hundreds of chemical compounds. But the main core generally creates only a very limited amount. So, for example, vanillin defines most of the aroma of natural vanilla, and if you don’t appreciate the subtle nuances too much, the remaining components can be neglected. mp3 uses the same principles, removing insignificant portions of the spectrum. Most people cannot tell the lossless formats by ear from the normally encoded 320kbps mp3s, which saves a lot of space when storing your media library.

Audio Coding: Secrets Revealed Part 2

Audio Coding: Secrets Revealed Part 2

Bit Depth

Bit depth

audio encoding

Along with the sample rate, there is the bit depth or depth of the sound. Bit depth is the number of bits of digital information to encode each sample. Simply put, the bit depth determines the “accuracy” of the input signal measurement. The larger the digit capacity, the smaller the error will be for each individual conversion from the magnitude of an electrical signal to a number and vice versa. With the smallest possible bit depth, there are only two options for measuring sound accuracy: 0 for full silence and 1 for full sound. If the bit width is 8 (16), then by measuring the input signal, 2 8 = 256 (2 16 = 65,536) different values ​​can be obtained.

Bit depth is fixed in the PCM codec, but for codecs that assume compression (eg MP3 and AAC), this parameter is calculated during encoding and may vary from sample to sample.

Bitrate
Bit rate is an indicator of the amount of information that one second of sound encodes. The higher it is, the less distortion and the closer the encoded composition is to the original. For linear PCM, the bit rate is very easy to calculate.

bitrate = sample rate × bit depth × channels

For systems like the Epiphan Pearl Mini that encode 16-bit (16-bit) linear PCM, this calculation can be used to determine how much additional bandwidth the PCM audio might require. For example, for stereo (two channels), the signal is digitized at 44.1 kHz at 16 bits and the bit rate is calculated as follows:

44.1 kHz × 16 bit × 2 = 1411.2 kbps

Meanwhile, audio compression algorithms like AAC and MP3 have fewer bits to transmit the signal (that’s their purpose), so they use low bit rates. Typically, the values ​​are in the range of 96 kbps to 320 kbps. For these codecs, the higher the bit rate you choose, the more audio bits you get per sample and the better the sound quality.

Sample rate, bit depth and bit rates in real life.
Audio CDs, one of the most popular early inventions for the general public for storing digital audio, used 44.1 kHz (20 Hz – 20 kHz, human ear range) and 16 bits. These values ​​were chosen to be able to save as much audio as possible to disk with good sound quality.

When video was added to audio and DVD and then Blu-ray discs came along, a new standard was created. DVD and Blu-Ray recordings typically use 48 kHz (stereo) or 96 kHz (5.1 surround) linear PCM format and 24-bit depth. These settings have been selected as ideal for keeping audio in sync with video while obtaining the best possible quality using the additional available disk space.

Our recommendations
CDs, DVDs, and Blu-Ray discs all have one goal: to provide the consumer with a high-quality playback engine. The goal of all developments was to provide high-quality audio and video without worrying about file size (if only it could fit on disk). Such quality could be provided by linear PCM.

In contrast, mobile media and streaming media have a completely different goal: to use the lowest bit rate, as low as possible, while still being sufficient to maintain acceptable quality for the listener. Compression algorithms are best suited for this task. You can follow the same principles for your records.

When recording audio from a video …
In case the record is used for the next on-ra-ki-bot, choose the 48 kHz PCM codec and the maximum bit depth (16 or 24) to provide the best audio quality. We recommend these parameters for Epiphan Pearl Mini.

When streaming audio from video …
With streaming or recording for later translation, good sound can be obtained with less bandwidth, using MP3 or AAC codecs with a frequency of 44.1 kHz and a bit rate of 128 kbit / s or higher. These parameters ensure that the sound is good enough without affecting the quality of the transmission.

Audio encoding: secrets revealed

Audio encoding: secrets revealed

Audio Encoding

Audio settings for video capture and transmission.

audio and video encoding

As people directly related to the AV sphere, we constantly talk about audio coding and audio codecs, but what is it? An audio codec is essentially a device or algorithm that can encode and decode a digital audio signal.

In practice, the audio waves that travel through the air are continuous analog signals. The signals are converted to digital form by a device called an analog-to-digital converter (ADC), and the reverse converter is called a digital-to-analog converter (DAC). The codec lies between these two functions and it is he who allows you to adjust some important parameters for the successful capture, recording and transmission of an audio signal: the codec algorithm, the sampling frequency, the bit width and the speed of the audio signal. data.

The three most popular audio codecs are Pulse-Code Modulation (PCM), MP3, and Advanced Audio Coding (AAC). The choice of codec determines the compression rate and the recording quality. PCM is a codec used by computers, CDs, digital phones, and sometimes SACD. The PCM signal source is sampled at regular intervals, and each sample is the digital amplitude of the analog signal. PCM is the simplest option for digitizing an analog signal.

With the correct parameters, this digitized signal can be completely converted back to analog without any loss. But this codec, which provides an almost complete identity with the original audio, is unfortunately not very cheap, which translates into very large file sizes, and such files are not suitable for streaming. We recommend using PCM to record digital images for your sources or when doing audio post-processing.

Fortunately, we always have the option of choosing a different codec that can compress digital data (rather than PCM) based on some helpful observations on the behavior of sound waves. But in this case, you have to make a compromise: all alternative algorithms are associated with “losses”, since it is impossible to completely restore the original signal, but nevertheless the result is still so good that most users will not be able to to catch the difference.

MP3 is an audio encoding format that uses a digital data compression algorithm that allows you to save the audio signal in smaller files. The MP3 codec is the most used by users to record and store music files. We recommend using MP3 to stream audio content as it requires less network bandwidth.

AAC is a newer audio encoding algorithm that is the successor to MP3. AAC has become the standard for MPEG-2 and MPEG-4 formats. In fact, this is also a digital data compression codec, but with less quality loss than MP3 when encoded with the same bit rate. We recommend using this codec for online streaming.

Sampling frequency (kHz, kHz)
Sample rate (or sample rate): the frequency with which the signal is digitized, stored, processed, or converted from analog to digital. Time sampling means that the signal is represented by several of its samples (samples) taken at regular intervals.

Measured in hertz (Hz, Hz) or kilohertz (kHz, kHz,) 1 kHz equals 1000 Hz. For example, 44,100 samples per second can be labeled 44,100 Hz or 44.1 kHz. The selected sample rate will determine the maximum playback frequency and, as follows from Kotelnikov’s theorem, to fully restore the original signal, the sample rate must be twice the highest frequency in the signal spectrum.

As you know, the human ear is capable of picking up frequencies between 20 Hz and 20 kHz. Given these parameters and the values ​​shown in the table below, you can understand why 44.1 kHz was chosen as the sampling frequency for CD and is still considered a very good frequency for recording.

There are several reasons for choosing a higher sample rate, although it may seem like a waste of time and effort to reproduce sound outside the range of human hearing. At the same time, 44.1 – 48 kHz will suffice for the average listener for a high-quality solution to most problems.

Audio encoding and processing

Audio encoding and processing

Encoding

Sound information.

ENCODING

Sound is a wave that travels through air, water, or other medium with a continuously changing intensity and frequency.

A person perceives sound waves (air vibrations) with the help of hearing in the form of sound of different volume and pitch. The higher the intensity of the sound wave, the louder the sound, the higher the frequency of the wave, the higher the pitch of the sound

The human ear perceives sound at a frequency of 20 vibrations per second (low sound) to 20,000 vibrations per second (high sound).

A person can perceive sound in a wide range of intensities, in which the maximum intensity is 10 14 times greater than the minimum (one hundred thousand billion times). A special unit “decibel” (dbl) is used to measure the volume of sound (Table 5.1). Decreasing or increasing the volume of the sound by 10 dB corresponds to a decrease or increase in the intensity of the sound by 10 times.

Table 5.1. Sound volume
Sonar Volume in decibels
Lower limit of human ear sensitivity 0
Whisper of Leaves 10
Conversation 60
Horn 90
Jet engine 120
Pain threshold 140
Sound time sampling. For a computer to process sound, a continuous audio signal must be converted to a discrete digital form using time sampling. A continuous sound wave is divided into separate small time sections, for each section a certain value of sound intensity is set.

Therefore, the continuous dependence of the loudness of the sound at time A (t) is replaced by a discrete sequence of loudness levels.

Sampling frequency.

A microphone connected to the sound card is used to record analog sound and convert it to digital format. The quality of the digital sound obtained depends on the number of measurements of the sound volume level per unit time, that is, the sampling frequency. The more measurements that are made in 1 second (the higher the sampling frequency), the more accurately the “ladder” of the digital audio signal repeats the curve of the dialogue signal.

The audio sample rate is the number of measurements of the volume of a sound in one second.

The audio sample rate can range from 8000 to 48000 sound volume measurements per second.

Audio encoding depth. Each “step” is assigned a specific value for the volume level of the sound. Loudness levels of sound can be viewed as a set of possible states N, for which a certain amount of information is needed to encode, which is called audio encoding depth.

Audio encoding depth is the amount of information required to encode the discrete volume levels of digital audio.

If the known encoding depth, the number of digital audio volume levels can be calculated using the formula N = 2 I. Let the sound encoding depth be 16 bit, then the number of sound volume levels is:

N = 2 I = 2 16 = 65 536.

During the encoding process, each sound volume level is assigned its own 16-bit binary code, the smallest sound level will correspond to the code 0000000000000000 and the highest, 1111111111111111.

The quality of digitized sound. The higher the sound sampling frequency and depth, the better the digitized sound will sound. The lowest quality of digitized sound, corresponding to the quality of telephone communication, is obtained at a sampling rate of 8000 times per second, a sampling rate of 8 bits, and by recording an audio track (“mono” mode). The highest quality digitized audio, corresponding to the quality of an audio CD, is achieved with a sampling rate of 48,000 times per second, a sampling rate of 16 bits, and the recording of two audio tracks (“stereo” mode ).

It should be remembered that the higher the quality of the digital sound, the greater the volume of information in the audio file. It is possible to estimate the information volume of a digital stereo sound file with a duration of 1 second with an average sound quality (16 bits, 24,000 measurements per second). To do this, the encoding depth must be multiplied by the number of measurements in 1 second and multiplied by 2 (stereo sound):

16 bits × 24,000 × 2 = 768,000 bits = 96,000 bytes = 93.75 KB.

Audio Coding: Secrets Revealed – Part 2

Audio Coding: Secrets Revealed – Part 2

AUDIO ENCODING

Audio settings for video capture and transmission.

AUDIO ENCODING

Sampling frequency (kHz, kHz)
Sample rate (or sample rate): the frequency with which the signal is digitized, stored, processed, or converted from analog to digital. Time sampling means that the signal is represented by several of its samples (samples) taken at regular intervals.

Measured in hertz (Hz, Hz) or kilohertz (kHz, kHz,) 1 kHz equals 1000 Hz. For example, 44,100 samples per second can be labeled 44,100 Hz or 44.1 kHz. The selected sample rate will determine the maximum playback frequency and, as follows from Kotelnikov’s theorem, to fully restore the original signal, the sample rate must be twice the highest frequency in the signal spectrum.

As you know, the human ear is capable of picking up frequencies between 20 Hz and 20 kHz. Given these parameters and the values ​​shown in the table below, you can understand why 44.1 kHz was chosen as the sampling frequency for CD and is still considered a very good frequency for recording.

There are several reasons for choosing a higher sample rate, although it may seem like a waste of time and effort to reproduce sound outside the range of the human ear. At the same time, 44.1 – 48 kHz will suffice for the average listener for a high-quality solution to most problems.

Bit depth
Along with the sample rate, there is the bit depth or depth of sound. Bit depth is the number of bits of digital information to encode each sample. Simply put, the bit depth determines the “accuracy” of the input signal measurement. The larger the digit capacity, the smaller the error for each individual conversion from the magnitude of an electrical signal to a number and vice versa. With the smallest possible bit depth, there are only two options for measuring sound accuracy: 0 for full silence and 1 for full sound. If the bit width is 8 (16), then by measuring the input signal, 2 8 = 256 (2 16 = 65,536) different values ​​can be obtained.

Bit depth is fixed in the PCM codec, but for codecs that assume compression (eg MP3 and AAC), this parameter is calculated during encoding and may vary from sample to sample.

Bitrate
Bit rate is an indicator of the amount of information that one second of sound encodes. The higher it is, the less distortion and the closer the encoded composition is to the original. For linear PCM, the bit rate is very easy to calculate.

bitrate = sample rate × bit depth × channels

For systems such as the Epiphan Pearl Mini that encode 16-bit (16-bit) linear PCM, this calculation can be used to determine how much additional bandwidth the PCM audio might require. For example, for stereo (two channels), the signal is digitized at 44.1 kHz at 16 bits and the bit rate is calculated as follows:

44.1 kHz × 16 bit × 2 = 1411.2 kbps

Meanwhile, audio compression algorithms like AAC and MP3 have fewer bits to transmit the signal (that’s their purpose), so they use low bit rates. Typically, the values ​​are in the range of 96 kbps to 320 kbps. For these codecs, the higher the bit rate you choose, the more audio bits you get per sample and the better the sound quality.

Sample rate, bit depth and bit rates in real life.
Audio CDs, one of the most popular early inventions for the general public for storing digital audio, used 44.1 kHz (20 Hz – 20 kHz, human ear range) and 16 bits. These values ​​were chosen to be able to save as much audio as possible to disk with good sound quality.

When video was added to audio and DVD and then Blu-ray discs came along, a new standard was created. DVD and Blu-Ray recordings typically use 48 kHz (stereo) or 96 kHz (5.1 surround) linear PCM format and 24-bit depth. These settings have been chosen as ideal for keeping the audio in sync with the video while obtaining the best possible quality using additional available disk space.

Our recommendations
CDs, DVDs, and Blu-Ray discs all have one goal: to provide the consumer with a high-quality playback engine. The goal of all developments was to provide high-quality audio and video without worrying about file size (if only it could fit on disk). Such quality could be provided by linear PCM.

By contrast, mobile media and streaming media have a completely different goal: to use the lowest bit rate possible, while still being sufficient to maintain acceptable quality for the listener.

Audio encoding: secrets revealed

Audio encoding: secrets revealed

audio encoding

Audio settings for video capture and transmission.

AUDIO ENCODING

As people directly related to the AV sphere, we constantly talk about audio coding and audio codecs, but what is it? An audio codec is essentially a device or algorithm that can encode and decode a digital audio signal.

In practice, the audio waves that travel through the air are continuous analog signals. The signals are converted to digital form by a device called an analog-to-digital converter (ADC), and the reverse converter is called a digital-to-analog converter (DAC). The codec lies between these two functions and it is he who allows you to adjust some important parameters for the successful capture, recording and transmission of an audio signal: the codec algorithm, the sampling frequency, the bit width and the speed of the audio signal. data.

The three most popular audio codecs are Pulse-Code Modulation (PCM), MP3, and Advanced Audio Coding (AAC). The choice of codec determines the compression rate and the recording quality. PCM is a codec used by computers, CDs, digital phones, and sometimes SACD. The PCM signal source is sampled at regular intervals, and each sample is the digital amplitude of the analog signal. PCM is the simplest option for digitizing an analog signal.

With the correct parameters, this digitized signal can be completely converted back to analog without any loss. But this codec, which provides an almost complete identity with the original audio, is unfortunately not very cheap, which results in very large file sizes, and such files are not suitable for streaming. We recommend using PCM to record digital images for your sources or when doing audio post-processing.

Fortunately, we always have the option of choosing a different codec that can compress digital data (rather than PCM) based on some helpful observations on the behavior of sound waves. But in this case, you have to make a compromise: all alternative algorithms are associated with “losses”, since it is impossible to completely restore the original signal, but nevertheless the result is still so good that most users will not be able to to catch the difference.

MP3 is an audio encoding format that uses a digital data compression algorithm that allows you to save the audio signal in smaller files. The MP3 codec is the most used by users to record and store music files. We recommend using MP3 to stream audio content as it requires less network bandwidth.

AAC is a newer audio encoding algorithm that is the successor to MP3. AAC has become the standard for the MPEG-2 and MPEG-4 formats. In fact, this is also a digital data compression codec, but with less quality loss than MP3 when encoded with the same bit rate. We recommend using this codec for online streaming.