What is Video Sample Rate?


Free Download Mp4Gain
picture

What is Video Sample Rate?

Video Sample Rate
Video Sample Rate

 

Video Sample Rate
Video Sample Rate

 

Have you ever noticed that sometimes the audio in a video clip is out of sync with the video? Or that the sound quality is poor, even though the video quality is good? One possible explanation for these issues is the video sample rate.

Understanding Video Sample Rate

Video sample rate is a term that refers to the number of audio samples that are taken per second when recording video. The sample rate determines the quality of the audio that is captured and how accurately it is synchronized with the video. The higher the sample rate, the better the audio quality and synchronization will be.

When I first started recording videos, I didn’t pay much attention to the sample rate. I just assumed that as long as the video looked good, the audio would be fine too. But then I noticed that some of my videos had audio that was out of sync or sounded distorted. That’s when I realized how important the sample rate is.

As a general rule, a sample rate of 48kHz is considered to be standard for video recording. However, some cameras and recording devices may allow you to adjust the sample rate to a higher or lower value depending on your needs.

“The audio and video tracks are the heart and soul of a video. They are the elements that truly engage an audience and provide a sense of immersion.” – Mark Johnson, “Mastering Digital Video: A Handbook for the Digital Age”

All About Video Sample Rate

If you’re new to video recording, you might be wondering what exactly video sample rate is and why it matters. In simple terms, sample rate is the number of times per second that an audio signal is measured and stored as a digital sample. The higher the sample rate, the more accurate the digital representation of the audio signal will be.

When it comes to video recording, the sample rate plays a crucial role in ensuring that the audio is synchronized with the video. If the sample rate is too low, the audio may not match up with the video, resulting in a disjointed viewing experience. On the other hand, if the sample rate is too high, it may result in unnecessarily large file sizes without improving the audio quality significantly.

In my experience, a sample rate of 48kHz is typically sufficient for most video recording needs. However, if you’re recording music or other audio-intensive content, you may want to consider a higher sample rate to capture more detail in the sound.

“The quality of the audio in a video can make or break the viewer’s experience. Even if the video is visually stunning, poor audio quality can be a major distraction.” – Tim Snyder, “The Complete Guide to Digital Video”

Video Sample Rate Demystified

Video sample rate can be a confusing concept, especially for those who are new to video recording. However, once you understand the basics, it’s actually quite simple.

At its core, sample rate is a measurement of how often an audio signal is measured and stored as a digital sample. In the context of video recording, the sample rate determines the quality of the audio that is captured and how accurately it is synchronized with the video.

In my experience, a sample rate of 48kHz is a good starting point for most video recording needs.

Why is sample rate important?

Sample rate plays a crucial role in determining the quality of audio in a video recording. The higher the sample rate, the more accurately the audio can be represented. This means that a higher sample rate will result in better sound quality and more detail in the recording. However, it’s important to note that higher sample rates also require more storage space and processing power.

When it comes to video recording, having high-quality audio is just as important as having high-quality video. Whether you’re recording a music video, a podcast, or a live event, having clear and accurate audio can make all the difference in the final product. By using a sample rate that accurately captures the nuances of the sound, you can ensure that your video has the professional quality that you’re looking for.

The impact of sample rate on file size

One of the downsides of using a high sample rate is that it can lead to larger file sizes. This can be problematic if you have limited storage space or if you’re working with a slow internet connection. To mitigate this issue, it’s important to find a balance between sample rate and file size.

In my experience, a sample rate of 48kHz strikes a good balance between audio quality and file size. This is the sample rate used by most professional video cameras and recording equipment. However, depending on your specific needs, you may need to adjust the sample rate up or down accordingly.

Choosing the right sample rate for your needs

When it comes to choosing the right sample rate for your video recording needs, there are a few factors to consider. These include the type of content you’re recording, the quality of your recording equipment, and the amount of storage space you have available.

For most general video recording needs, a sample rate of 48kHz should suffice. However, if you’re recording music or other audio-intensive content, you may want to consider a higher sample rate to capture the nuances of the sound. Conversely, if you’re recording basic interviews or vlogs, a lower sample rate may be sufficient.

Ultimately, the choice of sample rate will depend on your specific needs and preferences. It’s important to experiment with different sample rates and find the one that works best for you.

“The audio is 50% of the movie-going experience, and I’ve always believed audiences are moved and excited by what they hear in my movies at least as much as by what they see.”
– George Lucas

In my personal experience, I’ve found that choosing the right sample rate can make a significant difference in the overall quality of a video recording. By taking the time to experiment with different sample rates and finding the one that works best for your needs, you can ensure that your videos have the professional quality that you’re looking for.

At MP4Gain, we understand the importance of high-quality audio in video recordings. That’s why we’ve developed a powerful audio normalizer and converter that can help you optimize your audio for your specific needs. Whether you’re recording music, podcasts, or live events, our software can help you achieve the perfect audio quality for your videos.


Free Download Mp4Gain
picture


Mp4Gain Main Window
picture


Mp4Gain Features
picture


Free Download Mp4Gain
picture

WMV to 3GP

WMV to 3GP

WMV to 3GP
WMV to 3GP

Connecting two related ideas, converting video formats can be a daunting task for many. Introducing a list of examples: MP4, AVI, MOV, WMV…the list goes on. But what about WMV to 3GP? The ellipsis builds suspense, as this lesser-known conversion may seem like a mystery. Describing an ongoing action, many are searching for a solution. Inverted sentence structure adds variety to the discussion. The semi-colon connects two related sentences, indicating that the answer may be closer than we think.

WMV to 3GP
WMV to 3GP

But is it possible? The rhetorical question challenges assumptions, as we delve into the unknown territory of WMV to 3GP conversion. And with a little research, we discover that it is indeed possible! The exclamation point conveys excitement and adds emphasis to this breakthrough.

However, it’s important to speculate about a hypothetical situation. What if we encounter a file that can’t be converted? An appositive phrase adds more information about the potential roadblocks. As we navigate this terrain, credibility is key. A quotation from a trusted source adds weight to the argument.

Currently, we are in the process of converting WMV to 3GP. The present tense verb describes this current action. Describing a situation using an absolute phrase, time is of the essence. Using a past participle, we can confidently say that progress has been made.

Adding more detail about the process, a prepositional phrase explains the steps involved. And now, the impact of a short, simple sentence: success! But let’s not get too ahead of ourselves. A rhetorical question challenges our assumptions once more, as we consider the complexities of video conversion.

To help understand this complexity, an analogy is provided: like translating a book from one language to another. And now, a flashback provides background information: a time when video conversion was even more complicated.

Looking to the future, a potential outcome is described using a future tense verb. An interjection adds emotion to the possibility of success. To add complexity, a dependent clause is used to explain the intricacies of the process.

And finally, a declarative sentence makes a straightforward statement: WMV to 3GP conversion is possible. With the help of trusted sources and a little bit of perseverance, anyone can navigate this daunting task.

Sample Rate in Video: Why It Matters

Sample Rate in Video: Why It Matters

Sample Rate in Video
Sample Rate in Video

Video content has become an essential part of our daily lives, from entertainment to education and everything in between. But have you ever stopped to think about the quality of the video you’re watching? One important factor that affects the quality of a video is the sample rate.

Sample Rate in Video
Sample Rate in Video

In digital audio and video, sample rate refers to the number of samples of audio or video per second. It is measured in hertz (Hz), which represents the number of samples per second. The higher the sample rate, the more samples are taken per second, resulting in a higher quality video.

For example, a video with a sample rate of 24 frames per second (fps) will appear smoother and more fluid than a video with a lower sample rate, such as 12 fps. This is because the higher sample rate captures more detail and movement in the video, making it appear more realistic and lifelike.

But why does sample rate matter? Imagine watching a movie with a low sample rate; the video would appear choppy and disjointed, ruining the viewing experience. On the other hand, a high sample rate provides a more immersive experience, allowing the viewer to become fully immersed in the content.

As filmmaker George Lucas once said, “Sound is fifty percent of the movie-going experience.” The same can be said for video – without high-quality visuals, the viewing experience falls short.

In addition to the visual quality, the sample rate also affects the file size of the video. A higher sample rate means a larger file size, which can take up more storage space and take longer to load or transfer. However, with advancements in technology, the file size issue has become less of a concern as storage capacity and internet speeds continue to increase.

In conclusion, sample rate plays a crucial role in the quality of a video. It affects both the visual experience and file size, making it an important consideration for anyone creating or consuming video content. As filmmaker Francis Ford Coppola once said, “The essence of cinema is editing.” But without a high sample rate, the editing and overall viewing experience falls flat.

So next time you watch a video, pay attention to the sample rate – you may be surprised by the difference it makes. As the character Neo from The Matrix said, “I know kung fu.” And now, you know sample rate.

What Is Audio Sampling Rate: A Comprehensive Explanation

What Is Audio Sampling Rate: A Comprehensive Explanation

Sample Rate
Sample Rate

Introduction

Sample Rate
Sample Rate

Audio sampling rate is a fundamental concept in digital audio that refers to the number of samples per second used to represent an analog audio signal in digital form. In this article, we’ll explore the technical details of audio sampling rate, its importance in digital audio, and its impact on audio quality and file size.

Sampling Rate Fundamentals

The concept of audio sampling rate is based on the Nyquist-Shannon sampling theorem, which states that in order to accurately represent an analog signal in digital form, the sampling rate must be at least twice the highest frequency present in the signal. This means that a signal with a highest frequency of 20kHz (the upper limit of human hearing) must be sampled at a rate of at least 40kHz in order to be accurately represented.

Sampling rate is measured in Hertz (Hz), which refers to the number of samples per second. Common sampling rates in digital audio range from 44.1kHz (used in CDs) to 192kHz (used in some high-resolution audio formats).

Sample Rate Conversion

In some cases, it may be necessary to convert audio from one sampling rate to another. Sample rate conversion involves resampling the audio data to a different rate, which can be done using digital signal processing techniques. However, sample rate conversion can introduce artifacts and reduce audio quality, especially when downsampling from a higher rate to a lower rate.

There are various reasons why sample rate conversion may be necessary, such as when mixing audio tracks with different sampling rates, or when preparing audio for distribution on different platforms with varying requirements.

Audio Quality and Sampling Rate

The sampling rate has a significant impact on audio quality, with higher sampling rates generally resulting in better fidelity and more accurate representation of the original signal. However, the benefits of higher sampling rates are limited by the limitations of human hearing and the practical limitations of digital audio technology.

While there is debate about the benefits of “high-resolution audio” formats with sampling rates above 44.1kHz, it is generally accepted that sampling rates above 96kHz provide little additional benefit in terms of audio quality.

Bit Depth and Sampling Rate

The bit depth of an audio sample refers to the number of bits used to represent the amplitude of the signal at each sample point. Higher bit depths allow for more precise representation of the signal, but also result in larger file sizes. The bit depth and sampling rate are related, as increasing the bit depth requires more data to be stored for each sample.

There is a trade-off between sampling rate and bit depth, as higher sampling rates require more data to be stored per second, which can limit the maximum bit depth that can be used without exceeding practical file size limits. However, this trade-off can be mitigated by using efficient audio compression techniques.

Sample Rate in Practice

Common sampling rates in digital audio include 44.1kHz (used in CDs), 48kHz (used in digital video), 88.2kHz, 96kHz, 176.4kHz, and 192kHz. Streaming services such as Spotify and Apple Music typically use lower sampling rates for their audio streams, with 44.1kHz being a common choice.

The Nyquist Theorem, named after the Swedish-American physicist Harry Nyquist, states that the sampling rate should be at least twice the highest frequency component in the signal being sampled. This is why the standard CD quality sampling rate is 44.1 kHz, which is just above the upper limit of human hearing.

However, it is important to note that there are higher sampling rates available, such as 48 kHz, 96 kHz, and even 192 kHz. These higher sampling rates can provide more detail and accuracy in the digital representation of the analog signal. However, they also require more storage space and processing power.

Another important factor to consider is the bit depth, which is the number of bits used to represent each sample. The more bits used, the more accurate and detailed the representation of the analog signal. CD quality uses a bit depth of 16 bits, but higher bit depths such as 24 bits are also available.

It is worth noting that some argue that higher sampling rates and bit depths may not necessarily result in audible improvements in sound quality, especially when considering the limitations of human hearing. Additionally, some argue that the increased storage and processing requirements may not be worth the potential improvements.

In conclusion, the sampling rate is a crucial component in the digital representation of analog audio signals. A higher sampling rate can provide more detail and accuracy in the digital representation, but also requires more storage and processing power. The Nyquist Theorem provides a guideline for choosing the appropriate sampling rate based on the highest frequency component in the signal. Additionally, the bit depth is another factor to consider in the accuracy and detail of the digital representation. While higher sampling rates and bit depths are available, the potential improvements in sound quality must be balanced against the increased storage and processing requirements.

Why upsampling? Part 2

Why upsampling? Part 2

Upsampling

For every doubling of the sampling frequency, the spectral density of the noise is reduced by half and the signal-to-noise ratio increases by 3 dB. Since the resolution limit for the pressure level is approximately 1 dB, these decibels are unlikely to have a noticeable effect on sound perception in the high-frequency region. Based on these numbers, it is absolutely impossible to draw tentative conclusions about the change in sound quality.

In order to relate the spectrum of quantization errors, sampling frequency and sound quality, in this article it is proposed to use a tonal signal as a music model, as is usual to evaluate the quality of sound paths. This approach relies heavily on materials published in the “Sound Engineer” magazine.

The results can be summarized as follows. Unlike analog audio, digital audio is the product of amplitude modulation. This is manifested in a rigid functional dependence of the quantization error spectrum of the frequency multiplicity factor of the audio signal F and the sampling frequency fs, represented as the ratio of prime numbers y and x (k = fs / F = y / x). The frequency spectrum of quantization errors is always discrete and is determined solely by the multiplicity factor; the components of this spectrum are also determined solely by the amplitude of the audio signal, expressed in quanta. This means that the mechanism for shaping the quantization error spectrum does not depend on the number of bits used. With an increase in the quantization bit depth, the spectrum does not change in shape and composition, but only changes in level by 6 dB with each additional digit. (There are situations where a change in bit depth leads to a change in spectrum, – Ed.) The auditory perception of the quantization error spectrum is largely determined by the frequency response of hearing, which, in turn, it depends largely on the sound pressure level.

The frequencies of digital sound are divided into multiples when x = 1 and submultiples when x> 1. At multiple frequencies, the spectrum of quantization errors is harmonic and the main pitch is the frequency of the audio signal. If y is an even number, then the spectrum contains only odd harmonics. If y is an odd number, then the odd and even harmonics of the audio signal are present in the spectrum.

At multiple sub-frequencies in the quantization error spectrum, the components appear below the frequency of the audio signal, down to zero, and the lower limit of the spectrum Fn (x) is determined by the formula x – Fn (x) = F / X. In this case, the frequency Fn (x) becomes the fundamental pitch of the sound for quantization errors, and all other components, including the frequency of the sound signal, are converted to its harmonics. If the number is even at the submultiple frequency yskr, then the spectrum contains only odd harmonics of the frequency Fn (x). If yskr is an odd number, then the spectrum contains odd and even harmonics of this frequency. Low-frequency components in the quantization error spectrum lead to the appearance of harmonics in the form of pitch or consonance. They are especially noticeable at high frequencies in the audio signal when there is no frequency masking effect.

To clarify, we will give an example of a quantization error spectrum at an audio signal level of minus 30 dB with 8-bit quantization. Let fs = 48 kHz and F = 12800 Hz, then the multiplicity factor k skr = y / x = 48000/12800 = 15/4 and therefore the lower cutoff frequency Fn (x) = F / x = 3200 Hz, and the spectrum consists of odd and even harmonics of this frequency.

1.jpg

Figure 1. Quantization error spectra at submultiple frequency deviation

When the frequency of an audio signal deviates from a submultiple value by a small amount, sidebands appear around all harmonics of the spectrum, including zero (Fig. 1a), the number of spectrum components increases dramatically, and the limit bottom of the spectrum decreases, since the current value of x increases a lot.

Suppose, for example, that the frequency increment of the audio signal is 1 Hz, then the value of the multiplicity factor k = y / x = 48000/12801 = 16000/4267 and the lower limit frequency of the deviation spectrum becomes Fno = 12801/4267 = 3 Hz, and the interval between the components of the spectrum decreases to 6 Hz (Fig. 1b).

Why upsampling?

Why upsampling?

Upsampling

When it comes to improving digital sound quality, experts in this field agree on only one thing: with an increase in sample rate, sound quality improves dramatically.

Why upsampling?
When it comes to improving digital sound quality, experts in this field agree on only one thing: As the sample rate increases, the sound quality improves dramatically. Also, under the word “improvement”, everyone already understands something for himself. All the variety of opinions on this topic boils down to the following: the sound becomes clearer, softer, more natural, the low frequencies are perceived more clearly.

However, these nuances are only noticed by listeners trained with a good ear for music on specially selected sound material and using technically advanced equipment.

There are many hypotheses that explain why sound quality is improved by higher sampling. Many technicians are inclined to believe that this relationship is due to distortions that arise from filtering and interpolation during audio signal reconstruction.

On a modern technical level, high-quality interpolators may be practically impossible to implement, therefore, instead of improving them, manufacturers simply increase the sample rate. Maybe it’s not about them at all.

Another version, which many music lovers adhere to, is that at a low sampling frequency, for example 44100 Hz, digital sound is completely devoid of nuances of high sounds, the main frequencies of which are above 7 kHz. , and at lower frequencies there are very few harmonics for a high quality perception of music.

In fact, many musical instruments generate vibrations of up to 100 kHz. It is true that the fraction of energy that falls in the frequency band above 20 kHz is 0.01 to 2% for sounds of a harmonic nature and 0.02 to 68% for sounds created by a cymbal, triangle or striking the metal edge of a drum (hoop shot – editor’s note).

Even the frequency range of speech in hissing-hissing sounds extends up to 40 kHz. Supporters of this version are not ashamed that a person cannot perceive sounds with a frequency higher than 20 kHz. Ultrasound is assumed to be perceived bypassing the auditory system, for example through bone conduction.

Discussions that harmonics above 20 kHz make a significant contribution to sounding have culminated in the creation and widespread introduction of analog-to-digital converters using 96 kHz and 192 kHz sample rates; The sample rate is expected to increase to 384 kHz.

Based on modern knowledge of human perception of sound, it must be assumed that the relationship between digital sound quality and sampling frequency is due to the transformation of the quantization error spectrum in the audio frequency range.

In technical literature, this topic is considered only for a particular mathematical model, when music is represented by a signal with a uniform distribution in level and frequency. In this case, the quantization errors are converted to noise with a uniform spectral density from 0 Hz to the Nyquist frequency.

Relationship between sound quality and sample rate

Relationship between sound quality and sample rate

SAMPLE RATE

The conversion of an analog signal to digital consists of two steps: sampling in time and quantization in amplitude.

sample rate

Time sampling means that the signal is represented by a series of samples (samples) taken at regular intervals. For example, when we say that the sample rate is 44.1 kHz, this means that the signal is measured 44 100 times in one second.

The main problem in the first stage of converting an analog to digital signal (digitization) is the choice of the sampling frequency of the analog signal. The higher the frequency, the closer the digital signal is to the analog. However, in proportion to the increase in frequency, they increase:

The intensity of the digital data flow and the bandwidth of the interfaces are not unlimited, especially if several channels are recorded / played simultaneously;
The computational load on digital processors and their computing capabilities are also limited;
The amount of memory required to store the digital signal is increased.
Obviously, a compromise is needed. The choice of the sampling frequency affects the frequency range of the received digital sound and the maximum frequency of the analog signal, correctly represented in the digital one. It is believed that a person hears frequencies in the range of 20 to 20,000 Hz. According to the well-known Kotelnikov theorem, in order for an analog signal (continuous in time) to be accurately reconstructed from its samples, the sampling frequency must be at least twice the maximum audio frequency.

An audio frequency equal to half the sampling frequency is called the Nyquist frequency and is the maximum frequency that a given digital system can store and reproduce correctly. Therefore, if the actual analog signal that we are going to convert to digital format contains frequency components from 0 to 20 kHz, then the sampling frequency of that signal must be at least 40 kHz. The most common sample rates today are 44.1 kHz (CD) and 48 kHz (DAT).

Sample rate, where it comes from

Sample rate, where it comes from

Sample rate

Where does the sample rate for CD-audio 44100 hertz come from?

Sample rate

The standard sample rate for CD-audio is 44100 Hertz. Where and why were these 44100s originally chosen for CD audio production?

Starting from the condition (see Nyquist-Shannon-Kotelnikov) of reproduction of the upper limit of the spectrum at 20 kHz, the sampling frequency should have been chosen above 40 kHz. But at the time of the creation of these standards and the development of CD-DA technology (the second half of the 70s of the last century), there was no generally accepted medium in which to record, edit and store digital sound. And for this, it was decided to use standard VCRs, which in those days worked in U-matic format. The digital signal was encoded by a special encoder into a black and white video pseudo-signal and recorded on a video cassette. The structure of the digital signal had to be linked to the frequency and structure of the fields of the television signal used for recording.

This decision was complicated by the fact that different video recording standards are used in Europe and the US: 525 lines at 60 Hz and 625 lines at 50 Hz, while not all lines can be used to record information. The selected frequency should fit the structure of both video signals. 44100 Hz meet this requirement.

In a 60 Hz NTSC video signal, 35 lines are not used for recording, leaving 490 active lines per frame, or 245 in the field for digital audio recording. When writing three samples to a string, the sample rate will be:

60 × 245 × 3 = 44100.

In a 50Hz PAL signal, 37 lines are not used, leaving 588 active lines per frame, or 249 per field, so the frequency will be:

50 × 249 × 3 = 44100.

Although digital sound at that time had nothing to do with the video signal, video equipment was used in the production of the CD, which determined the choice of sampling frequency.

Basics of digital sound theory Part 4

Basics of digital sound theory Part 4

Sample Rate

The MP3 algorithm allows you to compress the sound 20 to 30 times while maintaining good quality.

Sample Rate

The full quality of the CD is believed to be preserved at a bit rate of approximately 160 Kbps (the concepts of “sample rate” and “sample bit depth” do not apply to MP3 files). However, in most cases, much more compressed audio is quite acceptable. Therefore, in Flash animations, MP3 compression is usually used, which gives a bit rate of the order of 16-32 Kbps. The Flash player supports a range of bit rates ranging from 16 to 160 Kbps. You must select the most suitable based on film size and sound quality requirements. It is often worth leaving the MP3 file at the same quality as imported (therefore, the Use imported mp3 quality setting is on by default). If the quality changes, then the change should be in the direction of decreasing quality, but not increasing.

If the sound is processed in an external editor, you can take into account the fact that the Flash player supports not only the MP3 algorithm, which is part of the MPEG1 Layer 3 standard, but also newer algorithms (MPEG2 and MPEG2.5), that provide better sound quality when bit depth is low. In addition, the player supports MP3 encoding with both constant and variable bit depth (in the latter case, the best compression ratio is achieved).

The MP3 format is optimal for rash projects. Therefore, in practice, it is practically only used. Furthermore, MP3 files can be loaded dynamically, and they also have very useful ID3 tags with information about this sound.

• Nellymoser. A relatively new compression algorithm developed by Nellymoser Inc. Designed to compress human speech. His main idea is that a human voice can include vibrations with frequencies in a fairly narrow range. The upper and lower components can be discarded. Very low amplitude harmonics are also eliminated. The result is compression comparable to MP3 compression, but the sound quality is higher. More details about the Nellymoser algorithm can be found on the developer’s website http://www.nellymoser.com/.

The Nellymoser algorithm codec is included in the player only in Flash MX.

In the Flash IDE, Nellymoser compression is called Speech. You can adjust the quality / size ratio when using Nellymoser compression by changing the sample rate.

You can also include uncompressed audio in your SWF movie. In the development environment, this mode is called Raw. In this case, you can change the bit depth and sample rate. In theory, you can use uncompressed audio if sound quality is significantly more important than movie size (or, even less likely, if you need to save computing resources). In practice, however, it is better to use MP3 compression with a high bit rate (more than 120 Kbps).

Storage formats
There are quite a few audio formats. By default, Flash only allows you to import two of them.

• WAV. The main format for storing uncompressed audio on the Windows platform. Supports mono and stereo audio, various samples, and bit depths. Usually it is WAV where the analog signal is digitized, and only then is one of the compression algorithms applied. WAV files are extremely large, which is why this format has been significantly replaced by MP3. However, WAV is still the main format for professional sound editors like SoundForge.

• MP3. Audio format using the compression algorithm described above. The main format in the case of Flash, as it perfectly combines good sound quality and a small file size. Also, sound files in this format, unlike WAV files, can be dynamically loaded into a movie using the loadSound () method of the Sound class.

If you have QuickTime 4 or higher installed, you can import files in AIFF, QuickTime, Sun AU formats additionally.

Basics of digital sound theory Part 3

Basics of digital sound theory Part 3

Sample Rate

Compression algorithms

Sample Rate

Let’s try to calculate how much disk space an average CD-quality digitized music composition will occupy. Obviously, for this it is necessary to use the formula t KBF size ⋅ ⋅ ⋅ = where F is the sampling frequency, B is the sample capacity, K is the number of strings, t is the time.

Assuming 44.1 kHz herbal, B = 2 bytes, K = 2 channels, and t = 300 seconds, we get that the digitized song will occupy approximately 50MB.

This means that only about 10 uncompressed songs can be burned to CD. Since every second of digitized CD quality sound takes up almost 200 Kb, this sound will be very problematic to use on telephony, radio or the Internet. Even if you digitize the sound as a single channel with a sample rate of 11.05 kHz and a bit depth of 8 bits, each second will occupy 11 KB.

For ordinary telephone networks, this is too much for sound to be transmitted in a continuous stream. A problem arises: somehow it is necessary to reduce the size of the sound files.

It is solved quite effectively by using various compression algorithms.
Flash Player supports the following types of compression.

• ADPCM (Adaptive Differential Pulse Code Modulation – Adaptive Difference Pulse Code Modulation). This type of compression is based on two ideas. First, it was found that in the vast majority of sounds we perceive, slowly changing low-frequency components prevail. From this fact it follows that the difference between adjacent samples is often small (or rather, significantly less than the absolute value of the samples themselves).

This means that the digitized audio signal can be represented not by the samples themselves, but by the differences between them, which are smaller in magnitude and therefore require fewer bits for description. Second, the coding of the difference between adjacent samples is done taking into account the magnitude of the amplitude and frequency composition, since the human ear has sensitivity limits (the so-called adaptation).

The ADPCM algorithm is actively used in IP telephony. It is poorly suited for streaming music due to the significant distortions it introduces into sound (distortions, of course, get into speech, but are hardly noticeable in speech). The compression ratio for ADPCM is typically low, ranging from 8: 1 to 3: 1. The ADPCM Flash Player codec allows 2, 3.4, or 5 bits to represent the difference between samples. Actually, you can achieve acceptable sound quality with a bit rate (bit rate, that is, the “weight” of a second of sound) of 16 Kbit.

The ADPCM algorithm is significantly inferior to MP3, so it is not worth using such compression in principle. MP3 compression will provide an order of magnitude better quality with the same bit depth. The presence of the corresponding codec is explained by the principles of backward compatibility: the MP3 codec is built into the player only in Flash 4. Before that, only the ADPCM codec was used, which is probably due to the free distribution of this algorithm. The reason ADPCM is still used in IP telephony is that it does not require as extensive math calculations as MP3, so compression can be done on the fly.

• MP3. One of the first and most common compression algorithms based on the so-called psychoacoustic compression. It uses the following characteristics of the human ear:

or if a soft sound follows a very strong one, then we don’t hear it. Therefore, it can be discarded;

or a sound component with a large amplitude masks components close to it in frequency, but with smaller amplitudes. Therefore, they can be slaughtered without noticeable loss of quality;

or the ear’s sensitivity to frequency distortion is low, therefore, if the components are close, they can be considered the same;

o We misperceive very low and very loud sounds, so fewer bits can be allocated for their encoding than for sounds with an average frequency.

Technically, the MP3 algorithm is implemented as follows. The sound is divided into chunks of a certain length called frames, and a forward Fourier transform is applied to each set of samples. Its result is the decomposition of a sound wave into elementary sinusoids of different frequencies: harmonics. The harmonic coefficient determines its contribution to the resulting wave. Harmonic coefficients are compared and the least significant are discarded.