To find the size of uncompressed sampled sound, multiply sample rate × bit depth × duration in seconds × channels, then divide by 8 for bytes. It appears whenever a question gives a recording length and quality settings and asks for the storage needed.
It follows the same idea as image size: count the items, multiply by the bits for each item, then convert.
How does a computer store sound?
Sound is a continuous wave. To store it, the computer measures the wave height many times each second. Each measurement is a sample, and the number of samples per second is the sample rate in hertz (Hz).
Each sample is stored as a binary number. The number of bits used is the sample resolution or bit depth. A higher sample rate and a higher bit depth both capture the wave more accurately, and both increase file size.
How do you calculate it, step by step?
- Convert the time to seconds (minutes × 60).
- Multiply the number of samples: sample rate × seconds.
- Multiply by the bit depth to get bits per channel.
- Multiply by the number of channels (mono is 1, stereo is 2).
- Divide by 8 for bytes, then convert units using the stated convention.
The method as Cambridge-style pseudocode:
INPUT Rate, Depth, Seconds, Channels
Bits ← Rate * Depth * Seconds * Channels
Bytes ← Bits / 8
OUTPUT Bytes
Trace with Rate = 44100, Depth = 16, Seconds = 10, Channels = 1: Bits = 44 100 × 16 × 10 × 1 = 7 056 000, so Bytes = 882 000.
Worked example
A voice note lasts 30 seconds. It is recorded mono at 8 000 Hz with an 8-bit sample resolution. Give the uncompressed size in bytes and in KiB (1 KiB = 1 024 bytes).
Step 1, samples: 8 000 × 30 = 240 000 samples.
Step 2, bits: 240 000 × 8 = 1 920 000 bits.
Step 3, bytes: 1 920 000 ÷ 8 = 240 000 bytes.
Step 4, KiB: 240 000 ÷ 1 024 = 234.375 KiB.
Check: each sample is 8 bits, which is exactly 1 byte, so bytes equal samples: 240 000. This matches.
Assumption: no compression, no metadata, one channel.
The mistake to watch for
A common slip is to leave the duration in minutes.
Mistaken answer: A 2-minute clip at 44 100 Hz, 16-bit, mono. Bits = 44 100 × 16 × 2 = 1 411 200.
The student multiplied by 2 minutes, but the sample rate counts samples per second.
The correction: 2 minutes = 120 seconds.
Bits = 44 100 × 16 × 120 = 84 672 000 bits = 10 584 000 bytes. The wrong answer was 60 times too small. Writing the unit beside each number as you go catches this.
Check yourself
Show every step, then open each answer.
1. A 1-minute mono recording is made at 22 050 Hz with 16-bit samples. How many bytes does it need?
Show answer
Seconds = 60. Bits = 22 050 × 16 × 60 = 21 168 000. Bytes = 21 168 000 ÷ 8 = 2 646 000 bytes.
2. A 5-second stereo clip is recorded at 48 000 Hz with 24-bit samples. How many bytes does it need?
Show answer
Bits = 48 000 × 24 × 5 × 2 = 11 520 000. Bytes = 11 520 000 ÷ 8 = 1 440 000 bytes.
3. A recording is 100 KiB at a sample rate of 40 000 Hz. The same recording is made again at 20 000 Hz with everything else unchanged. What is the new size?
Show answer
Size is proportional to sample rate, and the rate is halved, so the size is halved: 50 KiB.
Where this leads next
Large raw files are why compression exists, so move on to lossless and lossy compression. You can also check your arithmetic by writing the formula as code in the Python reasoning sandbox, then test everything in the practice set.
If you know the formula but still drop a factor under exam pressure, a teacher can watch your working and spot which quantity you tend to skip in online one-to-one Computer Science tuition.