The extra dimension a spectrogram adds
A waveform shows only two dimensions: time and amplitude. A spectrogram shows three: time on the horizontal axis, frequency on the vertical axis, and amplitude (loudness) at each frequency represented by color or brightness intensity — this is exactly the information a waveform structurally cannot display, since a waveform collapses all frequency content into a single amplitude value at each moment.
This makes a spectrogram the tool for questions a waveform simply can't answer: what pitch range is present, whether there's an unwanted hum at a specific frequency, or where the energy of a sound is actually concentrated across the frequency spectrum.
Spotting hums and unwanted tones by their fixed frequency
Electrical hum (from power lines or poorly shielded equipment) typically sits at a very specific, constant frequency, which shows up in a spectrogram as a thin, perfectly horizontal line running the entire length of the recording — a signature that's immediately obvious visually but can be much harder to isolate by ear alone, especially when it's masked by other sound. Identifying the exact frequency this way is also the first step toward removing it with a targeted filter.
Reading vocal range and musical content visually
Speech and singing show up as bands of energy that rise and fall in frequency corresponding to pitch changes, with harmonics (related frequencies above the fundamental note) appearing as fainter parallel bands above the main pitch line — musicians and audio engineers use this to visually confirm pitch, identify the frequency range a particular voice or instrument occupies, and spot problem frequencies that might clash with other elements in a mix.
When you actually need this over a waveform
For basic editing tasks like cutting, trimming, and joining, a waveform is usually all you need. A spectrogram becomes valuable specifically when you're diagnosing an unwanted tone or hum, analyzing the frequency content of a recording for mixing or mastering purposes, or trying to understand why a sound has a particular quality (harsh, muddy, thin) that a simple amplitude view can't explain.