Date: 2026-08-27
This experiment compared four models that can analyze audio directly—GPT Audio 1.5, Gemini 2.5 Flash, Gemini 3.1 Pro Preview, and Gemini 3.7 Flash—plus Whisper 1 as a transcription-only baseline. One speaker recorded 29 short samples covering literal and extended transcription, emotional tone, English and Spanish pronunciation, code-switching, context-sensitive names, human non-speech sounds, pauses, self-correction, contrastive emphasis, and an ordinary voice message.
The clearest result is that audio understanding is not the same capability as transcription. GPT Audio and Whisper produced the lowest intended-script word error rate on the initial corpus, but Gemini 3.1 Pro and Gemini 3.7 Flash were substantially better at locating contrastive stress, preserving meaningful hesitations, and describing non-speech sounds. The expanded Spanish tests were harder and more mixed: the sounded h was robustly detected by Gemini 3.1 and 3.7,