← BlogBlog · STT

Groq Whisper 502 service_unavailable on long audio chunks: it's the timestamps

Short answer: when you split long audio into OGG chunks with ffmpeg -f segment, every chunk except the first carries a non-zero start timestamp, and on those chunks we saw Groq hang for ~120 seconds and return 502 service_unavailable. The fix is one line: add -reset_timestamps 1 to the split command.

Published By Stanislav Shupilkin4 min read

The staircase: where Groq stops answering

60-second chunks, 0.16 MB each, cut with ffmpeg -f segmentwithout resetting timestamps. “Offset” is the time a chunk claims to start at (start_time in ffprobe).

Chunk start offsetAPI responseError bodyTime to response
0 s200—0.62–1.1 s
60 s200—0.58 s
120 s502service_unavailable120.12 s
180 s502service_unavailable120.12 s
240 s502service_unavailable120.12 s
300 s502service_unavailable120.12 s
480 s502service_unavailable120.12 s

The 0 s row is chunk zero from the same kind of split at 300 seconds; the 60-second staircase didn't log it separately. We kept the error code, not the full response body.

Before
# before: chunks 1..N keep a running offset → 502
ffmpeg -i input.opus -vn -c:a libopus -b:a 24k -ac 1 -ar 16000 \
  -f segment -segment_time 900 -segment_format ogg chunk_%04d.ogg
After
# after: every chunk starts at zero → 200
ffmpeg -i input.opus -vn -c:a libopus -b:a 24k -ac 1 -ar 16000 \
  -f segment -segment_time 900 -segment_format ogg \
  -reset_timestamps 1 chunk_%04d.ogg

How we found it

Our Telegram bot @smolevich_voice_bot transcribes voice messages and audio files through Groq, which serves the open model whisper-large-v3-turbo. Long recordings have to be chunked anyway: Groq caps files at 25 MB on the free tier.

The failure looked the same every time: chunk one works, everything after it dies. A daily probe logged 1 good chunk out of 6 almost every day for two months, and we honestly filed it as “Groq can't do long audio”. During the investigation we ruled out retries, a different IP, a different model and a fresh HTTP connection, one by one.

The test that settled it was boring. Take the same stretch of audio and cut it two ways. Segment #1 from -f segment (offset 300 s): 502 after 120 seconds. The same stretch via -ss 300 -t 300(offset 0, same size, 0.80 MB): 200 in 0.74 seconds. Same audio bytes, only the timestamp differs. So “the first chunk always passes” and “the first chunk always starts at zero” were the same observation all along.

How to reproduce

  1. Split any recording longer than 3 minutes with the “before” command, using -segment_time 60.
  2. Check offsets: ffprobe -v error -show_entries format=start_time -of csv=p=0 chunk_0002.ogg prints roughly 120.
  3. Send chunk_0001.ogg and chunk_0002.ogg to POST https://api.groq.com/openai/v1/audio/transcriptions with model whisper-large-v3-turbo. For us the first returned 200, the second 502.
  4. Re-split with -reset_timestamps 1: every chunk's start_time drops to about 0.

The bot now splits exactly like the “after” block: 900-second chunks, OGG/Opus 24 kbps, mono 16 kHz, -reset_timestamps 1. Two months of blaming the wrong component, basically

Questions

Is the Groq Whisper 502 caused by file size?

Not in our tests. 0.16 MB chunks with an offset failed just like 2.4 MB ones, and a 9.16 MB WAV went through. The threshold was the offset, somewhere between 60 and 120 seconds; we didn't narrow it further.

Will retries fix it?

No. The same offset chunk failed five times: different processes, different IPs, whisper-large-v3 and turbo, minutes apart. Each retry just adds another 2-minute wait.

Can I skip -reset_timestamps?

Splitting into MP3 or WAV without the flag worked for us. The cost is size: MP3 at 32 kbps came out 1.14 MB per 5 minutes versus 0.80 MB for Opus, WAV was 9.16 MB. The flag is simpler. It also fixes chunk durations in ffprobe: without it, ffprobe reports the chunk's end time in the source instead of its length, and timestamps in the merged transcript drift.

How we tested
  • Date: August 4, 2026.
  • Model: whisper-large-v3-turbo via https://api.groq.com/openai/v1/audio/transcriptions; whisper-large-v3 was tried only on one offset chunk (also 502).
  • Format: OGG/Opus 24 kbps, mono, 16 kHz, split with ffmpeg -f segment.
  • Main run: a 2 h 45 min recording (9,897 s), 900 s chunks with -reset_timestamps 1: 11 of 11 returned 200, zero errors, 17.5 s of API time in total. All 10 seams checked by hand: no duplicates, no gaps.
  • Staircase: 60 s chunks at offsets 60/120/180/240/300/480 s.
  • Requests used: 67 for the whole investigation.
  • ffmpeg docs: the segment muxer's reset_timestamps option: “Reset timestamps at the beginning of each segment, so that each segment will start with near-zero timestamps.”
  • Caveat:tested on one long recording, on that date. We don't know why Groq reacts to the offset; this is only what we observed. The API may behave differently now.