Audio0.36 credits / minOpenAI

Whisper Large v3

Speech to text

OpenAI's Whisper in 99 languages, or translated to English.

Model guide
# Token from `npx summer-engine login` (valid 30 days)
export SUMMER_TOKEN="$(cat ~/.summer/auth-token)"

curl -X POST https://www.summerengine.com/api/mcp/generate/audio \
  -H "Authorization: Bearer $SUMMER_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"capability":"speech_to_text","modelId":"whisper-large-v3","audioBuffer":"<the audio file, base64 encoded>","idempotencyKey":"a-new-uuid-per-run"}'

Input

Speech to text reads a recording. Upload or record one on the Audio page, or send it through the MCP server. Transcribe on the Audio page

Result

Your result shows here. Runs are charged in Summer credits and a failed run is refunded.