Transcribe audio
Sendmultipart/form-data to POST /audio/transcriptions with your STT API key in the Authorization header.
Request parameters
For predictable input quality, use mono PCM WAV at 16 kHz. Keep the recording clear and avoid clipping or background noise.
Real-time streaming
For live audio, connect towss://stt-rnnt-ar.sarj.ai/stream with Authorization: Bearer YOUR_STT_API_KEY in the WebSocket handshake headers. This endpoint is separate from the /openai/v1 HTTP base URL.
- Send
{"sample_rate": 16000}as JSON and wait for thesession_idresponse. - Send mono, signed 16-bit little-endian PCM audio in binary frames, without a WAV header. Receive partial and final transcript events while sending audio.
- Send
{"action": "end_stream"}as JSON when finished. Continue reading final events until the server closes the connection.
Transcript.Results. Each result has IsPartial and an Alternatives array; the recognized text is in Alternatives[0].Transcript when an alternative is present.
This WebSocket interface is not listed in the HTTP OpenAPI schema. For uploaded files, the separate stream=true form option returns transcription events over Server-Sent Events (SSE); it is not a live-audio upload interface.
API schema
See the STT OpenAPI schema for the full HTTP request and response schema, including thestream option for file transcription.
File transcription response
json.
