Skip to main content
Sarj STT transcribes Arabic speech. Its file-transcription endpoint supports an OpenAI SDK-compatible request format.

Transcribe audio

Send multipart/form-data to POST /audio/transcriptions with your STT API key in the Authorization header.

Request parameters

For predictable input quality, use mono PCM WAV at 16 kHz. Keep the recording clear and avoid clipping or background noise.

Real-time streaming

For live audio, connect to wss://stt-rnnt-ar.sarj.ai/stream with Authorization: Bearer YOUR_STT_API_KEY in the WebSocket handshake headers. This endpoint is separate from the /openai/v1 HTTP base URL.
  1. Send {"sample_rate": 16000} as JSON and wait for the session_id response.
  2. Send mono, signed 16-bit little-endian PCM audio in binary frames, without a WAV header. Receive partial and final transcript events while sending audio.
  3. Send {"action": "end_stream"} as JSON when finished. Continue reading final events until the server closes the connection.
Transcript events contain Transcript.Results. Each result has IsPartial and an Alternatives array; the recognized text is in Alternatives[0].Transcript when an alternative is present. This WebSocket interface is not listed in the HTTP OpenAPI schema. For uploaded files, the separate stream=true form option returns transcription events over Server-Sent Events (SSE); it is not a live-audio upload interface.

API schema

See the STT OpenAPI schema for the full HTTP request and response schema, including the stream option for file transcription.

File transcription response

The API returns the transcription in the requested response format. The default response format is json.
Last modified on September 15, 2026