Skip to main content
Sarj TTS converts text into speech using built-in voices or your own reference recording. The speech endpoint supports an OpenAI SDK-compatible request format.

Generate speech

Send a JSON request to POST /audio/speech with your TTS API key in the Authorization header. The response is binary audio.

Voices and samples

These five Saudi voices use the same sentence, generated with each voice’s production defaults.

يا مرحبا، تقدر تتابع طلبك من التطبيق، وإذا احتجت أي مساعدة حنا موجودين. قل لي وش تحتاج، ونرتبها لك خطوة بخطوة.

Use the ID next to each name as the voice value. For the current voice catalog, call GET /voices relative to the base URL above, with the same API key.

Request parameters

For example, add "speed": 1.1 to request a slightly faster delivery than 1.0. Other omitted generation controls use the deployment’s configured defaults.
Start with the defaults and change one control at a time. t_shift and the two temperatures can affect delivery, but they are not interchangeable expressiveness controls.See the TTS OpenAPI schema for the full request schema, including streaming and additional controls.

Voice cloning

For one-shot cloning, send multipart/form-data to POST /audio/speech/clone. Use a clean reference recording and its exact transcript. A 6-10 second reference is a useful starting point.
Replace ref_text with the recording’s transcript in its original language. This endpoint uses text, not input, and returns audio without registering a permanent voice.
Last modified on September 15, 2026