> ## Documentation Index
> Fetch the complete documentation index at: https://platform-docs.sarj.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Text-to-Speech

> Generate speech with Sarj TTS, choose a voice, and control its delivery.

export const VoiceSample = ({name, voice, src, duration: initialDuration, peaks}) => {
  const audio = useRef(null);
  const [playing, setPlaying] = useState(false);
  const [waiting, setWaiting] = useState(false);
  const [time, setTime] = useState(0);
  const [duration, setDuration] = useState(initialDuration);
  const [ready, setReady] = useState(false);
  const [error, setError] = useState(false);
  const formatTime = value => `${Math.floor(value / 60)}:${String(Math.floor(value % 60)).padStart(2, '0')}`;
  const toggle = async () => {
    const player = audio.current;
    if (!player) return;
    if (!player.paused) {
      player.pause();
      return;
    }
    if (error) player.load();
    setError(false);
    setWaiting(true);
    try {
      await player.play();
    } catch (cause) {
      if (cause.name !== 'AbortError') {
        setError(true);
        setWaiting(false);
        setPlaying(false);
      }
    }
  };
  return <div className="sarj-voice-player not-prose" data-playing={playing}>
      <audio data-sarj-voice-audio="true" hidden onEnded={() => {
    setPlaying(false);
    setWaiting(false);
  }} onError={() => {
    setError(true);
    setReady(false);
    setPlaying(false);
    setWaiting(false);
  }} onLoadedMetadata={event => {
    const actual = event.currentTarget.duration;
    if (Number.isFinite(actual) && actual > 0) {
      setDuration(actual);
      setReady(true);
    }
  }} onPause={() => {
    setPlaying(false);
    setWaiting(false);
  }} onPlay={event => {
    for (const other of document.querySelectorAll('[data-sarj-voice-audio]')) {
      if (other !== event.currentTarget) other.pause();
    }
    setPlaying(true);
  }} onPlaying={() => setWaiting(false)} onTimeUpdate={event => setTime(event.currentTarget.currentTime)} onWaiting={() => setWaiting(true)} preload="metadata" ref={audio} src={src} />
      <button aria-label={`${playing ? 'Pause' : error ? 'Retry' : 'Play'} ${name}`} aria-pressed={playing} className="sarj-voice-play" onClick={toggle} title={`${playing ? 'Pause' : error ? 'Retry' : 'Play'} ${name}`} type="button">
        <span aria-hidden="true" className={waiting ? 'sarj-voice-loading' : ''}>
          <Icon color="#ffffff" icon={waiting ? 'spinner' : playing ? 'pause' : error ? 'rotate-right' : 'play'} iconType="solid" size={16} />
        </span>
      </button>
      <div className="sarj-voice-body">
        <div className="sarj-voice-heading">
          <span className="sarj-voice-name">{name}</span>
          <code className="sarj-voice-id">{voice}</code>
        </div>
        <div className="sarj-voice-transport">
          <div className="sarj-voice-seek">
            <div aria-hidden="true" className="sarj-voice-wave">
              {peaks.map((height, index) => <span data-played={time > 0 && index / peaks.length < time / duration} key={index} style={{
    height: `${height}%`
  }} />)}
            </div>
            <input aria-label={`Seek ${name}`} aria-valuetext={`${formatTime(time)} of ${formatTime(duration)}`} disabled={!ready || error} max={duration} min="0" onChange={event => {
    const next = Number(event.target.value);
    if (audio.current && Number.isFinite(next)) {
      audio.current.currentTime = next;
      setTime(next);
    }
  }} step="0.1" type="range" value={time} />
          </div>
          <span aria-hidden="true" className="sarj-voice-time">{formatTime(time)} / {formatTime(duration)}</span>
        </div>
        {error ? <span className="sarj-voice-error" role="status">Audio unavailable</span> : null}
      </div>
      <a aria-label={`Download ${name} WAV`} className="sarj-voice-download" download={`${name.toLowerCase()}.wav`} href={src} title={`Download ${name} WAV`}>
        <Icon icon="download" size={16} />
      </a>
    </div>;
};

Sarj TTS converts text into speech using built-in voices or your own reference recording. The speech endpoint supports an OpenAI SDK-compatible request format.

| Base URL | Model |
| :- | :- |
| `https://sarj-omni-tts.sarj.ai/v1` | `sarj-tts` |

## Generate speech

Send a JSON request to `POST /audio/speech` with your TTS API key in the `Authorization` header. The response is binary audio.

```bash theme={null}
curl --fail-with-body https://sarj-omni-tts.sarj.ai/v1/audio/speech \
  -H "Authorization: Bearer $SARJ_TTS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "sarj-tts",
    "voice": "ars_male",
    "input": "يا مرحبا، كيف نقدر نساعدك اليوم؟",
    "response_format": "wav"
  }' \
  --output speech.wav
```

## Voices and samples

These five Saudi voices use the same sentence, generated with each voice's production defaults.

<p lang="ar" dir="rtl">يا مرحبا، تقدر تتابع طلبك من التطبيق، وإذا احتجت أي مساعدة حنا موجودين. قل لي وش تحتاج، ونرتبها لك خطوة بخطوة.</p>

<VoiceSample name="Abdullah" voice="ars_male" src="/audio/voices/abdullah.wav" duration={7.62} peaks={[8,87,87,38,37,49,71,72,80,72,58,59,47,48,31,58,40,75,58,79,65,84,54,62,62,43,42,8,8,8,8,8,52,100,71,55,66,22,57,63,62,79,29,69,67,44,36,8]} />

<VoiceSample name="Ghada" voice="ars_female" src="/audio/voices/ghada.wav" duration={9.33} peaks={[10,8,8,17,78,92,56,56,15,66,74,80,65,76,68,66,36,78,17,20,89,89,85,80,81,82,63,78,68,62,49,15,16,100,87,45,45,74,22,15,83,66,67,43,61,36,34,8]} />

<VoiceSample name="Nourah" voice="nourah" src="/audio/voices/nourah.wav" duration={8.28} peaks={[17,63,90,40,42,20,21,65,81,94,88,82,60,61,48,43,27,30,67,76,100,61,91,69,70,79,71,55,50,29,28,30,88,64,59,60,58,23,84,73,72,67,33,59,35,45,26,13]} />

<VoiceSample name="Ibrahim" voice="ars_ibrahim_studio_male" src="/audio/voices/ibrahim.wav" duration={7.42} peaks={[8,59,51,39,44,8,64,74,60,75,68,53,48,42,56,15,70,69,84,31,80,69,71,72,71,68,71,54,51,29,8,77,100,72,52,73,38,8,62,74,89,83,66,82,54,45,32,8]} />

<VoiceSample name="Fares" voice="ars_fares" src="/audio/voices/fares.wav" duration={9.4} peaks={[17,24,22,23,23,70,74,56,70,67,77,75,52,62,51,73,45,65,81,50,96,79,91,70,75,68,60,26,21,22,21,20,37,100,84,43,67,44,24,67,74,76,60,68,58,52,25,12]} />

Use the ID next to each name as the `voice` value. For the current voice catalog, call `GET /voices` relative to the base URL above, with the same API key.

## Request parameters

| Parameter | Description |
| :- | :- |
| `input` | Required. Text to speak, up to 50,000 characters. |
| `model` | `sarj-tts`, the only supported model and the default. |
| `voice` | A voice ID from the catalog. Defaults to `auto`; select an ID for a specific voice. |
| `response_format` | `mp3` (default), `wav`, `opus`, `aac`, `flac`, or `pcm`. |
| `speed` | Speaking rate, from `0.25` to `4.0`. Omit to use the voice's configured rate. An explicit value overrides that rate. |

For example, add `"speed": 1.1` to request a slightly faster delivery than `1.0`. Other omitted generation controls use the deployment's configured defaults.

<Accordion title="Advanced controls">
  Start with the defaults and change one control at a time. `t_shift` and the two temperatures can affect delivery, but they are not interchangeable expressiveness controls.

  | Parameter | Effect |
  | :- | :- |
  | `t_shift` | Changes the token-refinement schedule and can affect rhythm, prosody, and stability. Higher values do not necessarily mean more expression. Keep it above `0`; the maximum is `2`. |
  | `class_temperature` | Controls randomness in audio-token selection (`0` to `2`). Higher values add variation but can introduce artifacts; `0` uses greedy selection. |
  | `position_temperature` | Controls randomness in which token positions are refined next (`0` to `10`). It is separate from audio-token selection. |
  | `guidance_scale` | Conditioning strength (`0` to `10`). Increasing it is not a guarantee of better quality. |
  | `num_step` | Number of generation steps (`1` to `64`). Fewer steps can reduce quality; omit to retain the deployment default. |
  | `instructions` | Optional delivery instructions, up to 4,096 characters. Their effect depends on the voice and text. |
  | `language` | Optional language code, such as `ar` or `en`. This does not select a dialect-specific voice. |

  See the [TTS OpenAPI schema](https://sarj-omni-tts.sarj.ai/openapi.json) for the full request schema, including streaming and additional controls.
</Accordion>

## Voice cloning

For one-shot cloning, send `multipart/form-data` to `POST /audio/speech/clone`. Use a clean reference recording and its exact transcript. A 6-10 second reference is a useful starting point.

```bash theme={null}
curl --fail-with-body https://sarj-omni-tts.sarj.ai/v1/audio/speech/clone \
  -H "Authorization: Bearer $SARJ_TTS_API_KEY" \
  -F "ref_audio=@reference.wav" \
  --form-string "ref_text=The exact words spoken in reference.wav." \
  --form-string "text=يا مرحبا، كيف نقدر نساعدك اليوم؟" \
  -F "response_format=wav" \
  --output cloned-speech.wav
```

Replace `ref_text` with the recording's transcript in its original language. This endpoint uses `text`, not `input`, and returns audio without registering a permanent voice.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.