ModelRover

xAI Voice

Synthesize speech, transcribe audio, and list voices through xAI-compatible REST endpoints.

Endpoints

POST /v1/tts
POST /v1/stt
GET /v1/tts/voices
GET /v1/tts/voices/{voice_id}
Authorization: Bearer $API_KEY

Every voice call requires model. It is the TTS JSON body field, the STT multipart form field, or the model query parameter on voice list calls. The value is the complete public model ID and selects the platform route. The gateway removes it from TTS requests before they reach xAI, because the xAI TTS body has no model field. The model ID is the one shown on the model page; xAI's TTS API itself has no model name.

Text to speech

POST /v1/tts accepts the xAI TTS JSON body and forwards every field unchanged except model: text, voice_id, language, output_format, speed, text_normalization, optimize_streaming_latency, with_timestamps, and replace. The response is the raw audio with the upstream Content-Type. With with_timestamps=true, the response is xAI's JSON envelope.

curl "{{API_BASE_URL}}/v1/tts" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "xai/grok-tts",
    "text": "Welcome to the gateway.",
    "voice_id": "eve",
    "language": "en"
  }' \
  --output speech.mp3

Speech to text

POST /v1/stt takes multipart/form-data. Send the audio either as a file part or as a public url field, and never both. The other fields are the official xAI STT form fields and are forwarded unchanged. The gateway rewrites the form model to the channel's upstream model ID.

curl "{{API_BASE_URL}}/v1/stt" \
  -H "Authorization: Bearer $API_KEY" \
  -F model=xai/grok-voice-transcribe-2.0 \
  -F file=@meeting.wav

Uploads up to 50 MB are accepted by default. For larger recordings, host the audio at a public URL and send it as url; xAI downloads it on its side.

Voices

GET /v1/tts/voices lists voices for the model named in the model query parameter. GET /v1/tts/voices/{voice_id} returns one voice. These calls are not billed, and the response is the upstream JSON.

curl "{{API_BASE_URL}}/v1/tts/voices/eve?model=xai/grok-tts" \
  -H "Authorization: Bearer $API_KEY"

WebSocket endpoints

GET /v1/realtime, GET /v1/tts, and GET /v1/stt open WebSocket connections, so the request must be an Upgrade request. Authenticate with Authorization: Bearer $API_KEY; browser subprotocol tokens are not supported.

  • /v1/realtime?model=… is the realtime voice session. model is required, selects the platform route, and is sent to xAI as the upstream model ID. Other query parameters are forwarded unchanged.
  • /v1/tts?model=…&language=…&voice=…&codec=… is streaming TTS. model is required and selects the platform route, but nothing is sent upstream because xAI's streaming TTS has no model parameter.
  • /v1/stt?model=…&encoding=…&sample_rate=… is streaming STT. model is required and rewritten to the upstream model ID.

Frames are relayed unchanged in both directions, whether JSON text frames or binary audio frames. If xAI rejects the handshake, the HTTP error comes back as JSON and the connection is not upgraded. If the balance runs out during a session, the gateway sends an insufficient_quota error frame and closes the connection with code 1008.

Billing for WebSocket sessions:

  • Realtime: VAD sessions are billed by session duration in seconds, rounded up, and priced per minute. Push-to-talk sessions are billed by the seconds of audio sent and received, rounded up, and priced per minute. Each client conversation.item.create counts as a text input, except function results and items that carry audio.
  • Streaming TTS: per input character of text.delta.
  • Streaming STT: per audio second, rounded up, at the streaming rate.

Not supported: SIP call_id sessions, session resumption, the realtime file_search tool, ephemeral client secrets, and custom voices. A call_id is rejected before the upgrade with HTTP 400 stateful_not_supported. Resumption and file_search are rejected inside the session with an error frame.

Billing

  • TTS is billed per input character of text as sent.
  • STT is billed per audio second, rounded up. Admission reserves an upper bound from the upload size, or six hours for url input, and the final charge uses the duration xAI reports.

Restrictions

  • Custom voices, voice cloning, ephemeral tokens, and SIP are not exposed as routes. STT audio comes only from a file part or a public url.
  • Voice endpoints accept only audio models that are bound to an xAI Voice channel. Chat requests to these models are rejected.
  • Requests are subject to the platform body limit. Logs replace audio bytes with a placeholder that records the content type and size.

Help us improve this page

Found something unclear, outdated, or incorrect?

Last updated on