TTS API
Convert text to natural-sounding speech
Overview
The TTS API allows you to convert text into high-quality audio using various AI voice providers including ElevenLabs, OpenAI, Google, and Gemini. Each character of text costs 1 credit. The response is the generated audio itself, streamed back as an audio/mpeg (or audio/wav) file.
💰 Pricing: 1 character = 1 credit
/api/tts/generateGenerate audio from text. Returns audio file or base64 encoded audio.
Request Body
| Parameter | Type | Required | Description |
|---|---|---|---|
| text | string | Yes | Text to convert (max 5000 chars) |
| voiceId | string | Yes | Voice ID from voices endpoint |
| provider | string | No | Provider: eleven, openai, google, gemini (default: eleven) |
| voice_settings | object | No | ElevenLabs tuning object: model_id (see models below), stability 0-1, similarity 0-1, style 0-1, speaker_boost. |
| source | string | No | ElevenLabs voice engine: c, f or p. Omit to let the system auto-select the best available (recommended). |
ElevenLabs Models — voice_settings.model_id
Choose how your ElevenLabs voice is rendered by passing model_id inside voice_settings. Default is eleven_flash_v2_5. Longer text is automatically split into chunks and stitched into one complete file — you never get a cut-off clip.
| model_id | Best for |
|---|---|
| eleven_flash_v2_5 | Fastest, lowest latency — ideal for long narration (default) |
| eleven_turbo_v2_5 | Fast with strong quality — balanced choice |
| eleven_multilingual_v2 | Highest quality, 29 languages |
| eleven_v4_turbo | Newest v4 Turbo — fast & high quality (official route, 4x credits) |
| eleven_v3 | Most expressive — supports emotion audio-tags like [excited], [whispers], [sighs] |
voiceId decides who speaks; the model_id decides how. You always get exactly the voice you selected — the model only changes speed/expressiveness/quality.Example Request
curl -X POST "https://app.abhibots.com/api/tts/generate" \
-H "X-API-Key: your_api_key" \
-H "Content-Type: application/json" \
-d '{
"text": "Hello! Welcome to AbhiBots TTS API.",
"voiceId": "21m00Tcm4TlvDq8ikWAM",
"provider": "eleven",
"voice_settings": {
"model_id": "eleven_flash_v2_5",
"stability": 0.5,
"similarity": 0.75
}
}'Response
On success (200) the endpoint returns the audio file directly as a binary stream — Content-Type: audio/mpeg for ElevenLabs/OpenAI or audio/wav for Gemini. Write the response body to a file. On failure it returns JSON { "detail": "..." } with a 4xx/5xx status (e.g. 401 bad key, 402 no credits, 429 rate limit).
# Save the audio straight to a file
curl -X POST "https://app.abhibots.com/api/tts/generate" \
-H "X-API-Key: your_api_key" -H "Content-Type: application/json" \
-d '{"text":"Hi there","voiceId":"21m00Tcm4TlvDq8ikWAM","provider":"eleven","voice_settings":{"model_id":"eleven_flash_v2_5"}}' \
--output speech.mp3Python Example
import requests
API_KEY = "your_api_key"
API_URL = "https://app.abhibots.com/api/tts/generate"
response = requests.post(
API_URL,
headers={"X-API-Key": API_KEY},
json={
"text": "Hello from Python!",
"voiceId": "21m00Tcm4TlvDq8ikWAM",
"provider": "eleven",
"voice_settings": {"model_id": "eleven_flash_v2_5"}
}
)
if response.status_code == 200:
data = response.json()
print(f"Audio URL: {data['audio_url']}")
print(f"Credits used: {data['credits_used']}")
else:
print(f"Error: {response.json()}")JavaScript Example
const API_KEY = "your_api_key";
const response = await fetch("https://app.abhibots.com/api/tts/generate", {
method: "POST",
headers: {
"X-API-Key": API_KEY,
"Content-Type": "application/json"
},
body: JSON.stringify({
text: "Hello from JavaScript!",
voiceId: "21m00Tcm4TlvDq8ikWAM",
provider: "eleven",
voice_settings: { model_id: "eleven_flash_v2_5" }
})
});
const data = await response.json();
console.log("Audio URL:", data.audio_url);
console.log("Credits used:", data.credits_used);/api/tts/streamREAL-TIMEReal-time streaming TTS via Gemini Live API (WebSocket-backed). First audio chunk arrives in ~1 second — play progressively before the full clip is ready. Charged 2x (Gemini). Returns a live audio/wav stream.
Request Body
| Parameter | Type | Required | Description |
|---|---|---|---|
| text | string | Yes | Text to convert (up to 20,000 chars) |
| voiceId | string | No | Gemini voice (default Kore): Aoede, Puck, Zephyr |
| model | string | No | flash (fast) or pro (quality). Default flash |
| style | string | No | Style prompt e.g. cheerful, calm |
Example (curl)
curl -X POST "https://app.abhibots.com/api/tts/stream" -H "X-API-Key: your_api_key" -H "Content-Type: application/json" -d '{"text":"Hello in real time!","voiceId":"Kore","model":"flash"}' --output stream.wavPython (stream to file)
import requests
with requests.post(
"https://app.abhibots.com/api/tts/stream",
headers={"X-API-Key": "your_api_key"},
json={"text": "Streaming!", "voiceId": "Kore", "model": "flash"},
stream=True) as r:
with open("out.wav", "wb") as f:
for chunk in r.iter_content(8192):
f.write(chunk)Real-time playback (play as it streams)
# Play in real-time as chunks arrive (needs: pip install pyaudio)
import requests, pyaudio
pa = pyaudio.PyAudio()
stream = pa.open(format=pyaudio.paInt16, channels=1, rate=24000, output=True)
with requests.post(
"https://app.abhibots.com/api/tts/stream",
headers={"X-API-Key": "your_api_key"},
json={"text": "Plays as it streams!", "voiceId": "Kore"},
stream=True) as r:
first = True
for chunk in r.iter_content(4096):
if first: # skip 44-byte WAV header
chunk = chunk[44:]; first = False
stream.write(chunk)/api/tts/taskASYNCBest for long text and SaaS / batch use. Returns a task_id instantly (no waiting), then you poll the Get Task endpoint below for the result. Avoids sync-request timeouts on big generations. ElevenLabs voices. Note: the text field is input.
Request Body
| Parameter | Type | Required | Description |
|---|---|---|---|
| input | string | Yes | Text to convert (up to 50,000 chars) |
| voice_id | string | Yes | Voice ID from the voices endpoint |
| model_id | string | No | eleven_flash_v2_5 (default), eleven_multilingual_v2, eleven_v3, eleven_turbo_v2_5, eleven_v4_turbo |
| stability | float | No | 0-1 (default 0.5). Premium tuning |
| similarity | float | No | 0-1 (default 0.75). Premium tuning |
| source | string | No | Optional engine hint: c, f or p |
Example (curl)
curl -X POST "https://app.abhibots.com/api/tts/task" -H "X-API-Key: your_api_key" -H "Content-Type: application/json" -d '{"input":"Long text here...","voice_id":"21m00Tcm4TlvDq8ikWAM","model_id":"eleven_flash_v2_5"}'Response
{ "task_id": "a1b2c3d4" }/api/tts/task/{task_id}Poll this until status is completed, then download the audio from the result URL.
Response (completed)
{
"id": "a1b2c3d4",
"status": "completed", // processing | completed | failed
"result": "https://app.abhibots.com/api/tts/audio/12345_abcd.mp3",
"input": "Long text here...",
"voice_id": "21m00Tcm4TlvDq8ikWAM",
"model_id": "eleven_flash_v2_5"
}Python (create -> poll -> download)
import requests, time
H = {"X-API-Key": "your_api_key"}
task = requests.post("https://app.abhibots.com/api/tts/task", headers=H,
json={"input": "Long text...", "voice_id": "21m00Tcm4TlvDq8ikWAM"}).json()
tid = task["task_id"]
while True:
t = requests.get(f"https://app.abhibots.com/api/tts/task/{tid}", headers=H).json()
if t["status"] == "completed":
audio = requests.get(t["result"], headers=H).content
open("out.mp3", "wb").write(audio); break
if t["status"] == "failed":
raise Exception(t.get("error"))
time.sleep(2)/api/tts/historyReturns your account's most recent generations (up to 50, newest first) - text, voice, characters, duration and a downloadable audio URL.
Example (curl)
curl "https://app.abhibots.com/api/tts/history" -H "X-API-Key: your_api_key"Response
[
{
"id": "344079",
"provider": "eleven",
"voice": "21m00Tcm4TlvDq8ikWAM",
"text": "First chars of the text...",
"chars": 1483,
"duration": 98.8,
"createdAt": "2026-08-22T13:27:31Z",
"audioUrl": "https://app.abhibots.com/api/tts/audio/12345_abcd.mp3"
}
]