Grok TTS studio
Grok Text to Speech: Create Expressive AI Voice
Grok text to speech turns a written script into directed audio in your browser. Choose one of five voices, select from 20 languages or auto detection, add speech tags, and export a secure MP3, WAV, PCM, μ-law, or A-law file.
Live generation studio
XAI / GROK-TTS / READY
Write a line, shape its delivery, and generate through Miso One's authenticated server workflow. Your script stays in place if a request fails.
Model scope / 01
What is Grok text to speech?
Grok text to speech is xAI's text-to-audio capability; this page uses the narrower xai/grok-text-to-speech endpoint published on Replicate and exposes only the controls that endpoint documents.
The current Replicate endpoint accepts a script, one of five built-in voices, 20 named languages or automatic language detection, speech direction tags, output encoding, sample rate, and optional text normalization.
Miso One adds a browser workstation around that contract, then stores successful audio on owned storage so playback and downloading do not depend on a temporary provider result URL.
Review the Replicate model and API schemaVoice register / 02
Choose the right Grok TTS voice
Start with the delivery a script needs, then audition a short representative line. The five identifiers below are the built-in choices exposed by the current Replicate schema, not a claim about every xAI voice product.
Eve
Energetic and upbeat · narration
Ara
Warm and friendly · explainers
Rex
Confident and clear · product voiceover
Sal
Smooth and balanced · stories
Leo
Authoritative and strong · announcements
The tone phrases above follow Replicate's model page; each best-for label is Miso One editorial guidance, not an official provider claim. Voice fit still depends on script, language, pacing, and direction.
Verify the five voice descriptions on ReplicateWorkflow / 03
How to use Grok text to speech
Move from script to owned audio in four clear steps, from writing the script through previewing and downloading the stored result.
Write the spoken script
Enter the exact words to speak. The server applies the character limit associated with your signed-in account.
Direct voice and delivery
Choose Eve, Ara, Rex, Sal, or Leo, set a language, and insert concise speech tags where the performance should change.
Choose the output
Select MP3, WAV, PCM, μ-law, or A-law and a sample rate; MP3 requests can also set a bit rate.
Generate, listen, and download
Submit through the authenticated Miso One endpoint, preview the stored result, and download the same-origin audio file.
Speech tags that change delivery
Speech tags place compact performance cues inside the script. Insert them only where a listener should hear a pause, breath, laugh, sigh, or whispered delivery, then audition the result.
[pause]The result is ready. [pause] Let's hear it.
Create a deliberate beat between ideas.
[laugh]I thought the first take was final. [laugh]
Signal a brief laugh when it fits the script.
[sigh] / [breath][breath] We can begin again.
Add a small human transition without rewriting the line.
<whisper><whisper>Keep this part between us.</whisper>
Direct a quieter delivery for a selected phrase.
Grok text to speech formats for each delivery path
Choose the format your next tool expects. MP3 is compact for review, WAV or PCM suits editing pipelines, and μ-law or A-law targets compatible phone systems.
| Format | Useful when | Control note |
|---|---|---|
| MP3 | Sharing drafts, publishing, and lightweight previews | Includes a selectable bit rate |
| WAV | Editing and interchange with common audio tools | Bit rate is not sent |
| PCM | Raw downstream processing where a pipeline expects PCM | Confirm the receiving system's sample rate |
| μ-law / A-law | Compatible telephony and IVR workflows | Match the codec required by the phone platform |
Where Grok TTS fits a production workflow
The strongest use case is a workflow with a known script, a clear delivery decision, and a defined audio destination.

Case signal 01
Podcast and video voiceover
Turn intros, transitions, explainers, and revision lines into reviewable audio before the edit is locked.

Case signal 02
Multilingual localization
Keep one content structure while choosing the language and voice for each localized deliverable.

Case signal 03
IVR and support prompts
Prepare short phone-system lines in μ-law or A-law when those codecs match the destination platform.
Grok TTS vs ElevenLabs vs OpenAI TTS
Choose among these tools by workflow and documented controls, not by a universal winner. Product capabilities can differ by endpoint and change over time, so confirm the linked official documentation before a production decision.
| Decision | Grok TTS here | ElevenLabs TTS | OpenAI TTS |
|---|---|---|---|
| Try in this browser | Yes, through this signed-in Miso One studio | Use the ElevenLabs product or API workflow | Use an OpenAI API integration or supported product |
| Built-in voice choice | Five IDs in the current Replicate schema | Voice library and platform voice workflows | Built-in voices documented for the speech API |
| Language approach | 20 named languages plus auto detection here | Model-specific multilingual support | Model and voice instructions shape multilingual output |
| Delivery direction | Inline speech tags and text normalization | Controls vary by model and product workflow | Natural-language voice instructions on supported models |
| Output workflow | MP3, WAV, PCM, μ-law, or A-law stored by Miso One | Formats and delivery options documented per API | Speech API output formats documented by OpenAI |
| Choose it when | You want this five-voice, tag-directed browser workflow | Its voice platform and model options fit your wider voice workflow | You want speech inside an existing OpenAI application stack |
Choose Grok TTS here for a compact, tag-directed workflow with owned downloads; consider ElevenLabs when its documented voice platform matches the project; consider OpenAI TTS when speech belongs inside an OpenAI-based application. Validate the exact endpoint before committing.
Know the endpoint limits before production
This page promises only the current Replicate xai/grok-text-to-speech input contract. It does not imply that every feature in xAI's broader Voice offering is available through this endpoint.
xAI Voice overviewFree signed-in accounts can submit up to 120 characters per generation; active paid accounts can submit up to 1,000. The server is the authority for both limits.
Credits are calculated as one credit for each started block of 100 characters: ceil(characters / 100).
The built-in selection here is Eve, Ara, Rex, Sal, and Leo. Voice cloning, streaming, timestamps, and other broader product capabilities are not promised by this page.
Audio suitability depends on the script, selected voice, language, tags, format, and the requirements of the destination system.
Grok text to speech FAQ
Short answers about the exact browser workflow, voices, limits, credits, and endpoint scope.
What is Grok text to speech?
It is text-to-audio technology from xAI. This page connects to the xai/grok-text-to-speech model on Replicate and limits its controls and claims to that endpoint's published schema.
Can I use Grok TTS online?
Yes. Write and direct a script in this browser studio, sign in, generate through the protected server endpoint, then preview and download the stored result.
Which Grok TTS voices are available here?
The current Replicate schema exposes five voice identifiers here: Eve, Ara, Rex, Sal, and Leo.
Which languages and formats are supported?
The endpoint offers 20 named languages plus auto detection, with MP3, WAV, PCM, μ-law, and A-law output.
How do speech tags work?
Insert concise tags such as [pause], [laugh], [sigh], or [breath] at the point where delivery should change, or wrap a selected phrase in <whisper>...</whisper>, then generate a short audition.
How many credits does generation use?
Generation uses ceil(characters / 100) credits, so 1–100 characters use one credit and 901–1,000 use ten. The server calculates the final charge.
Is this the same as every feature in xAI Voice?
No. xAI's broader Voice product and this Replicate endpoint are not identical. This page promises only the five voices, languages, formats, tags, sample rates, MP3 bit rates, and normalization exposed by the endpoint used here.
Continue your voice workflow
Compare the focused Grok studio with Miso One's broader text-to-speech, voice, and planning resources.
Official sources
Endpoint facts and comparison boundaries are grounded in current first-party documentation.
Turn the next script into owned audio
Return to the studio, audition a short representative line, and choose the voice and output settings that match the real destination.
