Skip to main content

Grok Voice Transcribe 2.0 API Tutorial: Build a Real-Time Support Call Transcriber

Learn how to build a real-time support-call transcriber with the Grok Voice Transcribe 2.0 API in Python, then add diarization, Smart Turn, and 8 kHz phone audio.
Oct 5, 2026  · 12 min read

Explore with AI

ChatGPTClaudePerplexity

SpaceXAI's Grok Voice Transcribe 2.0 is a speech-to-text model. In this Grok Voice Transcribe 2.0 API tutorial, you send recordings over REST and live audio over a WebSocket. The API returns text, word timings, optional speaker IDs, and end-of-turn events; it does not answer the caller.

A support call is harder than one clean narrator. It has short gaps, unfamiliar names, several speakers, and contact details read over an 8 kHz line. Our project called Qivora Sync gives the tutorial one thread: a customer reports a failed file sync, the agent collects contact details, and an escalation engineer joins. The same Python client handles the recording first and live audio later.

For speech-to-speech, where the model answers the caller itself, see our Grok Voice Think Fast 2.0 tutorial. The code for this one is in the GitHub repository.

TL;DR

Short on time? Here is what the call showed.

  • POST /v1/stt handles recorded audio and wss://api.x.ai/v1/stt handles live audio, with shared controls for diarization, keyterms, fillers, and audio handling.

  • A keyterm fixed the invented product name, but strongly biased vocabulary pulled a faint echo toward that name in a live-speaker check.

  • Speaker labels stayed stable on the clean mix but became unreliable at 8 kHz.

  • The Arabic switch stayed in Arabic script, and format=true fixed the phone number, but only half-fixed the email.

  • On the long mid-number pause, Smart Turn crossed every tested threshold, so threshold tuning alone was not enough.

Associate AI Engineer

Train and fine-tune the latest AI models for production, including LLMs like Llama 4. Start your journey to becoming an AI Engineer today!
Explore Track

What Is Grok Voice Transcribe 2.0?

Grok Voice Transcribe 2.0 (grok-voice-transcribe-2.0) is SpaceXAI's speech-to-text model. The REST path transcribes a finished file, while the WebSocket path handles live audio.

SpaceXAI's Grok Voice Transcribe 2.0 announcement highlights phone calls, multiple speakers, credentials, and multilingual speech. For benchmark comparisons, see our Grok Voice Transcribe 2.0 overview.

Building a Real-Time Support Call Transcriber

The controlled Qivora Sync fixture stays fixed while the audio and API settings change. The call includes an invented product name, fillers, a language switch, spoken contact details, a pause during dictation, and a third speaker.

The Qivora Sync support call pipeline: three speakers into Grok Voice Transcribe 2.0, out as a diarized live transcript

Three speakers become one live transcript. Image by Author.

Creating the three-speaker call

The controlled fixture uses three distinct voices from the Grok Text to Speech API. Each language segment is synthesized separately and joined with ffmpeg so that the switch points stay fixed. The API also accepts language=auto; separate requests are an experiment design choice, not an API requirement.

Defining the expected transcript

Before the first request, define the expected text, speakers, product spelling, customer details, fillers, and pauses. Every setup then has the same target.

Setting Up Grok Voice Transcribe 2.0 in Python

Install the dependencies before sending audio.

Prerequisites

You need Python 3.10 or newer, an xAI API key, and ffmpeg for building the audio. The Python clients use requests, websockets, and python-dotenv.

The Speech to Text docs call 2.0 the default when you omit model, and grok-voice-transcribe-1.0 reached end of life on October 2, 2026. I would still pin the versioned ID.

Installing dependencies and building audio

Clone the repository, add your key to .env, and build the sample audio:

git clone https://github.com/KhalidAbdelaty/grok-voice-transcribe-2.0.git
cd grok-voice-transcribe-2.0
pip install -r requirements.txt
cp .env.example .env    # then paste your key into .env
python project/scripts/make_fixtures.py

The setup command creates the dialogue and the audio files used later. If you have your own recording, skip that command.

A .env written on Windows can leave a \r on the key, and requests rejects the header before anything reaches SpaceXAI. Strip the key before adding it to the authorization header.

Establishing a Batch Transcription Baseline

A baseline is the model with nothing turned on, so every later change has something to compare against. The first request sends the file and a pinned model:

import os
import requests
from dotenv import load_dotenv

load_dotenv()
api_key = os.environ["XAI_API_KEY"].strip()

with open("support_call.wav", "rb") as audio_file:
    response = requests.post(
        "https://api.x.ai/v1/stt",
        headers={"Authorization": f"Bearer {api_key}"},
        data=[("model", "grok-voice-transcribe-2.0")],
        files={"file": ("support_call.wav", audio_file, "audio/wav")},
    )

response.raise_for_status()
result = response.json()

The response carries text, detected language, duration, and a timed words array. The REST reference shows per-word confidence, but it did not appear in the batch responses for this fixture. I would treat the field as optional and check each API response before using it. Put option fields before file; later fields may be ignored.

The baseline stripped fillers, kept the Arabic in Arabic script, and left the spoken digits separate. It consistently misspelled the invented product name.

Adding Diarization, Keyterms, and Text Formatting

A support transcript needs speaker labels, the correct product spelling, and usable customer details. Each setting is one more form field:

data = [
    ("model", "grok-voice-transcribe-2.0"),
    ("diarize", "true"),         # a speaker id on every word
    ("keyterm", "Qivora Sync"),  # repeat the field for more terms
    ("language", "en"),          # required by format
    ("format", "true"),          # inverse text normalization
    ("filler_words", "false"),   # the default; true keeps "uh" and "um"
]

Add one option at a time to the same audio. Start with speaker labels.

Grouping words into speaker turns

Speaker diarization gives words numeric speaker IDs, not names. Group consecutive words with the same ID to build turns:

def group_turns(words):
    turns = []
    for word in words:
        if turns and turns[-1]["speaker"] == word.get("speaker"):
            turns[-1]["words"].append(word["text"])
            turns[-1]["end"] = word["end"]
        else:
            turns.append({"speaker": word.get("speaker"), "start": word["start"],
                          "end": word["end"], "words": [word["text"]]})
    for turn in turns:
        turn["text"] = " ".join(turn.pop("words"))
    return turns

On clean audio, each known turn stayed with a consistent speaker ID. Mapping names by order of first appearance only works when the call order is already known; production systems need their own speaker mapping.

Diarized transcript of the Qivora Sync call showing three separated speaker turns with timestamps

Clean audio keeps speaker labels consistent. Image by Author.

Using keyterm biasing for product names

Keyterm biasing is a per-request hint, not training. Pass keyterm=Qivora Sync (up to 100 terms, 50 characters each), and the model leans toward that spelling when the audio supports it.

The keyterm corrected the baseline's product-name error without changing the surrounding transcript.

In a separate live-speaker check, strongly biased vocabulary pulled a faint echo toward the keyterm. That result does not mean that keyterms create false text on their own; it means that ambiguous audio still needs an echo check.

Transcribing English-Arabic language switches

As the baseline showed, Khalid's Arabic stayed in Arabic script. The result was the same with automatic detection and with language=en, because language selects formatting rules rather than forcing an output language.

Formatting spoken phone numbers and emails

The baseline kept spoken digits separate. Inverse Text Normalization (ITN) turns those spoken forms into written ones. format=true switches it on and needs language, or the request fails with a 400.

The phone number became one continuous digit string. The email was only partly normalized: punctuation improved, but the spoken "at" and spelled domain still needed cleanup.

That uneven result is frustrating. ITN formats text; it does not validate contact data. I would validate both fields before storage.

ITN can also rewrite ordinary duration phrases as abbreviated quantities. In the formatted batch response for this fixture, only the top-level text was normalized; the words array kept the spoken form.

Keeping or removing filler words

As the baseline showed, fillers are stripped from text and words by default. filler_words=true brought Khalid's "uh" and "um" back where expected. Keep them off for support notes and on for a verbatim QA record.

Batch output includes speakers, vocabulary, formatting, and filler-word control. Next, send the same audio as a live stream.

Streaming Grok Voice Transcribe 2.0 Over WebSocket

The streaming path uses query parameters rather than a setup message. Wait for transcript.created, send raw binary audio (no base64), and close with {"type": "audio.done"}. Our GPT Live Transcribe tutorial uses the same pattern with another model.

Start with the events, then connect the client.

Batch uses the documented format=true with language=en. The streaming docs say that language turns on ITN, but in a live probe, language=en alone did not change the transcript. The WebSocket query list does not include format, so this tutorial treats streaming ITN as behavior to verify rather than depend on.

Reading partial and final events

Every transcription update is a transcript.partial event with two booleans. Interim text may still change. A chunk final (is_final=true) locks about 3 seconds of text while the turn stays open, and an utterance final (speech_final=true) closes the turn.

Streaming event flow from transcript created through interim, chunk-final, utterance-final, and transcript done states

Streaming states move text toward finalization. Image by Author.

Streaming 16 kHz PCM audio in Python

For streaming, resample the source to mono 16-bit PCM at 16 kHz first. The core client sends 100-millisecond chunks at a real-time pace while another task receives transcript events:

import asyncio, json, os, wave
import websockets
from dotenv import load_dotenv 

load_dotenv()

url = ("wss://api.x.ai/v1/stt?model=grok-voice-transcribe-2.0"
       "&sample_rate=16000&encoding=pcm&interim_results=true&diarize=true")
headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY'].strip()}"}

async def stream_call(path):
    async with websockets.connect(url, additional_headers=headers) as ws:
        assert json.loads(await ws.recv())["type"] == "transcript.created"

        async def send():
            with wave.open(path, "rb") as wf:
                assert wf.getframerate() == 16000
                assert wf.getnchannels() == 1
                assert wf.getsampwidth() == 2
                while chunk := wf.readframes(1600):
                    await ws.send(chunk)
                    await asyncio.sleep(0.1)
            await ws.send(json.dumps({"type": "audio.done"}))

        async def receive():
            async for raw in ws:
                event = json.loads(raw)
                if event["type"] == "transcript.partial":
                    print(event["text"])
                elif event["type"] == "transcript.done":
                    break

        await asyncio.gather(send(), receive())

Interim text grew about every half-second. This is a local measurement, not official latency.

Terminal showing a partial caption updating before it becomes a final transcript line

Partial captions settle into final transcript. Image by Author.

Chunk finals freeze text without closing the turn. Smart Turn controls when speech_final closes it.

Keeping transcript chunks in order

Showing only the active event makes earlier words vanish after each chunk's final, because the next interim starts over from the incoming audio.

Keep every locked chunk, append the current interim, and let the utterance final replace both.

The text can grow without losing earlier chunks. With display state handled, turn boundaries are the remaining streaming problem.

Using Smart Turn for End-of-Turn Detection

Smart Turn evaluates each silence and estimates whether the speaker has finished. It exists for Khalid's number, "zero one zero, five five five, [pause], one two three four," where silence alone can't tell a thinking pause from the end.

Testing the Smart Turn threshold

The threshold isn't transcription confidence or the VAD threshold. It is the end-of-turn probability a silence must clear before speech_final fires; below it, the turn stays open. Two query parameters set it:

params += [
    ("smart_turn", "0.7"),           # end-of-turn probability needed to close
    ("smart_turn_timeout", "3000"),  # close anyway after 3 s of silence
]

The docs call 0.5 balanced, 0.7 conservative for number sequences, and 0.9 very conservative. In this fixture, pauses shorter than the default endpointing window did not produce a useful Smart Turn decision. That is an observed result, not a documented timing rule.

In the streaming test, stopping audio frames did not advance the observed silence timer. Continuing to send digital silence lets Smart Turn close the utterance.

Extending the pause during number dictation makes the behavior visible. Short pauses stay inside one turn, while a long pause splits it at every threshold when the confidence exceeds all three settings.

Timeline of the phone-number utterance showing speech, pauses, and end-of-turn confidence at each threshold

Long pauses can split number dictation. Image by Author.

Human callers are less predictable. A short digit sequence can look finished. Then the caller continues.

If Smart Turn closes during number dictation, wait briefly and merge a continuation before replying.

Setting a Smart Turn timeout

smart_turn_timeout closes a turn after a fixed silence, even when Smart Turn is unsure. On the fast three-speaker stream, Smart Turn grouped several known turns before a timeout forced the close.

If you already know where turns end, send {"type": "finalize"} at each boundary; otherwise, pair Smart Turn with a timeout.

Once turn boundaries are under control, the same caller has to survive an 8 kHz line.

Transcribing 8 kHz Phone Audio

Phone-quality audio here is 8 kHz G.711 mu-law, made from the same call:

ffmpeg -i support_call.wav -ar 8000 -ac 1 -f mulaw support_call_8k.raw

Raw telephony audio has no container, so set audio_format=mulaw and sample_rate=8000 in the batch form, or encoding=mulaw&sample_rate=8000 on the socket. Check text and speaker labels separately.

Comparing clean and phone audio

The earlier keyterm, formatting, and language-switch findings changed little at 8 kHz.

Speaker labels became less reliable. The phone version introduced an extra speaker ID and assigned a closing turn to the wrong person. Counting segments alone hides both mistakes.

The flaky version band-limits the call to 300-3400 Hz, encodes it as 8 kHz mu-law, and drops each 20-millisecond packet with a probability of 0.03. A fixed random seed of 7 keeps the same gaps on every replay.

That packet loss did not change the English transcript much in this sample, and the spoken contact details stayed in order. This result applies only to this sample.

Phone simulation narrows audio, dropping packets. Image by Author.

Adjusting VAD for phone audio

Voice activity detection (VAD) decides whether audio is speech at all. The docs suggest lowering vad_threshold for quiet phone speech, at the risk of stray text from noise.

Lowering vad_threshold changed nothing on clean phone audio because there was no quiet speech to recover. The null result supports one rule: lower the threshold only when phone speech goes missing.

Using Multichannel Transcription for Separate Speakers

Use a fresh batch form without diarize:

data = [
    ("model", "grok-voice-transcribe-2.0"),
    ("multichannel", "true"),
]

The API detects the channel count from a WAV or other container. For raw multichannel audio, add ("channels", "3"); WebSocket multichannel input also requires an explicit channel count.

Send the form with the multichannel file through the REST request shown earlier, then read result["channels"]. Each item contains an index, transcript text, and timed words. In the controlled three-channel fixture, each channel contained only its assigned speaker. Streaming uses the same split and adds channel_index to its events.

I would use separate legs whenever the phone system provides them. Unlike diarization in the phone-audio section, a known split does not infer speakers.

Building the Complete Python Support Transcriber

The complete client exposes one group of settings, then builds the REST form or WebSocket URL separately. Shared settings cover diarization, keyterms, filler words, audio encoding, and turn handling; formatting follows the transport-specific rules covered earlier.

Apply the final settings to a phone-quality recording, then check product spelling, language switches, contact details, and speaker labels separately. In the controlled fixture, the text checks passed while one speaker label still required review. Save the settings and speaker mapping with each transcript so later comparisons use the same setup.

Exploring the full voice-agent demo

The support-transcription tutorial ends with that final check. The repository also contains a separate voice-agent extension with generated replies, spoken output, interruptions, and echo handling.

Transcribe keeps the same role in that demo: it produces text. A language model writes the replies, and Grok TTS speaks them.

Live call switches audio paths mid-conversation. Video by Author.

Grok Voice Transcribe 2.0 Limitations

Support transcripts can contain names, phone numbers, and emails. SpaceXAI's security FAQ says that it stores API data encrypted at rest for 30 days for abuse auditing. SpaceXAI also says that it does not train on the data without permission. Eligible teams can turn on Zero Data Retention at the team level.

Keep the API key on your server. The Speech-to-Text docs say to proxy the WebSocket through your backend.

One controlled call cannot represent every accent, room, or phone line. Check the settings with audio from the intended environment before using them in production.

Common Errors and Troubleshooting

Most failures here come from audio formatting or socket handling:

  • InvalidHeader ... return character(s) in header value is the Windows \r on the key.

  • A 400 can mean a missing file or url, an unsupported format, raw audio without sample_rate, or format=true without language.

  • In the streaming test, stopping audio frames did not advance the observed silence timer; continuing to send digital silence allowed the turn to close.

  • cannot call recv while another coroutine is already running recv means two coroutines read one socket. Give each connection a single reader.

  • In this Windows setup, the input-path audio processing chopped quiet syllables. Turning it off or using exclusive capture fixed the input.

If none of those cases fit, compare the raw events with the source audio to isolate the cause.

Grok Voice Transcribe 2.0 Pricing

SpaceXAI's pricing page lists transcription at $0.10 per hour over REST and $0.20 per hour streaming. The announcement says diarization, timestamps, and keyterms are included. Calculate cost from audio duration rather than request count.

Each open stream bills its own audio duration. A second listener adds streaming cost and doubles the STT minutes only when both streams receive the same full duration.

Final Thoughts

I would not evaluate a call transcriber on clean audio alone. The phone-audio section shows why.

The API returns transcription data; the client still owns conversation state and validation. Also, keep the versioned model ID. Treat the other settings as starting points, then check them with the target audio.

The next extensions are a SIP phone input, per-call vocabulary, and a CRM export. If you want an agent rather than a transcriber, our Grok Voice Agent API tutorial covers that path.

FAQs

Does Grok Voice Transcribe 2.0 support real-time transcription?

Yes, over the WebSocket, and not only as raw PCM. A client with limited bandwidth can stream encoding=opus, about 4 KB/s against 48 KB/s for 24 kHz PCM, as long as each frame carries one Opus packet. Opus is mono only, so it does not support streaming multichannel.

Does Grok Voice Transcribe 2.0 support speaker diarization?

Set diarize=true on either endpoint. In the diarized streaming response for this fixture, words also included an undocumented speaker_confidence field. I would not build application logic around it. Treat speaker IDs as request- or session-local labels, not persistent identity recognition.

Can Grok Voice Transcribe 2.0 transcribe multiple languages in one recording?

Automatic detection can preserve a mid-recording language switch without a hint. The language parameter controls formatting for 25 listed languages, including Arabic (ar), so test the relevant code against your own audio before relying on formatted output.

What is the difference between Smart Turn and VAD?

VAD asks whether audio is speech; Smart Turn asks whether the speech is finished. vad_threshold defaults to 0.5 in batch and 0.08 on the stream. endpointing defaults to 400 milliseconds and sets the silence needed before an utterance can close.

Can I transcribe a recording from a URL instead of uploading a file?

Use the batch endpoint's url field instead of file. SpaceXAI downloads the recording server-side, and a failed download returns a 502.


Khalid Abdelaty's photo
Author
Khalid Abdelaty
LinkedIn

I’m a data engineer and community builder who works across data pipelines, cloud, and AI tooling while writing practical, high-impact tutorials for DataCamp and emerging developers.

Topics
Artificial Intelligence
AI Agents

Learn AI With DataCamp!

Track

Associate AI Engineer for Developers

29 hr
Learn how to integrate AI into software applications using APIs and open-source libraries. Start your journey to becoming an AI Engineer today!
See DetailsRight Arrow
Start Course
See MoreRight Arrow
Related

blog

Grok Voice Transcribe 2.0: Accuracy, Features, Pricing, and API Access

SpaceXAI's new speech-to-text model ranks first for streaming accuracy on the public leaderboard at an unchanged hourly price. We cover what changed from 1.0, the benchmarks, pricing, and how to call it.
Tom Farnschläder's photo

Tom Farnschläder

9 min

Tutorial

Grok Voice Think Fast 2.0 API Tutorial: Build a Real-Time Voice Agent in Python

Learn how to use Grok Voice Think Fast 2.0 to build a real-time voice agent that handles spoken conversations, calls tools, manages interruptions, and resumes disconnected sessions.
Khalid Abdelaty's photo

Khalid Abdelaty

15 min

Tutorial

Grok Voice Agent Builder: A Hands-On Guide in Python

Build a Python voice agent with the same API used by Grok Voice Agent Builder: WebSocket setup, audio streaming, tool calling, cost tracking, and a FastAPI endpoint.
Khalid Abdelaty's photo

Khalid Abdelaty

11 min

Tutorial

GPT Live Transcribe API Tutorial: Build Real-Time Captions in Python

Learn how to use OpenAI’s gpt-live-transcribe API to stream microphone audio, generate multilingual live captions, improve domain-specific accuracy, and balance latency against transcription quality.
Khalid Abdelaty's photo

Khalid Abdelaty

15 min

Tutorial

Grok Imagine API: A Complete Python Guide With Examples

Learn how to generate videos using the Grok Imagine API. This Python guide covers everything from image animations to video editing with the new xAI video model.
François Aubry's photo

François Aubry

8 min

Tutorial

GPT-Live-1 API Tutorial: Build a Full-Duplex Voice Assistant

Follow this GPT-Live-1 API tutorial to build a full-duplex voice learning assistant with browser WebRTC, backend delegation, web search, and confirmed actions.
Khalid Abdelaty's photo

Khalid Abdelaty

15 min

See MoreSee More