feat(xai): configurable TTS output sample_rate#6509
Open
randyyu13 wants to merge 1 commit into
Open
Conversation
The xAI TTS plugin hardcoded the output to 24000 Hz PCM, so pipelines that run at a different rate (e.g. 16000 Hz for telephony) always incur an extra resampling stage. Add a sample_rate constructor option (default 24000, one of the xAI-supported rates) that is threaded into the websocket request and the audio emitter, letting callers match their pipeline rate and avoid the resample. Codec stays PCM (LiveKit-native). Validated against the supported-rate set with a clear error.
|
Randy Yu seems not to be a GitHub user. You need a GitHub account to be able to sign the CLA. If you have already a GitHub account, please add the email address used for this commit to your account. You have signed the CLA already but the status is still pending? Let us recheck it. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The xAI TTS plugin hardcodes its output to 24000 Hz PCM (
SAMPLE_RATEin_connect_wsand the audio emitter). This PR adds asample_rateconstructor option so callers can pick any xAI-supported rate.Why
When the agent pipeline runs at a different rate than 24 kHz — e.g. 16 kHz for telephony — the fixed 24 kHz output forces an extra resampling stage (24k → pipeline rate), which adds latency and can introduce artifacts. Matching the TTS output to the pipeline rate avoids that stage entirely. The xAI API already supports configurable output sample rate; the plugin just wasn't exposing it.
Changes
sample_rate: int = 24000added toTTS.__init__and_TTSOptions, threaded intosuper().__init__(...), the websocket request params, and theAudioEmitterinit.(8000, 16000, 22050, 24000, 44100, 48000)with a clearValueError.pcm(LiveKit-native); this PR is scoped to sample rate only.Tests
Added to
tests/test_plugin_xai_tts.py:ValueErrorruff check+ruff formatclean.