MOSS-TTS-v1.0-Pro
LegacySpeech generationMOSS-TTS is our latest flagship text-to-speech model. It delivers high-fidelity voice cloning that faithfully reproduces a speaker's identity, tone, and prosody, producing natural, vivid, and expressive speech.
ENDPOINT
POST /v1/audio/speech
Current request path
Input
Text + voice_id
Content accepted by the request
Output
Audio
Content returned on success
Model overview
Single-speaker text-to-speech model. Use the stable model ID moss-tts-1.0-pro by default; to reproduce a fixed version, pass moss-tts-1.0-pro-2026-02-07 in model.
Billing
This page does not include placeholder pricing. Refer to console pricing and bills for actual units and rates.
ItemUnit priceDescription
Model invocation
See pricingUsage is measured from actual requests and outputs. Refer to console bills for the final amount.Modality support
TTextInput only
AAudioOutput only
IImageNot supported
VVideoNot supported
Supported endpoints
Single-speaker speech
/v1/audio/speechSupportedMulti-speaker dialogue speech
/v1/audio/speech/speakersNot supportedAudio transcription
/v1/audio/transcriptionsNot supportedVoice design
/v1/audio/voice/generationsNot supportedMultimodal understanding
/v1/responsesNot supportedCapabilities
PerformanceHigh
Inference speedMedium
Streaming Not supported
Async Supported
- 65,536-token context window
- Up to 45,000 output tokens
- 15+ languages supported
- Phoneme-level pronunciation control & token-level duration control
Versions and snapshots
Model ID
moss-tts-1.0-proSnapshot
moss-tts-1.0-pro-2026-02-07Latest