MOSS-TTS-v1.5-Flash
RecommendedSpeech generationMOSS-TTS-v1.5-Flash is designed for low-latency streaming speech generation. It provides high-fidelity voice cloning and speech control while continuously returning audio chunks with stream=true.
ENDPOINT
POST /v1/audio/speech
Current request path
Input
Text + voice_id
Content accepted by the request
Output
Audio
Content returned on success
Model overview
Single-speaker text-to-speech model. Use the stable model ID moss-tts-1.5-flash by default; to reproduce a fixed version, pass moss-tts-1.5-flash-2026-06-26 in model.
Billing
This page does not include placeholder pricing. Refer to console pricing and bills for actual units and rates.
ItemUnit priceDescription
Model invocation
See pricingUsage is measured from actual requests and outputs. Refer to console bills for the final amount.Modality support
TTextInput only
AAudioOutput only
IImageNot supported
VVideoNot supported
Supported endpoints
Single-speaker speech
/v1/audio/speechSupportedMulti-speaker dialogue speech
/v1/audio/speech/speakersNot supportedAudio transcription
/v1/audio/transcriptionsNot supportedVoice design
/v1/audio/voice/generationsNot supportedMultimodal understanding
/v1/responsesNot supportedCapabilities
PerformanceHigher
Inference speedMedium
Streaming Supported
Async Supported
- 32,768-token context window
- Up to 7,500 output tokens
- Cantonese supported
- 30+ languages supported
- Phoneme-level pronunciation control & token-level duration control
Versions and snapshots
Model ID
moss-tts-1.5-flashSnapshot
moss-tts-1.5-flash-2026-06-26Streaming version