MOSS-TTSD
LatestSpeech generationMOSS-TTSD (text to spoken dialogue) is an open-source bilingual (Chinese-English) spoken dialogue synthesis model that turns a two-speaker dialogue script into natural, expressive conversational speech. It supports zero-shot two-speaker voice cloning and long single-pass speech generation, making it ideal for podcasts, interviews, chat, and other dialogue scenarios.
ENDPOINT
POST /v1/audio/speech/speakers
Current request path
Input
Text
Content accepted by the request
Output
Audio
Content returned on success
Model overview
Multi-speaker dialogue speech synthesis model. Use the stable model ID moss-ttsd-1.0 by default; to reproduce a fixed version, pass moss-ttsd-1.0-2026-03-20 in model.
Billing
This page does not include placeholder pricing. Refer to console pricing and bills for actual units and rates.
See pricingUsage is measured from actual requests and outputs. Refer to console bills for the final amount.Modality support
Supported endpoints
/v1/audio/speechNot supported/v1/audio/speech/speakersSupported/v1/audio/transcriptionsNot supported/v1/audio/voice/generationsNot supported/v1/responsesNot supportedCapabilities
- Up to 40 minutes of audio per generation
- Zero-shot two-speaker cloning
Versions and snapshots
moss-ttsd-1.0moss-ttsd-1.0-2026-03-20Latest