Console

Model ListMOSS-TTSD

MOSS-TTSD

LatestSpeech generation

MOSS-TTSD (text to spoken dialogue) is an open-source bilingual (Chinese-English) spoken dialogue synthesis model that turns a two-speaker dialogue script into natural, expressive conversational speech. It supports zero-shot two-speaker voice cloning and long single-pass speech generation, making it ideal for podcasts, interviews, chat, and other dialogue scenarios.

ENDPOINT

POST /v1/audio/speech/speakers

Current request path

Input

Text

Content accepted by the request

Output

Audio

Content returned on success

Model overview

Multi-speaker dialogue speech synthesis model. Use the stable model ID moss-ttsd-1.0 by default; to reproduce a fixed version, pass moss-ttsd-1.0-2026-03-20 in model.

Billing

This page does not include placeholder pricing. Refer to console pricing and bills for actual units and rates.

ItemUnit priceDescription
Model invocationSee pricingUsage is measured from actual requests and outputs. Refer to console bills for the final amount.

Modality support

TTextInput only
AAudioOutput only
IImageNot supported
VVideoNot supported

Supported endpoints

Single-speaker speech/v1/audio/speechNot supported
Multi-speaker dialogue speech/v1/audio/speech/speakersSupported
Audio transcription/v1/audio/transcriptionsNot supported
Voice design/v1/audio/voice/generationsNot supported
Multimodal understanding/v1/responsesNot supported

Capabilities

PerformanceHigh
Inference speedMedium
Streaming Not supported
Async Supported
  • Up to 40 minutes of audio per generation
  • Zero-shot two-speaker cloning

Versions and snapshots

Model IDmoss-ttsd-1.0
Snapshotmoss-ttsd-1.0-2026-03-20Latest