MOSS-Transcribe-Diarize-Pro
RecommendedSpeech recognitionMOSS-Transcribe-Diarize-Pro outputs high-accuracy transcripts together with speaker identities and timestamps, built for multi-speaker scenarios such as meetings, interviews, podcasts, and film or TV content.
ENDPOINT
POST /v1/audio/transcriptions
Current request path
Input
Audio
Content accepted by the request
Output
Text
Content returned on success
Model overview
A multi-speaker transcription model supporting audio up to 60 minutes long. Use the model ID moss-transcribe-diarize-pro; the server always provides the latest version.
Billing
Billing methodCNY priceActual deduction
By input audio or video duration
¥2.00 / hour40 credits / hourEach request is billed by actual usage. CNY prices are for reference; the final charge follows the credits bill. View full pricing
Modality support
TTextOutput only
AAudioInput only
IImageNot supported
VVideoNot supported
Supported endpoints
Single-speaker speech
/v1/audio/speechNot supportedMulti-speaker dialogue speech
/v1/audio/speech/speakersNot supportedAudio transcription
/v1/audio/transcriptionsSupportedVoice design
/v1/audio/voice/generationsNot supportedMultimodal understanding
/v1/responsesNot supportedCapabilities
PerformanceHigh
Inference speedMedium
Streaming Supported
Async Supported
- Supports audio up to 60 minutes
- Extra-long context — 32k
- Up to 16k output length
- Robust handling of noise and overlapping speech
- Keyterm boosting for domain-specific terms, names, brands, and business terminology
Versions and snapshots
Model ID
moss-transcribe-diarize-pro