Console

Model ListMOSS-Transcribe-Diarize-Pro

MOSS-Transcribe-Diarize-Pro

RecommendedSpeech recognition

MOSS-Transcribe-Diarize-Pro outputs high-accuracy transcripts together with speaker identities and timestamps, built for multi-speaker scenarios such as meetings, interviews, podcasts, and film or TV content.

ENDPOINT

POST /v1/audio/transcriptions

Current request path

Input

Audio

Content accepted by the request

Output

Text

Content returned on success

Model overview

A multi-speaker transcription model supporting audio up to 60 minutes long. Use the model ID moss-transcribe-diarize-pro; the server always provides the latest version.

Billing

Billing methodCNY priceActual deduction
By input audio or video duration¥2.00 / hour40 credits / hour

Each request is billed by actual usage. CNY prices are for reference; the final charge follows the credits bill. View full pricing

Modality support

TTextOutput only
AAudioInput only
IImageNot supported
VVideoNot supported

Supported endpoints

Single-speaker speech/v1/audio/speechNot supported
Multi-speaker dialogue speech/v1/audio/speech/speakersNot supported
Audio transcription/v1/audio/transcriptionsSupported
Voice design/v1/audio/voice/generationsNot supported
Multimodal understanding/v1/responsesNot supported

Capabilities

PerformanceHigh
Inference speedMedium
Streaming Supported
Async Supported
  • Supports audio up to 60 minutes
  • Extra-long context — 32k
  • Up to 16k output length
  • Robust handling of noise and overlapping speech
  • Keyterm boosting for domain-specific terms, names, brands, and business terminology

Versions and snapshots

Model IDmoss-transcribe-diarize-pro