MOSS-Transcribe
LatestSpeech recognitionMOSS-Transcribe-1.0 delivers high-accuracy transcription across speakers of different ages and voice characteristics — from children to the elderly. It fits meetings, interviews, podcasts, and film or TV content, producing clear and accurate transcripts even in complex acoustic environments.
ENDPOINT
POST /v1/audio/transcriptions
Current request path
Input
Audio
Content accepted by the request
Output
Text
Content returned on success
Model overview
Standard audio transcription model. Use the stable model ID moss-transcribe-1.0 by default; to reproduce a fixed version, pass moss-transcribe-1.0-2026-07-22 in model.
Billing
This page does not include placeholder pricing. Refer to console pricing and bills for actual units and rates.
ItemUnit priceDescription
Model invocation
See pricingUsage is measured from actual requests and outputs. Refer to console bills for the final amount.Modality support
TTextOutput only
AAudioInput only
IImageNot supported
VVideoNot supported
Supported endpoints
Single-speaker speech
/v1/audio/speechNot supportedMulti-speaker dialogue speech
/v1/audio/speech/speakersNot supportedAudio transcription
/v1/audio/transcriptionsSupportedVoice design
/v1/audio/voice/generationsNot supportedMultimodal understanding
/v1/responsesNot supportedCapabilities
PerformanceHigh
Inference speedMedium
Streaming Not supported
Async Supported
- High-accuracy transcription with high-quality text output
- Adapts to speakers with different voices, tones, and pronunciation
- Covers meetings, interviews, podcasts, film, and TV scenarios
Versions and snapshots
Model ID
moss-transcribe-1.0Snapshot
moss-transcribe-1.0-2026-07-22Latest