Console

Model ListMOSS-Transcribe

MOSS-Transcribe

LatestSpeech recognition

MOSS-Transcribe-1.0 delivers high-accuracy transcription across speakers of different ages and voice characteristics — from children to the elderly. It fits meetings, interviews, podcasts, and film or TV content, producing clear and accurate transcripts even in complex acoustic environments.

ENDPOINT

POST /v1/audio/transcriptions

Current request path

Input

Audio

Content accepted by the request

Output

Text

Content returned on success

Model overview

Standard audio transcription model. Use the stable model ID moss-transcribe-1.0 by default; to reproduce a fixed version, pass moss-transcribe-1.0-2026-07-22 in model.

Billing

This page does not include placeholder pricing. Refer to console pricing and bills for actual units and rates.

ItemUnit priceDescription
Model invocationSee pricingUsage is measured from actual requests and outputs. Refer to console bills for the final amount.

Modality support

TTextOutput only
AAudioInput only
IImageNot supported
VVideoNot supported

Supported endpoints

Single-speaker speech/v1/audio/speechNot supported
Multi-speaker dialogue speech/v1/audio/speech/speakersNot supported
Audio transcription/v1/audio/transcriptionsSupported
Voice design/v1/audio/voice/generationsNot supported
Multimodal understanding/v1/responsesNot supported

Capabilities

PerformanceHigh
Inference speedMedium
Streaming Not supported
Async Supported
  • High-accuracy transcription with high-quality text output
  • Adapts to speakers with different voices, tones, and pronunciation
  • Covers meetings, interviews, podcasts, film, and TV scenarios

Versions and snapshots

Model IDmoss-transcribe-1.0
Snapshotmoss-transcribe-1.0-2026-07-22Latest