Copy MarkdownModel OverviewMoss API platform models are grouped by speech generation, speech recognition, and multimodal understanding. Click any model card to open its details on the right.Speech GenerationMOSS-TTS-1.5-FlashSingle-speaker text-to-speech modelMOSS-TTS-1.0-ProSingle-speaker text-to-speech modelMOSS-TTSD-1.0Multi-speaker dialogue speech synthesis modelMOSS-Voice-Generator-1.0Generate speech from text and voice instructionsSpeech RecognitionMOSS-Transcribe-1.0General audio transcription modelMOSS-Transcribe-Diarize-ProMulti-speaker transcription modelMultimodal UnderstandingMOSS-VL-1.0Image and video understanding model