MOSS-VL
LatestMultimodal understandingMOSS-VL accepts text instructions with images, video, or multiple images for image understanding, video understanding, OCR, document parsing, visual question answering, and cross-media content analysis. Each request supports up to 5 images or 1 video; images and video cannot be mixed.
ENDPOINT
POST /v1/responses
Current request path
Input
Text + images / video
Content accepted by the request
Output
Text
Content returned on success
Model overview
Multimodal understanding model. Use the stable model ID moss-vl-1.0 by default; to reproduce a fixed version, pass moss-vl-1.0-2026-07-08 in model.
Billing
This page does not include placeholder pricing. Refer to console pricing and bills for actual units and rates.
ItemUnit priceDescription
Model invocation
See pricingUsage is measured from actual requests and outputs. Refer to console bills for the final amount.Modality support
TTextInput and output
AAudioNot supported
IImageInput only
VVideoInput only
Supported endpoints
Single-speaker speech
/v1/audio/speechNot supportedMulti-speaker dialogue speech
/v1/audio/speech/speakersNot supportedAudio transcription
/v1/audio/transcriptionsNot supportedVoice design
/v1/audio/voice/generationsNot supportedMultimodal understanding
/v1/responsesSupportedCapabilities
PerformanceHigh
Inference speedMedium
Streaming Not supported
Async Not supported
- Image description, OCR, and document content understanding
- Video summaries, event understanding, and question answering
- Up to 5 images or 1 video per request
- Supports file ID and URL input
Versions and snapshots
Model ID
moss-vl-1.0Snapshot
moss-vl-1.0-2026-07-08Latest