Console

Model ListMOSS-VL

MOSS-VL

LatestMultimodal understanding

MOSS-VL accepts text instructions with images, video, or multiple images for image understanding, video understanding, OCR, document parsing, visual question answering, and cross-media content analysis. Each request supports up to 5 images or 1 video; images and video cannot be mixed.

ENDPOINT

POST /v1/responses

Current request path

Input

Text + images / video

Content accepted by the request

Output

Text

Content returned on success

Model overview

Multimodal understanding model. Use the stable model ID moss-vl-1.0 by default; to reproduce a fixed version, pass moss-vl-1.0-2026-07-08 in model.

Billing

This page does not include placeholder pricing. Refer to console pricing and bills for actual units and rates.

ItemUnit priceDescription
Model invocationSee pricingUsage is measured from actual requests and outputs. Refer to console bills for the final amount.

Modality support

TTextInput and output
AAudioNot supported
IImageInput only
VVideoInput only

Supported endpoints

Single-speaker speech/v1/audio/speechNot supported
Multi-speaker dialogue speech/v1/audio/speech/speakersNot supported
Audio transcription/v1/audio/transcriptionsNot supported
Voice design/v1/audio/voice/generationsNot supported
Multimodal understanding/v1/responsesSupported

Capabilities

PerformanceHigh
Inference speedMedium
Streaming Not supported
Async Not supported
  • Image description, OCR, and document content understanding
  • Video summaries, event understanding, and question answering
  • Up to 5 images or 1 video per request
  • Supports file ID and URL input

Versions and snapshots

Model IDmoss-vl-1.0
Snapshotmoss-vl-1.0-2026-07-08Latest