MOSS-VL-Realtime
BetaMultimodal understandingMOSS-VL-Realtime continuously receives JPEG, PNG, or WebP image frames over one WebSocket connection and returns incremental text answers for ongoing camera or screen understanding.
ENDPOINT
WebSocket /v1/realtime
Current request path
Input
Continuous image frames + text
Content accepted by the request
Output
Incremental text
Content returned on success
Model overview
Realtime video understanding model, Beta. Use moss-vl-realtime-1.0 in the model connection query parameter; to pin a version, use moss-vl-realtime-1.0-2026-09-09.
Billing
Free for a limited time.
0 credits View full pricing
Modality support
TTextInput and output
AAudioNot supported
IImageInput only
VVideoNot supported
Supported endpoints
Single-speaker speech
/v1/audio/speechNot supportedMulti-speaker dialogue speech
/v1/audio/speech/speakersNot supportedAudio transcription
/v1/audio/transcriptionsNot supportedVoice design
/v1/audio/voice/generationsNot supportedMultimodal understanding
/v1/responsesNot supportedRealtime video understanding
/v1/realtimeSupportedCapabilities
PerformanceHigh
Inference speedRealtime interaction
Streaming Supported
Async Not supported
- Bidirectional WebSocket event protocol
- JPEG, PNG, and WebP image frames
- Silence, interruption, and multiple answer segments
- Session-level token usage reporting
- No complete video files, audio input, or speech output
Versions and snapshots
Model ID
moss-vl-realtime-1.0Snapshot
moss-vl-realtime-1.0-2026-09-09Latest