Console

Model ListMOSS-VL-Realtime

MOSS-VL-Realtime

BetaMultimodal understanding

MOSS-VL-Realtime continuously receives JPEG, PNG, or WebP image frames over one WebSocket connection and returns incremental text answers for ongoing camera or screen understanding.

ENDPOINT

WebSocket /v1/realtime

Current request path

Input

Continuous image frames + text

Content accepted by the request

Output

Incremental text

Content returned on success

Model overview

Realtime video understanding model, Beta. Use moss-vl-realtime-1.0 in the model connection query parameter; to pin a version, use moss-vl-realtime-1.0-2026-09-09.

Billing

Free for a limited time.

0 credits View full pricing

Modality support

TTextInput and output
AAudioNot supported
IImageInput only
VVideoNot supported

Supported endpoints

Single-speaker speech/v1/audio/speechNot supported
Multi-speaker dialogue speech/v1/audio/speech/speakersNot supported
Audio transcription/v1/audio/transcriptionsNot supported
Voice design/v1/audio/voice/generationsNot supported
Multimodal understanding/v1/responsesNot supported
Realtime video understanding/v1/realtimeSupported

Capabilities

PerformanceHigh
Inference speedRealtime interaction
Streaming Supported
Async Not supported
  • Bidirectional WebSocket event protocol
  • JPEG, PNG, and WebP image frames
  • Silence, interruption, and multiple answer segments
  • Session-level token usage reporting
  • No complete video files, audio input, or speech output

Versions and snapshots

Model IDmoss-vl-realtime-1.0
Snapshotmoss-vl-realtime-1.0-2026-09-09Latest