Realtime video understanding
Use MOSS-VL-Realtime to continuously send JPEG, PNG, or WebP frames over one WebSocket connection, ask questions in the same session, and receive incremental text answers for camera or screen understanding.
/v1/realtimemoss-vl-realtime-1.0Create an API Key
Create a key in API Key management and use it only in server-side environment variables.
Environment variable MOSS_API_KEY=<your_api_key_here>Connect and authenticate
Connect with a server-side WebSocket client. Specify the model in the query parameters and pass the API Key in the handshake request header.
FieldTypeRequirementDescriptionModel ID. Use moss-vl-realtime-1.0, or moss-vl-realtime-1.0-2026-09-09 to pin a snapshot. Pass it in the query parameters of the WebSocket URL.
Bearer
<API Key>; it must be sent in the WebSocket handshake request header.Connection URL wss://api.mosi.cn/v1/realtime?model=moss-vl-realtime-1.0 Authorization: Bearer $MOSS_API_KEYConfigure the session
After connecting, receive
session.created, then sendsession.configureonce. Wait forsession.configuredandsession.readybefore sending frames or questions.FieldTypeRequirementDescriptionMust be session.configure.
Initial user prompt. Defaults to an empty string.
Generation budget reset after each new input. Defaults to 4096; it is not the output limit for the entire session.
Whether to emit session.usage observation events. Defaults to false; this does not affect final usage reporting.
session.configure { "type": "session.configure", "prompt": "Keep observing the scene and answer my questions.", "max_new_tokens": 128, "include_usage": false }Send realtime frames
Convert camera or screen content into image frames. Start with one frame to verify the connection, then repeat the same process for subsequent frames when continuous observation is needed. For each frame, send
input.framemetadata, send image binary data afterinput.frame.ready, and wait forinput.frame.acceptedbefore sending the next frame. Frames alone do not generate answers; send questions separately in the next step.FieldTypeRequirementDescriptionMust be input.frame.
Frames and questions share a continuous sequence starting at 0.
Frame capture time in seconds; it must not precede the timestamp of a previously accepted frame.
Use image/jpeg, image/png, or image/webp, matching the actual encoding.
input.frame · JSON metadata { "type": "input.frame", "seq_no": 0, "timestamp": 0.0, "mime_type": "image/jpeg" }After receiving
input.frame.readyfor the frame, send the complete binary bytes of the JPEG file. Do not put the image in JSON, Base64, or multipart. Once the image is accepted, send another frame or proceed to the next step to ask a question.Send a question
After sending the frame, ask a question with a separate
input.prompt. You do not need to attach the image again; the server generates an answer based on the realtime frames it has received. A question withfinal=falsedoes not trigger an answer by itself: keep sending frames as in the previous step, and the answer appears after a later frame arrives.FieldTypeRequirementDescriptionMust be input.prompt.
Use the next sequence number after the frame.
Non-empty question text.
Defaults to false. Set to true on the last question to end the session.
input.prompt { "type": "input.prompt", "seq_no": 1, "prompt": "Please describe the details in the scene.", "final": false }After asking, keep sending frames with the next consecutive
seq_no, following the same flow as the previous step. This example sends one more frame after the question; the model usually starts answering after that frame is processed. If it stays silent, keep sending frames.input.frame · next frame after the question { "type": "input.frame", "seq_no": 2, "timestamp": 1.0, "mime_type": "image/jpeg" }Receive and display answers
Continuously receive server events and concatenate the
deltafromresponse.text.deltain order. Useresponse_idto distinguish answer segments. Process text, errors, and close events even while waiting for input confirmation.Server events {"type":"response.text.delta","delta":"There is a car in the scene.","response_id":"resp_xxx","response_seq":0} {"type":"response.done","response_id":"resp_xxx","response_seq":0,"finish_reason":"stop"}End the session
To keep asking questions, continue with the next sequence number. When ready to finish, set
final=trueon the lastinput.promptand keep receiving events untilsession.done. To stop early, sendsession.abortinstead. Choose one of these two ways to end the session.Final input { "type": "input.prompt", "seq_no": 3, "prompt": "Please summarize what you just saw.", "final": true }Stop early { "type": "session.abort" }Verify success
For the question in this example, confirm that you received and concatenated non-empty text after sending the follow-up frame, followed by
session.donewithreason=completed. The model may remain silent for some frames or questions, so an absence of text does not always mean the call failed; if no text arrives at all, first check that you kept sending frames after the question.
Next steps
View API: WSS /v1/realtime
Continuously send image frames and questions over WebSocket and receive incremental text events. This is a visual understanding protocol; it is incompatible with audio Realtime and is not SSE.
View model overview
View model IDs, capabilities, and use cases.