Console

Realtime video understanding

Use MOSS-VL-Realtime to continuously send JPEG, PNG, or WebP frames over one WebSocket connection, ask questions in the same session, and receive incremental text answers for camera or screen understanding.

Endpoint
WSS/v1/realtime
Model ID
moss-vl-realtime-1.0
  1. Create an API Key

    Create a key in API Key management and use it only in server-side environment variables.

    Environment variable
    MOSS_API_KEY=<your_api_key_here>
  2. Connect and authenticate

    Connect with a server-side WebSocket client. Specify the model in the query parameters and pass the API Key in the handshake request header.

    FieldTypeRequirementDescription
    model: query stringRequired

    Model ID. Use moss-vl-realtime-1.0, or moss-vl-realtime-1.0-2026-09-09 to pin a snapshot. Pass it in the query parameters of the WebSocket URL.

    Authorization: HTTP headerRequired

    Bearer <API Key>; it must be sent in the WebSocket handshake request header.

    Connection URL
    wss://api.mosi.cn/v1/realtime?model=moss-vl-realtime-1.0
    Authorization: Bearer $MOSS_API_KEY
  3. Configure the session

    After connecting, receive session.created, then send session.configure once. Wait for session.configured and session.ready before sending frames or questions.

    FieldTypeRequirementDescription
    type: stringRequired

    Must be session.configure.

    prompt: stringOptional

    Initial user prompt. Defaults to an empty string.

    max_new_tokens: integerOptional

    Generation budget reset after each new input. Defaults to 4096; it is not the output limit for the entire session.

    include_usage: booleanOptional

    Whether to emit session.usage observation events. Defaults to false; this does not affect final usage reporting.

    session.configure
    {
      "type": "session.configure",
      "prompt": "Keep observing the scene and answer my questions.",
      "max_new_tokens": 128,
      "include_usage": false
    }
  4. Send realtime frames

    Convert camera or screen content into image frames. Start with one frame to verify the connection, then repeat the same process for subsequent frames when continuous observation is needed. For each frame, send input.frame metadata, send image binary data after input.frame.ready, and wait for input.frame.accepted before sending the next frame. Frames alone do not generate answers; send questions separately in the next step.

    FieldTypeRequirementDescription
    type: stringRequired

    Must be input.frame.

    seq_no: integerRequired

    Frames and questions share a continuous sequence starting at 0.

    timestamp: numberRequired

    Frame capture time in seconds; it must not precede the timestamp of a previously accepted frame.

    mime_type: stringRequired

    Use image/jpeg, image/png, or image/webp, matching the actual encoding.

    input.frame · JSON metadata
    {
      "type": "input.frame",
      "seq_no": 0,
      "timestamp": 0.0,
      "mime_type": "image/jpeg"
    }

    After receiving input.frame.ready for the frame, send the complete binary bytes of the JPEG file. Do not put the image in JSON, Base64, or multipart. Once the image is accepted, send another frame or proceed to the next step to ask a question.

  5. Send a question

    After sending the frame, ask a question with a separate input.prompt. You do not need to attach the image again; the server generates an answer based on the realtime frames it has received. A question with final=false does not trigger an answer by itself: keep sending frames as in the previous step, and the answer appears after a later frame arrives.

    FieldTypeRequirementDescription
    type: stringRequired

    Must be input.prompt.

    seq_no: integerRequired

    Use the next sequence number after the frame.

    prompt: stringRequired

    Non-empty question text.

    final: booleanOptional

    Defaults to false. Set to true on the last question to end the session.

    input.prompt
    {
      "type": "input.prompt",
      "seq_no": 1,
      "prompt": "Please describe the details in the scene.",
      "final": false
    }

    After asking, keep sending frames with the next consecutive seq_no, following the same flow as the previous step. This example sends one more frame after the question; the model usually starts answering after that frame is processed. If it stays silent, keep sending frames.

    input.frame · next frame after the question
    {
      "type": "input.frame",
      "seq_no": 2,
      "timestamp": 1.0,
      "mime_type": "image/jpeg"
    }
  6. Receive and display answers

    Continuously receive server events and concatenate the delta from response.text.delta in order. Use response_id to distinguish answer segments. Process text, errors, and close events even while waiting for input confirmation.

    Server events
    {"type":"response.text.delta","delta":"There is a car in the scene.","response_id":"resp_xxx","response_seq":0}
    {"type":"response.done","response_id":"resp_xxx","response_seq":0,"finish_reason":"stop"}
  7. End the session

    To keep asking questions, continue with the next sequence number. When ready to finish, set final=true on the last input.prompt and keep receiving events until session.done. To stop early, send session.abort instead. Choose one of these two ways to end the session.

    Final input
    {
      "type": "input.prompt",
      "seq_no": 3,
      "prompt": "Please summarize what you just saw.",
      "final": true
    }
    Stop early
    {
      "type": "session.abort"
    }
  8. Verify success

    For the question in this example, confirm that you received and concatenated non-empty text after sending the follow-up frame, followed by session.done with reason=completed. The model may remain silent for some frames or questions, so an absence of text does not always mean the call failed; if no text arrives at all, first check that you kept sending frames after the question.

Next steps