Console

Image and Video Understanding

Use MOSS-VL to understand images or video and return text in a Responses API object. Every request must include both a text instruction and media.

Endpoint
POST/v1/responses
Model name
moss-vl-1.0
  1. Create an API Key

    Open the API Key Management Platform, create an API Key, and save it as an environment variable. Do not write it into client-side code or commit it to a repository.

    Environment variable
    MOSS_API_KEY=<your_api_key_here>
  2. Prepare the request input

    Prepare a non-empty text instruction and select an image or video file_id, or a public or signed URL accessible to the service. Images and videos cannot be mixed in one request.

    FieldTypeRequirementDescription
    model: stringRequired

    moss-vl-1.0; use moss-vl-1.0-2026-07-08 to pin a snapshot

    input: arrayRequired

    Must contain one user message whose content includes both input_text and valid media

    max_output_tokens: integerOptional

    Controls the maximum number of output tokens

    Request body
    {
      "model": "moss-vl-1.0",
      "input": [
        {
          "role": "user",
          "content": [
            { "type": "input_text", "text": "Describe the main content of this image." },
            { "type": "input_image", "file_id": "<image_file_id>" }
          ]
        }
      ],
      "max_output_tokens": 1024
    }
  3. Send your first request

    Send a JSON request to /v1/responses. The endpoint synchronously returns a Responses API object with text output.

    请求示例
    curl https://api.mosi.cn/v1/responses \
      -H "Authorization: Bearer $MOSS_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
      "model": "moss-vl-1.0",
      "input": [
        {
          "role": "user",
          "content": [
            { "type": "input_text", "text": "Describe the main content of this image." },
            { "type": "input_image", "file_id": "<image_file_id>" }
          ]
        }
      ],
      "max_output_tokens": 1024
    }'
  4. Read the response

    After the request completes, read output[].content[].text. If the response returns status=incomplete, the output may have been truncated by max_output_tokens.

    Core result fields
    {
      "status": "completed",
      "output": [
        {
          "type": "message",
          "role": "assistant",
          "content": [
            {
              "type": "output_text",
              "text": "There are four nuts in the image."
            }
          ]
        }
      ]
    }
  5. Verify success

    Success means the response status is completed and output[].content[].text contains text generated by the model.

Next steps