Create Response
/v1/responsesUse MOSS-VL to understand images or videos and return text results as a Responses API object. Each request must include both a text instruction and media.
Request parameters
Model ID. Use the stable model ID for regular calls, or a snapshot model ID when you need to reproduce results from a fixed version.
The input message array. It must currently contain exactly one message with role=user.
Currently fixed to user.
The content array must include at least one non-empty input_text part and either 1–5 input_image parts or 1 input_video part. Images and video cannot be mixed in one request.
Content type.
The image or video understanding instruction. It must be a non-empty string.
An image URL reachable by the service; mutually exclusive with file_id in the same content item. Each image can be up to 30 MB. Supported formats: PNG, JPG, JPEG, WebP, BMP, GIF, and TIFF.
The ID of an uploaded image; mutually exclusive with image_url in the same content item. Each image can be up to 30 MB. Supported formats: PNG, JPG, JPEG, WebP, BMP, GIF, and TIFF.
A video URL reachable by the service; mutually exclusive with file_id in the same content item. Each video can be up to 200 MB, with no current duration limit. Supported formats: MP4, M4V, AVI, MOV, WebM, and MKV.
The ID of an uploaded video; mutually exclusive with video_url in the same content item. Each video can be up to 200 MB, with no current duration limit. Supported formats: MP4, M4V, AVI, MOV, WebM, and MKV.
The maximum number of output tokens, from 1 to 8192.
Return value
The endpoint returns a Responses API object. Model-generated text is available at output[].content[].text; use status to determine whether the output is complete.
The response status, currently completed or incomplete.
The reason for incomplete output; max_output_tokens when the output limit is reached.
The output item list.
The output content list.
Currently output_text for text results.
The text generated by the model.
The input tokens used for this request.
The output tokens generated for this request.
The total input and output token count.
The response also contains common metadata such as the ID, object type, timestamps, and model. See the response example on the right for the full structure. Clients should ignore unused additional fields; usage only reports the token usage for this request.
Error codes
| HTTP | error.code | Trigger | Suggested action |
|---|---|---|---|
| 400 | missing_required_field | A text instruction or media content is missing | Ensure input[].content contains both non-empty input_text and valid media. |
| 400 | media_count_exceeded | More than 5 images or more than 1 video is provided | Keep each request to 1–5 images or 1 video. |
| 400 | mixed_media_not_supported | Images and video are provided in the same request | Send image understanding and video understanding as separate requests. |
| 400 | invalid_media_source | A media item has both a URL and file_id, or neither | Set only the matching URL field or file_id on each media item. |
| 400 | unsupported_media_format | An uploaded or referenced media format is unsupported | Use a documented media format and keep the file within the size limit. |