Console

Create Response

POST/v1/responses

Use MOSS-VL to understand images or videos and return text results as a Responses API object. Each request must include both a text instruction and media.

Request parameters

FieldTypeRequiredDescription
model: stringRequired

Model ID. Use the stable model ID for regular calls, or a snapshot model ID when you need to reproduce results from a fixed version.

input: arrayRequired

The input message array. It must currently contain exactly one message with role=user.

input[].role: stringRequired

Currently fixed to user.

input[].content: arrayRequired

The content array must include at least one non-empty input_text part and either 1–5 input_image parts or 1 input_video part. Images and video cannot be mixed in one request.

input[].content[].type: stringRequired

Content type.

input_text.text: stringRequired

The image or video understanding instruction. It must be a non-empty string.

input_image.image_url: stringConditional

An image URL reachable by the service; mutually exclusive with file_id in the same content item. Each image can be up to 30 MB. Supported formats: PNG, JPG, JPEG, WebP, BMP, GIF, and TIFF.

input_image.file_id: stringConditional

The ID of an uploaded image; mutually exclusive with image_url in the same content item. Each image can be up to 30 MB. Supported formats: PNG, JPG, JPEG, WebP, BMP, GIF, and TIFF.

input_video.video_url: stringConditional

A video URL reachable by the service; mutually exclusive with file_id in the same content item. Each video can be up to 200 MB, with no current duration limit. Supported formats: MP4, M4V, AVI, MOV, WebM, and MKV.

input_video.file_id: stringConditional

The ID of an uploaded video; mutually exclusive with video_url in the same content item. Each video can be up to 200 MB, with no current duration limit. Supported formats: MP4, M4V, AVI, MOV, WebM, and MKV.

max_output_tokens: integerOptional

The maximum number of output tokens, from 1 to 8192.

Return value

The endpoint returns a Responses API object. Model-generated text is available at output[].content[].text; use status to determine whether the output is complete.

FieldTypeDescription
status: string

The response status, currently completed or incomplete.

incomplete_details.reason: string | null

The reason for incomplete output; max_output_tokens when the output limit is reached.

output: array

The output item list.

output[].content: array

The output content list.

output[].content[].type: string

Currently output_text for text results.

output[].content[].text: string

The text generated by the model.

usage.input_tokens: integer

The input tokens used for this request.

usage.output_tokens: integer

The output tokens generated for this request.

usage.total_tokens: integer

The total input and output token count.

The response also contains common metadata such as the ID, object type, timestamps, and model. See the response example on the right for the full structure. Clients should ignore unused additional fields; usage only reports the token usage for this request.

Error codes

HTTPerror.codeTriggerSuggested action
400missing_required_fieldA text instruction or media content is missingEnsure input[].content contains both non-empty input_text and valid media.
400media_count_exceededMore than 5 images or more than 1 video is providedKeep each request to 1–5 images or 1 video.
400mixed_media_not_supportedImages and video are provided in the same requestSend image understanding and video understanding as separate requests.
400invalid_media_sourceA media item has both a URL and file_id, or neitherSet only the matching URL field or file_id on each media item.
400unsupported_media_formatAn uploaded or referenced media format is unsupportedUse a documented media format and keep the file within the size limit.