Image and Video Understanding
Use MOSS-VL to understand images or video and return text in a Responses API object. Every request must include both a text instruction and media.
/v1/responsesmoss-vl-1.0Create an API Key
Open the API Key Management Platform, create an API Key, and save it as an environment variable. Do not write it into client-side code or commit it to a repository.
Environment variable MOSS_API_KEY=<your_api_key_here>Prepare the request input
Prepare a non-empty text instruction and select an image or video
file_id, or a public or signed URL accessible to the service. Images and videos cannot be mixed in one request.FieldTypeRequirementDescriptionmoss-vl-1.0; use moss-vl-1.0-2026-07-08 to pin a snapshot
Must contain one user message whose content includes both input_text and valid media
Controls the maximum number of output tokens
Request body { "model": "moss-vl-1.0", "input": [ { "role": "user", "content": [ { "type": "input_text", "text": "Describe the main content of this image." }, { "type": "input_image", "file_id": "<image_file_id>" } ] } ], "max_output_tokens": 1024 }Send your first request
Send a JSON request to
/v1/responses. The endpoint synchronously returns a Responses API object with text output.请求示例 curl https://api.mosi.cn/v1/responses \ -H "Authorization: Bearer $MOSS_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "moss-vl-1.0", "input": [ { "role": "user", "content": [ { "type": "input_text", "text": "Describe the main content of this image." }, { "type": "input_image", "file_id": "<image_file_id>" } ] } ], "max_output_tokens": 1024 }'Read the response
After the request completes, read
output[].content[].text. If the response returnsstatus=incomplete, the output may have been truncated bymax_output_tokens.Core result fields { "status": "completed", "output": [ { "type": "message", "role": "assistant", "content": [ { "type": "output_text", "text": "There are four nuts in the image." } ] } ] }Verify success
Success means the response status is
completedandoutput[].content[].textcontains text generated by the model.