Direct
Direct upstream connection — best when you need native behavior and the full context window.
| Input context | Input | Output | Cache read |
|---|---|---|---|
| ≤ 32K | 0.74/M | 3.24/M | 0.18/M |
| > 32K | 1.03/M | 3.82/M | 0.26/M |

GLM-5V-TurboGLM-5V-Turbo is Zhipu's first multimodal Agent foundation model, deeply optimized for visual programming and complex task scenarios. It supports multimodal inputs including images, videos, text, and files, with enhanced visual understanding, long-horizon planning, and action execution capabilities. Compared with general-purpose multimodal models, it is better suited for integration into Agent workflows, completing the full closed loop of “environment perception → task planning → execution,” enabling multimodal capabilities to move from “being able to understand” to “being able to act.”
The same model is available through multiple service channels — choose based on latency, reliability and cost.
Prices in $ / 1M tokensprovider field to the request body, for example "provider": { "channel": "direct" }. Valid values are direct / stable / economical; omit it to use the default channel.Direct upstream connection — best when you need native behavior and the full context window.
| Input context | Input | Output | Cache read |
|---|---|---|---|
| ≤ 32K | 0.74/M | 3.24/M | 0.18/M |
| > 32K | 1.03/M | 3.82/M | 0.26/M |
GLM-5V-Turbo is Zhipu's first multimodal Agent foundation model, deeply optimized for visual programming and complex task scenarios. It supports multimodal inputs including images, videos, text, and files, with enhanced visual understanding, long-horizon planning, and action execution capabilities.
Compared with general-purpose multimodal models, GLM-5V-Turbo is better suited for integration into Agent workflows, completing the full loop of "environment perception → task planning → action execution" — taking multimodal capabilities from "being able to understand" to "being able to get things done."
SeaWhale AI provides GLM-5V-Turbo through an OpenAI-compatible interface, supporting multimodal inputs, tool calling, and streaming output.
Get API Key · Model ID:
GLM-5V-Turbo
GLM-5V-Turbo is deeply optimized for visual programming scenarios: writing frontend from design mockups, locating UI issues from screenshots, and generating data processing code from charts. This is the most direct engineering value of multimodal capabilities.
As an Agent foundation model, its strength lies in converting visual understanding into concrete actions: recognizing UI elements, assessing the current state, and deciding the next step.
In visual tasks that require multiple steps, GLM-5V-Turbo maintains planning toward the overall goal rather than restarting judgment at each step.
Images, videos, text, and files are processed uniformly in the same context, making it suitable for scenarios that require cross-modal associative reasoning.
| Scenario | Description |
|---|---|
| Visual programming | Design mockup to frontend code, screenshot issue localization |
| UI automation | Vision-based UI operation and testing |
| Multimodal Agent | Agents that need to understand the environment before deciding and executing |
| Video content analysis | Temporal understanding and content extraction |
| Document understanding | Complete parsing of mixed image-text documents |
| Quality inspection and patrol inspection | Image-based judgment and subsequent action triggering |
| Capability | GLM-5V-Turbo | GLM-5 Turbo | Gemini 3.5 Flash |
|---|---|---|---|
| Model ID | GLM-5V-Turbo |
GLM-5-Turbo |
gemini-3.5-flash |
| Positioning | Multimodal Agent foundation | Fast tier for text-only agents | Efficient multimodal workhorse |
| Input modalities | Images, videos, text, files | Text | Text, images, videos, audio, PDF |
| Context window | 131K token | 131K token | 105K token |
| Max output | 32K token | 131K token | 64K token |
| Focus | Visual programming and action execution | Long execution chain stability | Coding and parallel agents |
For specific pricing, refer to the real-time price card at the top of the page.
1. Create a SeaWhale AI API Key
Generate a key in the console and add credits.
2. Treat vision as part of the context
GLM-5V-Turbo's strength is not in simple image description, but in "what to do after seeing." Clearly stating the target action in the prompt yields the best results.
3. Call the API
curl -X POST https://api.seawhaleai.com/v1/chat/completions \
-H 'Authorization: Bearer YOUR_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
"model": "GLM-5V-Turbo",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://example.com/design.png"}},
{"type": "text", "text": "Implement the corresponding Vue component based on this design mockup, paying attention to spacing and font sizes."}
]
}],
"stream": true
}'
How is it different from regular multimodal models?
Regular multimodal models aim to "understand," while GLM-5V-Turbo aims to "get things done after understanding" — it integrates visual understanding, long-horizon planning, and action execution into a closed loop, making it better suited for integration into Agent workflows.
What input modalities are supported?
Images, videos, text, and files. The output is text.
What exactly does the visual programming capability mean?
It includes converting design mockups to code, locating issues from UI screenshots, and code-based processing of chart data — directly converting visual information into actionable engineering output.
What are the context and output limits?
131,000-token context, with a maximum output of 32,000 tokens.
Is it suitable for UI automation?
Yes. It is optimized for the "environment perception → task planning → action execution" loop, making it a suitable choice for vision-based automation scenarios.
Can it process videos?
Yes. It supports video input and temporal understanding, suitable for content analysis scenarios.
GLM-5V-Turbohttps://api.seawhaleai.com/v2/"provider": { "channel": "direct" }SeaWhale AI is compatible with the OpenAI API protocol, so you can call it with the OpenAI SDK or plain HTTP requests. Streaming is enabled by default.
About the provider parameter (optional, a SeaWhale AI extension): most models are served over several channels that differ slightly in price and reliability. Add a provider field to the request body to pick one; omit it and the system selects the default channel — normal calls are unaffected.
provideris not part of the official OpenAI protocol — it is a SeaWhale AI extension that only takes effect on this platform. The OpenAI SDK allows custom fields like this to pass through; see the examples below.
| Value | Channel | Best for |
|---|---|---|
direct | Direct | The official upstream link, for native behavior and the full context window |
stable | Preferred | Balanced availability and speed — a good default for production traffic |
economical | Economy | Cost first, well suited to batch processing and price-sensitive workloads |
Available channels and their prices are listed under "Pricing" above (channels vary by model). Additional notes:
"provider": { "channel": "direct" }.extra_body; in Node.js put it directly on the request object and it passes through. In TypeScript projects, add a // @ts-expect-error line to skip the type check.curl https://api.seawhaleai.com/v2/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <API_KEY>" \
-d '{
"model": "GLM-5V-Turbo",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"provider": { "channel": "direct" },
"stream": true
}'
# provider is optional — remove this line to use the default channelfrom openai import OpenAI
client = OpenAI(
base_url="https://api.seawhaleai.com/v2",
api_key="<API_KEY>",
)
stream = client.chat.completions.create(
model="GLM-5V-Turbo",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"},
],
stream=True,
# Optional: pick a service channel; omit to use the default
extra_body={"provider": {"channel": "direct"}},
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)import OpenAI from 'openai'
const client = new OpenAI({
baseURL: 'https://api.seawhaleai.com/v2',
apiKey: '<API_KEY>',
})
const stream = await client.chat.completions.create({
model: 'GLM-5V-Turbo',
messages: [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'Hello!' },
],
stream: true,
// Optional: pick a service channel; omit to use the default
// @ts-expect-error provider is a SeaWhale AI extension, not in the OpenAI SDK types
provider: { channel: 'direct' },
})
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}