Skip to content
Sign in

GLM-5.3 is Zhipu AI's latest flagship model, with comprehensive advancements in complex software engineering and Agent tasks. It uses the same base model as GLM-5.2—all improvements come from post-training. Compared with GLM-5.2, it performs even better on complex programming and long-range tasks.

Context window1.0M
Provider智谱AI
Released2026/09/01

Playground

Pricing

The same model is available through multiple service channels — choose based on latency, reliability and cost.

Prices in $ / 1M tokens
To pick a channel, add a provider field to the request body, for example "provider": { "channel": "direct" }. Valid values are direct / stable / economical; omit it to use the default channel.

Direct

Direct upstream connection — best when you need native behavior and the full context window.

InputOutputCache read
1.25/M4.40/M0.30/M

Overview

Input
Text
Output
Text

GLM-5.3 API: The Open-Weight Model with the Strongest Coding Capabilities

GLM-5.3 is the flagship model released by Z.ai on August 14, 2026, with weights opened on August 28. It doesn't change the base architecture—it continues to use GLM-5.2's MoE base (753 billion total parameters, ~40 billion activated per token), with all improvements coming from post-training scaling: a tenfold increase in long-horizon task environments, richer environment types, and significantly extended post-training time.

The result is a 50% improvement in coding capability over GLM-5.2 on Z.ai's self-developed Z.ai Code Bench, and the top ranking among open-weight models on public benchmarks such as Terminal-Bench 3.0 and Agents' Last Exam. Another direction the official team didn't anticipate was cybersecurity: CyberGym 84.5, ExploitBench 54.4, with vulnerability discovery and white-box code audit capabilities significantly exceeding expectations.

SeaWhale AI provides GLM-5.3 through an OpenAI-compatible interface, supporting tool calls, reasoning_effort reasoning intensity control, and streaming output.

Get API Key · Model ID: GLM-5.3


Why Choose GLM-5.3

  • Strongest coding among open-weight models — 50% improvement over GLM-5.2 on Z.ai Code Bench
  • #1 among open models on agent benchmarks — Terminal-Bench 3.0 jumped from 4.6 to 28.3, Agents' Last Exam 28.5
  • Outstanding cybersecurity capabilities — CyberGym 84.5, ExploitBench 54.4, with white-box audit and vulnerability discovery as key strengths
  • 1 million token context — maximum output of 128K tokens
  • Shorter execution paths — uses more than half fewer output tokens than closed-source flagships for similar tasks

Core Capabilities

01 Project-Level Coding and Engineering Agent

This generation focuses post-training on long-horizon task environments, with cross-file feature implementation, refactoring, and debugging all significantly more stable than GLM-5.2. On Z.ai Code Bench's High tier, GLM-5.3 achieves 31.4% accuracy with about 50K tokens, while Claude Opus 4.8 needs about 120K tokens at the highest tier to reach 29.5%—for the same task, GLM-5.3's execution path is significantly shorter.

  • Cross-file, cross-module feature development and refactoring
  • Long-chain defect localization and fixing
  • Agent-style autonomous engineering workflows

02 Cybersecurity and Vulnerability Discovery

The official team describes this capability as "developing faster than expected." In white-box code audit, vulnerability discovery, and exploit chain construction, GLM-5.3 is in the top tier of current open-weight models, with CyberGym 84.5 being the SOTA for this benchmark.

  • White-box code security audit
  • Vulnerability discovery and validation
  • Automated security engineering workflows

03 1 Million Token Long Context

With a 1 million token context and 128K maximum output, it can read an entire medium-to-large code repository, hundreds of pages of PDFs, or a complete long conversation history in one pass, then output a complete analysis or change plan.

  • Complete code repository-level analysis
  • Extra-long technical document processing
  • Full preservation of long conversation history

04 Three-Level Reasoning Intensity

GLM-5.3's thinking mode is always on, providing low, high, and max levels via reasoning_effort (default max). Drop to a lower level for simple tasks to save tokens, or use max for complex engineering tasks to get the most complete reasoning chain.

  • low — lightweight tasks, cost control
  • high — routine development and analysis
  • max — complex engineering and agent tasks (default)

Best Use Cases

Scenario Description
Project-level coding Coding SOTA among open-weight models, short execution paths
Coding agents #1 among open models on Terminal-Bench 3.0, Agents' Last Exam
Security audit White-box code audit and vulnerability discovery, CyberGym SOTA
Long-context analysis One-pass processing of entire repositories, hundreds of pages of documents
Domestic replacement Open weights, can be privately deployed
Cost optimization Significantly fewer output tokens for equivalent results

Differences Between GLM-5.3, GLM-5.2, and GLM-5.3-Flash

Capability GLM-5.3 GLM-5.2 GLM-5.3-Flash
Model ID GLM-5.3 GLM-5.2 GLM-5.3-Flash
Total params / activated 753B / ~40B Same base 320B / 18B
Context window 1M tokens 1M tokens 1M tokens
Max output 128K tokens 128K tokens 128K tokens
Input modality Text Text Text + Image
Terminal-Bench 3.0 28.3 (#1 open) 4.6 Below GLM-5.3
Weight license GLM-5.3 License (commercial use allowed) MIT MIT
Reasoning intensity low / high / max high / xhigh low / high / max
Positioning Flagship: strongest coding and security Previous-gen flagship Lightweight multimodal, ~1/10 price

Specific billing is subject to the real-time price card at the top of the page.

FAQ

When was GLM-5.3 released? Released on August 14, 2026, API launched on August 19, and model weights opened on August 28.

What improvements does it offer over GLM-5.2? The base model hasn't changed; all improvements come from post-training scaling. Coding capability improved 50% on Z.ai Code Bench, Terminal-Bench 3.0 went from 4.6 to 28.3, Agents' Last Exam went from 23.8 to 28.5, and a new strong cybersecurity capability emerged.

How does it compare to closed-source flagships? Coding and agent capabilities are close to Claude Fable 5. On Terminal-Bench 3.0, 28.3 still trails GPT-5.6 Sol's 34.6, but it has the highest Agents' Last Exam score among the same batch of models, and requires significantly fewer output tokens to complete similar tasks.

What is the open-source license? The weights use Z.ai's custom GLM-5.3 License (no longer MIT like GLM-5.2), allowing commercial use, with ultra-large-scale MaaS scenarios requiring security review; the code portion is Apache 2.0.

What are the context and output limits? 1 million token context, maximum output of 128K tokens, text-only input. For image input, use GLM-5.3-Flash.

Can thinking mode be turned off? No. GLM-5.3's thinking mode is always on; you can only adjust the depth via reasoning_effort between low/high/max.

Does it support tool calls? Yes. Function calling, streaming output, and JSON structured output are all supported.


Why Use GLM-5.3 API with SeaWhale AI

  • Call directly without deployment — the 753B parameter model requires no self-built GPU cluster
  • OpenAI-compatible interface — connect with just two lines of changes to existing code
  • Side-by-side comparison with closed-source models — directly compare effectiveness and cost with Claude and GPT under the same account
  • Settlement in RMB — domestic models are priced in RMB, with clear accounting

API

API integration ​

Model IDUse this value as the model in inference requests
glm-5.3
API KeyBearer token used to authenticate inference requests
Base URLOpenAI compatible · /chat/completions
OpenAIhttps://api.haijingai.com/v2/
provider OptionalSelects a service channel; omit it and the system picks the default
"provider": { "channel": "direct" }

glm-5.3 usage examples ​

SeaWhale AI is compatible with the OpenAI API protocol, so you can call it with the OpenAI SDK or plain HTTP requests. Streaming is enabled by default.

js
curl https://api.haijingai.com/v2/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <API_KEY>" \
  -d '{
    "model": "glm-5.3",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Hello!"}
    ],
    "provider": { "channel": "direct" },
    "stream": true
  }'
# provider is optional — remove this line to use the default channel
js
from openai import OpenAI

client = OpenAI(
    base_url="https://api.haijingai.com/v2",
    api_key="<API_KEY>",
)

stream = client.chat.completions.create(
    model="glm-5.3",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello!"},
    ],
    stream=True,
    # Optional: pick a service channel; omit to use the default
    extra_body={"provider": {"channel": "direct"}},
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
js
import OpenAI from 'openai'

const client = new OpenAI({
  baseURL: 'https://api.haijingai.com/v2',
  apiKey: '<API_KEY>',
})

const stream = await client.chat.completions.create({
  model: 'glm-5.3',
  messages: [
    { role: 'system', content: 'You are a helpful assistant.' },
    { role: 'user', content: 'Hello!' },
  ],
  stream: true,
  // Optional: pick a service channel; omit to use the default
  // @ts-expect-error provider is a SeaWhale AI extension, not in the OpenAI SDK types
  provider: { channel: 'direct' },
})

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}
Contact support