Pin GLM-5.3-Flash for cost-effective coding
GLM-5.3-Flash is in the AshnaAI catalog as the low-cost coding default: a 320B MoE, 1M-context, vision-capable model people pin in chat when the job is code, tools, or UI screenshots. Ashna-X1 still covers mixed everyday routing.
AshnaAI

What shipped in the catalog
AshnaAI lists GLM 5.3 Flash (catalog id glm-5.3-flash) in the Tops catalog. The row is served through Baseten as zai-org/GLM-5.3-Flash. Creator is Zhipu AI. Tools, web search, image query, and deep think are enabled. Context is the 1M-token GLM-5.3-Flash window.
Z.ai introduced the model on 26 August 2026 as the first natively multimodal GLM-5, after a short anonymous preview as Ox Alpha. The official scoreboard is in the table below, copied from z.ai/blog/glm-5.3-flash. Those are Z.ai’s published numbers. The same model is the one people pin on AshnaAI next to GPT, Claude, and Kimi rows.
Z.ai published benchmark table
Read the Flash column against Opus 4.8 and GPT-5.6 Terra. Flash is within a point of Opus 4.8 on Terminal Bench 2.1 (84.3 vs 85.0), ahead on DeepSWE v1.1 (63.4 vs 58.0), and cheaper than the GPT-5.6 rows on the same catalog. That is why people treat it as the coding default—not because it tops every cell.
The table below is the competitive scoreboard from Z.ai’s 26 August 2026 launch post.
| Benchmark | GLM-5.3-Flash | GLM-5.2 | DeepSeek-V4-Vision-Exp | Opus 4.8 | GPT-5.6 Terra | Gemini 3.7 Flash |
|---|---|---|---|---|---|---|
| Coding | ||||||
| Terminal Bench 2.1 | 84.3 | 81.0 | 83.9 | 85.0 | 87.4 | 85.8 |
| DeepSWE v1.1 | 63.4 | 46.2 | 59.3 | 58.0 | 69.6 | 65.3 |
| NL2Repo | 56.3 | 48.9 | 57.7 | 69.7 | - | - |
| Agentic | ||||||
| Toolathlon Verified | 78.4 | 59.9 | 75.9 | 76.2 | 74.9 | - |
| AutomationBench v1.0.6 | 48.8 | 26.2 | 38.8 | 41.0 | 37.2 | 52.3 |
| Agents' Last Exam | 26.3 | 20.4 | 27.3 | 27.0 | 28.0 | - |
| HLE w/ Tools | 55.3 | 54.7 | 55.1 | 57.9 | - | - |
| GDPval-AA v2 | 1773 | 1504 | 1675 | 1582 | 1571 | 1527 |
| Vision | ||||||
| OfficeQA Pro | 62.4 | - | 57.9 | 48.9 | - | - |
| CharXiv Reasoning w/ Tools | 89.4 | - | 80.4 | 89.9 | 88.0 | 88.7 |
| Chartography w/ Tools | 78.0 | - | 64.3 | 75.0 | 68.0 | 65.0 |
| BabyVision | 53.4 | - | 35.1 | 46.8 | 61.6 | 70.9 |
| MVbench | 77.8 | - | 69.4 | 67.1 | 75.0 | 82.2 |
| MMVU | 80.5 | - | 72.7 | 67.4 | 75.8 | 82.3 |
Base-model table from the same Z.ai post
The same launch note also published base-model scores. GLM-5.3-Flash-Base uses 18B activated parameters of 320B total, and Z.ai reports it ahead of GLM-4.5-Base overall, including LiveCodeBench-Base at 37.6.
| Benchmark | GLM-4.5-Base | GLM-5-Base | DeepSeek-V4-Flash-Base | GLM-5.3-Flash-Base |
|---|---|---|---|---|
| Activated Params | 32B | 40B | 13B | 18B |
| Total Params | 355B | 744B | 284B | 320B |
| MMLU | 86.1 | 88.3 | 88.5 | 88.1 |
| BBH | 86.2 | 87.4 | 84.9 | 86.6 |
| HellaSwag | 87.1 | 88.1 | 85.3 | 87.1 |
| LiveCodeBench-Base | 28.1 | 34.4 | 29.9 | 37.6 |
| SimpleQA | 30.0 | 36.0 | 31.2 | 33.5 |
Why it is the coding default
On AshnaAI, new coding, debug, and repo agents typically start on GLM-5.3-Flash rather than Kimi K2.7 Code, GLM-5.2, GPT-5.6 Terra, GPT-5.6 Sol, or a Claude row. That is a cost-and-capability default, not a claim that Flash wins every reasoning contest against GPT-5.6 Sol or Claude Fable 5.
Flash fits when the job is write or repair code, clone a UI from a screenshot, or run many tool calls. A premium GPT or Claude row fits when a customer mandates that vendor, or when a team already measured a gap on that repo.
Cost class, not a mystery price
The catalog marks Flash as comparative cost low. The Zhipu-via-Baseten list rates on the platform are about $0.15 / 1M input and $0.50 / 1M output. GPT-5.6 Terra is about $2 / $12, GPT-5.6 Sol about $5 / $30, Kimi K3 about $0.80 / $3.20. Anthropic’s public list for Claude Opus 5 is $5 / $25 and Claude Fable 5 is $10 / $50. Those are vendor list rates, not your invoice line if you are on credits.
For everyday coding volume, Flash is the row that stays near frontier coding benches without premium-token spend. Pair-by-pair write-ups live on GLM-5.3-Flash vs GPT-5.6 Terra, vs GPT-5.6 Sol, vs Claude Fable 5, vs Claude Opus 5, and vs Kimi K3.
How people run Flash in practice
Reasoning stays on. Effort is low, high, or max. Teams start on low for small edits and raise it when the patch needs more thinking.
Image input is on. Screenshot-to-UI and frontend stills are the common vision jobs on this row.
Try GLM-5.3-Flash now
Someone trying Flash inside the product opens GLM-5.3-Flash in AshnaAI chat. That URL pins catalog id glm-5.3-flash, so the first message already runs on Zhipu’s GLM-5.3-Flash. New accounts start at app.ashna.ai/signup. If the next task is not coding, people switch the catalog back to Ashna-X1 instead of leaving Flash pinned out of habit.
Teams who want the same model from their own product use the OpenAI-compatible AshnaAI API. They request access at apply for API access, create a key in Account → API, then POST chat completions with model set to glm-5.3-flash. The same model field works for every catalog row. Walkthrough: How to call any catalog model through the API. Reference: list models and chat completions.
Frequently asked questions
- What is GLM-5.3-Flash?
- Zhipu AI’s first natively multimodal GLM-5 model: a 320B-parameter Mixture-of-Experts with about 18B active parameters, 1M context, vision, and tool calling. On AshnaAI the catalog id is glm-5.3-flash.
- How do I open GLM-5.3-Flash in AshnaAI?
- Use the pinned chat link https://app.ashna.ai/chat?agent=glm-5.3-flash or pick GLM 5.3 Flash in the model catalog. New accounts start at https://app.ashna.ai/signup
- Is GLM-5.3-Flash cheaper than GPT-5.6 and Claude for coding?
- On list rates used by the catalog, Flash is the low-cost coding row. GPT-5.6 Terra and Sol, Claude Fable 5, Claude Opus 5, and Kimi K3 sit in a higher comparative cost class. Pair-by-pair write-ups are linked from this page.
- Does Flash replace Ashna-X1?
- X1 remains the default for mixed agent work. People pin Flash when the job is coding, frontend, heavy tools, or vision-plus-code.
- Can GLM-5.3-Flash see images?
- Yes. Image input is on. Teams use screenshots and stills for UI clones and frontend work.
- Does GLM-5.3-Flash keep reasoning on?
- Reasoning stays on. Effort is low, high, or max. Higher effort increases latency and tokens, so people start on low for small edits.
Tags
Related
- AshnaAI for Work
- AshnaAI vs ChatGPT
- AshnaAI vs Claude
- Large Language Model (LLM)
- Foundation Model
- glm 5 3 flash vs gpt 5 6 terra
- glm 5 3 flash vs gpt 5 6 sol
- glm 5 3 flash vs claude fable 5
- glm 5 3 flash vs claude opus 5
- glm 5 3 flash vs kimi k3
- how to call any catalog model through the api
- how to use ashna x1 instead of picking models yourself
- ashna x1 task aware model routing
Try this in AshnaAI. Create a free account.
Found this article helpful? Share it with your network.