# GLM-5.3-Flash Benchmark Scores & Performance

> GLM-5.3-Flash by Z.AI scores 57.36/100 overall, ranking #55 out of 889 AI models.

Canonical page: https://benchlm.ai/models/glm-5-3-flash

Last updated: October 10, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Z.AI |
| Source Type | Open Weight |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Official model card | [GLM-5.3-Flash model card](https://huggingface.co/zai-org/GLM-5.3-Flash) |
| Overall Score | 57.36/100 |
| Overall Rank | #55 of 889 |

## Family & Coverage

- Family: GLM-5
- Variant: flash (5.3 Flash)
- Benchmarks covered: 39 of 625
- Sibling models: [GLM-5.3](/models/glm-5-3), [GLM-5.2](/models/glm-5-2), [GLM-5.1](/models/glm-5-1), [GLM-5-Turbo](/models/glm-5-turbo), [GLM-5](/models/glm-5), [GLM-5V-Turbo](/models/glm-5v-turbo), [GLM-5 (Reasoning)](/models/glm-5-reasoning)
- Related earlier model: [GLM-4.7-Flash](/models/glm-4-7-flash)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score | Source |
|-----------|-------|--------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 84.3% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 78.4% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [AutomationBench](/benchmarks/automationbench) | 48.8% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 26.3% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [HLE w/ tools](/benchmarks/hlewithtools) | 55.3% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1773 | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [AA Tau3 Banking](/benchmarks/aatau3banking) | 47.2% | [Artificial Analysis: tau3-banking leaderboard](https://artificialanalysis.ai/evaluations/tau3-banking) |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/terminalbench21) | 62.9% | [Vals AI: Terminal-Bench 2.1 leaderboard](https://www.vals.ai/models/zai_glm-5.3-flash) |
| [AA Briefcase](/benchmarks/aabriefcaseelo) | 1454 | [Artificial Analysis: aa-briefcase leaderboard](https://artificialanalysis.ai/evaluations/aa-briefcase) |
| [AA AutomationBench](/benchmarks/aaautomationbench) | 60.4% | [Artificial Analysis: automationbench-aa leaderboard](https://artificialanalysis.ai/evaluations/automationbench-aa) |
| [AA EnterpriseOps-Gym](/benchmarks/aaenterpriseopsgym) | 33.2% | [Artificial Analysis: enterprise-ops-gym-aa leaderboard](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa) |
| [AA ITBench](/benchmarks/aaitbench) | 51.2% | [Artificial Analysis: itbench-aa leaderboard](https://artificialanalysis.ai/evaluations/itbench-aa) |
| [aaTerminalBench21](/benchmarks/terminalbench21) | 84.3% | [Artificial Analysis: terminalbench-v2-1 leaderboard](https://artificialanalysis.ai/evaluations/terminalbench-v2-1) |
| [AA Terminal-Bench 4.0](/benchmarks/terminal-bench-4) | 32.8% | [Artificial Analysis: terminalbench-v4-0 leaderboard](https://artificialanalysis.ai/evaluations/terminalbench-v4-0) |
| [GDP.pdf](/benchmarks/aagdppdf) | 15.4% | [Artificial Analysis: gdp-pdf leaderboard](https://artificialanalysis.ai/evaluations/gdp-pdf) |

## Coding Benchmarks

| Benchmark | Score | Source |
|-----------|-------|--------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 84.3% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [DeepSWE](/benchmarks/deepswe) | 63.4% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [NL2Repo](/benchmarks/nl2repo) | 56.3% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 80.5% | [Vals AI: LiveCodeBench leaderboard](https://www.vals.ai/models/zai_glm-5.3-flash) |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 92.0% | [Vals AI: SWE-bench leaderboard](https://www.vals.ai/models/zai_glm-5.3-flash) |
| [OpenHarmony Bench](/benchmarks/openharmonybench) | 57.3% | [OpenHarmony Bench official leaderboard](https://bench.matrix.openharmony.cn/) |
| [FrontierSWE v2](/benchmarks/frontierswev2) | 18.1% | [Proximal: FrontierSWE v2 leaderboard](https://www.frontierswe.com/) |
| [AA-SciCode](/benchmarks/aascicode) | 51.6% | [Artificial Analysis: scicode leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| [Bug Hunt Bench](/benchmarks/bug-hunt-bench) | 17.7 fixes | [Bug Hunt Bench public data](https://bughunt.productcompass.pm/data/benchmark.json) |

## Multimodal & Grounded Benchmarks

| Benchmark | Score | Source |
|-----------|-------|--------|
| [OfficeQA Pro](/benchmarks/officeqapro) | 62.4% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [CharXiv](/benchmarks/charxiv) | 89.4% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [Chartography (tools)](/benchmarks/chartographywithtools) | 78.0% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [BabyVision](/benchmarks/babyvision) | 53.4% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [MMVU](/benchmarks/mmvu) | 80.5% | [Z.AI GLM-5.3-Flash launch post](https://z.ai/blog/glm-5.3-flash) |
| [Design Arena Website](/benchmarks/designarenawebsite) | 1278 | [OpenRouter model benchmarks](https://openrouter.ai/z-ai/glm-5.3-flash/benchmarks) |

## Reasoning Benchmarks

| Benchmark | Score | Source |
|-----------|-------|--------|
| [MLCR-AA](/benchmarks/aamlcr) | 51.1% | [Artificial Analysis: mlcr-aa leaderboard](https://artificialanalysis.ai/evaluations/mlcr-aa) |
| [AA-LCR](/benchmarks/lcr) | 80.0% | [Artificial Analysis: artificial-analysis-long-context-reasoning leaderboard](https://artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning) |
| [CritPt](/benchmarks/critpt) | 15.4% | [Artificial Analysis: critpt leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

## Knowledge Benchmarks

| Benchmark | Score | Source |
|-----------|-------|--------|
| [GPQA Diamond (Vals)](/benchmarks/gpqa-diamond) | 86.4% | [Vals AI: GPQA Diamond leaderboard](https://www.vals.ai/models/zai_glm-5.3-flash) |
| [MMLU-Pro (Vals)](/benchmarks/mmlu-pro) | 86.1% | [Vals AI: MMLU Pro leaderboard](https://www.vals.ai/models/zai_glm-5.3-flash) |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 41.8% | [Artificial Analysis: artificial-analysis-intelligence-index leaderboard](https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index) |
| [AA-GPQA Diamond](/benchmarks/gpqa-diamond) | 91.2% | [Artificial Analysis: gpqa-diamond leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| [AA-HLE](/benchmarks/aahle) | 39.9% | [Artificial Analysis: humanitys-last-exam leaderboard](https://artificialanalysis.ai/evaluations/humanitys-last-exam) |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 7.5% | [Artificial Analysis: omniscience leaderboard](https://artificialanalysis.ai/evaluations/omniscience) |

## Other Z.AI Models

- [GLM-5.3](/models/glm-5-3) - Score: 68.64
- [GLM-5.2](/models/glm-5-2) - Score: 61.55
- [GLM-5.1](/models/glm-5-1) - Score: 56.48
- [GLM-5-Turbo](/models/glm-5-turbo) - Score: 54.15
- [GLM-5](/models/glm-5) - Score: 54.14
- [GLM-5V-Turbo](/models/glm-5v-turbo) - Score: 50.1
- [GLM-4.7](/models/glm-4-7) - Score: 48.28
- [GLM-4.7-Flash](/models/glm-4-7-flash) - Score: 39.36
- [GLM-4.6](/models/glm-4-6) - Score: 37.37
- [GLM-4.5-Air](/models/glm-4-5-air) - Score: 34.99
- [GLM-4.5](/models/glm-4-5) - Score: not computed
- [GLM-5 (Reasoning)](/models/glm-5-reasoning) - Score: not computed
