Review Machine · Master Review · AI models
Z.ai GLM-5.3
The open-weights model everyone benchmarked this month
Across four Z.ai pages, two independent sources, two dated reports and two community threads read on 17 September 2026, GLM-5.3 is first among open-weights models on the one independent index we could read, its coding gains come from post-training on the GLM-5.2 base as its maker says, and the licence stays permissive for almost everyone, while no independent run of its own sixteen-row benchmark table was readable.
Weights 753B total, 40B activeLicence GLM-5.3 LicenseContext 1M in, 128K outRead 17 September 2026Skip to the verdictRead as text

Verdict age
ProvisionalEvaluations in · developers thinThe launch material is from 14 August 2026 and the weights from 28 August, so the model is between three weeks and a month old. The independent index reading is from 15 to 16 September and was taken after the index itself was revised, so an earlier reading of 60 is not comparable. The vendor's table has no independent run, and the developer evidence is one thread and a handful of repository discussions.
- Early · editors only
- Provisional · editors in, owners thin
- Settled · owner reviews read over months
From 11 sources to one verdict
Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.
What we read
11 sources, 14 Aug – 17 Sep 2026. News, forum and other tiers thin. Buyer and video tiers empty.
By tier, heaviest first
The launch post, the model card, the licence file and the API documentation, read on 17 September 2026.
One evaluator's index page and one analysis of the release, 2 to 17 September 2026.
Two dated reports, 18 and 28 August 2026.
No buyer platform sells a model you host yourself. The only platform figure is the repository's own download count, reported by a third party rather than read here.
The launch thread and the repository's open discussions, both read as ordinary pages.
One tracker mirroring an evaluator's figures with observation dates, 16 September 2026.
Date window
14 Aug 2026 to 17 Sep 2026
Editorial reviews 2 Sep – 17 Sep 2026 · News 18 Aug – 28 Aug 2026Dated material runs from 14 August to 16 September 2026; the official pages were read on 17 September 2026.
Couldn’t read, so not used
- Hugging Face model pagepre-release snapshot, a countdown to 28 August
- Hacker News commentsthread figures read, no comment list paged
The model page as fetched still showed the release as upcoming, so the parameters and the licence were taken from the repository's own card and licence file and from the evaluator's specification. No review list, issue list or comment thread was paged.
Read, then rejected
- BenchLM model page · still lists the weights as unpublished and the parameters as undisclosed, older than the release
- Morph's GLM-5.3 page · a vendor serving the model, and its figures are the evaluator's reprinted
What they measured
753B total, 40B active, 1M tokens in and 128K out. First of the open-weights models on the independent index at 45.
Parameters
753B total40B active per token
Artificial Analysis · read 17 SepContext
1Mtokens in128K maximum output
Z.ai · read 17 SepIntelligence Index v4.3
45indexfirst of the open-weights models ranked; the page calls it expensive, slower than average and very verbose
Artificial Analysis · read 17 SepDownload count, 2 September
94,403downloads1,478 likes on the same day
noze · 2 SepTerminal-Bench 2.1, vendor table
88.2%the same table lists 88.3 for Kimi K3 and 88.8 for GPT-5.6 Sol
Z.ai · read 17 SepCyberGym, vendor table
84.5%first in the table; it lists 83.8 for Mythos 5 and 83.6 for GPT-5.6 Sol
Z.ai · read 17 SepExploitBench, vendor table
54.4%against 78.0 for Mythos 5 and 76.5 for GPT-5.6 Sol in the same table
Z.ai · read 17 SepTerminal-Bench 3.0, vendor table
Z.ai · read 17 SepEvery vendor-table figure is Z.ai's own, from the model card. No independent run of those tests was readable on 17 September.
Where they agree
First of the open-weights models on the independent index we could read, and a licence that stays permissive for almost everyone, in 2 of the 2 independent sources.
One tick per editorial review, in order of publication:NO nozeAA Artificial Analysis
First of the open-weights models on the independent index
1/2Artificial Analysis scores GLM-5.3 (max) 45 on Intelligence Index v4.3 and puts it ahead of Kimi K3 at 44, at roughly a fifth of the price per million tokens. T2Editorial
The coding gain comes from post-training alone
1/2Z.ai's card and documentation both say the base model is the previous generation's, and the launch post calls scaling post-training the only change. Terminal-Bench 3.0 rises 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9 in the same table. T1Official T2Editorial
The licence stays permissive for almost everyone
2/2The licence file grants use, copying, modification, fine-tuning, distribution, sublicensing and sale, with one condition that needs a Model-as-a-Service business above ten billion dollars of trailing twelve-month revenue. T1Official T2Editorial
The repository's own counters
94,403 downloads and 1,478 likes by 2 September, reported by a third party. The download count says how many took the files, not how many ran them. T2Editorial
Hugging Face's own counters, quoted by an analysis and read here as one page load.
What keeps coming up against it
The two-week hold before the weights, and the size of what arrived, in 1 of the 2 independent sources each.
The weights came two weeks after the model
Announced at launch · delivered on 28 AugustThe launch post says the weights would follow two weeks after launch, once safety evaluation and hardening were complete, and says cyber capability developed faster than expected.
Z.ai · 14 Aug
1/2The licence text and the benchmark footnotes arrived with the weights rather than with the announcement, and the footnotes say how the cyber numbers were produced.
Reports put the weights on Hugging Face on 28 August, with the repository's last modification at 15:22 UTC that day.
Mervin Praison · 28 Aug
The launch thread carries 1,171 points and 584 comments, the largest discussion of the model we could read.
Hacker News · 14 Aug
Verbose, slow, and priced against the closed frontier
1/2Artificial Analysis reports about 210 million output tokens for the evaluation suite against a median near 120 million for open-weights models. The documentation says reasoning cannot be switched off and max effort is the default. T1Official T2Editorial
Self-hosting the whole checkpoint needs a rack
1/2The repository ships fp8 weights of several hundred gigabytes, the reports put a multi-GPU node as the realistic minimum, and the download count by 2 September measures interest rather than use. T2Editorial T2News
Safety behaviour changed with the model
0/2A discussion on the repository, opened in the week before this reading, reports more refusals than the previous generation. One report, not a pattern, and Z.ai has published nothing about it. Forum
Where they split
Four places the evidence pulls two ways, starting with the day the weights landed.
The day the weights landed
Dates the weights to 25 August, three days ahead of the deadline Z.ai set itself.
Against that: Dates them to 28 August, with the repository modified at 15:22 UTC that day.
StagedBoth fit a staged push: the weights and the licence file need not arrive in the same commit. We date the release 28 August.
The independent index moved under it
Carries the first independent reading, 60 on Intelligence Index v4.1.1, tying Kimi K3.
Against that: Records 44.9 for the same model under v4.3, 8.5 points behind Claude Fable 5.1.
Both trueThe index was revised between the two readings, so this is not a fall in the model. The two numbers are not the same measurement.
Where the cyber lead stops
First on CyberGym at 84.5, ahead of the 83.8 and 83.6 the same table lists for Mythos 5 and GPT-5.6 Sol. That benchmark is finding flaws.
Against that: The two-hour and six-hour ExploitGym budgets that sit beside it are renormalised on each model's own tokens per second, not hours on a clock.
SoThe lead is real and narrow in scope: the same table has ExploitBench at 54.4, against 78.0 for Mythos 5.
The vendor's table against the index
Sixteen rows, with GLM-5.3 first among open weights on several of them.
Against that: An independent composite on its own infrastructure, with the model 8.5 points behind the closed frontier on 16 September.
Not comparableOne is a vendor running its own tests at best effort, the other an index of ten evaluations. Both are honest about what each measured.
Who it's for, who should pass
For long agentic coding you can host, and for anyone who wants the top of the open-weights index. Pass if it has to fit one machine.
It suits you if
- You run coding agents over long tasks and want weights you can host.The licence grants fine-tuning and redistribution, and the model card documents serving in SGLang, vLLM, TokenSpeed and Transformers.
- You pay per token and want the strongest open-weights option an independent index scores.Artificial Analysis ranks it first among them, and Z.ai's list price is unchanged from the previous generation.
- You want to study how much of a capability jump comes from post-training.The base model is the one the previous generation used, and Z.ai says so on the card and in the documentation.
Pass if
- You want frontier-class output from a machine in the corner.The fp8 repository runs to several hundred gigabytes and the reports assume a multi-GPU node.
- You need images in or out.The documentation and the model card both say text only.
- You need predictable bills.Reasoning is always on, max effort is the default, and the independent evaluation measured about 210 million output tokens against a median near 120 million.
This is a reading of published reviews, not medical, financial or legal advice. For a decision about your health, your money or your rights, a qualified professional is the right next step, and not a review.
The verdict
The open-weights leader on the one independent index we could read, two weeks late and heavy to run.
The strongest open-weights model on the one independent index we could read, with its coding gains coming from post-training on an unchanged base, a licence that stays permissive for almost everyone, and a two-week hold before the weights arrived.
Confidence, by tier
The launch post, model card, licence file and documentation were read directly on 17 September.
One evaluator's index, run on its own infrastructure, plus one dated analysis of the release.
Two reports, 18 and 28 August, both about the release rather than the model in use.
Nothing to buy: the weights are downloaded, and no store rates them.
One launch thread and a handful of repository discussions, none counting how many developers use it.
One tracker mirroring the evaluator's figures, dated 16 September.
No named channel reviewed this model.
- Editorial evidence
- under a month old
- Newest report
- 28 Aug 2026
- Read
- 17 Sep · month 1
Free to download. The multi-GPU rack is the subscription.
Rests onThe licence file grants free use, modification and resale, and the fp8 repository runs to several hundred gigabytes; the dated reports put a multi-GPU node as the realistic minimum.
Sources
11 sources, heaviest tier first. Every figure above comes from one of these.
- T1OfficialGLM-5.3 launch post14 Aug
- T1OfficialGLM-5.3 model cardread 17 Sep
- T1OfficialGLM-5.3 License fileread 17 Sep
- T1OfficialGLM-5.3 API documentationread 17 Sep
- T2EditorialThe weights, the licence and the benchmark method2 Sep
- T2EditorialArtificial Analysis on GLM-5.3 (max)index score, specification and the site's own summary, one page loadread 17 Sep
- T2NewsFirst index reading, 60 on v4.1.118 Aug
- T2NewsGLM-5.3 ships open weights28 Aug
- ForumGLM-5.3 launch threadpoints and comment count from one Hacker News search call14 Aug
- Forumzai-org/GLM-5.3 discussionsopen thread list read as an ordinary pageread 17 Sep
- OtherGLM 5.3 benchmarks and pricing, observationsbenchmark rows with their observation dates, one page load16 Sep
MethodOn 17 September 2026 we read four official Z.ai pages, two independent sources, two dated reports, one benchmark tracker and two community sources, published between 14 August and 16 September 2026, by web search and plain page fetch, and synthesised them with AI; no model was run, no library installed and nothing was tested.
Dates without a year are 2026.
The review as text
Z.ai's GLM-5.3 is a 753-billion-parameter mixture-of-experts model, with 40 billion parameters active per token, a one-million-token context window and text as its only input. It is the open-weights release this reading picked: its weights landed on 28 August 2026, inside the last thirty days, and it carries more dated evaluations than any other model whose weights shipped in that window.
We read four Z.ai pages (the launch post, the model card, the licence file and the API documentation), two independent sources (an evaluator's index page and an analysis of the release), two dated reports, one benchmark tracker and the two community sources a plain page fetch could open, the Hacker News thread from launch day and the repository's own discussions. The dated material runs from 14 August to 16 September 2026, and the pages themselves were read on 17 September 2026. Nothing was run, installed, prompted or tested.
Consensus
The evidence agrees on what GLM-5.3 is: a long-horizon coding and agentic model made by taking the GLM-5.2 base and spending more post-training on it, rather than a new pretraining run. Z.ai's launch post puts it plainly, and its model card and API documentation repeat it, that every gain over GLM-5.2 comes from post-training.
The one independent index we could read has it first among open-weights models. Artificial Analysis scores GLM-5.3 (max) 45 on Intelligence Index v4.3, ahead of Kimi K3 at 44, and calls it amongst the leading models in intelligence while noting that it is particularly expensive beside open-weights models of similar size, slower than average and very verbose. On Z.ai's own table it is first among open weights on Terminal-Bench 3.0, at 28.3 against 4.6 for GLM-5.2, on Agents' Last Exam at 28.5 and on CyberGym at 84.5. Both readings agree that coding is where the model moved, and they part company on how much of that reaches a reader who is not running agent harnesses.
Confidence is uneven. The official pages are strong: the model card, the licence file and the documentation were read directly. The independent side is one evaluator's index and one analysis of the release, both dated. Developer evidence is one long thread and a handful of repository discussions. The vendor's table is a claim, and no independent run of its sixteen benchmarks was readable on 17 September.
Recurring strengths
It is the open-weights leader on the independent index. Artificial Analysis places GLM-5.3 (max) first among the open-weights models it ranks, at roughly a fifth of Kimi K3's list price per million tokens. That is the index reading in this sample.
The coding gain is large, and it comes from post-training alone. Z.ai reports Terminal-Bench 3.0 rising from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9 and Agents' Last Exam from 23.8 to 28.5 against GLM-5.2 on the same base model. Its own model card is the source, and its own model card is a claim.
The licence stays permissive for almost everyone. The licence file grants use, copying, modification, fine-tuning, distribution, sublicensing and sale, and adds one condition: a licensee running a Model-as-a-Service business whose revenue passes ten billion dollars over any twelve consecutive months must pass Z.AI's security review before commercial use. For nearly every reader the file is MIT in everything but its name.
Recurring complaints
The weights arrived two weeks after the model did. The launch post said the weights would follow two weeks after launch, once safety evaluation and hardening were complete, and attributed the hold to cyber capability developing faster than expected. Reports date the weights to 28 August, with the repository's last modification at 15:22 UTC that day. One analysis writes down 25 August instead.
It is verbose, and it is priced against the closed frontier. Artificial Analysis reports about 210 million output tokens across its evaluation suite against a median near 120 million for open-weights models, and calls the model slower than average. Z.ai's own documentation explains part of it: reasoning cannot be switched off, and the default effort level is max.
Running the whole checkpoint is a rack, not a workstation. The repository ships fp8 weights of several hundred gigabytes, and the analysis we read puts a multi-GPU node as the realistic minimum. The repository's download count reached 94,403 by 2 September with 1,478 likes, which measures interest rather than use.
Where reviewers split
How long the hold was. One analysis dates the weights to 25 August, three days ahead of the deadline Z.ai had set itself. Reports date them to 28 August, and the repository's last modification is that day. Both are consistent with a staged push, since the weights and the licence file need not arrive together.
The independent index moved under it. One report carried Artificial Analysis's first reading, 60 on Intelligence Index v4.1.1 on 18 August, tying Kimi K3. A tracker records 44.9 for the same model under v4.3 on 16 September. This is not a fall in the model: the index itself was revised between the two readings, so the two numbers are not the same measurement.
Where the cyber lead stops. Z.ai's table is first on CyberGym, at 84.5 against the 83.8 and 83.6 it lists for Mythos 5 and GPT-5.6 Sol, which is detecting and locating flaws. The same table has ExploitBench at 54.4 against 78.0 for Mythos 5 and 76.5 for GPT-5.6 Sol, which is further up the exploitation chain. The analysis adds that the two-hour and six-hour ExploitGym budgets are not hours on a clock: they are renormalised on each model's own tokens per second.
The vendor's table against the index. Z.ai's sixteen rows put GLM-5.3 at the top of the open-weights field on several benchmarks. Artificial Analysis's composite, which runs its own suite on its own infrastructure, places the model 8.5 points behind Claude Fable 5.1 on a 16 September reading. The two are not comparable, and each is honest about what it measured.
Who it suits
You run coding agents over long tasks and want weights you can host. The licence grants fine-tuning and redistribution, and the model card documents serving in SGLang, vLLM, TokenSpeed and Transformers.
You pay per token and want the strongest open-weights option an independent index scores. Artificial Analysis ranks it first among them, and Z.ai's list price is unchanged from the previous generation.
You want to study how much of a capability jump comes from post-training. The base model is the one the previous generation used, and the vendor says so on the card.
Who should pass
You want frontier-class output from a machine in the corner. The fp8 repository runs to several hundred gigabytes and the reports assume a multi-GPU node, so the download is the cheap part.
You need images in or out. The documentation and the model card both say text only, which several rivals at this size do not.
You need predictable bills. Reasoning is always on, max effort is the default, and the independent evaluation measured about 210 million output tokens where open-weights models of that size median near 120 million.
This is a reading of published reviews, not medical, financial or legal advice. For a decision about your health, your money or your rights, a qualified professional is the right next step, and not a review.
One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it. Get the weekly verdict.
Sources
- Z.ai launch post, official, 14 August 2026.
- GLM-5.3 model card, official, read 17 September 2026.
- GLM-5.3 License, official, read 17 September 2026.
- GLM-5.3 API documentation, official, read 17 September 2026.
- Artificial Analysis on GLM-5.3 (max), editorial, read 17 September 2026.
- noze, the weights, the licence and the benchmark method, editorial, 2 September 2026.
- Unite.AI report on the first index reading, news, 18 August 2026.
- Mervin Praison on the weights release, news, 28 August 2026.
- AI Atlas model page and benchmark observations, other, observed 16 September 2026.
- Hacker News launch thread, forum, 14 August 2026.
- zai-org/GLM-5.3 discussions, forum, read 17 September 2026.
On 17 September 2026 we read four official Z.ai pages, two independent sources, two dated reports, one benchmark tracker and two community sources, published between 14 August and 16 September 2026, by web search and plain page fetch, and synthesised them with AI; no model was run, no library installed and nothing was tested.
One verdict a week.
Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.