How good is the AI?
Lighter doesn't mean weaker.
The model behind our text replies is GLM-4.5-Air - an open MoE model with 106B parameters, only 12B of them active per token. The technical report benchmarks it against frontier models. It lands within about a point of Claude Opus 4 overall, and slightly ahead on agentic and reasoning work.
GLM-4.5-Air 12B active
Claude Opus 4
Overall 12 benchmarks
59.8
60.9
Agentic tool use, browsing
55.7
54.6
Reasoning maths, science, code
66.1
65.1
Coding SWE-bench, Terminal-Bench
43.8
55.5
Every bar is scored out of 100, on the same axis - nothing is zoomed.
Opus 4 is ahead on agentic coding. For chat, questions, research and everyday reasoning, GreenAI's model uses a fraction of the energy but is doing frontier-class work.
Source: the GLM-4.5 technical report. Group figures are plain means of the report's published scores (overall = all 12 benchmarks, TAU-bench counted once); Opus 4's overall and reasoning means are computed from the same tables the same way.