The Green AI assistant

Energy efficient models. They only think harder when your question needs it.

Most AI products answer a "what's the capital of France" with the same enormous model they'd use to debug your codebase. GreenAI runs lighter, more efficient open models - a fraction of the compute per reply - and turns its reasoning on only when the question earns it. Every reply shows what it cost.

Power efficiency and energy accounting is built into the product.

SmallmodelsNo huge gas guzzling models
2modesReasoning only when it is earned
£0To start - no premium tier
1clickTo read your own footprint

The problem

Large language models are enormously power-hungry. Ignoring that isn't an option.

A single frontier model can take tens of gigawatt-hours to train - comparable to the annual electricity use of a small town - and every message you send afterwards draws a lot more. And most AI companies are using a mix of dirty power. We measure our power consumption, minimise it using efficient models, and over-correct by investing in renewables until the numbers point the other way.

Estimated energy for 100 replies of about 500 tokens each Each bar is the top of the estimated range, on one axis - nothing is zoomed.
GreenAI Air GLM-4.5-Air, 12B active
2-5 Wh · ~0.5 g CO2
Frontier Model no extended thinking
30-100 Wh · 10-40 g CO2
Frontier Model extended thinking, ~2k tokens
100-600 Wh · 40-200 g CO2

Frontier models are the largest general-purpose models - Claude Opus, OpenAI's GPT, xAI's Grok, Google Gemini - and none of their vendors publish per-query energy figures, nor the model sizes you would need to derive one. Each bar is drawn at the high end of its range; the low end is in the figures beside it. GreenAI Air is interpolated from measured benchmarks for Llama 3.1 70B (~0.04 Wh per reply) and 405B (~0.21 Wh); the frontier rows assume 405B-class active compute, but it is likely much larger. Figures include host CPU, cooling and PUE - accelerator-only numbers run roughly half. CO2 at the UK grid's ~125 g/kWh; a US data centre is nearer 350-400 g. Batching and how many thinking tokens get generated move the result more than the choice of model does.

  • Training is the big upfront debt - and we don't add to it. We train nothing. GreenAI runs open weights that already exist, so the largest single energy cost in AI is one we don't repeat. Reuse beats retraining.
  • Inference is the daily drip. Small per message, but it adds up across millions of chats - so how hard the model thinks, and where it runs, is where the real savings live.
  • Grids are still mixed. "Renewable" electricity often means certificates, not clean electrons on the wire. We fund genuinely new capacity, not just paperwork.

Our approach

Three commitments, in the order they matter.

01

Use less

The greenest kilowatt-hour is the one you never spend. Compact open models, a small share of its parameters active per token, reasoning off unless the question needs it, and context trimmed rather than padded. Efficiency ships before any offset is counted.

02

Move to a clean grid

Where the electrons come from changes the footprint tenfold. Our own hardware is moving to Helsinki, on Nordic hydro and wind with free-air cooling.

03

Give back more

Once subscriptions cover the running costs, a fixed share funds new renewable capacity - community solar, wind, storage. The goal is net-positive, not net-zero. We'll report the kilowatt-hours funded against the kilowatt-hours used, starting from zero.

The models we run

Open, low energy models, run well - and a reasoning switch.

GreenAI runs a few energy-efficient, open-weight models and vary how hard they think. That makes the energy per reply predictable, the setup inspectable, and power usage greener.

Lighter models, run well

open weights

A variety of mixture-of-experts models that activate only a small share of their parameters for each token they write, so it answers like a big model while using far less power.

Quick

reasoning off · the default

Chat, lookups, rewrites, summaries. The model answers directly without generating a reasoning trace first. Most messages land here, and it is by far the cheapest way to get a good answer.

Deep

reasoning on · when earned

Code, maths, multi-step problems and long documents. The model works through the problem before answering. Better results, several times the tokens - which is exactly why it isn't the default.

How good is the AI?

Lighter doesn't mean weaker.

The model behind our text replies is GLM-4.5-Air - an open MoE model with 106B parameters, only 12B of them active per token. The technical report benchmarks it against frontier models. It lands within about a point of Claude Opus 4 overall, and slightly ahead on agentic and reasoning work.

Opus 4 is ahead on agentic coding. For chat, questions, research and everyday reasoning, GreenAI's model uses a fraction of the energy but is doing frontier-class work.

Source: the GLM-4.5 technical report. Group figures are plain means of the report's published scores (overall = all 12 benchmarks, TAU-bench counted once); Opus 4's overall and reasoning means are computed from the same tables the same way.

Ask something. See what it cost.

Try it out at chat.greenai.org.uk. We'd love you to join the GreenAI revolution!

Free tier, no card. Sign up with an email address and start asking.