AI model comparison

GLM 4.5 Air vs Llama 4 Maverick

Pricing, context window, benchmarks, and features compared side by side. Or skip the guesswork and send one prompt to both.

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications.

Input / 1M
$0.13
Output / 1M
$0.85
Context
131K

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward

Input / 1M
$0.19
Output / 1M
$0.65
Context
1.05M

At a glance

Quick verdict

How the two models stack up on the things people ask about most.

Lower price

z-ai logo

GLM 4.5 Air

1.0x cheaper for a typical chat

Higher intelligence score

z-ai logo

GLM 4.5 Air

11.1 vs 10 on Artificial Analysis

Larger context window

meta-llama logo

Llama 4 Maverick

1.05M vs 131K tokens

Newer release

z-ai logo

GLM 4.5 Air

Released July 25, 2025

Benchmarks

Benchmark scores

Independent scores from Artificial Analysis. Higher is better.

Intelligence index

z-ai logoGLM 4.5 Air11.1
meta-llama logoLlama 4 Maverick10

Coding index

z-ai logoGLM 4.5 AirNot available
meta-llama logoLlama 4 Maverick16.3

Agentic index

z-ai logoGLM 4.5 AirNot available
meta-llama logoLlama 4 Maverick0.6

Pricing

GLM 4.5 Air vs Llama 4 Maverick API pricing

Per-token API rates. Cheaper option highlighted.

Metricz-ai logoGLM 4.5 Airmeta-llama logoLlama 4 Maverick
Input tokensPer 1M tokens$0.13$0.19
Output tokensPer 1M tokens$0.85$0.65
Cached input (read)Per 1M tokens$0.03$0.05
Cache writePer 1M tokensNot availableNot available

What it costs in practice

Estimated cost per 1,000 requests at standard rates. Coding and long-document figures also show the cost when the input is already cached.

Chat message

2K in, 500 out

z-ai logoGLM 4.5 Air$0.69
meta-llama logoLlama 4 Maverick$0.70

Coding task

30K in, 4K out

z-ai logoGLM 4.5 Air$7.30$4.15 cached
meta-llama logoLlama 4 Maverick$8.24$4.11 cached

Long document summary

150K in, 2K out

z-ai logoGLM 4.5 Air$21.20$5.45 cached
meta-llama logoLlama 4 Maverick$29.43$8.81 cached

Specs

Context window and capabilities

How much each model can read, how much it can write, and what it accepts as input.

Metricz-ai logoGLM 4.5 Airmeta-llama logoLlama 4 Maverick
Context window131K tokens1.05M tokens
Max output98K tokens16K tokens
Input typesTextImage, Text
Output typesTextText
Reasoning effort levelsNot availableNot available
Default reasoning effortNot availableNot available
Release dateJuly 25, 2025April 5, 2025

Features

Supported features

API features available for each model.

Metricz-ai logoGLM 4.5 Airmeta-llama logoLlama 4 Maverick
Tool callingSupportedSupported
Structured outputsNot supportedSupported
JSON modeNot supportedSupported
ReasoningSupportedNot supported
TemperatureSupportedSupported
Stop sequencesSupportedSupported
Deterministic seedSupportedSupported
Verbosity controlNot supportedNot supported

Our take

GLM 4.5 Air is 1.0x cheaper for a typical chat. GLM 4.5 Air scores higher on the Artificial Analysis Intelligence Index (11.1 vs 10). Llama 4 Maverick has the larger context window (1.05M vs 131K tokens). The right pick depends on your workload, so the quickest way to settle it is to send the same prompt to both and compare.

FAQ

Frequently asked questions

Is GLM 4.5 Air or Llama 4 Maverick cheaper?

GLM 4.5 Air is cheaper for a typical chat (2K input tokens and 500 output tokens). GLM 4.5 Air costs $0.13 per 1M input tokens and $0.85 per 1M output tokens, while Llama 4 Maverick costs $0.19 per 1M input tokens and $0.65 per 1M output tokens.

Which has a bigger context window, GLM 4.5 Air or Llama 4 Maverick?

Llama 4 Maverick supports up to 1.05M tokens of context, compared with 131K for GLM 4.5 Air.

Is GLM 4.5 Air smarter than Llama 4 Maverick?

On the Artificial Analysis Intelligence Index, GLM 4.5 Air scores 11.1 and Llama 4 Maverick scores 10, putting GLM 4.5 Air ahead. Benchmarks don't capture everything, so the best test is running your own prompts through both.

Can I use GLM 4.5 Air and Llama 4 Maverick at the same time?

Yes. Shortcut Chat sends one prompt to multiple models at once, so you can see GLM 4.5 Air and Llama 4 Maverick answer side by side and keep the better response.

z-ai logometa-llama logo

Why choose? Ask both.

Send one prompt to GLM 4.5 Air and Llama 4 Maverick at the same time. Compare the answers side by side and keep the best one.