SylphxModels

Z.ai

GLM Flash family · v4.7

GLM-4.7-Flash

Retiring in 2d
z-ai/glm-4.7-flash

GLM-4.7-Flash on Sylphx Models · 200K of context. One Responses contract, with list prices published per million tokens.

Released
Jan 19, 2026 · 249 days old
Retires
Sep 26, 2026 · 2 days left
Provider
z-ai
Context
200,000 tokens
Open in the playgroundSee a request sampleRead the docs
Input $0.0693 / MOutput $0.462 / MCached input $0.01155 / MFull rate sheet
Context window200,000tokens per request
Max output131,072tokens per response
Input$0.0693 / Mper million tokens
Output$0.462 / Mper million tokens
Cached input$0.01155 / M83% off input
ReleasedJan 19, 2026249 days old

Pricing

USD per million tokens
Input
$0.0693
Output
$0.462
Cached input
$0.01155

$0.17 for 750,000 input and 250,000 output tokens.

Assumption: one batch of 750,000 input plus 250,000 output tokens at the published list prices — recompute for your own mix. Blended $0.1675 per million at a 3:1 mix.

Chat replies · 3:1 input:output
$0.1675 / M
Long documents · 9:1 input:output
$0.1086 / M
Heavy output · 1:1 input:output
$0.2656 / M
Cache write
Not published
Currency and unit
USD / million tokens

Cached input reads cost 83% less than fresh input tokens.

Full rate sheet and billing rules.

Specification

Context window
200,000 tokens
Max output
131,072 tokens
Input modalities
Not published
Output modalities
Not published
Family
GLM Flash
Version
v4.7
Provider
z-ai
Released
Jan 19, 2026 · 249 days old
Model id
z-ai/glm-4.7-flash

Data posture

Standard

By default, prompts and outputs sent to this model are not used to train models.

Training on what you send · default
No
Zero Data Retention · default
Yes
curl https://api.sylphx.ai/v1/responses \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "z-ai/glm-4.7-flash",
    "input": "Summarise this incident report in three bullets.",
    "stream": false
  }'

How to call it

Send the official Responses document to the base URL above with your organization key and this exact model id. Streaming arrives as ordered events with one terminal event.

  • Retries: send an Idempotency-Key and a retry replays the original response instead of billing twice.
  • Errors: typed envelopes tell you whether to retry, wait, or fix the request.
  • Keys: mint and revoke organization keys from the console.

Independent evaluation

We do not publish benchmark scores for this model — we have not verified any. These external leaderboards run their own tests, so check them against your own workload.

Not our numbers