GLM Coding Plans Are Back, Starting at RMB 118

Zhipu has reopened subscriptions to the GLM Coding Plan, replacing vague quotas with a transparent points-based system. The new rules make costs easier to calculate, but developers should still carefully compare point multipliers, peak and off-peak pricing, and the experience offered by the high-speed version.
GLM Coding Plan Returns: Zhipu Is Not Just Selling Subscriptions, but Reworking the Compute Ledger
On July 31, Zhipu reopened subscriptions to the GLM Coding Plan. The new plans start at RMB 118 per month, with input tokens, output tokens, cache hits, different models, and MCP capabilities all converted into credits according to publicly disclosed rules. From now through August 15, 2026, the new annual and quarterly plans are available for a limited time at 30% and 20% discounts, respectively.
The key point of this return is not a price cut, but the transformation of an AI coding bill that was previously difficult to calculate into a relatively transparent credit account.
For developers who have agents read repositories, modify files, and run tests every day, this matters more than simply receiving a few extra calls. What makes AI coding products unsettling is often not their high prices, but the fog surrounding their usage limits: How much does a single conversation consume? How are long contexts, model switching, and tool calls billed? Users have little ability to predict any of this in advance. This time, Zhipu is at least attempting to clear away the fog.

From Restricted Availability to Reopening Sales, It Still Comes Down to Compute
GLM Coding Plan is Zhipu's subscription service for AI coding scenarios, covering natural-language programming, code debugging and repair, repository Q&A, and automated task processing. It is not aimed at light users who occasionally ask a model to complete a function, but at developers who want to integrate models into terminals, IDEs, and agent workflows.
Previously, Zhipu temporarily restricted the number of available subscriptions due to rapid growth in AI coding demand. The company says today's reopening follows continued infrastructure expansion. Related reports also state that Zhipu has begun building a 1 GW domestic AI computing data center powered entirely by domestically produced AI chips.
The 1 GW figure is certainly eye-catching, but data center scale cannot be directly equated with the effective inference capacity currently available to the Coding Plan. Whether server racks come online as scheduled, how efficiently domestic chip clusters are utilized, whether inference frameworks have been specifically optimized, and how efficiently workloads are scheduled during peak periods will all affect the speed and stability developers ultimately experience.
In other words, reopening subscriptions indicates that Zhipu believes supply pressure has eased, but the real test will begin only after users start flooding back in. AI coding has a pronounced tidal pattern: request volumes rise rapidly on weekday mornings, before releases, and during intensive debugging periods. Agents also make multiple consecutive model and tool calls within a single task, producing a steeper concurrency curve than ordinary chat products.
Zhipu's decision to classify entire weekends as off-peak hours also indirectly reflects the core of this business: it is not just selling model capabilities, but also using pricing mechanisms to steer demand toward periods of idle compute capacity.
A Transparent Credit System Answers: "How Much Longer Can I Use It?"
The new plans use a credit system, with the following types of consumption converted according to publicly disclosed rules:
- Input tokens, including user instructions, project context, and code read by the model;
- Output tokens, including explanations, patches, code, and task-planning results;
- Cache-hit tokens, meaning consumption generated when repeated context is reused;
- Calls to different models, with different capabilities assigned different credit multipliers;
- MCP capabilities and related tool calls.
Credits do not make billing itself simpler. On the contrary, they add another unified settlement unit on top of tokens. But as long as the conversion rules remain stable and usage records are sufficiently detailed, developers can account for these complex forms of consumption in a single ledger.
Tokens can be thought of as water meter readings, while credits represent the final water bill. Different models are like water sources with different prices, cache hits resemble discounted water rates, and MCP tool calls are like additional service fees. What developers truly care about is not how many characters flowed through the system, but how many code reviews, rounds of repository-level Q&A, and automated repair tasks a month's worth of credits can support.
| Use Case | Main Sources of Consumption | Potential Impact on Credits | | --- | --- | --- | | Single-file completion | Small numbers of input and output tokens | Usually low and easy to estimate | | Repository-level Q&A | Large amounts of code context | Input and caching strategies matter more | | Automated agent repair | Multiple rounds of reasoning, tool use, and MCP calls | Consumption per task can vary significantly | | Generating and repeatedly running tests | Output tokens and consecutive tool calls | Can easily result in cascading consumption | | Batch refactoring on weekends | Long contexts and multiple rounds of changes | Can benefit from off-peak multipliers |
This design is more honest than vague promises of "several times the base quota," but transparency does not necessarily mean affordability. Whether the credit system is ultimately useful depends on three details: first, whether model multipliers change frequently; second, whether the usage dashboard can provide breakdowns at the conversation, task, or even tool-call level; and third, whether reliable warnings are provided as a plan approaches exhaustion.
If users can only see their total credits decrease each day without knowing which repository scan or MCP service consumed the most, credits merely replace a black box measured in "calls" with one measured in "points." Zhipu now needs to prove that transparency means more than a single rules table—that it extends across a complete experience encompassing usage queries, alerts, and retrospective analysis.
Starting at RMB 118, the Price Is Reasonable, but the Monthly Fee Is Not the Whole Story
The new GLM Coding Plan starts at RMB 118 per month. Within China's developer tools market, this price is not aggressive, but neither is it cheap enough to make the decision inconsequential.
Light users who only occasionally complete code or ask for error explanations may get better value from usage-based APIs, free quotas built into IDEs, or general-purpose chat products. Coding Plan is better suited to people with stable call volumes who are willing to keep models embedded in their development workflows—for example, those who conduct multiple rounds of repository Q&A every day or assign agents to test generation, refactoring, and bug fixing.
When evaluating a plan, developers should not compare only monthly fees. They should examine the cost per task:
- How many credits does it take to index or scan a medium-sized repository?
- How much does a complete agent task—from locating a problem to submitting a patch—consume?
- Can long contexts consistently hit the cache, and is the discount after a cache hit significant?
- How much do peak and off-peak multipliers affect their working hours?
- When the plan runs out, does the service stop, slow down, or allow users to purchase additional credits?
The 30% discount on annual plans and 20% discount on quarterly plans can reduce the nominal monthly cost, but long-term lock-in also carries risks at a time when AI coding products are evolving rapidly. Model capabilities, IDE integrations, agent toolchains, and competitors' pricing can all change within a few months. Unless a team has already tested the service on real projects and established stable usage patterns, it is generally safer to estimate costs with a short-term plan before committing to an annual subscription solely for the discount.
Existing User Benefits Remain Largely Unchanged, While v1 Users Must Still Wait for a Migration Option
Zhipu says users who currently hold plans, including team-plan users, will not be affected by this adjustment in terms of pricing, benefits, quotas, or calculation methods. They can continue using, renewing, and upgrading their original plans.
The situation is more complicated for the earliest group of users on v1 plans without weekly rate limits. Because Zhipu discontinued its earliest plan offerings in February this year, these users currently cannot renew or upgrade their original plans. According to the announcement, their benefits will remain unchanged during the current validity period, and they will be able to purchase and continue using the service long-term at the v2 price that applied before this adjustment, as long as they do so before their current plans expire. The corresponding feature is expected to launch within two weeks.
This is a relatively gentle migration plan, but "launching within two weeks" still means the feature is not yet actually available. Existing users should monitor when the purchase option becomes available, what subscription periods can be purchased, and whether the promised long-term usage comes with any additional restrictions, so they do not face a service gap as their original plans approach expiration.
In April this year, Zhipu stopped automatic renewals for legacy plans without weekly usage limits, citing the need to unify its plan structure and provide stable support for future feature and benefit upgrades. Today's launch of the credit system shows that the earlier plan restructuring was not an isolated move, but an effort to make room for a new resource-scheduling and billing framework.
Applications Are Open for the High-Speed Version, but Speed Should Not Be Measured Only by Time to First Token
Coding Plan users can now apply to try the high-speed version, with users of legacy plans receiving priority access.
"High speed" is certainly valuable for AI coding, but it cannot be measured solely by time to first token. Coding agents typically go through multiple stages, including reading context, planning steps, calling tools, modifying files, running tests, and making further corrections. Even if a model outputs dozens more tokens per second, the overall task can still take a long time if tool queues are congested, MCP calls are slow, or concurrency is restricted.
A genuinely useful high-speed version must provide stable performance in at least the following areas:
- Time to first token and output speed should not fluctuate significantly during peak periods;
- Long-context processing should not frequently time out or fall back to degraded modes;
- Consecutive calls in multi-round agent tasks should not be subject to strict rate limits;
- MCP tool failures should provide clear error messages and retry mechanisms;
- Faster performance should not come at the cost of significantly higher credit consumption.
Giving legacy-plan users priority access is a reasonable staged-rollout strategy. These users tend to have higher usage frequency and are more likely to expose problems under real workloads. There is still insufficient information about whether the high-speed version will become a separately billed benefit or what credit multiplier it will use, so developers should not form expectations based solely on the words "high speed."
This Adjustment Is an Improvement, but the Outcome Will Not Be Decided on the Billing Page
The transparent credit system is a sound product adjustment. AI coding has evolved from autocomplete plugins into agents capable of continuously executing tasks, and the traditional model of "a certain number of conversations per month" can no longer accurately describe resource consumption. Bringing models, tokens, caching, and MCP into a single queryable system at least gives users the ability to calculate costs.
For developers, however, clear billing is merely the baseline, not a reason to buy. Whether GLM Coding Plan can retain users over time will still depend on how its models perform in real repositories: whether they can understand cross-file dependencies, comply with existing project constraints, autonomously correct failures after tests, and produce patches that are both minimal and trustworthy.
It also faces a competitive market reality. Developers can choose native coding subscriptions, or they can use platforms compatible with the OpenAI API to access different models—including GPT, Claude, Gemini, DeepSeek, and GLM—and integrate them into their existing IDEs or agent tools. OpenAI Hub already supports this type of unified multi-model access. Fixed plans offer stable budgets and out-of-the-box usability, while multi-model APIs allow users to switch capabilities and costs based on the task without being locked into a single model.
The value of GLM Coding Plan therefore ultimately depends on two things: whether the models included in the plan are good enough, and whether the credit economics truly add up. If RMB 118 can reliably cover most daily tasks for a high-frequency developer, the plan will be competitive. If just a few repository-level agent runs noticeably deplete the credit balance, even the most transparent rules will only expose its poor value for money more quickly.
Reopening subscriptions is only the first step. Peak-period stability, credit breakdowns, and the performance of the high-speed version over the next month or two will be the real stress test of Zhipu's capacity expansion.
References
- ITHome: Zhipu GLM Coding Plan Subscriptions Return With a Transparent Credit System, Starting at RMB 118 per Month—An overview of the new plan pricing, limited-time discounts, credit rules, high-speed version trial, and arrangements for existing users.
- Linux.do: Discussion of the GLM Coding Plan's Return—A compilation of developer community discussions about the new subscriptions, credit system, and pricing changes.



