← Journal

Z.ai is holding GLM 5.3's weights back for two weeks

Z.ai launched GLM 5.3 on August 14 but held its weights back, on a day DeepSeek and an anonymous provider shipped models only as API endpoints.

Z.ai launched GLM 5.3 on August 14 and then did not publish the weights. The model is described as open-weight, but for about two weeks the only way to reach it is through the company's own GLM Coding Plan and ZCode agent, plus a set of vetted security partners. The reason given is a number Z.ai produced itself and has not had verified: 84.5 percent on the CyberGym vulnerability-finding benchmark, against 77.2 percent for GLM 5.2 (BetaNews).

A model good enough at finding vulnerabilities is exactly the case where a lab would rather hand out calls than files. Three more items in today's ledger show the same layer thickening from different directions.

What ships is an endpoint

DeepSeek put DeepSeek-V4-Flash-Vision-Exp on its API platform, an experimental multimodal build the company says holds V4-Flash's text, reasoning and agent scores while bringing multimodal agent performance close to Anthropic's Opus 4.8 (The Times of India). Experimental is the load-bearing word. A build that lives on an API can be revised, rate-limited or withdrawn on the provider's schedule, and the person who benchmarked it on Tuesday cannot prove on Friday that they measured the same thing. A downloaded checkpoint is a fixed object; a served one is a service.

OpenRouter is hosting the other end of that trade. Ox Alpha appeared last Thursday, free to use, with a context window just over a million tokens, from a provider that will not say who it is and that retains every prompt and completion sent to it. The open-source agent OpenCode puts its capacity at 100 trillion tokens a day (TNW).

Free, a million tokens of context, and no name on the receipt. Nothing about that is hidden: the retention is stated. It is simply the shape the bargain takes once the model is a socket rather than a file. You cannot audit the weights, you cannot pin the version, and the traffic you send is the consideration.

The traffic has to land somewhere

Cursor launched Origin, a Git forge that hosts repositories and syncs projects from GitHub, built on a storage system it calls Continuity that runs on S3 and write-ahead logging. The pitch is explicit: code hosting rebuilt for the agent-driven traffic GitHub's architecture is straining under (The Stack). Anysphere, which builds Cursor, agreed in June to sell itself to SpaceX for $60B in stock (Yahoo Finance), so this is not a side project competing for attention.

Cerebras introduced CS-4 on the same day, claiming 30 times the performance of CS-3 and up to 10 times the throughput per watt (The Stack).

A forge rebuilt for machine-rate commits and a system built for throughput per watt are answers to the same question the endpoints are asking. If models are consumed by agents through APIs rather than pulled down and run, then the scarce things are serving capacity and a version-control layer that does not fall over when the committer is a program.

What today's record shows

Four launches, and the artifact is missing from all of them. Z.ai's weights exist and are not downloadable yet. DeepSeek's build is experimental and lives on a platform. Ox Alpha has no published provider at all. Cursor's answer to GitHub is a hosted service with its own storage engine underneath.

The word open-weight is still being used, and today it described a model nobody outside a partner list can run. That is not a contradiction the industry is hiding; it is a default quietly changing. Access is becoming something granted per call and revocable, and every figure above was measured against a thing that may not be the same thing tomorrow.

Built from the digest of 2026-08-23.