An engineering team sees a new model with published weights, a long context window and strong benchmark claims. The first reaction is understandable: perhaps the team can run it privately, reduce API costs and depend less on a large US model provider.
That conclusion may be right. It is not automatic.
DeepSeek V4.1 Flash, Z.ai’s GLM 5.3 family and Moonshot AI’s Kimi K3 show how quickly Chinese labs are advancing open-weight AI. They also expose a gap that model announcements often hide. Being able to download a model is different from being able to operate it safely, economically and reliably.
First, open weight is not the same as open source
An open-weight model makes its trained parameters available for download. This can let a team inspect, adapt and host the model without sending every request to the developer’s API.
The term does not guarantee that the complete training data, training code or training process is available. It also does not guarantee unrestricted commercial use. Each model comes with its own licence and usage conditions.
That distinction matters here. DeepSeek publishes V4.1 Flash checkpoints and technical details.[1] Z.ai publishes GLM 5.3 and GLM 5.3 Flash weights.[3] Moonshot publishes Kimi K3 weights under its own model licence.[5] A procurement or legal review still needs to examine the exact version a team plans to use.
Accessible text alternative for Figure 1
A six-layer stack showing that published model weights require licence review, infrastructure, inference serving, governance and workload evaluation before becoming a useful AI service.
Three models, three different operating questions
The headline features do not answer the same business question.
| Model | Documented design signal | What a team should ask |
|---|---|---|
| DeepSeek V4.1 Flash | 552 billion backbone parameters, with 16 billion active during decoding and 8 billion during prompt processing. Supports up to one million tokens. | Can its compressed context design lower serving cost for long, input-heavy agent workloads? |
| GLM 5.3 Flash | 320 billion total parameters and 18 billion active parameters. Native multimodal support and an efficiency-focused hybrid attention design. | Does the faster model meet the team’s coding and multimodal needs without requiring the larger GLM 5.3? |
| Kimi K3 | A 2.8 trillion parameter multimodal model with a one-million-token context window and published weights. | Is its capability worth the much larger hosting and operational footprint? |
These are vendor-documented specifications, not proof that one model is universally better. Independent testing also changes with the provider, reasoning setting, quantisation method and evaluation harness.[7] A leaderboard rank should start an evaluation, not finish one.
Mixture of experts helps compute, not model storage
All three model families use a mixture-of-experts, or MoE, design. Instead of using every parameter for every token, the model activates only part of the network.
This can reduce the computation required for each response. It does not make the inactive weights disappear. The system may still need to store and move a very large model across GPU memory, host memory and storage.
This is why an active-parameter number can be misleading when viewed alone. A model that activates 16 billion parameters may still have hundreds of billions of parameters to load and coordinate. Network links, memory bandwidth, inference kernels, cache management and parallel serving become part of the real product.
These are not ordinary laptop models. Smaller quantised variants may become available, but a quantised community build should be evaluated as a separate artefact. It may not behave exactly like the vendor’s hosted model.
Long context is capacity, not guaranteed understanding
One-million-token context windows are now a visible part of this competition. They can help with large codebases, document collections and long agent sessions.
Yet context capacity does not prove that a model will find the right detail, preserve instructions or reason consistently across the entire window. Longer prompts also increase processing time, cache requirements and cost.
DeepSeek V4.1 Flash is technically interesting because it targets this operating problem directly.[1] Its paper describes compressed attention and smaller key-value caches for input-heavy workloads. That is an architectural response to a real bottleneck. A team still needs to test retrieval accuracy, time to first token and end-to-end task completion with its own data.
API price and self-hosting cost are different calculations
A low API price is easy to compare. Self-hosting cost is spread across GPUs, power, engineering time, monitoring, security, upgrades and spare capacity.
A busy service with predictable demand may justify dedicated infrastructure. A smaller team with irregular traffic may pay more to keep GPUs available than it would spend on API calls. Hosting can still be the right choice when privacy, data location, customisation or provider independence matters more than the lowest immediate cost.
The comparison should therefore use cost per successful task, not only cost per token. A cheaper model that needs more retries, produces longer answers or fails tool calls may cost more in the workflow that matters.
A practical evaluation order
Do not begin by buying hardware. Begin with the workload.
- Select twenty to fifty real tasks that represent normal work and difficult edge cases.
- Test the hosted model or a trusted inference provider before building a private cluster.
- Measure task success, latency, tool-use reliability, output length and human correction time.
- Review the model licence, data path, logging policy and regional requirements.
- Estimate self-hosting with realistic utilisation, redundancy and engineering support.
- Pilot private hosting only when it solves a requirement the API route cannot meet.
For most teams, the immediate lesson is not that they should self-host DeepSeek, GLM or Kimi. It is that they now have more credible models to evaluate and more leverage when choosing how AI work is delivered.
The Chinese open-weight model race is expanding choice. The winning model for an enterprise will not be the one with the largest parameter count or the strongest launch chart. It will be the one that completes the required work within the team’s cost, control and operational limits.
References and further reading
- DeepSeek V4.1 Flash technical paper, DeepSeek AI, September 2026. Architecture, context and cache-compression details. ↩
- DeepSeek V4.1 Flash model repository, DeepSeek AI. Checkpoints, model card and licence for the released version. ↩
- GLM 5.3 release, Z.ai, August 2026. Vendor explanation of the model and its intended workloads. ↩
- GLM 5.3 Flash release, Z.ai, August 2026. Architecture and efficiency claims for the smaller active model. ↩
- Kimi K3 technical article, Moonshot AI, July 2026. Model design, evaluations and release context. ↩
- Kimi K3 model repository, Moonshot AI. Published weights, model card and licence. ↩
- Kimi K3 independent model analysis, Artificial Analysis. Provider, performance and cost comparisons using its published methodology. ↩
