Kimi K3's open weights are live — but at 2.8 trillion parameters, 'open' doesn't mean you can run it
TL;DR: Moonshot published Kimi K3’s full weights on Hugging Face on July 26 (~7:30 PM EDT, a day early) — the largest open-weight release ever at 2.8 trillion parameters, and now irreversibly public (which is exactly why US policy targets chips, not weights). The number that matters: ~594 GB in BF16, needing 4–8 H100 GPUs minimum — and no consumer hardware (RTX 4090, Mac Studio) can load it even quantized (~300–400 GB at Q4). What “open” actually means here: a new cheap hosted option (Together AI and Modal went live day-0) and downward price pressure — not a model you personally run. What this means for you: unless you have a GPU cluster, treat K3 as a cheaper API to rent, not weights to self-host — and factor in the ongoing Moonshot sanctions cloud.
What dropped
The most-anticipated open-weight release of the year is done. On July 26, 2026, around 7:30 PM EDT — a day ahead of its July 27 target — Moonshot published the full weights for Kimi K3 on its Hugging Face organization. Per VentureBeat and community trackers:
- 2.8 trillion parameters — the largest open-weight release in history, and Moonshot’s “first open 3T-class model.” It’s a Mixture-of-Experts design that activates 16 of 896 experts per token.
- New architecture — built on Kimi Delta Attention and Attention Residuals, with native agentic capabilities (tool calling, browsing, multi-step planning) and an extended context window aimed at repository-scale code understanding.
- Size: roughly 594 GB in BF16; community Q4 GGUF quantizations are expected to bring that to ~300–400 GB.
- Hardware floor: a minimum of about 4–8× H100 80GB GPUs for experimental use. Consumer hardware cannot load it — not an RTX 4090, not a Mac Studio, even quantized.
- Ecosystem: 505 people were queued on Hugging Face before it went live; Together AI and Modal both launched day-0 hosted access, shifting competition to whoever serves K3 fastest and cheapest.
And the fact that overshadows all of it: once these weights are public, they can’t be recalled.
Why this matters
1. “Open weights” and “you can run it” are not the same thing — and at 2.8T, the gap is a canyon. This is the single most important correction to how the release is being covered. “Free to download” sounds like democratization, but a 594 GB model needing 4–8 H100s is not something a developer runs on a laptop, a workstation, or even a single high-end GPU. When we covered K3’s launch, we flagged that self-hosting a 2.8T model was a real option for almost nobody. The weights being public confirms it with hard numbers: for the overwhelming majority of people and companies, “open” means someone else will host it cheaply, not “it’s yours to run.” The democratization is of price and access, not of self-hosting.
2. The real winners are the hosted providers, and the real benefit to you is cheaper inference. Together AI and Modal going live on day zero is the actual mechanism by which open weights help you: they turn a 594 GB file into an API endpoint at commodity margins, and they compete on price to serve it. For a buyer, that means K3-class capability becomes available at hosted-provider prices, likely well below frontier closed-model rates — the same dynamic DeepSeek created, now at the top of the open-weight size range. If K3 is genuinely strong on your tasks, the move isn’t to buy 8 H100s; it’s to price out Together, Modal, and other hosts.
3. Self-hosting is now a real option — but only for the well-resourced, and mostly for data control. For organizations that can field the hardware, the open weights unlock something the API never could: full data sovereignty and air-gapped deployment. No data leaves your infrastructure, no dependency on Moonshot’s servers, no China-routing concern. Given the sanctions cloud over Moonshot, that insulation is genuinely valuable — a self-hosted model is far less exposed to vendor risk than a hosted Chinese API. But this benefit is gated behind serious capital; it’s a Fortune-500-and-labs option, not an indie-developer one.
4. It’s the concrete proof of yesterday’s policy logic. We argued that the US response to K3 is chip export controls, not an open-source ban, because you can’t un-release open weights. Twelve hours later, the weights are on Hugging Face, mirrored, and downloading worldwide — un-bannable in practice. The policy debate is now purely academic on the model: it exists, it’s out, and no rule can claw it back. All remaining leverage is on the inputs (chips) and the vendor (sanctions on Moonshot), exactly as the confirmed legislation targets.
5. The capability question is still open — and worth waiting on. K3 took #1 on a frontend-code leaderboard at launch, but Moonshot itself conceded it trails Claude Fable 5 and GPT-5.6 Sol overall, and a UK AISI/CAISI assessment found it significantly below frontier on cyber. Now that the weights are public, independent benchmarks on the actual model — not Moonshot’s own numbers — will start landing. Those are the results to trust. Until they do, treat K3 as “near-frontier on some tasks, unproven overall,” and test on your own workload before committing.
What this means for you
- Don’t buy hardware to run K3 — rent it. Unless you already operate a multi-H100 cluster, use a hosted provider (Together AI, Modal, and others racing in). That’s where the price benefit of “open weights” actually reaches you.
- If data sovereignty is a hard requirement and you can afford it, self-hosting the open weights is now genuinely possible and side-steps both the China-server data path and vendor-transaction risk. Budget for 4–8 H100s minimum.
- Wait for independent benchmarks before migrating anything important. Moonshot’s own numbers are favorable; the neutral ones on the released weights are what matter. Compare against your current tool using the best AI coding tools guide.
- Keep the vendor/sanctions overhang in view. The weights are legal to run, but building a commercial dependency on Moonshot carries the Entity-List risk we’ve flagged. Self-hosting mitigates it; a hosted Moonshot API does not.
- If you were considering open weights for cost, right-size the model to your task. K3’s 2.8T scale is overkill (and over-budget) for most work. Smaller open models — DeepSeek V4, Qwen, GLM-5.2 — are far cheaper to self-host and often good enough. Reach for K3 specifically when you need its peak capability on a task where it genuinely leads, and let a host carry the hardware.
The honest caveats
- Independent benchmarks on the released weights are still emerging. As of this writing, the strongest capability claims remain Moonshot’s own. The value of the open release is that neutral testing can now happen — but it hadn’t fully happened yet.
- Hardware figures are early community estimates. The ~594 GB (BF16), ~300–400 GB (Q4), and 4–8 H100 numbers come from community analysis in the first hours after release; expect refinement as more people actually deploy it.
- Read the license — don’t assume. K3’s license terms may differ from the earlier K2-line terms, and quantization/redistribution rights matter for commercial use. Confirm from Moonshot’s own repo, and verify you’re downloading from the genuine
moonshotaiorganization (copycat repos exist). - “Open” is doing a lot of work. Open weights aren’t the same as open training data or a fully permissive license. Judge the openness by the actual license, not the headline.
- The sanctions situation is unresolved. BIS is investigating and Entity-List action remains possible. That wouldn’t delete the weights, but it could complicate commercial use, hosting, and support. Track it.
The grounded summary: the largest open-weight model ever is now a free, un-recallable download — a real milestone, and a real irreversibility that reshapes the policy fight. But the practical truth for buyers is smaller and more useful than the headline: at 2.8 trillion parameters, “open” means a cheaper model to rent, not one you’ll run, unless you own a GPU cluster. Rent it from a day-0 host, wait for independent benchmarks, and keep the Moonshot sanctions cloud on your risk register.
Frequently asked questions
Are Kimi K3's weights actually available now?
Yes. Moonshot published the full model weights on its Hugging Face organization (huggingface.co/moonshotai) on July 26, 2026, roughly 7:30 PM EDT — a day ahead of the July 27 target. It's the largest open-weight release in history at 2.8 trillion parameters, and once public, it can't be recalled. Always confirm you're downloading from Moonshot's verified organization and read the actual K3 license file before using it commercially.
Can I run Kimi K3 on my own computer?
No. At ~594 GB in BF16 format, K3 requires a minimum of about 4–8 H100 80GB GPUs for experimental use. Community Q4 quantizations may cut the footprint to ~300–400 GB, but even those are far beyond consumer hardware — an RTX 4090 or a Mac Studio cannot load the model even quantized. Realistically, self-hosting K3 is for organizations with serious GPU clusters.
So what does 'open weights' actually get me?
For most people, a cheaper hosted option and downward price pressure — not a model you personally run. Hosted providers like Together AI and Modal went live with day-0 K3 hosting, and once weights are public, competition shifts to whoever serves them fastest and cheapest. The self-host benefit (full data control, air-gapped deployment) is real but only accessible to well-resourced organizations that can afford the hardware.
Is it safe to build on Kimi K3 given the US allegations?
Weigh two separate things. Technically, the weights are downloadable and legal to run today. But Moonshot is under a US export-control/distillation cloud — the White House alleged distillation of Anthropic's Fable and banned-chip access, and BIS is investigating. That doesn't restrict the model itself, but it creates vendor/sanctions risk if you depend on Moonshot commercially. Self-hosting the open weights actually insulates you somewhat, since you're not transacting with the company or routing data to its servers.
How does Kimi K3 compare to running DeepSeek or Qwen?
K3 is dramatically larger (2.8T vs. DeepSeek's and Qwen's smaller footprints), so it's far harder and costlier to self-host — the hardware bar is much higher. For most self-hosting use cases, smaller open models like DeepSeek V4 or Qwen are more practical. K3's edge is peak capability on specific tasks (it took #1 on a frontend-code leaderboard), best accessed through a hosted provider unless you have datacenter-scale compute.
Sources
- Kimi K3 Open Weights Arrive Sunday: Self-Hosting Cuts China Data Risk the API Never Can (Tech Times)
- Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community (Hugging Face)
- Kimi K3: The open-weights escalation (Interconnects, Nathan Lambert)
- China's Moonshot AI releases Kimi K3, the largest open-source model ever (VentureBeat)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.