CO/AI Subscribe
Thursday · July 23, 2026 · Issue No. 934
The Week Kimi K3 Broke the Internet: Inside Moonshot AI’s Biggest Launch Yet
Essay

The Week Kimi K3 Broke the Internet: Inside Moonshot AI’s Biggest Launch Yet

On July 17, Chinese AI lab Moonshot AI released Kimi K3, and by every measure available, the internet did not see it coming. Within 48 hours, demand outstripped capacity so badly that Moonshot suspended new subscriptions. Chip stocks briefly dipped into what one Bloomberg anchor called a “bare market” before rebounding. And multiple outlets independently reached for the same comparison: this was “the DeepSeek moment,” again.

The numbers that started it

Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model, activating 16 of 896 experts per token, with a 1-million-token context window priced flat across the entire window, no long-context tax. Pricing lands around $3 input / $15 output per million tokens, Sonnet-tier pricing for output that competes with Claude Fable 5 quality. Per gate_ventures, it ranked first on Program Bench (77.8) and scored 42.0 on SWE Marathon, ahead of Claude Opus 4.8 (40.0), GPT-5.6 Sol (39.0), Claude Fable 5 (35.0), and GPT-5.5 (14.0).

Artificial Analysis, one of the more trusted independent benchmarking outfits, put it in more measured terms: Kimi K3 is second only to Fable 5 on their agentic knowledge-work benchmark, a +727 Elo jump over Kimi K2.6, but it costs more than Opus 4.8 to run and averages nearly an hour per task. That tension, frontier-level output at a fraction of the sticker price, but slower and more token-hungry underneath, ended up defining most of the week’s real debate.

The moment that actually went viral

Benchmarks are one thing. What spread was a screenshot. @Virexontic posted a side-by-side: same prompt, an armory bay scene, lighting, props, atmosphere, fed to both Kimi K3 and Claude Opus 4.8. Kimi came back with textured walls, weapon racks, ammo crates, alarm lighting. Opus came back with an empty room. The task cost three cents on Kimi. It cost Opus thirty-eight dollars.

But the same account followed up days later with the other side of that story: someone actually timed the work. Three developers gave Kimi K3, Fable 5, and GPT-5.6 Sol the same front-end task through Chase AI. Kimi used 21.5 million tokens and took an hour and 33 minutes; Fable used a fraction of that. Kimi topped the front-end leaderboard, ahead of Fable 5, on a post that pulled 18 million views, but David Sacks reportedly called the result “concerning.” Cheap and capable is not the same as fast.

Not everyone’s convinced

The most useful pushback came from actual users, not benchmarks. One thread collected Chinese developers’ real-world complaints: slow responses, weaker agent capabilities, high quota consumption, and inconsistent results on real projects, one user reported Kimi spending several hours evaluating a SaaS project without finishing.

Reddit’s reaction split along similar lines. On r/LocalLLaMA, the open-weights release reopened the national-security-risk argument that follows every major Chinese open model, met by a 409-upvote counter from u/bornlasttuesday: “Sure, if it were about defending defenders. But no, it’s about defending profits.” Someone else pointed out that gatekeeping open weights is largely theater, since “anyone with an internet connection can download it.”

And in the middle of a very serious week, the internet still found time for a viral office photo. A picture of Moonshot’s workspace, taken two days before launch, racked up 2,795 upvotes on r/singularity, mostly for jokes about monitor size and posture: “20B valuation deserves some flash for the employees. Get them bigger monitors,” followed by, simply, “bro in the back got my posture.”

Musk, markets, and money

Elon Musk called Kimi K3 “impressive,” then said xAI’s upcoming 2-trillion-parameter model “might beat” it. Moonshot’s response, posted to Weibo, was a one-line invitation: “Welcome to the 2 trillion+ club.”

The market took it more seriously than a meme exchange. Bank of America flagged the release as bullish for memory and Micron specifically, since open-weight downloads create new client-side memory demand that closed models don’t. Micron rose roughly 6% pre-market on the news. Bloomberg’s own coverage described chip stocks briefly falling into a “bare market” on the Friday before launch over questions the release raised, before rebounding as the NASDAQ gained nearly a point the following week. Reports also surfaced that Moonshot is seeking a valuation as high as $50 billion ahead of a potential IPO, a direct line from this launch’s reception to the company’s next funding round.

Worth noting: this wasn’t Moonshot’s only headline of the week. Just days earlier, Anthropic formally accused Alibaba’s Qwen lab of running the largest known distillation campaign against Claude, roughly 25,000 fake accounts generating 29 million exchanges over six weeks. That’s a separate company and a separate dispute, but it’s the backdrop Kimi K3 landed against: a week where the US-China AI rivalry was already the story before Moonshot even shipped.

What happens July 27

Kimi K3 launched as an API-accessible model, but the full open weights are promised by July 27, and Moonshot has already open-sourced Kimi Code CLI, a free terminal agent powered by K3 with video input support. Polymarket traders are currently pricing the July 27 weights release at 94%, about as close to a sure thing as prediction markets get.

Whether Kimi K3 is actually “better” than Fable 5 or Opus 4.8 depends entirely on which task you ask about, and the internet spent this week arguing exactly that. What’s harder to argue with is the reaction itself: a suspended waitlist, a rattled chip sector, a public jab at Elon Musk, and 2,795 Reddit upvotes for a photo of an office. That’s not a benchmark score. That’s a moment.

By Anthony Batt — 20+ years building software and digital media products at scale. Podcasting host at Future-Proof Podcast by CO/AI.

Share: X LinkedIn Email
Essays

More like this

All essays →
The Real AI Race Isn’t Model vs. Model Anymore. Here’s Where Poolside Actually Fits.
Essay

The Real AI Race Isn’t Model vs. Model Anymore. Here’s Where Poolside Actually Fits.

The closed frontier is down to two names, and one of them just changed Start with the giants,...

Sam Altman Told Stanford What He’d Actually Teach About Startups Now
Essay

Sam Altman Told Stanford What He’d Actually Teach About Startups Now

On May 21, 2026, Sam Altman walked into a Stanford lecture hall for the first time in a...

Google Shipped Three New Gemini Models This Week, and Still No Sign of the One Everyone Actually Wants
Essay

Google Shipped Three New Gemini Models This Week, and Still No Sign of the One Everyone Actually Wants

On July 21, Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, and Google VP...

CONSULTING

Outsider
Labs.

A management consulting team focused on AI transformations for executives and business owners.

Work with us →