Lilith Lilith.
CS EN PL

Zvi Mowshowitz reads Kimi K3 soberly: 2.8 trillion parameters, excellent benchmarks and potentially the strongest open model if the weights ship. He also warns about spotty access, token hunger and practical performance lagging the charts.

Kimi K3 looks like the largest open model of its moment

Zvi’s core thesis has two parts. Kimi K3 is, in his view, a very good model with excellent benchmarks and, if the weights ship as planned, could be the strongest open model in raw capability.

The numbers are large. Kimi K3 is described in the announcement he quotes as a 2.8 trillion parameter model with native vision capabilities and a 1 million token context window. Zvi argues that size explains part of why the model looks strong.

At the same time, he rejects the simple story that China has caught the closed frontier. His estimate places Kimi K3 several months behind the best closed models, with post-training closer and pre-training farther out.

Developers should care more about workflow fit than leaderboard rank

Kimi K3 may be practically useful where its profile fits: agentic coding, long tasks and some knowledge work. That is exactly where open weights can matter even without beating the best closed model outright.

For teams, the question is whether the model fits the workflow. If it is slow, token hungry or expensive in real use, part of the benchmark shine disappears. Performance per million tokens is not the same as performance per developer workday.

Maximum-effort benchmarks can sell a different model than production sees

Zvi notes that Kimi K3 benchmarks are often scored at maximum effort and with more tokens than comparable tests for rival models. That does not make them fake. It means the comparison may mix model quality, thinking time and willingness to spend inference.

He also warns about jagged performance. The model may be excellent at some tasks and weaker elsewhere. That matters more to buyers than one big table, because production punishes unevenness harder than Twitter does.

The first weeks will expose price, latency and blind spots

Independent evals and user reports outside demos will decide the first round. Watch latency, cost per completed task, quality in agentic coding and behavior in less popular domains.

The second signal is safety. Zvi explicitly flags the need for bio and cyber testing of open models. If the open-weight frontier moves upward, knowing that a model can write code is not enough. Teams need to know which doors it opens, and for whom.

Lilith's verdict

Kimi K3 is a large animal in a cage with the door open. Before the applause, someone has to check whether it can run a marathon or just pose nicely under the lights.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗