2026-08-07 · ← News
ByteDance Reportedly Training a Model With Ten Trillion Parameters
A Chinese response to the scale of US models
ByteDance, the owner of TikTok, has reportedly embarked on training a massive AI model aimed at reaching 10 trillion parameters, according to non-public sources. If these figures are accurate, it would be three times the size of Moonshot’s Kimi K3, which is currently the largest announced model out of China.
The information comes from leaks and remains officially unconfirmed, but it aligns with the strategy of major Chinese players. They are striving to keep pace with US labs like OpenAI or Anthropic and, despite export restrictions, are sparing no expense on hardware infrastructure.
US sanctions have not halted scaling yet
This move is particularly significant in the context of US sanctions on the export of advanced AI chips. The sanctions were intended to slow down Chinese AI development by restricting access to state-of-the-art hardware from Nvidia.
However, it appears that companies like ByteDance can bypass or compensate for these restrictions. They may be utilizing vast reserves of older, less efficient chips connected into massive clusters, or relying on domestic alternatives. The ability to launch a training run with 10 trillion parameters demonstrates that compute limitations are not an insurmountable hurdle for large Chinese tech giants.
Massive parameters do not guarantee an immediate advantage
The sheer size of a model does not guarantee top-tier performance. Even with 10 trillion parameters, ByteDance will face the same challenges as Western labs: the quality of training data, the stability of massive clusters, and the costs of inference.
If the model reaches production, the question will be its application. ByteDance could integrate it directly into TikTok for hyper-personalization of content, or attempt an expansion into enterprise AI services, where other players currently dominate.
Real impact will be determined by cost and availability
The proof of success will not just be the completion of the training run. The key signal to watch is whether ByteDance can operate such a massive model economically, and whether its deployment yields a measurable difference compared to smaller, more efficient competitor models.
Lilith's verdict
The US tried to limit China’s compute with export sanctions. ByteDance just handed them the receipt for a cluster proving that with enough money, you can scale even on subpar hardware.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗