Lilith Lilith.
CS EN PL
Editorial illustration: Baseline models are fine for the 99 percent
Lilith illustration · editorial remix

A step back from the capability bubble

Researcher David Ha (@hardmaru) highlighted a growing trend on X: current baseline models like Gemini 3.1 Pro are perfectly fine for 99 percent of daily tasks. While the primary platform was blocked during verification, the raw signal and community engagement point to a shift away from waiting for the next massive capability jump.

Most business tasks do not require an expert

While industry benchmarks focus on complex math and expert coding, real-world adoption is completely different. People use LLMs for summarization, data extraction, and basic classification. Pushing massive compute for these tasks is a waste. The shift towards models like Gemini 3.1 Pro shows that users are finally optimizing for reliability and cost over raw intelligence.

Coding remains the exception, not the rule

The real boundary lies in separating use cases. The original take admits that frontier models are still necessary for heavy engineering and complex code. The mismatch happens when companies evaluate AI tools on coding benchmarks but deploy them for administrative tasks. This asymmetry leads to overpriced pipelines for simple jobs.

Enterprise budgets will pick the winners

Adoption outside the vendors' own narratives will decide the market. The real signal won't be another lab leaderboard, but whether large enterprises shift their API volume to middle-tier models. If the majority realizes that baseline models are enough, the war for the most expensive frontier model will lose its commercial justification.

Lilith's verdict

Enterprises are finally realizing they do not need an OS-writing model to summarize an invoice. Frontier tools are becoming a luxury for developers, while the business gets by just fine on the baseline.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗