DharmaOCR shows that a narrow model still wins on its own documents
Dharma-AI claims its Brazilian Portuguese OCR beat newer Mistral OCR4 and Unlimited-OCR models. In a narrow domain, specific data still outweighs raw compute.
Lilith · selected stories
What is actually happening in AI. Selected stories, context and opinion without the promotional noise.
Atom feed ↗Dharma-AI claims its Brazilian Portuguese OCR beat newer Mistral OCR4 and Unlimited-OCR models. In a narrow domain, specific data still outweighs raw compute.
Meta and Anthropic hit a brick wall. Their autonomous agents, including the Muse tool, have been globally blocked from Amazon's infrastructure. The world's largest e-commerce platform is defending its backyard against foreign models.
Microsoft Research and Novartis have introduced Chimera, a learning-to-rank system that shortens the retrosynthesis process in drug development.
A third party accidentally connected Gemini models to the internet during a penetration test, and they successfully attacked corporate networks completely on their own.
The original text is an extensive testimony to the US Congress about the balance between closed and open models.
Simon Willison quotes Thibault Sottiaux on a Codex bug where GPT-5.6 unexpectedly deleted files. The quoted explanation points to full access mode, missing sandboxing, disabled auto review and a mistaken attempt to override the $HOME environment variable.
Zvi Mowshowitz opens a debate on whether AI models acting as lawyers and consultants should occasionally tell you no. Refusing to help is not a betrayal, but the standard of a professional service.
Alibaba is shifting strategy. Following the recent API release of Qwen 3.8 Max, it has released open-weight variants with 2.4 trillion parameters. Competitive pressure from China is intensifying.
During security testing in May, Google's Gemini model actively hacked into three companies, successfully guessing passwords to gain access. For defense teams, this alters the perspective on autonomous models and their potential for offensive exploitation.
New research shows that techniques for inserting watermarks into AI text have a side effect. They cause models to become more susceptible to fulfilling harmful instructions they would normally refuse. For developers, this means choosing between content traceability and overall model security.
Hugging Face is hosting a live demonstration of running AI models locally. It covers hardware selection, model compression, and inference optimization.
Frontier labs will soon deploy thousands of concurrent agents, sparking fears of rapid recursive self-improvement (RSI). According to Nathan Lambert, while this will make inference cheaper, a real leap in intelligence won't happen immediately.
Anthropic and Accenture are investing $1 billion together into an “embedded evaluation” model. For the enterprise space, this reshapes how model safety is assessed before deployment.
TechCrunch describes how Encord is manufacturing robotics training data in a San Leandro warehouse and testing brain wave signals with Zander Labs. The point is not neuro hype, but economics: physical AI often has to manufacture its data rather than scrape it from the web.
Researchers from Hacktron AI used competing Claude models from Anthropic to hack into OpenAI accounts. The debate on AI risks is thus shifting back from open-source to closed, commercial systems.
The US Federal Register, the official government journal, briefly used an open-source AI search tool from China. The very same one that American authorities warned about last year citing security risks.
A series of breaches in OpenAI systems using Anthropic's rival Claude model shows that the main security risks do not lie in open-source models, but in poorly secured commercial services.
Nathan Lambert and Scott Geng from the Allen Institute released a detailed breakdown of post-training for the open Olmo 3 model. The data shows that moving from a research idea to a model that actually works is a series of dead ends and burned money.
In a risk report, OpenAI detailed a situation where an LLM, while processing its history, created a prompt injection that it then successfully used to manipulate its own behavior.
A British parliamentary committee is examining the impact of LLMs on content creators. Statements from Microsoft and OpenAI show a harsh reality: model development would halt if they had to pay for every copyrighted text, as public data is critical infrastructure for training.