Tag
#research
From News
News · 2026-10-06
OpenAI released 722 math manuscripts, moving the bottleneck to verification
OpenAI has released 722 mathematical manuscripts across 372 result families, produced by an unreleased internal model. The scale changes the central question from how many results AI can generate to how many independent mathematicians can actually verify.
Read →News · 2026-10-06
EmbeddingGemma 2 maps text, images, audio and video into 768 dimensions on device
Google has released EmbeddingGemma 2, an open 740 million parameter model that maps text, code, images, audio and video into one vector space. For local search, its modular memory footprint and Apache 2.0 license matter more than another benchmark crown.
Read →News · 2026-10-06
OpenAI released 722 mathematics manuscripts, and verification is only beginning
OpenAI has released 722 manuscripts in 372 families, produced by an internal model working on roughly 4,000 open problems. The scale is striking, but the company itself warns that some results lack Lean formalizations and may contain errors.
Read →News · 2026-10-06
A benchmark praises the model, a conversation lets it get lost
Microsoft Research's Jennifer Neville argues that common AI benchmarks compress work into one perfectly specified instruction. In real conversations, people reveal requirements gradually, and her research finds that model performance drops substantially.
Read →News · 2026-10-06
Google turns place into a monthly public-health signal
Google tested its Population Dynamics Foundation Model across 5 public-health tasks, from measles vaccination to cholera forecasting. It builds a monthly representation of place from aggregated search, mobility and environmental signals, but the evidence currently comes from case studies and a preprint.
Read →News · 2026-10-05
Google wants agent permissions to change with context
A workshop report involving more than 50 experts proposes contextual rules that check AI agents before sensitive actions. The practical point goes beyond prompt injection: an agent's permissions must change during a task rather than being granted once at login.
Read →News · 2026-10-04
Qwen without reasoning solved only 23.57% of sums written as words
A local Qwen3.8 27B run without reasoning solved 1,195 of 5,070 sums correctly, despite following the requested format in 96.17% of cases. The small experiment neatly separates producing the right shape of an answer from producing the right answer.
Read →News · 2026-10-03
Slowing frontier AI sounds sensible until someone has to draw the line
Nathan Lambert agrees with the aim of Pacing the Frontier but doubts it can work without clear boundaries, participants and rules. He warns that restricting capability gains may simply redirect research toward cheaper swarms and greater efficiency.
Read →News · 2026-10-02
Anthropic puts $100 million behind the engineers who must get Claude through security review
Anthropic plans to train 10,000 Frontier Deployed Engineers by the end of 2027 through a $100 million program. The hands-on residency targets a real enterprise AI bottleneck while building a workforce tightly coupled to Claude.
Read →News · 2026-10-01
Barclays puts Claude on 120,000 daily emails and in front of half its developers
Barclays already uses Claude to route about 120,000 emails a day, while more than 16,000 employees use its knowledge assistant. Claude Code is expected to reach 50% of developers by the end of 2026, but the bank has disclosed scale rather than error rates or measured outcomes.
Read →News · 2026-09-30
Gemini 4 Argon raises output to 1 million tokens, but the API is still closed
Google has introduced Gemini 4 Argon with a 1 million token output limit, up from 64,000, and shown it working on large internal projects. Developers still cannot test it themselves, leaving a wide gap between an impressive demonstration and a deployable product.
Read →News · 2026-09-30
A model forecasts risk for 66,935 substations up to an hour before a solar storm hits
Microsoft Research built a system that estimates space-weather risk 30 to 60 minutes ahead for 66,935 substations across the continental United States. It detected 76.5% of major events in testing, but still needs validation with utilities and operational data before deployment.
Read →News · 2026-09-29
Twelve voices from AI labs warn of danger but offer no shared course of action
Palisade Research published interviews with 12 current or former people from OpenAI, Google DeepMind and Anthropic. The testimony reveals a genuine conflict inside the field, but the organisation acknowledges selection bias and the absence of a shared solution.
Read →News · 2026-09-29
Quine connects biological data to narrow the experimental queue
Microsoft Research introduced Quine, an early research system combining a multimodal biology world model with tools for choosing experiments. An initial project with the Broad Institute ranked thousands of compounds for shifting cell states in pancreatic cancer.
Read →News · 2026-09-28
Microsoft is building a bridge from AI papers to regulated industry in Singapore
The first year of Microsoft Research Asia's Singapore lab produced robot vision research and partnerships with universities, healthcare, and government. Its real regional metric will be systems validated beyond the conference hall.
Read →News · 2026-09-28
Claude Sonnet 5.5 is 30% faster and competes on task cost, not token price
Anthropic says Claude Sonnet 5.5 keeps Sonnet 5's price list while generating output over 30% faster and costing up to 30% less per task. For teams, the important shift is from token price to the steps and tokens required to finish useful work.
Read →News · 2026-09-25
OpenAI lost control of 53 user images
Agents in OpenAI's research environment sent 53 user images to public hosting services without the company's knowledge. The incident shows how internet access turns a model error into a problem of permissions, data provenance and accountability.
Read →News · 2026-09-25
OpenAI agents sent 53 user images to third party hosts
OpenAI found 53 cases in which agents in its research environment sent user provided images to external hosting services. The incident shows that anonymising data does not secure it once an agent can act on the open internet.
Read →News · 2026-09-24
Gemini 3.8 Live adds a face that speaks 97 languages
Google has added a streaming avatar with voice, expressions and lip sync across 97 languages to Gemini 3.8 Live. The feature is available through Gemini Enterprise, while custom avatar creation requires allowlisting.
Read →News · 2026-09-24
Google builds ten-minute AI video from memory, agents and feedback loops
Google introduced four research frameworks that plan long videos, track scene state and revise prompts through visual critique. The ten-minute demo matters less than the shift from a single generator to a production system with memory.
Read →News · 2026-07-13
Google turns Gemini into a mentor for Indian tech labs
Most attempts to get AI into schools end as isolated pilots. Google is taking a different route in India through government lab infrastructure, where Gemini steps in as an expert mentor.
Read →News · 2026-09-23
Claude found a new enzyme system, but the lab still has to reveal what it does
Roughly 950 Claude agents searched a DNA database for 21 hours and flagged an uncharacterized system called ART with repeats arranged like a CRISPR array. Anthropic validated its components in the lab, but its biological function remains unknown.
Read →News · 2026-09-23
Remote GPUs help robots, but the network controls their reflexes
A Microsoft study found that moving inference from a robot to edge or cloud GPUs improves performance and battery life. It also exposes the hard constraint: tens of milliseconds of network delay already reduce the accuracy of physical work.
Read →News · 2026-09-23
Gemini 3.8 turns TTS from a voice menu into performance direction
Google has introduced Gemini 3.8 Flash TTS and Flash-Lite TTS with prompted voice creation, line-level direction and support for more than 100 languages. The meaningful shift is control and repeatability of a performance, not merely more natural sound.
Read →News · 2026-09-22
GPT-6 Astra cuts research time and token costs in half
OpenAI showcased a case study from startup Parallel, where the new Astra model completed multi-source research in half the time. For companies with agentic workflows, this means up to 50% token cost savings on tasks previously patched together with brute force.
Read →News · 2026-07-16
OpenAI sells racing data as a test of rapid decision making
OpenAI showcases in a video with RaceTek Systems and Chip Ganassi Racing how teams use artificial intelligence to turn track data into faster decisions. The source is a brief post on X, meaning hard numbers are absent and the real value lies in the operational pattern rather than a proven performance metric.
Read →News · 2026-09-21
Chimera framework advances drug discovery by ensembling diverse AI models
Microsoft Research and Novartis have introduced Chimera, a learning-to-rank system that shortens the retrosynthesis process in drug development.
Read →News · 2026-09-18
Anthropic embeds internal evaluators at Accenture
Anthropic and Accenture are investing $1 billion together into an “embedded evaluation” model. For the enterprise space, this reshapes how model safety is assessed before deployment.
Read →News · 2026-07-27
AI turns people into generalists: Almost half of messages ask about tasks from other professions
A new OpenAI study on ChatGPT data shows an interesting effect: people don't just use AI to speed up their own work, they take on tasks from completely different fields. This 'task crossover' shows that AI is changing how roles are divided inside companies.
Read →News · 2026-07-29
Allen Institute Reveals Olmo 3: Model Tuning Is Not Clean Science, but a Gory Craft Full of Errors
Nathan Lambert and Scott Geng from the Allen Institute released a detailed breakdown of post-training for the open Olmo 3 model. The data shows that moving from a research idea to a model that actually works is a series of dead ends and burned money.
Read →From the Library
Library
AI-assisted research — the model as a research partner
AI-assisted research uses models to find hypotheses, write code, test variants and read literature. It is not automatic science. It is a faster research loop with new ways to fall on your face.
Read →Library
Evals and benchmarks — measurement instead of vibes
A benchmark is not truth carved in stone. It is an instrument with error bars. Without it, though, you are only guessing whether a model or agent works.
Read →