Tag
#reasoning
From Radar
Radar · 2026-07-31
DeepSeek V4 Flash 0731 pushes agent performance down to sub-$0.30 per million tokens
DeepSeek shipped official DeepSeek-V4-Flash-0731: MIT weights, 304B parameters on Hugging Face (Artificial Analysis lists 284B total / 13B active), stronger agentic post-training, and API pricing around $0.14 per million input tokens and $0.27-$0.28 per million output tokens. Willison and Artificial Analysis place it among the best value-per-intelligence open-weight models.
Read →Radar · 2026-08-03
OpenAI rebuilt the voice stack so GPT-Live can listen while it speaks
OpenAI detailed the GPT-Live architecture: a full-duplex voice model with no turn detector on the audio path, async delegation to a frontier model such as GPT-5.5, and transport tuned for continuous audio. Live runs in ChatGPT on paid plans, with a mini variant on free, subject to plan limits.
Read →Radar · 2026-07-29
GPT-5.6 jumps in intelligence tests by turning on background reasoning
New API settings for GPT-5.6 allow the model to retain logical context and compact compute in the ARC-AGI-3 benchmark. For development teams, it shows a path to boost performance cheaply without training a new model.
Read →Radar · 2026-07-23
ChatGPT Health opens to every US adult. Clinician-level claims meet a lawsuit
OpenAI is rolling Health in ChatGPT to all signed-in US users 18+ across free through Pro plans. People can connect Apple Health and hospital records while the company cites 300 million weekly health queries and faces a lawsuit over dangerous advice.
Read →Radar · 2026-07-24
Fugu Ultra claims gains of up to 7.9 points over v1.0 but has not shown the full methodology
Hardmaru says Fugu Ultra v1.1 dynamically combines frontier models and improves performance by up to 7.9 points over v1.0. The announcement supports the case for orchestration, but its comparison with Fable 5 remains a vendor claim without a public evaluation.
Read →Radar · 2026-07-01
BAIR shows where the next wave of AI talent is flowing
BAIR published a showcase of 33 Ph.D. graduates from its 2026 class. The celebration doubles as a map of people moving into robotics, LLM agent systems, AI safety and AI for science.
Read →Radar · 2026-06-24
Google shows reasoning can pull out plain facts too
Google Research examines why chain-of-thought helps LLMs answer simple factual questions. The study on Gemini 2.5 and Qwen3-32B points to two mechanisms: extra computation in generated tokens and factual priming.
Read →Radar · 2026-06-03
GPT-Rosalind moves from benchmarks toward governed science
OpenAI updated GPT-Rosalind for life sciences and is offering it in research preview to selected organizations globally. The more important move is not the scorecard, but the attempt to connect a model, Codex and bioinformatics tools into an auditable workflow.
Read →