Lilith Lilith.
CS EN PL

From Radar

Radar · 2026-07-31

DeepSeek V4 Flash 0731 pushes agent performance down to sub-$0.30 per million tokens

DeepSeek shipped official DeepSeek-V4-Flash-0731: MIT weights, 304B parameters on Hugging Face (Artificial Analysis lists 284B total / 13B active), stronger agentic post-training, and API pricing around $0.14 per million input tokens and $0.27-$0.28 per million output tokens. Willison and Artificial Analysis place it among the best value-per-intelligence open-weight models.

Read

Radar · 2026-08-03

OpenAI rebuilt the voice stack so GPT-Live can listen while it speaks

OpenAI detailed the GPT-Live architecture: a full-duplex voice model with no turn detector on the audio path, async delegation to a frontier model such as GPT-5.5, and transport tuned for continuous audio. Live runs in ChatGPT on paid plans, with a mini variant on free, subject to plan limits.

Read

Radar · 2026-07-29

GPT-5.6 jumps in intelligence tests by turning on background reasoning

New API settings for GPT-5.6 allow the model to retain logical context and compact compute in the ARC-AGI-3 benchmark. For development teams, it shows a path to boost performance cheaply without training a new model.

Read

Radar · 2026-07-23

ChatGPT Health opens to every US adult. Clinician-level claims meet a lawsuit

OpenAI is rolling Health in ChatGPT to all signed-in US users 18+ across free through Pro plans. People can connect Apple Health and hospital records while the company cites 300 million weekly health queries and faces a lawsuit over dangerous advice.

Read

Radar · 2026-07-24

Fugu Ultra claims gains of up to 7.9 points over v1.0 but has not shown the full methodology

Hardmaru says Fugu Ultra v1.1 dynamically combines frontier models and improves performance by up to 7.9 points over v1.0. The announcement supports the case for orchestration, but its comparison with Fable 5 remains a vendor claim without a public evaluation.

Read

Radar · 2026-07-01

BAIR shows where the next wave of AI talent is flowing

BAIR published a showcase of 33 Ph.D. graduates from its 2026 class. The celebration doubles as a map of people moving into robotics, LLM agent systems, AI safety and AI for science.

Read

Radar · 2026-06-24

Google shows reasoning can pull out plain facts too

Google Research examines why chain-of-thought helps LLMs answer simple factual questions. The study on Gemini 2.5 and Qwen3-32B points to two mechanisms: extra computation in generated tokens and factual priming.

Read

Radar · 2026-06-03

GPT-Rosalind moves from benchmarks toward governed science

OpenAI updated GPT-Rosalind for life sciences and is offering it in research preview to selected organizations globally. The more important move is not the scorecard, but the attempt to connect a model, Codex and bioinformatics tools into an auditable workflow.

Read