Tag
#Open source
From News
News · 2026-10-03
ThinkingBox grades agents by the database, not by confident answers
Microsoft and Hugging Face have made ThinkingBox available, a benchmark of 507 stateful business tasks that runs each task 20 times and inspects the actual backend outcome. It shows why a successful tool call or polished answer does not mean the job was completed.
Read →News · 2026-10-02
AI writes 60% of Airbnb code, but handoffs are the bigger change
Airbnb CTO Ahmad Al-Dahle says AI authors 60% of company code and average pull request throughput per engineer is up 1.6 times. The deeper change is a move from documents and handoffs to prototypes, the Everest context layer and use-case-specific evals.
Read →News · 2026-10-02
AstaBrief writes cited research reports in 51 seconds with 8 billion parameters
Ai2 has released the weights for AstaBrief 8B, which generates cited scientific reports in Fast mode in an average of 51.1 seconds versus 178.5 seconds for Thinking mode. Its speed and self-hosting option are compelling, but much of the evaluation dates from 2025 and the small human study covered only 14 questions.
Read →News · 2026-09-28
Holo4 spans screens, code and APIs, but its benchmarks run on different tracks
H Company released the Holo4 27B and 35B-A3B models, combining GUI control, code execution and MCP or API calls in one agent. Open weights and published trajectories improve auditability, while results from different harnesses and task sets still complicate direct comparisons with closed models.
Read →News · 2026-07-12
White House tests the waters for banning frontier open-source models
The US government is considering how to hit the brakes on open AI models before they catch up to the closed frontier. The official narrative is about security risks from China, but the reality will hit everyone.
Read →News · 2026-09-24
DSpark makes a vision model up to 3.13 times faster, but cannot speed up the first token
Liquid AI has added an experimental 279.5-million-parameter speculative-decoding drafter to LFM2.5-VL-3B. Decoding became up to 3.13 times faster and end-to-end performance improved by up to 2.62 times because image encoding and prefill remain unchanged.
Read →News · 2026-07-16
DharmaOCR shows that a narrow model still wins on its own documents
Dharma-AI claims its Brazilian Portuguese OCR beat newer Mistral OCR4 and Unlimited-OCR models. In a narrow domain, specific data still outweighs raw compute.
Read →News · 2026-09-21
Long read without a conclusion: Models, Congress and the balance of power
The original text is an extensive testimony to the US Congress about the balance between closed and open models.
Read →News · 2026-09-19
Why an immediate intelligence explosion won't happen even with thousands of agents
Frontier labs will soon deploy thousands of concurrent agents, sparking fears of rapid recursive self-improvement (RSI). According to Nathan Lambert, while this will make inference cheaper, a real leap in intelligence won't happen immediately.
Read →News · 2026-09-15
Agents ace the demo. Why do they fail in production?
If you set an AI agent on the same production task ten times, it might choose a different path each time and hit different problems. New research from IBM and Hugging Face points out that traditional benchmarks ignore this inconsistency.
Read →News · 2026-09-11
Open Models Close the Gap, But the West Relies on Eastern Innovation
The performance gap between closed and open models has shrunk to a few months, with Asian labs driving the acceleration using data distillation. While regulators explore restrictions, corporations are already replacing expensive domestic models with Asian open-weights alternatives for cost reasons.
Read →News · 2026-09-09
IBM released an open-source time-series model that beats the giants
IBM released PatchTST-FM-r2, a new foundation model for time series with 385 million parameters. The model holds the top spot in zero-shot forecasting among openly licensed competitors on the GIFT-Eval benchmark.
Read →News · 2026-08-09
Lessons from the hacks: How AI safety shifted from model alignment to containers
The AI safety debate has undergone a quiet pivot. Following a series of agent breakouts from testing environments, labs are no longer just solving internal model tuning, but the security architecture of the infrastructure where the model runs.
Read →News · 2026-09-08
The open source community expands: Motif-3, GLM-5.3 and fading commercial barriers
The latest open models overview by Interconnects shows an important trend: top-tier models like Motif-3 and GLM-5.3-Flash are being released under the permissive MIT license. This creates viable alternatives to the Llama series without complex commercial restrictions.
Read →News · 2026-09-08
Blanket blocking as an alibi: Hugging Face analyzes AI safety failures
A recent security incident at Hugging Face has sparked debate about how model creators handle risks. Instead of surgical filtering, platforms apply a blunt blanket ban on certain topics. Safety thus often serves to protect the operator, not the user.
Read →News · 2026-08-10
Meta Releases Local Muse Glimmer with Clean License for Agent Work
Meta is back in the open weights game. Their new 30B vision model Muse Glimmer comes with an Apache 2.0 license and is built from the ground up for running local agents.
Read →News · 2026-09-03
RL shifts the perception of beauty from human to model
An open experiment shows that RL (Reinforcement Learning) can be used not only to solve math and code but also to train aesthetic appreciation. The model itself writes JavaScript that simulates watercolor painting.
Read →News · 2026-08-17
Nvidia doesn't want you buying AI from others. It wants you building it
Nvidia continues to invest heavily in open-source models to boost its ecosystem. The goal is to teach companies how to train their own AI, ensuring sustained demand for its compute hardware.
Read →News · 2026-08-20
Liquid AI speeds up local agents with DSpark drafts for LFM2.5
Local models on laptops received a substantial speed boost. Liquid AI is releasing DSpark draft models that let the LFM2.5 family reach up to 2.87 times the on-device speed on Apple hardware through speculative decoding while preserving output quality.
Read →News · 2026-08-21
Model scores no longer reflect voice AI capabilities
Speech recognition models boast great scores on public benchmarks, but a new study shows they are learning to cheat. Hugging Face researchers found that top-rated models reproduce errors from test data instead of transcribing what they actually hear.
Read →News · 2026-08-26
Hugging Face Shows How to Train Your Own Multi-Vector Model in Hours on a Single GPU
A new update to the Sentence Transformers library adds support for ColBERT-style late interaction. Developers can fine-tune models to their own domain regardless of the general limits of existing models.
Read →News · 2026-08-25
The 4-bit model that forgot its worse state and plays in a better league
Hugging Face and Multiverse Computing show Quantization-Aware Healing (QAH), letting a smaller 4-bit model beat its uncompressed bfloat16 version in benchmarks. For teams saving GPU memory, this shifts the perspective: compression is no longer a loss, but a real path to more accurate answers.
Read →News · 2026-08-10
Lambert drops post-training textbook detailing open model engineering
Nathan Lambert announced the release of his textbook on post-training AI models, condensing years of practical engineering experience. For developer teams, this could mean a critical shift from reading dense academic papers to following battle-tested fine-tuning guides.
Read →News · 2026-07-17
NeMo Automodel moves Diffusers from notebook to cluster
NVIDIA and Hugging Face show a NeMo Automodel integration with Diffusers for large scale fine-tuning of image and video diffusion models. The practical point is simple: fewer checkpoint conversions, more scaling paths and a cleaner route from Hub model to training.
Read →News · 2026-07-23
Nunchaku in Diffusers cuts image model memory to 12 GB
Hugging Face added Nunchaku 4-bit diffusion inference support to Diffusers: the blog reports a 1024x1024 image on an RTX 5090 in about 1.7 seconds and 12 GB of memory instead of 24 GB in BF16. The practical point is less friction between research code and normal Python workflows.
Read →News · 2026-07-27
Amodei rejects an open-weights ban. He wants chips, distillation limits, and mandatory tests
Dario Amodei wrote that Anthropic has never pushed a ban on open-weights models. Instead of a blanket ban he presses chip export controls, action against industrial-scale distillation, and mandatory safety tests for sufficiently capable open and closed systems.
Read →News · 2026-07-22
Open models are becoming a geopolitical strategy, not a charity project
Interconnects ties Kimi K3, Qwen 3.8, GLM 5.2 and Xi Jinping’s WAIC speech into a map of the open model fight. The point is that open weights are becoming an industrial and political lever, not just a technical preference.
Read →News · 2026-07-15
Shippy shows that production agents start with boundaries and audits
Ai2 describes Shippy, a maritime AI agent for Skylight that works with live satellite and vessel signals across more than 70 countries. The lesson is not the model. It is isolation, deterministic tools and evals for the whole agent.
Read →News · 2026-07-02
Nemotron 3.5 turns content safety from a filter into a policy engine
NVIDIA released Nemotron 3.5 Content Safety on Hugging Face, a 4B multimodal model for safety verdicts with custom policies, reasoning traces and broad language coverage.
Read →News · 2026-07-02
Cohere sends a 30B coding model into agentic harnesses
Cohere is releasing North Mini Code, a 30B Mixture of Experts model with 3B active parameters under the Apache 2.0 license. The interesting signal is not only the benchmark chart, but the focus on robustness across harnesses, because coding agents often fail at the interface, not just in the code.
Read →From the Library