Lilith.
⌕
Editorial illustration: Open d1 makes edge decisions in 16 ms, but multimodal evidence is still missing
Lilith illustration · editorial remix

Liquid AI has released the open-weight d1-3B and experimental d1-omni-600M. Both return structured choices, scores or yes-or-no answers in one forward pass rather than generating text tokens. On a Jetson AGX Thor, d1-3B answered one question in 16 ms.

Two small models return decisions instead of prose

d1-3B is derived from LFM2.5-VL-3B and accepts text and images. d1-omni-600M combines a 350-million-parameter bidirectional encoder with vision and audio encoders. It handles text plus image or text plus audio, and Liquid AI labels it an early research release.

Across seven public datasets covering reading comprehension, toxicity, intent classification, medical QA and cross-lingual understanding, d1-3B averaged 82.9 and d1-omni-600M scored 78.4. d1-3B answered one question in 8 ms on an RTX 4090 and 50 ms on a Jetson Orin Nano.

On-device classification can remove the generative detour

Many applications use a generative LLM only to extract a label, score or yes-or-no answer at the end. A decision model removes text generation and parsing. That is useful for routing, filtering, moderation and on-device control where stable output, latency and privacy matter.

Open weights also allow images or audio to stay off a third-party API. Developers must accept a different System One interface and, in the published example, load model-specific code with trust_remote_code=True.

The multimodal promise currently lacks a public benchmark

Vision and audio at the edge are the strongest product claim, yet the authors report no vision or audio benchmark in this release. They say Decision Index v0.3 has only a private vision split and that audio decision benchmarks remain an open problem. The published table therefore validates mostly text tasks.

The speed figures also apply only to d1-3B. The team reports none for the experimental d1-omni-600M. Quality comparisons come from the model maker, and an average across seven datasets does not show how well probabilities remain calibrated in a factory inspection or safety filter.

Real hardware must confirm both quality and calibration

The next step is independent vision and audio testing, memory and energy measurement, and probability calibration under data shift. Long-state latency matters too: a 3,400-token input took 1,640 ms on Jetson Orin Nano, far longer than a short question.

If the small d1 models retain accuracy on multimodal data from real devices, they could replace oversized generative models in narrow decision loops. Sixteen milliseconds alone does not settle that case.

Lilith's verdict

d1-3B clears a short question in 16 ms, while the camera and microphone are still waiting outside the test room. An edge model earns its name in field results, not on an empty straightaway.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗