Multimodal open d1 decision models for the edge
Today, two open decision models from the d1 decision model family were released: d1-3B and d1-omni-600M (experimental). These models, built on Liquid Foundation Models (LFMs), differ from generative models as they provide single forward-pass answers instead of token generation. The d1-3B and d1-omni-600M models were trained on various public datasets, including reading comprehension, toxicity detection, intent classification, medical QA, and cross-lingual understanding.
Upon benchmarking, d1-3B achieved a mean score of 82.9, the highest among tested models, while d1-omni-600M scored 78.4, outperforming Decider 2B (77.1) despite having only a quarter of its parameters. It was also confirmed that d1-3B retains vision capabilities from its LFM2.5-VL-3B backbone on standard vision benchmarks, and d1-omni-600M supports all three modalities.
For further validation, d1-3B was evaluated on the NVIDIA stack across several devices, with edge inference showing that it answers a single question in under 50 ms on all measured devices. GPU inference demonstrated that d1-3B answers a question in under 10 ms and processes a 384px image in under 18 ms. The d1 decision models are suitable for fast, structured decisions, including multimodal inputs.
When high decision quality is required at a smaller size, d1-3B is recommended, while d1-omni-600M is ideal for situations where footprint is a concern. The models can be installed using their dependencies (requiring transformers =5.14) and loaded with trust_remote_code=True. The d1-3B model card provides instructions on how to run d1-omni-600M.
Both decision models are open-weight and available on Hugging Face today. If this work is used, please cite the release blog.
Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.