Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size
Just two weeks after Thinking Machines released Inkling , its first open source AI language model, the well-funded startup led by former OpenAI chief technology officer Mira Murati today introduced Inkling-Small without sacrificing much of any performance — and in fact, the new model surpasses its larger predecessor on several benchmarks. Inkling Small is a 276-billion-parameter multimodal…
Just two weeks after Thinking Machines unveiled Inkling, its inaugural open-source AI language model, the venture, spearheaded by former OpenAI CTO Mira Murati, unveiled Inkling-Small. Remarkably, this new model mirrors its larger counterpart's performance while being significantly smaller. With a 276-billion-parameter multimodal reasoning model, Inkling-Small operates under an Apache 2.0 license and meets the Artificial Analysis Intelligence Index score of its bigger sibling, Inkling, which boasts 975 billion parameters and 41 billion active parameters, albeit with a single-point difference.
Despite its reduced size, Inkling-Small accepts text, image, and audio inputs, generating text, and supports up to a one-million-token context window. Compared to Inkling's 41 billion active parameters per token, Inkling-Small uses 12, achieving coding, reasoning, and multimodal performance without a major trade-off. The appeal for enterprises lies in the fact that developers can compromise on capability while cutting down on compute requirements, inference costs, and deployment footprint.
Although too large for a laptop or traditional workstation, Inkling-Small is more manageable to operate, making it ideal for enterprises with some, but not extensive, GPU resources. The model is available for free on Hugging Face and supports fine-tuning via the Tinker model training API, offering a 50% discount on API pricing for the first 64K-context Inkling-Small model.
With scores of 80.2% on SWE-bench Verified, 64.7% on Terminal Bench 2.1, and better performance on SciCode, Humanity's Last Exam, GPQA Diamond, and CritPt, Inkling-Small outperforms Inkling in several evaluations. However, it still lags behind in factual knowledge and some agentic tasks. Despite its compact size, Inkling-Small necessitates substantial GPU memory, requiring at least 600 GB of aggregate GPU memory or about 180 GB with quantized NVFP4.
Consequently, it is primarily designed for enterprise GPU servers, cloud clusters, and specialized inference providers, rather than consumer-grade devices.
Written by urgent.news from VentureBeat's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

