Welcome Inkling by Thinking Machines
Thinking Machines Lab has introduced Inkling-Small, a smaller variant of the Inkling model, designed for easier deployment and better performance. Inkling, an open-source large language model (LLM) of 1 trillion parameters, can process and understand text, images, and audio inputs simultaneously. It supports multimodal reasoning, a capability that enables it to reason across different modalities such as audio, images, and text.
Inkling utilizes a unique architecture called "relative attention" to encode positional information, replacing the commonly used RoPE method. This approach allows attention layers to learn position information directly from the attention logits. Additionally, Inkling incorporates "hybrid attention," which alternates between global and sliding window attention layers, as well as "short convolution" for local attention.
The model also features a "MoE with shared experts sink," where the router selects top-k experts and always includes two shared experts. Inkling's vision understanding module employs a hierarchical MLP patchifier and discretized mel spectrograms for audio inputs. It utilizes a simple multimodal tower design, with image and audio embeddings passed through their respective towers.
Inkling is available on Hugging Face and supports deployment through serverless inference routers and major inference engines like SGLang and vLLM. The BF16 checkpoint requires 2TB of VRAM, while the NVFP4 version demands 600GB of VRAM. Local deployment options include llama.cpp and quantized checkpoints. Inkling can be used with transformers via the Auto classes and provides example snippets for different modalities in its model card.
Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.