{
  "id": 17388,
  "title": "Welcome Inkling by Thinking Machines",
  "url": "https://urgent.news/2026/07/15/welcome-inkling-by-thinking-machines",
  "topic": "ai",
  "section": "AI",
  "published": "2026-07-15T00:00:00.000Z",
  "source": {
    "name": "Hugging Face",
    "slug": "hugging-face",
    "url": "https://huggingface.co/blog/thinkingmachines-inkling"
  },
  "original_language": "en",
  "account": "Thinking Machines Lab has introduced Inkling-Small, a smaller variant of the Inkling model, designed for easier deployment and better performance. Inkling, an open-source large language model (LLM) of 1 trillion parameters, can process and understand text, images, and audio inputs simultaneously. It supports multimodal reasoning, a capability that enables it to reason across different modalities such as audio, images, and text.\n\nInkling utilizes a unique architecture called \"relative attention\" to encode positional information, replacing the commonly used RoPE method. This approach allows attention layers to learn position information directly from the attention logits. Additionally, Inkling incorporates \"hybrid attention,\" which alternates between global and sliding window attention layers, as well as \"short convolution\" for local attention.\n\nThe model also features a \"MoE with shared experts sink,\" where the router selects top-k experts and always includes two shared experts. Inkling's vision understanding module employs a hierarchical MLP patchifier and discretized mel spectrograms for audio inputs. It utilizes a simple multimodal tower design, with image and audio embeddings passed through their respective towers.\n\nInkling is available on Hugging Face and supports deployment through serverless inference routers and major inference engines like SGLang and vLLM. The BF16 checkpoint requires 2TB of VRAM, while the NVFP4 version demands 600GB of VRAM. Local deployment options include llama.cpp and quantized checkpoints. Inkling can be used with transformers via the Auto classes and provides example snippets for different modalities in its model card.",
  "summary": null,
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}