Urgent.News

What's breaking now, across thousands of outlets.

Tech

Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

Janus is a single Go binary that enables running .gguf models on your machine, whether GPU or CPU, without the need for Python, Docker, or Ollama. It offers an OpenAI-compatible API and a built-in web UI. Users can run it via command line, integrate it into various OpenAI clients, or utilize the web UI. The design principle is to allow the model to determine the appropriate actions while Go handles inference, request routing, tool execution, and local operation.

The required disk space depends on the model size (2-8 GB typically) plus an additional ~50 MB for Janus and llama.dll. Optional Tesseract OCR is available for scanned-document OCR tasks. To build Janus, download pre-built llama.cpp Vulkan DLLs via the included build script or manually from Hugging Face. After the first startup, Janus will load the model into VRAM, which may take 10-60 seconds depending on model size and disk speed.

Linux users need libllama.so in the same directory or specified in LD_LIBRARY_PATH. Janus can be accessed via http://127.0.0.1:8990, and its status will be indicated in the browser. The terminal where Janus is running can be stopped by pressing Ctrl+C. Additional documentation includes the user manual, tools reference, and operator guide.

Janus operates independently, without any restrictions or paywalls, and can be used from the command line, scripts, Cursor, the web UI, or any OpenAI-compatible client. It can also be extended via the API.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at github.com →

More in Tech

More from Thursday 1 October →