This AI Model Has Native Text, Image, and Video Capabilities: Here's What You Should Know
Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-GGUF is a locally runnable, text-to-text GGUF release with native text, image, and video capabilities.
Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-GGUF is a locally runnable AI model that offers native text, image, and video capabilities. Developed by LuffyTheFox, it is built upon the uncensored HauhauCS Qwen3.6-35B-A3B base and incorporates the Genesis tensor-repair process. The model incorporates 2,000 blocks from two FFN expert tensors through a Hermes fine-tune, adding Hermes-agent behavior.
The model features a massive 35 billion parameters, with around 3 billion active per forward pass. It boasts a 262K-token native context window and utilizes a hybrid MoE architecture that combines Gated DeltaNet linear attention with full softmax attention. To optimize its performance, the model is designed to run in llama.cpp, LM Studio, koboldcpp, and other GGUF-compatible runtimes.
For optimal use, it is recommended to employ NVFP4 or APEX quantization, with APEX Compact being suitable for systems with 8 GB or 12 GB of GPU memory. The model adheres to the Apache-2.0 license, though commercial deployment should consider the terms of upstream components.
Best suited for local coding, precise instruction-following, and tool-oriented assistant workflows, the model's Hermes-derived agent behavior and large context window make it a strong candidate for tasks such as code explanation, multi-file planning, and structured technical work. However, without benchmark results provided in the README, users should test the model against their specific coding tasks before relying on it for production code.
The model is particularly valuable for long-document analysis, as its 262K-token context window allows for the examination of extensive reports, codebases, or document collections. However, users should keep in mind the potential impact on runtime memory and context-cache costs.
While the model is described as uncensored and capable of creative writing and role-play, the maintainer's claims of 0 refusals on a 465-prompt test should not be considered an independent safety or quality evaluation. The model card offers a non-thinking creative profile and optional creative system prompts for users seeking such functionality.
Multimodal local experiments are also feasible with this model, as it is designed to support text, image, and video inputs natively. However, the required mmproj file for vision capabilities and the lack of provided information on supported image or video formats, resolution limits, and multimodal benchmark results necessitate validation in the chosen runtime before implementing a multimodal workflow.
Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.