Introduction

A new open-source AI model has emerged with native support for text, image, and video processing, aiming to bring multimodal capabilities to local runtimes without relying on cloud APIs. The model, built on a 35-billion-parameter mixture-of-experts architecture, is distributed in GGUF format for compatibility with popular local inference engines.

What Happened

The model Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-GGUF combines a base uncensored Qwen3.6 architecture with a Hermes fine-tune and a custom Genesis tensor-repair process. It features 35 billion total parameters, of which approximately 3 billion are active per forward pass, and a 262K-token native context window that can be extended to 1M tokens using YaRN. The architecture employs a hybrid MoE design mixing Gated DeltaNet linear attention with full softmax attention, using 256 experts with 8 routed and 1 shared per token.

Why This Matters

For developers and researchers running AI locally, this release matters because it offers a rare combination of large-scale multimodal support, extremely long context, and fine-grained sampling profiles for both coding and creative tasks. The Hermes-agent behavior and recommended function-calling prompt pattern make it relevant for agent workflows, while the long context window opens possibilities for analyzing entire codebases or lengthy documents in a single session.

Key Takeaways

  • The model runs in GGUF-compatible runtimes such as llama.cpp, LM Studio, and koboldcpp, with recommended quantization options including NVFP4 and APEX Compact for 8GB and 12GB GPUs.
  • Vision and video support require a separate mmproj file alongside the main model; the README does not specify supported formats or resolution limits.
  • Two sampling profiles are provided: a thinking-mode Hermes agent profile and a coding/precise task profile, each with distinct temperature, top-p, and repeat-penalty settings.
  • The uncensored base model means it does not inherently refuse harmful prompts; users must implement their own safety guardrails for commercial or public deployments.
  • No benchmark results are published, so performance testing on target hardware is recommended before relying on the model for production workloads.

Conclusion

Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-GGUF represents a significant step toward truly local, multimodal AI assistants that can handle long contexts and diverse input types. While the lack of published benchmarks and the uncensored base require careful evaluation, the model's architecture, quantization flexibility, and extensive context window make it a compelling option for developers exploring large-scale local AI workflows.