---
url: "https://www.reddit.com/r/Qwen_AI/comments/1vhix48/here_is_my_local_setup_for_running/"
title: here is my local setup for running Qwen3.6-35B-A3B-uncensored-MTP-I-Quality and KAT-Coder-V2.5-Dev-MTP-I-Quality
source_kind: reddit
subreddit: Qwen_AI
author: ntaybak
score: 10
captured: "2026-08-17T00:35:18+00:00"
comment_tree: false
topics: [llm-frontend]
summary: User shares a local setup for running Qwen and KAT-Coder models on a single RTX 3090 Ti with MoE offloading and MTP speculative decoding, reporting token speeds and test accuracy.
status: ok
---

models :
SC117/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-APEX-I-Quality.gguf 21.8gb
gbuzhf/Kwaipilot_KAT-Coder-V2.5-Dev-MTP-APEX-I-Quality.gguf 22GB
UNSLOTH mmproj-F16.gguf for vision
Both models run on a single RTX 3090 Ti 24 GB · Ryzen 9 9950X  · 96 GB  DDR5-5600 via llama.cpp b10223 with MoE expert offloading and MTP speculative decoding.
SC117/Qwen3.6-35B-A3B-uncensored decodes at 118–130 tok/s (96K context, temp 0.6)
gbuzhf/KAT-Coder-V2.5-Dev decodes  at 82–95 tok/s (128K context, temp 1.0),
GSM8K test accuracy of 80% / 86%
two bat files for running every model, one with vision support and one without it, i have tested them in openchamber and Reasonix desktop apps and they both were fast , but i did not tested them in my main work so when i do i will update the post comparing their result.
i will keep pushing Qwen and DeepSeek to try deferent settings and re benchmark until they get the best thinking quality of them.
All of this was done by Qwen 3.8 Max and DeepSeek V4 Flash 0731, I didn't actually do anything myself, but I wanted to share the setup. Maybe someone can suggest some improvements for better thinking quality, or hopefully others will find it useful.