All discussions

Model Buzz · Text & code

Qwen

Recent community conversations about using Qwen.

7 selected discussionsReviewed Sep 9, 2026
Latest discussionsLatest first
LINUX DO分区:开发调优Writing

Which model writes Chinese with less of an AI voice?

Laplace9 is building a writing agent and asks for a model that produces more natural Chinese prose. Replies name Claude Opus 4.6, Gemini and Qwen 3.8 Max, with different preferences for wording and style.

In the repliesSeveral participants prefer Opus 4.6, while another reports good essay writing with Gemini 3.8 Flash. Szou says even their preferred models still need manual editing to remove familiar AI phrasing.

Chinese prose · English summary of a Chinese-language thread

LINUX DO分区:开发调优Hands-on

Qwen 3.8 27B on a 24 GB GPU: long-context praise meets speed questions

zhiqing reports stable long tasks and a full 262K context with a specific Qwen 3.8 27B quantization. In follow-up replies, they identify an RTX 3090 setup with DFlash2 enabled.

In the repliesslothboy reports slow output on their own setup, and others ask about the runtime and hardware. One reply questions the long thinking time; another suggests testing beyond the pelican example. The reported result belongs to this particular setup.

Qwen3.8-27B-DFlash2-EXL3-5.0bpw · Local inference

LINUX DO分区:国产替代Comparison

Qwen, GLM or DeepSeek Flash for coding? Waiting time divides users

yooinsung asks which Flash model people choose for writing code in IDEs such as Qoder and Trae. Replies compare Qwen 3.8 Flash, GLM 5.3 Flash and DeepSeek V4 Flash on waiting time and instruction following.

In the repliesOne user finds DeepSeek fast for routine work but still prone to hallucinations. Others describe long thinking delays in Qwen or GLM. A GLM user reports better instruction following at the lowest thinking setting, showing why settings matter to the comparison.

IDE coding · Different products and thinking settings

Hacker NewsHacker NewsLocal inference

Qwen Mac Studio speed reports draw questions about the software setup

Submitter speckx shared an external writeup reporting local benchmark numbers for running Qwen 3.8 27B on an Apple Mac Studio, drawing immediate scrutiny over performance measurements that trailed earlier model releases.

In the replieskgeist questioned the writeup’s architecture explanation, while Infernal asked whether missing software optimizations could explain the reported slowdown. Atreiden described similarly sluggish output across several quantizations and was waiting for additional acceleration support. pwython preferred a different local model on an M4 Max. The comments offer leads for investigating a slow setup, but they do not isolate a cause or establish that the model is inherently slower across Macs.

Qwen 3.8 27B · Apple Silicon inference; runtime and quantization vary.

Historical discussion; added to the archive on September 8, 2026. Selected replies, not a community-wide verdict.

Hacker NewsHacker NewsLocal inference

Qwen Flash Next sparks conflicting estimates of local memory needs

In the Qwen3.8-Flash-Next discussion, simonw shared SVG pelican samples generated on a DGX Spark at several reasoning settings. They used Unsloth’s UD-IQ1_S GGUF quantization and liked the results less than their Qwen 3.8 27B sample, suggesting quantization as a possible explanation.

In the repliesReaders disagreed about how much memory the model would need. andy99 questioned whether a 128 GB system could fit a useful quantization; a_humean was more optimistic and was waiting for runtime support. Other replies proposed offloading parts of the model. Those estimates depend on quantization, runtime and context settings, so the thread should not be read as a verified shopping specification.

Qwen3.8-Flash-Next · DGX Spark sample with UD-IQ1_S GGUF; memory estimates in replies are unverified.

Historical discussion; added to the archive on September 8, 2026. Selected replies, not a community-wide verdict.

Hacker NewsHacker NewsLocal inference

Qwen 3.8 27B users look for ways to rein in long reasoning

Submitter bilsbie shared an external post discussing Qwen 3.8 27B's local execution and its tendency toward excessive internal reasoning chains during inference.

In the repliesCommenters celebrated consumer hardware handling large local models smoothly in everyday setups. In response to verbose reasoning steps, practitioners detailed interim runtime workarounds. One developer shared a custom fork injecting text at specific thresholds, warning that forcing out-of-distribution patterns might degrade output quality. Another user tested dedicated reasoning effort parameters across multi-turn prompts, while others guided execution using rigid planning prompts. Commenters also debated whether overthinking stems from reinforcement learning incentives optimized for benchmark evaluation.

Qwen 3.8 27B · Local inference and experimental llama.cpp controls.

Historical discussion; added to the archive on September 8, 2026. Selected replies, not a community-wide verdict.

Redditr/LocalLLaMAComparison

Qwen or Gemma? Local users get different answers from different tasks

The author favored Gemma 4 26a4B QAT over Qwen 3.6 35a3B when running both models at Q4 quantization, highlighting superior instruction adherence and coherence. In a follow-up clarification, the author noted the evaluation was based on non-coding tasks like following a manual.

In the repliesCommunity responses showed divergent results across different workloads and configurations. Bulky-Priority6824 reported improved performance at higher quantization levels within their own setup. For an ongoing assistant role, o0genesis0o preferred Qwen for following instructions, noting issues with Gemma entering repetitive loops. Additionally, FilterJoe observed that chat template implementations alongside quantization choices make direct head-to-head comparisons difficult to standardize.

Qwen 3.6 35a3B and Gemma 4 26a4B QAT · Q4 comparison and other user setups; templates vary.

Historical discussion; added to the archive on September 8, 2026. Selected replies, not a community-wide verdict.

Summaries reflect individual posts and replies, not a community-wide verdict.