Skip to main content

Qwen3.6-35B-A3B-FP8

Description​

"Qwen3.6-35B-A3B-FP8" is a Mixture-of-Experts (MoE) language model by Alibaba with 35 billion total parameters, of which approximately 3 billion are active per forward pass. It is designed for efficient, high-quality chat and agentic workflows with reasoning and vision capabilities, suitable for long-document analysis and extended multi-turn conversations.

This model has the lifecycle status Stable. If we shut it down or replace it, we announce that at least 3 months in advance.

It supports and is suitable for:

  • Text generation within a chat completion (text to text)
  • Tool-calling for agentic workflows
  • Image understanding (vision)
  • Thinking / reasoning for step-by-step problem solving
  • Processing long documents and extended contexts

The following limitations apply:

  • Maximum context length: 256,000 tokens
  • Thinking mode requires at least 128,000 tokens of remaining context to function properly
  • Images must be submitted as Base64-encoded data URLs (no remote URLs)

Thinking mode is enabled by default. See Disabling thinking mode below for the correct way — the parameter must be nested inside chat_template_kwargs.

Using this model from n8n? The built-in OpenAI Chat Model node can't set chat_template_kwargs — see Reasoning models and thinking mode for a community-node workaround.

Sending this to /v1/responses instead of /v1/chat/completions? chat_template_kwargs is not part of the Responses API and is dropped there unless it is named in allowed_openai_params — see Turning the thinking phase off, which also covers the endpoint's own reasoning.effort field.

Disabling thinking mode​

from openai import OpenAI

client = OpenAI(
base_url="https://llm.aihosting.mittwald.de/v1",
api_key="sk-your-api-key-here",
)

response = client.chat.completions.create(
model="Qwen3.6-35B-A3B-FP8",
messages=[{"role": "user", "content": "What is 2 + 2?"}],
temperature=0.7,
top_p=0.8,
max_tokens=32768,
extra_body={
"chat_template_kwargs": {"enable_thinking": False},
# ^^^^^^^^^^^^^^^^^^^^^^^^
# Must be nested here — passing enable_thinking at the top level
# is silently ignored by the API.
},
)

print(response.choices[0].message.content)

Reading the response​

When thinking mode is enabled (default), the model returns two separate fields:

FieldContents
choices[0].message.reasoning_contentInternal chain-of-thought (may be very long)
choices[0].message.contentFinal answer

If content is empty, the model placed its answer inside the reasoning block — disable thinking mode to ensure content is always populated.

print(response.choices[0].message.reasoning_content) # internal chain-of-thought
print(response.choices[0].message.content) # final answer

The model has different recommended settings depending on the use case. Do not use greedy decoding (temperature 0) - it can cause performance degradation and repetitions.

Thinking mode (default)​

General tasks:

ParameterValue
temperature1.0
top_p0.95
top_k20
presence_penalty1.5

Precise coding / web development:

ParameterValue
temperature0.6
top_p0.95
top_k20
presence_penalty0.0

Non-thinking mode (enable_thinking: false)​

General tasks:

ParameterValue
temperature0.7
top_p0.8
top_k20
presence_penalty1.5

Reasoning / math / complex problem solving:

ParameterValue
temperature1.0
top_p1.0
top_k40
presence_penalty2.0

Output length​

Set max_tokens according to task complexity to control cost and latency:

Task typeRecommended max_tokens
Standard queries32,768
Complex problems (math, programming contests)81,920

Tips for specific tasks​

Vision (image to text)​

Always disable thinking mode for vision tasks - thinking adds latency without improving image understanding:

extra_body={"chat_template_kwargs": {"enable_thinking": False}}

Recommended parameters for vision:

ParameterValue
temperature0.7
top_p0.8
top_k20
max_tokens512–2048 depending on task

For accurate text extraction (OCR) or data reading, use temperature=0.1 instead.

Always resize images to a maximum of 1024 px on the longest edge before encoding as Base64 - large images significantly increase time to first token (TTFT). The first request for a new image will have a longer TTFT while the image encoder warms up; subsequent requests with the same image benefit from caching. See the Python examples or JavaScript examples for a ready-to-use helper.

Math problems​

For best results on mathematical tasks, append the following instruction to your prompt:

Please reason step by step, and put your final answer within \boxed{}.

Multiple-choice questions​

To get consistent, parseable output on multiple-choice tasks, add this to your prompt:

Please show your choice in the 'answer' field with only the choice letter, e.g., 'answer': 'C'.

Terms of use and licensing​

The model is provided by Alibaba under the Apache 2.0 License, and reuse of the generated content is not subject to any additional restrictions.