LMSYS Org Adds Day Zero SGLang Support for Qwen3.6-27B Reasoning Model

LMSYS OrgLMSYS Org

Ā· Updated

LMSYS Org integrated immediate support for Qwen3.6-27B into its SGLang inference framework, enabling high-speed serving of the new 27-billion parameter model. The model outperforms the massive Qwen3.5-397B-A17B on coding benchmarks and introduces native thinking modes for complex reasoning.

LMSYS Org, the organization behind the SGLang inference framework (a high-performance engine for serving large models), announced day-zero support for Qwen3.6-27B. This release mirrors the launch of Alibaba's dense multimodal model, which introduces native thinking modes for deliberative reasoning and supports text and vision reasoning.
Parameter count
27B
Reasoning modes
Thinking and non-thinking
Framework support
SGLang (day 0)
Modality
Text and multimodal
Coding performance
Beats Qwen3.5-397B-A17B
Availability
Open weights, self-hostable

This release marks an efficiency shift, following the pattern of the Qwen3.5-397B-A17B but surpassing it on coding benchmarks. While the previous generation used a massive Mixture-of-Experts architecture (specialized sub-networks for efficiency), Qwen3.6 achieves superior results with a smaller footprint. It continues the trajectory of the Qwen3-Coder-Next series by prioritizing agentic coding.

You can deploy Qwen3.6-27B immediately via SGLang to build high-throughput coding agents. This integration is part of a broader shift toward day-zero inference support across major serving frameworks. The model is available for self-hosting, providing a cost-effective alternative to larger frontier models for teams requiring local, high-performance environments.

LMSYS Org
LMSYS Org
@lmsysorg
X

šŸš€ Qwen3.6-27B is here, and we have day 0 support on SGLang āœ… 27B params, beats Qwen3.5-397B-A17B across major coding benchmarks → Agentic coding → Text + multimodal reasoning → Thinking / non-thinking modes Smaller model. Bigger results. Try it on SGLang now šŸ”„ https://t.co/x7tVVdvoFm

5retweets46likes
View on X

Still wondering? A few quick answers below.

Qwen3.6-27B is a 27-billion parameter large language model developed by Alibaba. It is designed for high-efficiency performance in technical tasks, specifically focusing on agentic coding and multimodal reasoning. Despite its smaller size compared to previous generations, it is built to handle complex logic and vision-based inputs for developers and researchers.

Qwen3.6-27B outperforms the larger Qwen3.5-397B-A17B model on coding benchmarks. While the older model uses a Mixture-of-Experts architecture, which routes tasks to specialized sub-networks to save compute, the new 27B model achieves better results in agentic coding tasks where AI autonomously writes and debugs code across multiple steps.

Qwen3.6-27B features dual operational modes that allow users to toggle between fast responses and deeper deliberation. The thinking mode enables the model to perform internal reasoning before generating a final answer, which is ideal for complex logic or math problems. The non-thinking mode provides quicker, direct outputs for standard conversational or creative tasks.

Developers can run Qwen3.6-27B on SGLang with day-zero support from LMSYS Org. SGLang is a high-performance serving framework designed to speed up model inference, which is the process of running a trained model to generate outputs. Users can follow the official cookbook guide in the SGLang documentation to set up autoregressive serving.

Yes, Qwen3.6-27B supports multimodal reasoning, allowing it to process text and visual data like images at the same time. This capability is essential for building advanced AI agents that need to understand screen content or analyze visual documents while following complex text instructions to complete multi-step tasks in a digital environment.

Every HeadsUpAI update is written based on its original source and reviewed before it's published. Read our editorial standards →

Share this update