Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use
Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM), has been unveiled by Alibaba's Qwen team, claiming to outperform GPT-5.6 Sol Max and Fable 5 on agentic computer use benchmarks, with a score of 86.1 on the OSWorld-Verified benchmark. The model targets autonomous software engineering and long-horizon enterprise work, and its release may signal a strategic shift for Alibaba, with open weights to be released next week. This could substantially reshape enterprise adoption, but the licensing terms remain undisclosed. The practical implication for engineers building AI systems is the potential for more advanced autonomous execution capabilities in their workflows.
⚡ Key Takeaways
- Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark, surpassing GPT-5.6 Sol Max (83.2) and Fable 5 (85.0).
- The model is a 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM).
- Qwen3.8-Max can autonomously complete software projects lasting more than 10 days and reproduce research papers involving thousands of lines of code.
- The model's release may be under a permissive license, allowing for self-hosted deployment, but the licensing terms are not yet disclosed.
- The benchmark suite released alongside Qwen3.8-Max reflects a shift towards evaluating long-horizon execution.
The release of Qwen3.8-Max and its potential for autonomous execution capabilities could significantly impact the development of AI systems for enterprise automation, allowing for more advanced and efficient workflows. The model's performance on benchmarks such as OSWorld-Verified demonstrates its ability to compete with leading proprietary models.
✅ Practical Steps
- Evaluate Qwen3.8-Max's performance on specific benchmarks, such as OSWorld-Verified, to determine its suitability for autonomous software engineering and long-horizon enterprise work.
- Consider the potential implications of Qwen3.8-Max's release on the development of AI systems for enterprise automation, including the potential for self-hosted deployment.
- Monitor the release of Qwen3.8-Max's open weights and licensing terms to determine the feasibility of integrating the model into existing workflows.
Want the full story? Read the original article.
Read on VentureBeat AI ↗