← Back
VentureBeat AI

Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use

9 min read
#llm#agents#enterprise
Qwen3.8-Max arrives with a bold claim: it outperforms GPT-5.6 Sol Max and Fable 5 on agentic computer use
Level:Advanced
For:AI Engineers
TL;DR

Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM), has been unveiled by Alibaba's Qwen team, claiming to outperform GPT-5.6 Sol Max and Fable 5 on agentic computer use benchmarks, with a score of 86.1 on the OSWorld-Verified benchmark. The model targets autonomous software engineering and long-horizon enterprise work, and its release may signal a strategic shift for Alibaba, with open weights to be released next week. This could substantially reshape enterprise adoption, but the licensing terms remain undisclosed. The practical implication for engineers building AI systems is the potential for more advanced autonomous execution capabilities in their workflows.

⚡ Key Takeaways

  • Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark, surpassing GPT-5.6 Sol Max (83.2) and Fable 5 (85.0).
  • The model is a 2.4-trillion-parameter mixture-of-experts (MoE) multimodal large language model (LLM).
  • Qwen3.8-Max can autonomously complete software projects lasting more than 10 days and reproduce research papers involving thousands of lines of code.
  • The model's release may be under a permissive license, allowing for self-hosted deployment, but the licensing terms are not yet disclosed.
  • The benchmark suite released alongside Qwen3.8-Max reflects a shift towards evaluating long-horizon execution.
💡 Why It Matters

The release of Qwen3.8-Max and its potential for autonomous execution capabilities could significantly impact the development of AI systems for enterprise automation, allowing for more advanced and efficient workflows. The model's performance on benchmarks such as OSWorld-Verified demonstrates its ability to compete with leading proprietary models.

✅ Practical Steps

  1. Evaluate Qwen3.8-Max's performance on specific benchmarks, such as OSWorld-Verified, to determine its suitability for autonomous software engineering and long-horizon enterprise work.
  2. Consider the potential implications of Qwen3.8-Max's release on the development of AI systems for enterprise automation, including the potential for self-hosted deployment.
  3. Monitor the release of Qwen3.8-Max's open weights and licensing terms to determine the feasibility of integrating the model into existing workflows.

Want the full story? Read the original article.

Read on VentureBeat AI

More like this

From weeks to minutes: How Formula 1® uses agentic AI on AWS to accelerate data operations

AWS ML Blog#agents

Prompt, Context, Loop: The Three Engineering Layers Every RAG System Is Built On

Towards Data Science#rag

Decoding Strategies and Output Control

Machine Learning Mastery#llm

Daniela Rus receives Bavarian Minister-President's High-Tech Prize

MIT News AI#llm

EXPLORE AI NEWS

Daily hand-picked stories on LLMs, RAG, agents and production AI — curated for engineers who ship.

BROWSE NEWS

GET THE WEEKLY DIGEST

Join engineers getting the Monday signal-over-noise AI breakdown. No spam, unsubscribe anytime.

LEARN AI ENGINEERING

Curated courses, research papers, repos and tutorials built for engineers leveling up in AI.

START LEARNING