How Qwen3.8-Max Matches Claude’s Unsupervised AI Work: A Deep Dive

16

Alibaba is making noise again.

This time, the claim is bigger than just raw processing speed. It’s about endurance. The Chinese tech giant says its newest model, Qwen3.8-Mai, can work alone—unsupervised—for days. No humans hovering. No oversight. Just the AI churning through tasks until it’s done.

For a while, this was Anthropic’s territory. Their AI models, like Claude, were touted for their ability to run long, complex workflows without constant human intervention. Now, Alibaba says they’ve caught up. Or maybe passed them.

The “Unsupervised” Claim Explained

Let’s break down what “working alone” actually looks like for an AI. It’s not sitting there staring at a screen. It’s executing a sequence of coded actions, writing reports, or debugging software. And it does it over hours, maybe days.

According to Alibaba, one specific test ran for 125 continuous hours. That’s five full days.

What did the AI do with all that time? It rebuilt a mathematical research paper’s experiment from scratch. Then, it didn’t just copy it. It improved on the original methodology. That’s not just recall. That’s synthesis and execution.

In another test, Qwen3.8-Max entered a live online coding contest. The deadline? 24 hours. It competed against 526 human engineering teams. The result? It beat everyone but 68.

Is that impressive? Sure. But it’s also a self-reported metric. Independent verification is scarce. We have to take their word for it for now.

The US Rival: Claude and Anthropic

Why does this matter? Because the US and China are in a race. And the track is narrowing.

Anthropic, the US-based firm behind Claude, has been leading the pack in the “unsupervised AI” narrative. Their flagship model, Claude Opus (often referred to as Claude Fable in recent coverage contexts), is designed to handle unfamiliar problems with minimal human input.

They even sell a product called Claude Code (and previously Claude Cowork). These tools are built for handoffs. You give the AI a research project, a spreadsheet, or a draft. You let it run. You come back later to check the work. It’s meant to replace the “busy work” layer of knowledge jobs.

Anthropic claims their newest version handles long stretches of work better than any predecessor. They say it can read software architecture, plan fixes, test the code, and verify the results—all on its own.

It’s the same pitch. Same promise. Just different brands.

Why This Race Matters Now

This isn’t just a tech bragging match. It’s geopolitical.

Washington has spent the last two years trying to slow China down. The strategy? Restrict the export of advanced AI chips. Keep Chinese firms like Alibaba a generation behind by cutting off their hardware access.

The logic is simple: No chips, no cutting-edge models.

Alibaba’s response? Show them up.

By releasing the weights of Qwen3.8-Max for free, Alibaba is doing something bold. It’s inviting the entire global developer community to audit the model. No black box. No hidden tricks. Just code for anyone to inspect.

If the claims hold up, it proves two things:
1. Algorithmic efficiency can offset hardware deficits.
2. Chinese AI development is not lagging behind the US. It’s competing on the same track.

The Verdict: Skepticism is Healthy

Here’s the reality check.

Neither Alibaba nor Anthropic has published third-party audits. Both are presenting results that make them look good. It’s in their interest to show their models as capable, autonomous, and superior.

The question isn’t whether these claims are true or false. The question is: How close is the gap really?

Investors and governments are watching this closely. If Alibaba can deliver high-quality, unsupervised AI work using less powerful hardware than US firms, the entire narrative of “hardware supremacy” starts to crack.

It’s a fascinating shift. We’re moving from a race for faster chips to a race for smarter code. And the finish line keeps moving.

Will you trust a five-day AI run without a human eye on it? Or is that exactly what you want?

The tools are here. The testing hasn’t stopped.