Alibaba Ships Qwen 3.7 Preview as Researchers Allege Censorship in Earlier Model

The same morning Alibaba began previewing its newest flagship model, researchers said they had extracted evidence of political censorship from the weights of the one before it. The coincidence has reignited a familiar argument about whether content restrictions can be rooted out at the model-weight level — and whether open weights actually mean open. Community users reported Qwen 3.7 options appearing in Alibaba’s Qwen Chat interface on Monday, with Max and Plus preview variants having surfaced on chat.qwen.ai and the LM Arena leaderboard around mid-May.

The previews run with deep thinking enabled by default and with web search and code interpreter disabled. A full unveiling is expected at Alibaba’s cloud summit in Hangzhou this week. The rollout follows the playbook Alibaba used with Qwen 3.6 Max in April: validate the model on public leaderboards first, then commercialize it. Early third-party measurements of the Max preview have already circulated — one review cited a hallucination rate of 22.9%, among the lowest measured in comparable frontier models, achieved in part because the model declines to answer roughly half of broad-recall queries, a behavior developers have taken to calling the “abstention tax.”

The 3.7 family comes in two tiers, following the pattern Alibaba established with Qwen 3.6: a Max flagship focused on reasoning and a Plus variant with vision and video understanding. Alibaba has framed the release around improvements in reasoning and multilingual capability, and early previews on LM Arena — which ranks models by blind, crowd-sourced comparison — put the Max preview in the top tier of text models, with strong showings in math and software tasks.

What is notable is what has not been released. Alibaba has not published open weights for the 3.7 family on Hugging Face, and the company has not confirmed whether it will. Every Qwen flagship through the 3.5 series shipped under permissive licenses with downloadable weights; the pattern broke with Qwen 3.6 Max, which was API-only, and the 3.7 release so far continues that direction.

The controversy centers on the previous generation. Researchers said they had extracted evidence of political censorship from Qwen 3.5’s weights, a finding that feeds a long-running debate on Hacker News, X and arXiv about whether content restrictions embedded during training can be detected and audited in open-weight models. For a family that has built its reputation on openness — more than 940 million cumulative downloads by March 2026, more than 200,000 derivative models, and a global open-weight download share above 50% by some counts — the question is whether downloadable weights are the same as transparent ones.

The argument cuts both ways. Open weights make censorship detectable in a way that closed models never can be; the same transparency that lets researchers inspect a model’s behavior is the transparency that lets them find the restrictions in the first place. The practical stakes are concrete for enterprises: companies that self-host Qwen to control their data inherit whatever behavioral restrictions are baked into the weights, and if those restrictions are discoverable by researchers, they are also discoverable by users, regulators and competitors. The finding gives ammunition to both sides of the open-weight debate — it shows open models can be audited, and it shows what an audit can turn up. The finding also complicates the pitch that open-weight models are simply neutral software: the weights carry policy choices encoded in training data, and those choices are not always visible to developers who self-host.

The timing matters. Alibaba is moving its most capable models behind a paywall just as it is asking the developer community to keep building on Qwen. The company killed the free tier of Qwen Code last month, and the 3.7 release confirms the direction: developers who want the top Qwen model will increasingly be API customers rather than weight users. The open community’s standard response is already forming — fork the last open checkpoint, which remains fully available on Hugging Face, and continue development independently. The two-tier architecture — open smaller models to seed adoption, close the flagship to capture enterprise revenue — mirrors the strategy OpenAI and Anthropic have used for years, and it is a notable direction for the lab that built its ecosystem by giving away its best work.

Export controls add pressure. U.S. restrictions on advanced Nvidia accelerators constrain how Chinese labs train and serve frontier models, which makes the API business more important and self-hosted open weights less central to Alibaba’s economics. The company has reorganized around AI this year, with Chief Executive Eddie Wu coordinating a new task force after the departure of the Qwen division’s head, and the Qwen team’s cadence — 3.6 in April, 3.7 preview in May — suggests it is trying to keep pace with Western labs while spending fewer compute resources per model. The company has acknowledged the tension publicly.

Qwen 3.7’s preview arrives with its two audiences pulling in opposite directions — developers who built on Qwen because the weights were open, and an Alibaba that needs the flagship to pay for itself. The censorship finding adds a third force: if open weights come with invisible restrictions, then “open” becomes a weaker claim for every lab that makes it.

Related Posts

  • September 6, 2026
  • 6 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 6 views
OpenAI Quietly Revises GPT-6 Astra Scores After Launch

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often…