Alibaba’s Qwen group has made Qwen3.8-Max broadly accessible and confirmed that its open weights ship subsequent week. A second checkpoint, Qwen3.8-27B, can be going open-weights. Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts mannequin. It accepts textual content, picture and video as enter and returns textual content.
Is it deployable
Sure, however the deployable floor depends upon which artifact you might be making use of.
The hosted API is deployable in the present day by any firm dimension. It’s OpenAI- and DashScope-compatible, so integration is a base-URL and model-ID change. The open weights are a distinct matter. At 2.4T complete parameters, the checkpoint is a multi-node datacenter artifact. Alibaba has not disclosed the activated-parameter depend. Serving value subsequently can not but be modeled. Qwen3.8-27B is the checkpoint that matches atypical on-premise GPU {hardware}.
The printed characteristic set maps cleanly onto 4 industries. These are software program engineering, authorized and monetary doc overview, media and e-commerce operations, and design.
Functions embrace repository-scale coding brokers and long-document information bases. Lengthy-video indexing, structured information extraction and multi-step analysis assistants additionally match.
Interactive explainer
“+row[0]+’‘+(vals[qi]!==null&&vals[fi]!==null?(vals[qi]>vals[fi]?’above’:vals[qi]
‘;
C.m.forEach(perform(mn,j));
el.innerHTML=h;L.appendChild(el)});
q(‘#bmw’).textContent=”Qwen3.8-Max above Fable5 on “+win+’ of ‘+cmp+’ comparable rows’;
setTimeout(perform(){qa(‘#bml .f’).forEach(perform(f){f.type.width=f.dataset.w+’%’})},40);
setTimeout(rs,140);}
/* RL curve */
var RL=[[0,.474],[500,.469],[1000,.487],[1500,.586],[2000,.606],[2500,.622],[3000,.616],[3500,.647],[4000,.725],[4500,.719],[5000,.689]];
var rlDone=false;
perform paintRl(){
if(rlDone)return;rlDone=true;
var g=q(‘#rlg’),X=perform(v){return 44+v/5000*570},Y=perform(v){return 210-(v-.44)/(.76-.44)*180};
perform mk(t,a){var e=d.createElementNS(‘http://www.w3.org/2000/svg’,t);
for(var okay in a)e.setAttribute(okay,a[k]);g.appendChild(e);return e}
[.45,.5,.55,.6,.65,.7,.75].forEach(perform(v){
mk(‘line’,{x1:44,y1:Y(v),x2:614,y2:Y(v),stroke:’#1d1830′,’stroke-width’:1});
var t=mk(‘textual content’,{x:6,y:Y(v)+3,class:’lbl’});t.textContent=v.toFixed(2)});
mk(‘line’,{x1:44,y1:Y(.474),x2:614,y2:Y(.474),stroke:’#4a4170′,’stroke-width’:1,’stroke-dasharray’:’4 4′});
var sb=mk(‘textual content’,{x:470,y:Y(.474)+14,class:’lbl’});sb.textContent=”SFT baseline 0.474″;
var pts=RL.map(perform(p){return X(p[0])+’,’+Y(p[1])}).be part of(‘ ‘);
var pl=mk(‘polyline’,{factors:pts,fill:’none’,stroke:’#7C5CFF’,’stroke-width’:2.6,’stroke-linejoin’:’spherical’,’stroke-linecap’:’spherical’});
var len=pl.getTotalLength?pl.getTotalLength():2000;
pl.setAttribute(‘stroke-dasharray’,len);pl.setAttribute(‘stroke-dashoffset’,len);
var an=d.createElementNS(‘http://www.w3.org/2000/svg’,’animate’);
an.setAttribute(‘attributeName’,’stroke-dashoffset’);an.setAttribute(‘from’,len);an.setAttribute(‘to’,0);
an.setAttribute(‘dur’,’1.5s’);an.setAttribute(‘fill’,’freeze’);pl.appendChild(an);
RL.forEach(perform(p,i){
var pk=p[1]===.725;
var c=mk(‘circle’,{cx:X(p[0]),cy:Y(p[1]),r:pk?6:3.6,fill:pk?’#B49CFF’:’#7C5CFF’,opacity:0});
var a2=d.createElementNS(‘http://www.w3.org/2000/svg’,’animate’);
a2.setAttribute(‘attributeName’,’opacity’);a2.setAttribute(‘from’,0);a2.setAttribute(‘to’,1);
a2.setAttribute(‘dur’,’.3s’);a2.setAttribute(‘start’,(i*0.13)+’s’);a2.setAttribute(‘fill’,’freeze’);c.appendChild(a2);
var t=mk(‘textual content’,{x:X(p[0]),y:Y(p[1])-12,’text-anchor’:’center’,class:’lbl’,fill:pk?’#B49CFF’:’#8b81ab’,opacity:0});
t.textContent=p[1].toFixed(3);
var a3=d.createElementNS(‘http://www.w3.org/2000/svg’,’animate’);
a3.setAttribute(‘attributeName’,’opacity’);a3.setAttribute(‘from’,0);a3.setAttribute(‘to’,1);
a3.setAttribute(‘dur’,’.3s’);a3.setAttribute(‘start’,(i*0.13)+’s’);a3.setAttribute(‘fill’,’freeze’);t.appendChild(a3);
if(ipercent2===0){var xl=mk(‘textual content’,{x:X(p[0]),y:228,’text-anchor’:’center’,class:’lbl’});xl.textContent=p[0]}});
var pk=mk(‘textual content’,{x:X(4000),y:Y(.725)-26,’text-anchor’:’center’,class:’lbl-b’,fill:’#B49CFF’});
pk.textContent=”peak”;
setTimeout(rs,200)}
/* cross-harness */
var HZ=[
{n:’CoWorkBench’,b:[[‘Fable5 (OpenClaw)’,75.9,0],[‘Opus4.8 (OpenClaw)’,72.3,0],[‘Qwen3.7-Max (OpenClaw)’,64.6,0],[‘QwenWork’,73.2,1],[‘Claude Code’,74.6,1],[‘Codex’,75.8,1],[‘OpenClaw’,74.8,1],[‘Hermes’,73.3,1]]},
{n:’WorkspaceBench’,b:[[‘Fable5 (OpenClaw)’,68.7,0],[‘Opus4.8 (OpenClaw)’,66.8,0],[‘Qwen3.7-Max (OpenClaw)’,61.4,0],[‘QwenWork’,67.0,1],[‘Claude Code’,67.3,1],[‘Codex’,67.6,1],[‘OpenClaw’,67.7,1],[‘Hermes’,71.2,1]]},
{n:’JobBench’,b:[[‘Fable5 (OpenCode)’,57.4,0],[‘Opus4.8 (OpenCode)’,48.4,0],[‘Qwen3.7-Max (OpenCode)’,31.3,0],[‘QwenWork’,58.0,1],[‘Claude Code’,59.0,1],[‘Codex’,57.5,1],[‘OpenClaw’,59.8,1],[‘Hermes’,59.6,1],[‘OpenCode’,53.4,1]]}];
HZ.forEach(perform(h){
var c=d.createElement(‘div’);c.className=”hc”;
var s=”
“+h.n+’
‘;
var mx=Math.max.apply(null,h.b.map(perform(x){return x[1]}));
h.b.forEach(perform(b){s+=’
‘+
‘‘+b[0]+’‘+
‘‘+b[1].toFixed(1)+’
‘});
c.innerHTML=s;q(‘#hzl’).appendChild(c)});
/* deploy */
var DEP=[
‘Path: hosted API only. Call model ID qwen3.8-max through the DashScope or OpenAI-compatible endpoint. Self-hosting 2.4T weights is out of reach. Best fits: coding agents, document and video ingestion, research assistants. Owner: the founding engineer or a solo AI engineer.’,
‘Path: API first, 27B for anything sensitive. Use the hosted flagship for long-horizon agent work and keep Qwen3.8-27B on your own GPUs for PII-bearing or high-volume routing. Best fits: support automation, analytics copilots, internal tooling. Owner: an ML platform team.’,
‘Path: both, with an evaluation gate. Multi-node clusters can host the open weights once activated parameters and license are published. Until then run the API behind your own harness. The cross-harness numbers suggest portability across Claude Code, Codex and OpenClaw. Owner: an applied AI group reporting to the CTO.’,
‘Path: wait, then take 27B. No license file has been published for Qwen3.8-Max, so procurement cannot clear it yet. Air-gapped deployment realistically means Qwen3.8-27B. Best fits: healthcare, defense, public sector. Owner: security architecture plus a data science lead.’];
qa(‘#who button’).forEach(perform(b){b.onclick=perform(){
qa(‘#who button’).forEach(perform(x){x.classList.take away(‘on’)});
b.classList.add(‘on’);q(‘#dep’).innerHTML=DEP[+b.dataset.w];setTimeout(rs,60)}});
q(‘#dep’).innerHTML=DEP[0];
/* proof */
var EV=[[
‘Full benchmark table across coding, agent, reasoning, multimodal, document, video’,
‘2.4 trillion total parameters, mixture-of-experts’,
‘1M context; 991K max input, 131K max output, 262K max reasoning’,
‘Text, image and video input; text output’,
‘$2.00 input, $6.00 output, $0.25 implicit cache read per 1M tokens’,
‘Built-in code_interpreter, web_search, web_extractor, t2i_search, i2i_search’,
‘RL scaling curve, including the decline past the 4,000-environment peak’,
‘Cross-harness results on QwenWork, Claude Code, Codex, OpenClaw, Hermes’,
‘Open weights for Qwen3.8-Max and Qwen3.8-27B announced for next week’
],[
‘Model card’,
‘License file’,
‘Activated parameters per token’,
‘Independent evaluation from Artificial Analysis or LMArena’,
‘Harness, attempt count and reasoning settings behind each score’,
‘Note: the multimodal table compares against Qwen3.7-Plus, not Qwen3.7-Max’
]];
var evMode=0;
perform paintEv(){var L=q(‘#evl’);L.innerHTML=”;
EV[evMode].forEach(perform(t,i){var e=d.createElement(‘div’);e.className=”q-ev”;
e.type.animationDelay=(i*55)+’ms’;
e.innerHTML=’‘+(evMode?’?’:’✓’)+’‘+t+’‘;
L.appendChild(e)});setTimeout(rs,120)}
qa(‘#evp button’).forEach(perform(b){b.onclick=perform(){
qa(‘#evp button’).forEach(perform(x){x.classList.take away(‘on’)});
b.classList.add(‘on’);evMode=+b.dataset.e;paintEv()}});
paintEv();paintBm();
window.addEventListener(‘load’,rs);setTimeout(rs,300);setTimeout(rs,900);
})();
” type=”width:100%;peak:600px;border:0;overflow:hidden;show:block” scrolling=”no” loading=”lazy” title=”Qwen3.8-Max Interactive Explainer”>
What’s Technically Obtainable
The mannequin web page lists a 1M-token context window. Most enter is 991K tokens, dropping to 983K when considering is enabled. Most output is 131K tokens in each modes, and the utmost reasoning finances is 262K tokens. Charge limits are 2M tokens per minute and 15K requests per minute.
Pricing is $2.00 per 1M enter tokens and $6.00 per 1M output tokens. Implicit cache reads value $0.25 per 1M tokens. Specific cache creation is $2.50 and specific cache reads are $0.17 per 1M tokens. Cached enter is eight instances cheaper than contemporary enter. Prefix stability subsequently drives value greater than immediate size does.
Supported capabilities embrace perform calling, structured outputs, batches, prefix completion and fine-tuning. 5 built-in instruments ship on the Responses API: code_interpreter, web_search, web_extractor, t2i_search and i2i_search.


Efficiency
Alibaba printed a full benchmark desk with this launch. Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1, forward of Claude Opus 4.8 and Claude Fable 5 at 84.6, behind GPT-5.6 Sol (max) at 88.8. It stories 67.7 on SWE-bench Professional towards Fable 5’s 80.0, and 73.5 on FrontierSWE towards Fable 5’s 88.8. It leads PaperBench at 93.0 and IFBench at 82.8. GPQA Diamond lands at 92.6, up marginally from Qwen3.7-Max’s 92.4. The clearest positive factors are multimodal and agentic, not reasoning. It tops most imaginative and prescient rows, together with OSWorld-Verified 86.1, Parametric CAD Bench 91.5, and OmniDocBench 1.5 at 92.1. Towards its personal predecessor the leap is giant: DeepSWE 1.1 strikes from 21.6 to 56.6, FrontierSWE from 40.7 to 73.5, JobBench from 31.3 to 53.4. Two caveats belong in any sincere learn. The multimodal desk benchmarks towards Qwen3.7-Plus, not Qwen3.7-Max, which flatters the generational delta. And Alibaba’s personal RL scaling curve peaks at 0.725 close to 4,000 coaching environments, then declines to 0.719 and 0.689.
Key Takeaways
- Qwen3.8-Max is a 2.4T-parameter MoE mannequin with 1M context, now typically accessible.
- Pricing is $2 enter, $6 output and $0.25 cached enter per 1M tokens.
- Open weights for Qwen3.8-Max and Qwen3.8-27B are promised subsequent week.
- No benchmark desk, license, or activated-parameter depend has been printed.
- The 27B checkpoint, not the flagship, is the practical on-premise deployment path.
Take a look at the Technical particulars, API and Qwen Studio. Additionally, be happy to comply with us on Twitter and don’t neglect to affix our 150k+ML SubReddit and Subscribe to our Publication. Wait! are you on telegram? now you’ll be able to be part of us on telegram as effectively.
Must companion with us for selling your GitHub Repo OR Hugging Face Web page OR Product Launch OR Webinar and many others.? Join with us
Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is dedicated to harnessing the potential of Synthetic Intelligence for social good. His most up-to-date endeavor is the launch of an Synthetic Intelligence Media Platform, Marktechpost, which stands out for its in-depth protection of machine studying and deep studying information that’s each technically sound and simply comprehensible by a large viewers. The platform boasts of over 2 million month-to-month views, illustrating its recognition amongst audiences.
