跳到正文
原文
Mistral AI·精选AI 评分78

Mistral 发布 Mistral Large 4 公开预览

Introducing Mistral Large 4

官方原文核对 · 关键事实加强核对

依据原始来源与逐字段引文;不代表独立实测,厂商性能声明仍是厂商自述。

  • 原文依据:Introducing Mistral Large 4
  • 原文依据:Le Chonk Today, we’re launching a public preview of Mistral Large 4.
  • 原文依据:Unofficially ML4, very officially: le Chonk .
  • 原文依据:ML4 pushes the frontier of open-weight performance.
  • 原文依据:You can try the preview API today on Mistral Studio .
  • 原文依据:Weights drop end of this month.
  • 原文依据:Frontier performance ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters.
  • 原文依据:It is our largest and most capable model to date, and it continues to improve rapidly as we refine it.
  • 原文依据:The model demonstrates exceptional performance across coding, agentic workflows, and multimodal understanding.
  • 原文依据:It already achieves performance competitive with the strongest open-source models globally, while significantly outperforming any open-weight model developed in the US or Europe.
  • 原文依据:On critical enterprise workloads, including cybersecurity, finance and law, we find it to be state-of-the-art among open models.
  • 原文依据:In some domains such as visual grounding, it goes further still, surpassing even frontier closed models.
  • 原文依据:We will release the weights by the end of the month.
  • 原文依据:ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.
  • 原文依据:The public preview is served on that same infrastructure.
  • 原文依据:Fun fact: a significant share of ML4’s training data was multilingual, spanning more than 160 languages, including every official language of the European Union.
  • 原文依据:Capabilities deep-dive Cybersecurity ML4 is one of the world's strongest AI models for cybersecurity.
  • 原文依据:On the Artificial Analysis Cyber Index, an independent evaluation of how well AI models find and fix security flaws in real software, it ranks among the top five models globally and leads open-weight models developed outside China by a wide margin.
  • 原文依据:On one of the index's tests, which asks a model to reproduce a real vulnerability in open-source software and then patch it, ML4 scores 82%, the highest of any model.
  • 原文依据:It also solves 93% of the challenges in Cybench, a set of 40 exercises drawn from security competitions, one of the highest scores reported for an open-weight model.
  • 原文依据:ML4 Preview ranked second of five models (3.74), ahead of Kimi K3 (3.59), GLM-5.3 (3.60) and GLM-5.2 (3.40), and behind only Claude Opus 5 (4.22).
  • 原文依据:On AutomationBench — 657 business workflows across apps like Gmail, Google Sheets, Slack, and Salesforce — it scores 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro, and DeepSeek V4 Pro.
  • 原文依据:On AA-Briefcase, which evaluates long-horizon knowledge work, it reaches 1,393 Elo, ahead of DeepSeek V4 Pro.
  • 原文依据:On visual grounding particularly, we find ML4 to be one of the most capable models we tested, for instance surpassing GPT-6-Astra on Dense 200 (42% vs 41%).
  • 原文依据:In benchmarks, ML4 is state of the art on SciCode-Verified among open-weight models.
  • 原文依据:Notably, we evaluated ML4 through third party evaluators ( vals.ai ) on representative tasks for both legal and financial tasks, finding the model exceeds GPT-6-Astra in both cases.
  • 原文依据:On HarveyAI’s Legal Agent benchmark, ML4 outperforms all open-source models.
  • 原文依据:On Lakera’s public B3 AI Security Benchmark , ML4 resists 93.3% of attacks – we see no higher scores among competitors.
  • 原文依据:We highlight our results on the KORA Benchmark , where ML4 again sits at our highest measured score among OSS models (1.691, with 2 being the maximum denoted as “ Exemplary ”).
  • 原文依据:Despite strong performance on Cyber benchmarks, the average refusal rate of the model on cyber prompts from JailbreakBench , StrongREJECT , and AgentHarm is higher than all OSS models.
  • 原文依据:ML4 is the first milestone on the roadmap funded by our €3 billion Series D — the largest equity round ever raised by a European technology company.
  • 原文依据:We will release the weights by the end of the month, along with more details on the architecture, additional benchmarks, and our post-training methodology.
AI 导读

Mistral 官方发布 Mistral Large 4 公开预览,称其为 1 万亿参数、49B 激活参数的原生多模态模型,可在 Mistral Studio 试用预览 API,权重将于本月底放出。

推荐理由

官方给出 1 万亿参数、49B 激活的开放权重路线与多模态、智能体编码基准,可据此判断开源前沿模型的当前水位。

来源:Mistral AI · mistral.ai