Search...
Ctrl K
Models
Providers
Rankings
Chat
Models
Providers
Rankings
Chat
Search...
Ctrl K
Sign In
Sign In
SWE Bench Multilingual - Benchmark Leaderboard & Model Performance | AI Stats
SWE Bench Multilingual
Overview
Overview
Code
Recorded Results
2
Average Score
0.76
Score Range
0.76 - 0.76
Leading Model
0.76 - Claude Opus 4.5
Scores Over Time
Individual benchmark scores plotted by date.
Models Using This Benchmark
Organisation
Model
Reported
Top Score
Info
Self Reported
Source
Anthropic
Claude Opus 4.5
24 Nov 2025
0.76
Avg@5
Yes
Source