Stepfun Step 5 Preview (LLM): On AA Pareto frontier

GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index

Artificial Analysis Intelligence Index v4.2

Qwen3.8-Flash-Next Intelligence, Performance and Price Analysis

Benchmarking Pocket-Scale Inference

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

Qwen3.8 27B scores 52 on Artificial Analysis

GLM-5.3-Flash Intelligence, Performance and Price Analysis

GLM-5.3 Artificial Analysis Benchmarks

Grok 4.6 (High) Intelligence, Performance and Price Analysis

Qwen3.8 Max now ranked as the best overall model by agentic index

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

Kimi K3: second only to Fable 5 on AA-Briefcase

DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

Grok Build 0.1: Intelligence, Performance and Price Analysis

Claude Sonnet 5 – benchmark results

GLM-5.2 is the new leading open weights model on Artificial Analysis

GLM 5.2 Performance Benchmarks

LLM leaderboard – Comparing models from OpenAI, Google, DeepSeek and others

Benchmarks and comparison of LLM AI models and API hosting providers