AI Agent Evaluations
Performance results of AI coding agents on Nuxt code generation tasks, measuring success rate and execution time.
Agent Performance Results
| Model | Agent | Avg Duration | Avg List Cost | Success Rate | First-Try Rate | |
|---|---|---|---|---|---|---|
Kimi K3 | OpenCode | 412.71s | $0.311 | 100% | 97% | |
Claude Fable 5 | Claude Code | 309.87s | $1.194 | 100% | 97% | |
GPT 5.6 Sol (xhigh) | Codex | 322.11s | $0.425 | 100% | 93% | |
GPT 5.3 Codex (xhigh) | Codex | 291.15s | $0.193 | 100% | 90% | |
Claude Opus 4.8 | Claude Code | 266.89s | $0.542 | 100% | 90% | |
Kimi K2.7 Code | OpenCode | 328.95s | $0.091 | 100% | 80% | |
Claude Opus 5 | Claude Code | 417.04s | $1.873 | 97% | 97% | |
Claude Sonnet 5 | Claude Code | 307.80s | $0.511 | 97% | 97% | |
GPT 5.5 Pro | Codex | 700.25s | $8.630 | 97% | 93% | |
Cursor Composer 2.0 | Cursor | 286.92s | $0.086 | 97% | 93% | |
Cursor Composer 2.5 | Cursor | 273.15s | $0.090 | 97% | 87% | |
Claude Opus 4.7 | Claude Code | 222.55s | $0.381 | 97% | 87% | |
MiniMax M3 | OpenCode | 233.70s | $0.042 | 97% | 83% | |
Claude Opus 4.6 | Claude Code | 243.58s | $0.341 | 97% | 83% | |
Gemini 3.1 Pro Preview | OpenCode | 297.97s | $0.270 | 93% | 80% | |
Claude Sonnet 4.6 | Claude Code | 254.27s | $0.321 | 90% | 80% | |
Kimi K2.6 | OpenCode | 296.14s | $0.080 | 90% | 77% | |
GPT 5.4 (xhigh) | Codex | 310.34s | $0.368 | 90% | 73% | |
Claude Sonnet 4.5 | Claude Code | 239.83s | $0.187 | 57% | 47% | |
MiniMax M2.7 | OpenCode | 204.39s | $0.009 | 47% | 30% |
Each evaluation is attempted up to 4 times. Success Rate is the percentage of evals that passed on at least one attempt; First-Try Rate is the percentage that passed on the first attempt, used to break ties between models with the same success rate. Avg Duration is the mean time an agent took per eval. Avg List Cost is the mean cost per eval, estimated from the tokens each run used at the provider's public list price, so it's a relative guide and not a bill: your rate depends on caching, discounts and subscription plans. Expand a row to see per-eval results, where a 1/3 badge means the eval failed twice before passing.