What to look for
Arena compares 21 combinations across Claude Code, Codex CLI, and Pi. The study shows why model selection also involves the software managing its tools and context: similar success rates can come with different costs.
Keep in mind
The experiment samples 30 tasks each from SWE-bench Lite and Terminal-Bench 2.0, with three attempts per combination and task. Its pricing snapshot and harness settings matter. These results do not establish a universal winner for other workloads.
Read the original work
HarnessTax: How Much Does the Harness Matter for Coding Agents? ↗By Arena Team · Arena
Published Sep 16, 2026. Source checked Oct 5, 2026.
Reading note by SaveMyToken · Added Oct 5, 2026. This page is a short editorial introduction; the full article belongs to its original publisher.
