What to look for

Arena compares 21 combinations across Claude Code, Codex CLI, and Pi. The study shows why model selection also involves the software managing its tools and context: similar success rates can come with different costs.

Keep in mind

The experiment samples 30 tasks each from SWE-bench Lite and Terminal-Bench 2.0, with three attempts per combination and task. Its pricing snapshot and harness settings matter. These results do not establish a universal winner for other workloads.

Read the original work

HarnessTax: How Much Does the Harness Matter for Coding Agents? ↗

By Arena Team · Arena

Published Sep 16, 2026. Source checked Oct 5, 2026.

Reading note by SaveMyToken · Added Oct 5, 2026. This page is a short editorial introduction; the full article belongs to its original publisher.