Skip to content

Harness · anthropic (proprietary)

Claude Code

The strongest proprietary coding harness in the ring, and the honest ceiling to measure against. Included because a benchmark that only ranks the harnesses we can beat is not a benchmark.

AEQI builds one of the harnesses measured here. This board is in method preview and pre-registration — no performance scores are published yet, and we lead with our own losses, not our wins. Read the method & charter →

Result: not yet measured

The facts

Maintainer
Anthropic
Source
anthropic (proprietary)
License
Proprietary
Category
Coding harness
Headless run
Yes — `claude -p` (print mode, `--output-format json`)
Native OpenRouter
— can be pinned to one shared model
Determinism notes
Speaks the Anthropic wire format rather than the OpenAI one, but OpenRouter serves that format too — so it runs on the same pinned model and the same account as every other entrant, with no proxy in between.

Where it sits

Performance scores land here only with published traces, under the pre-registered method. Until then, this profile is the fair-run readiness check: can we start Claude Code headless and pin it to one shared model to isolate the scaffold?

Compare

← Back to the roster

Cookies for sign-in and analytics. No third-party tracking.