Skip to content

CompanyBench

Everyone benchmarks the model. We benchmark the harness that runs the Company.

One model pinned across every entrant. Every call measured at one boundary.

AEQI builds one of the harnesses measured here. This board is in method preview and pre-registration — no performance scores are published yet, and we lead with our own losses, not our wins. Read the method & charter →

4
Tasks
1 pinned
Model
0
LLM judges
100%
Solved
0
Tests edited
14×
Cost spread

coding-v1

z-ai/glm-5.24 tasks × 1 trial · one model pinned across every entrant2026-07-29

coding-v1 results by harness, one model pinned, every call measured at a single boundary.
HarnessSolvedStepsToolsTokens inCachedTokens outEmbedDurationScoreCost
aeqihost
Company (daemon + memory)
4/437331,516,2071,432,064 · 94%6,2708 · 1,0334m 1s15$0.3988provider-reported
Claude Code
`claude -p`
4/42018690,032538,813 · 78%5,13203m 15s26$0.2003list-price
aeqihost
`aeqi run` (one-shot, no memory)
4/4181840,8770 · 0%3,235059s100$0.0257provider-reported
Pi
`pi -p`
4/41939,67924,421 · 62%2,991057s96$0.0277self-reported
Codex
runs on the pinned model · full run not yet spent
not yet run
Hermes Agent
exits silently on a headless prompt
not yet run
Kilo Code
IDE-native; CLI path not yet smoke-tested
not yet run
OpenCode
not yet installed
not yet run
OpenClaw
not yet installed
not yet run

Score is resource-efficiency at fixed success (cost 0.50, steps 0.25, wall-time 0.25), anchored so the leanest row is 100. AEQI in Company mode scores last: 14× Pi’s cost and 2× Claude Code’s for the same four tasks, carrying 94% of its 1.5M input tokens as cache reads. That reproduces the Stage-1 finding on better instrumentation rather than overturning it — the memory and company scaffold AEQI carries is real spend, and this board prices it. Cost sources differ and are labelled per row: provider-reported is OpenRouter’s own figure, list-price is computed from metered tokens with cache reads priced as cache reads, self-reported is the harness’s own usage block.

Tasks

fix-failing-suite
runs the tests, reads the failures, iterates to green
hidden-edge-cases
implements the spec, not the two visible cases
crash-only-at-runtime
executes the code to find a fault reading cannot reveal
no-regression
fixes the reported bug without breaking what already worked

Method

  • Held-out tests the agent never saw. No LLM judge.
  • Protected test files are hash-checked — editing one scores as a failure.
  • Each task must fail from its seed and pass from a reference solution, or it is cut.
  • Every harness routed through one metering boundary; cells the provider never served are void.
  • AEQI is measured as a Company, not as a one-shot CLI — the one-shot row is shown for contrast only.

Method & conflict-of-interest charter →

The ring

Cookies for sign-in and analytics. No third-party tracking.