Skip to content

Open any AI leaderboard and you hit the same wall: math, physics, language, code. MMLU, GSM8K, GPQA, HumanEval, SWE-bench, Terminal-Bench. We have gotten very good at measuring whether a model can solve a competition problem or fix a GitHub issue.

The economy does not run on competition problems. It runs on operations — reconciling the ledger, fulfilling the order, answering the ticket, closing the month. The unglamorous machine of a company, kept running one more day. And almost nobody measures whether an agent can do that.

Why the frontier is empty

It's empty for one reason: business is hard to grade. A coding task passes or fails a test suite — clean, objective, cheap. “Did the company do well?” is a judgment call, and the field's instinct is to hand judgment calls to another model. An LLM grading an LLM is a hall of mirrors; nobody serious cites it.

So the people who build benchmarks — mostly from math and computer science — benchmark what grades cleanly, and operations gets skipped. Which leaves the most economically important question in AI right now — can an agent actually run the back office of a business — with no scoreboard at all.

That absence isn't a gap to lament. It's the moat. Whoever makes business operations gradeable to the cent owns the category — because the reason it doesn't exist is the reason it's hard to copy.

What we're building

CompanyBench is our answer, and it starts narrow on purpose. Not “business” in the abstract — operations, the part that grades. A harness runs a small store's back office over a run of shifts: process orders, issue refunds to a frozen policy, reconcile a double-entry ledger, manage inventory, resolve tickets. Every consequential action goes through one typed command, so the outcome is a structured event, not a paragraph. Then a plain program grades it against a golden — no model judging another model. Double-entry books and integer inventory give conservation laws that check to the cent: the balance reconciles, stock never goes negative.

And because whatever state the agent leaves is the next shift's opening state, mistakes compound. That lets us measure the thing a single-shot test never can: drift — how long a harness keeps a business coherent before it loses the plot, how often it needs a human, what it costs to run a correct day. We hold the model constant across a sweep and vary the harness, so the score is the scaffolding, not the raw IQ of the LLM inside it. The full method is public, pre-registered before a single number exists.

The honesty tax

We run this benchmark and we compete in it. There is exactly one way that isn't a scam: we publish our own losses and lead with them, we pin every competitor to its own recommended config, we publish the traces, and we never rank ourselves first without an outside party reproducing it. It's early — method-preview, first runs rolling out, and our own harness is not at the top. We're publishing the method before the numbers on purpose. A benchmark you can trust is one whose rules were frozen before anyone knew the scores.

The long game: authority, co-authored

A benchmark becomes a standard when the people who define the field help write it. So we're not building this alone. We want to build it with universities — business schools to define what “running a company well” actually means, computer-science labs to keep the grading rigorous and the reproduction honest. Not a vendor's scoreboard: a co-authored exam. The tools people already trust — HAL, Terminal-Bench — earned their authority exactly this way, in the open, with academic hands on the method.

If you teach or research any of this — operations, agent evaluation, the economics of autonomous work — come build it with us. Co-author a scenario, contribute the reference oracle, reproduce the board and tell us where we're wrong. We reply to every serious proposal.

Everyone benchmarks the model. We're benchmarking the Company — the one job the whole economy actually pays for, and the one nobody has bothered to measure. See the board. Then help us make it the exam that matters.

Cookies for sign-in and analytics. No third-party tracking.