← AgentMeasure

Position note · Why we build

Outcome-based AI needs a yardstick

Roy Tong · September 2026

Something quiet happened in software pricing. In August 2024, Zendesk started charging per automated resolution — $1.50 to $2.00 each. Salesforce, Intercom, and Sierra followed: per-resolution billing, per-outcome contracts, money-back guarantees on AI performance. The seat is giving way to the outcome.

We think this is the right direction, and we think it is missing a piece: nobody defined the outcome.

What counts as one resolution? One operation? One “successfully completed task”?

The edges of that question are exactly where the money sits. A customer who stays silent is counted as resolved by some vendors — and billed. A ticket the customer reopens three days later was, apparently, resolved twice. A retry chain inflates an API bill that was never really three requests. Every one of these is a billing dispute waiting to happen, and both sides of the dispute are currently armed with nothing but their own dashboards.

“Just measure the final result” doesn’t work

The obvious answer — measure the business result at the end — runs into four gaps that don’t close by themselves:

Advertising went through this exact sequence. “Just count the sales” did not kill the measurement industry — it built one: Nielsen made audience a currency, DoubleVerify and IAS verified delivery, AppsFlyer attributed installs. When money attaches to a measured quantity, a measurement layer attaches to the money. Per-outcome AI is now at the start of that same road.

What we’re building, and what we refuse

AgentMeasure is an attempt to be the open yardstick for that road — usage and outcomes, one semantics, evidence you can audit. The spec, the SDK, conformance checks, and the local dashboard are free and open, and they stay that way. Our first external milestone came the way we think standards actually arrive: an upstream project merged one of our measurement invariants into their main branch.

Three rules we wrote down before anyone asked us to bend them:

  1. Free, forever. The free list is written into governance, not revocable later. Adoption is the product.
  2. Local-first, opt-in. Raw data never leaves your machine by default. Only aggregated, anonymized evidence returns — only if you choose.
  3. The referee doesn’t open a store. No paid rankings, no recommendation slots, no custody of funds. The moment a referee sells shelf space, the measurements stop meaning anything.

That last rule is not a moral pose. Every measurement business that lasted — Nielsen for a hundred years, FICO through institutional adoption — kept the referee out of the storefront. Neutrality is not a constraint on the business; it is the business.

Where this goes

When enough interactions run through a shared yardstick, some questions become answerable for the first time: which tool actually works, for which kind of task, at what real cost. Those answers belong to whoever runs the measurement — which is why the measurement has to be open, and why we’d rather it be governed in public than owned in private.

If you run agent-facing software or bill per outcome: send us a trace, get a measurement check. If you maintain a telemetry library: our conformance vectors are designed to be merged, not installed. And if you think any of the above is wrong, the discussions tab is open — we publish our nulls.