For CX, finance and procurement

Review your AI support bill.

Put the effective terms, invoice and business records together. Get reviewable billing differences, a finance one-pager and a clear account of what the evidence cannot prove.

GitHub v0.5.0 · 450 Healthcheck tests verified on 2026-10-03. Use the pinned Git install while PyPI publishing is pending. Customer payment and repeat use remain unverified.

Developer sample · conformance --fixture retry-chain.jsonl
✓ PASS execution-grain1 op · 3 attempts
✓ PASS retry-reconciliationdeclared 3 = rows 3
✓ PASS cost-preservation13 units conserved
✗ FAIL operation-graingrain drift
declared grainoperation
actual grainassignment
expected2.5 over 2
observed5.0 over 1
? UNPROVABLE cache-distinction
no cache-hit signal in evidence —
unsafe inference refused
found by an external contributor · fixed with their fixture as the regression guard UNPROVABLE is a finding, not a zero
Open sourceMIT
SpecDraft 0.4.5
CLIv0.5.0 · Git release
SDKv0.1.1
Ecosystem audit110+ tools · 45+ verified bugs
Upstream fixes23 merged or accepted
110+repositories deep-read
45+verified accounting bugs
23fixes merged or accepted upstream
136Kevents in a third-party re-review

We audited the AI usage-tool ecosystem.
The same five bugs kept appearing.

One assistant message billed per content block (2–5×). Re-emitted Codex events counted twice. Cache writes priced at 1.25× instead of 2×. Price tables 5× off — including in the industry's pricing source of record. Sessions wiped on resume. Every finding: pinned commit, runnable repro, public filing — 23 merged or accepted upstream, an external maintainer's 136k-event corpus review challenged an initial finding and led to a correction.

Read the Token-Accounting Bug Report →

External methodology correction: the author revised 436k to 54,154 tokens; not a dollar invoice →

Maintain a usage tool? Audit it in 10 minutes with our fixtures →

AI is starting to charge by the outcome.
Nobody defined the outcome.

When money rides on a measured unit, someone has to define the unit. That is what a measurement standard is for.

The question

What counts as one resolution, one operation, one outcome? Retries, reopens, and silent “assumed resolutions” all change the number — and the bill.

A yardstick, or an argument
$1.50–2.00 per resolution

AI support vendors charge for outcomes under different conditions. Intercom currently includes several outcome types; the effective terms and billing period determine which rules apply.

Per-outcome billing is live
$0.99 per resolution

Intercom Fin bills per resolution with a money-back guarantee. Buyers dispute what was counted; sellers refund what cannot be proven.

Both sides need a referee
The buyer’s side of the bill

Billed per resolution, per task, or per usage — re-run the bill on your own machine and get an AMS-1 settlement statement: what was charged, what the vendor’s own rules would charge, and what the evidence cannot prove. Data never leaves your machine.

One semantics, three billing units
2026-05 · “Verified” only

Zendesk now bills only “Verified Resolutions.” When vendors compete to define what “solved” means, the buyer needs an independent count.

The unit itself is contested
20 vendors checked

Not one publishes a dispute or refund process for miscounts. Intercom’s own docs: “We cannot guarantee the limit will be 100% accurate.”

No appeals process, anywhere

AgentMeasure is the open yardstick for both sides of that bill —
usage and outcomes, one semantics, evidence you can audit.

Read the note: the yardstick for outcome-based AI →

A billing outcome needs
the applicable rules and evidence.

Available≠Presented
Presented≠Selected
Selected≠Used
Used≠Useful
Useful≠Incremental Value
Measured Usage≠Billable Usage

Three attempts of one operation are not three billable operations, unless the metering policy says so.

Agent systems collapse these states into vague “tool usage.”
AgentMeasure keeps them separate.

CX and support operations

Your team knows which cases reopened, needed handover or remained unresolved.

Bring the business records
Finance and procurement

Review the actual outcome charges and effective terms before a renewal or payment decision.

Bring the invoice and decision
Implementation partners

Help identify the authorized exports and the person already doing a billing review.

Start with one qualified account

Similar events. Different facts.

01

Attempt≠Operation

Attempts are execution facts. Operations are logical intents. Retries are multiple attempts of one operation, not “2 operations with 50% success.”

1 intent
↓
Attempt 01 ✕
Attempt 02 ✓
↓
= 1 Operation
02

Available≠Influential

A tool result appearing in the next context proves availability. Influence — that it changed the agent's behavior — needs counterfactual evidence. Reference is a weak estimator, not a state. Unknown is a valid, first-class answer.

Tool returned result
↓
?
Did the caller actually use it?
RETURNEDobservable
CONSUMEDneeds evidence
03

Observed≠Inferred

A selection is not a preference. A model, router, policy, workflow, or user can make it. AgentMeasure records what happened before inferring why.

selected_by:
model / router / policy
workflow / user / platform
↓
what happened ≠ why

Measurement model

From reach to value.

Metric families, not a universal KPI. The same semantics map onto search, booking, and compute.

01 Reach Did the capability enter the choice set? Spec · Defined
02 Choice Was it selected? Spec · Defined
03 Use Was it actually executed? Spec · Defined
04 Utility Did the result contribute to the task? Spec · Partial Consumption: partial
Effect confirmation: planned
05 Value Did it improve the outcome? Spec · Draft Incremental value: formula defined

Defined does not mean observable in every runtime. Semantics are spec status; observation depends on the harness. See the observability matrix ↓

Provider-side measurement.
No agent-side install.

Agent Claude / Codex / any runtime
MCP / API / CLI▼
Capability Provider
Business handler
AgentMeasure SDK
observations · JSONL▼
Local Collector
▼
Local analyticsno cloud required
▼
Hosted analyticsopt-in · planned
01

No agent-side install

Third-party agents never need AgentMeasure. Measurement happens where the capability is delivered.

02

Provider-side measurement

The SDK wraps your tool handlers and emits canonical observations using one schema and six payload classes.

03

Not on the critical path

Non-blocking, durable spool with loss accounting. If the SDK fails, your server doesn’t.

04

Your data can stay local

JSONL on your machine, analytics generated locally. Hosted ingestion is explicit opt-in, later.

MCP is the first reference surface, not a requirement. Closed-source software is fine.

Telemetry records events. AgentMeasure defines what those events mean.

What works today.

Shipped, with tests, conformance vectors, and published profiles.

Provider SDK @agentmeasure/mcp · release tarball v0.1.1
Registry publish npm: pending scope/token Pending
Canonical observations schemas/observation.schema.json · 6 payload classes Shipped
Local analytics product/local-analytics.py Shipped
Conformance vectors CI · conformance workflow Passing
Last verified CI runs Traceable
Harness profiles codex / claude-code / dsh / pydantic-ai / otel 5
Hosted ingestion · dashboard next on the product roadmap In development
21 SDK tests Deterministic E2E fixture No cloud required Data stays local

Numbers track releases and the roadmap.

Start with one scoped review.

Check the effective terms, export fields and upcoming decision before agreeing the review. The open-source tools and AMS-1 standard remain free.

Open source
$0 forever
  • Local tools, conformance pack and CI action
  • AMS-1 statements and public rule library
  • Reproducible checks with missing evidence visible
  • Run in your environment
Run it yourself
First-look review
$990 one-time pilot quote
  • One vendor and one complete billing period
  • Data and rule feasibility checked first
  • Finance one-pager, findings and evidence bundle
  • 30-minute walkthrough; processing scope agreed with you
Check the review scope

Design partners are selected by data feasibility and feedback. Public testimonials require separate consent.

Recurring or renewal review
By scope
  • Monthly, quarterly or before a renewal
  • Effective rule versions and period comparison
  • Vendor count, record volume and manual review agreed in advance
  • Repeat need established after the first review
Discuss the next review

Billing-rule differences, stricter buyer standards and actual credits stay separate. Every delivery includes coverage and unresolved evidence. Findings do not establish a vendor-confirmed refund; negotiation stays with the parties. The verification rules stay public and free; paying does not change what the evidence says.

See the model in two minutes.

sh / demo
$ git clone https://github.com/roy-tong/AgentMeasure
$ cd AgentMeasure
$ ./examples/demo-e2e.sh
✓ 42 calls captured - mock MCP server, 3 caller classes
✓ 84 canonical observations · 0 rejected
✓ operation correlation - strict, no fallback counting
✓ caller attribution: claude 14 / codex 14 / unknown 14
✓ strict qualified 0% - synthetic traffic excluded
✓ deterministic replay - same fixture, same result
requires node ≥ 18 + python3, no external deps, isolated workspace

What the demo proves

  • ● Canonical observations - one portable schema
  • ● Operation / attempt correlation, fail-closed
  • ● Caller attribution: correlated / declared / unknown
  • ● Local analytics without any cloud
  • ● Deterministic replay - reproducible numbers
Open Quickstart →

Conformance in CI

Turn measurement assumptions into checks.

One workflow step reads your telemetry fixture and reports PASS / FAIL / UNPROVABLE per invariant. The evidence contract is explicit; what the data cannot decide is a finding, not a zero.

Workflow - uses: roy-tong/AgentMeasure@<sha> 1 step
Input fixtures/telemetry.jsonl (FMT-002) yours
Verdicts execution-grain · retry-reconciliation · cost-preservation · operation-grain · evidence-boundary 3-state

External provider trial

Find out what your agent traffic actually is.

Start with one sample. Scale to 7 days when it pays. Your data stays on your machine.

Billed per AI “resolution” (Intercom Fin, Zendesk AI, Salesforce)? — get an evidence-based check of that bill. Send a sanitized conversations/tickets export, or run our free CLI locally — data never leaves your machine. You get back a two-line settlement statement (AMS-1): Tier 1, everything the vendor’s own published rules would bill — including resolutions only its own system attested; Tier 2, only resolutions an affected party actually confirmed. Every line traces to a named rule, and anything the export cannot prove is listed as unprovable, never guessed. No charge — we are validating the method on real bills.

Not on per-resolution billing? Start smaller — a zero-install measurement check. Send 20–100 anonymized trace or log rows (or point to a public export). We map them locally and send back a short check: what safely counts as attempts vs operations, where retries may inflate usage, and what the telemetry cannot prove. No SDK, no integration, no 7-day commitment. Raw data stays local by default — a sanitized sample is only shared if you explicitly authorize it. A sample check is not a full audit and claims no causal effect. Production samples and synthetic fixtures enter through separate intakes.

Then, if the check pays for itself — the full 7-day audit. One capability, one week of real traffic, a Measurement Report and a 45-minute review call, every business conclusion carrying its evidence class:

Retry inflation

How many calls are actually one logical operation?

Execution economics

What does one successful operation really cost?

Traffic quality

How much “production” traffic is CI, synthetic, or unknown?

You get: a Measurement Report plus a 45-minute review call. Every business conclusion includes its evidence class.

Three providers, one capability each. If the method breaks on your data, we publish that. We already applied this discipline to six public claims in Benchmark Run #001.

One measurement language
across agent runtimes.

A public record of observation blind spots: what each runtime can and cannot observe.

Runtime Operation Attempt Consumption Delegation
Claude Code ● ● ◐ ◐
Codex ● ● ○ ○
DeepSeek Harness ● ● ○ ●
Pydantic AI ● ● ○ ○
OTel GenAI ◐ ● ○ ○
● observable ◐ partial / derived ○ unavailable

Harness profiles, Draft 0.4.4: a fixed questionnaire per runtime, with the mapping rules and blind spots in writing. Read the profiles →

Measure usage. Then measure causality.

AgentMeasure

What happened?

Observation / correlation / evidence-graded facts

AgentMeasure Lab

Did the change actually improve it?

Preregistered experiments / effect sizes / honest nulls

Explore Lab →

Software is becoming callable.

Human Software Economy
Human↓ UI↓ App↓ Seat / Month
Agent Capability Economy
Agent↓ Capability↓ Execution↓ Outcome↓ Usage → Value → Transaction

Before capabilities can be priced, compared, or purchased by agents,
they need a shared measurement layer.

That layer is the foundation for Capability as a Service. It is a direction, not a product sold today. The standard defines economic facts; money movement stays with payment rails.

The referee doesn’t open a store.

Rules we wrote down before anyone asked us to bend them.

Free, forever

The spec, the engine, the SDK, conformance, and the local dashboard stay open. The free list is written into governance — not revocable later.

Adoption is the product
Local-first, opt-in

Raw data never leaves your machine by default. Only aggregated, anonymized evidence returns to the network — only if you opt in.

Your traces stay yours
No rankings for sale

No paid placement, no recommendation slots, no custody of funds. The moment a referee sells shelf space, the measurements stop meaning anything.

Neutrality is the product