For CX, finance and procurement
Review your AI support bill.
Put the effective terms, invoice and business records together. Get reviewable billing differences, a finance one-pager and a clear account of what the evidence cannot prove.
GitHub v0.5.0 · 450 Healthcheck tests verified on 2026-10-03. Use the pinned Git install while PyPI publishing is pending. Customer payment and repeat use remain unverified.
We audited the AI usage-tool ecosystem.
The same five bugs kept appearing.
One assistant message billed per content block (2–5×). Re-emitted Codex events counted twice. Cache writes priced at 1.25× instead of 2×. Price tables 5× off — including in the industry's pricing source of record. Sessions wiped on resume. Every finding: pinned commit, runnable repro, public filing — 23 merged or accepted upstream, an external maintainer's 136k-event corpus review challenged an initial finding and led to a correction.
Read the Token-Accounting Bug Report →
External methodology correction: the author revised 436k to 54,154 tokens; not a dollar invoice →
Maintain a usage tool? Audit it in 10 minutes with our fixtures →
AI is starting to charge by the outcome.
Nobody defined the outcome.
When money rides on a measured unit, someone has to define the unit. That is what a measurement standard is for.
What counts as one resolution, one operation, one outcome? Retries, reopens, and silent “assumed resolutions” all change the number — and the bill.
A yardstick, or an argumentAI support vendors charge for outcomes under different conditions. Intercom currently includes several outcome types; the effective terms and billing period determine which rules apply.
Per-outcome billing is liveIntercom Fin bills per resolution with a money-back guarantee. Buyers dispute what was counted; sellers refund what cannot be proven.
Both sides need a refereeBilled per resolution, per task, or per usage — re-run the bill on your own machine and get an AMS-1 settlement statement: what was charged, what the vendor’s own rules would charge, and what the evidence cannot prove. Data never leaves your machine.
One semantics, three billing unitsZendesk now bills only “Verified Resolutions.” When vendors compete to define what “solved” means, the buyer needs an independent count.
The unit itself is contestedNot one publishes a dispute or refund process for miscounts. Intercom’s own docs: “We cannot guarantee the limit will be 100% accurate.”
No appeals process, anywhereAgentMeasure is the open yardstick for both sides of that bill —
usage and outcomes, one semantics, evidence you can audit.
A billing outcome needs
the applicable rules and evidence.
Three attempts of one operation are not three billable operations, unless the metering policy says so.
Agent systems collapse these states into vague “tool usage.”
AgentMeasure keeps them separate.
Your team knows which cases reopened, needed handover or remained unresolved.
Bring the business recordsReview the actual outcome charges and effective terms before a renewal or payment decision.
Bring the invoice and decisionHelp identify the authorized exports and the person already doing a billing review.
Start with one qualified accountSimilar events. Different facts.
Attempt≠Operation
Attempts are execution facts. Operations are logical intents. Retries are multiple attempts of one operation, not “2 operations with 50% success.”
Available≠Influential
A tool result appearing in the next context proves availability. Influence — that it changed the agent's behavior — needs counterfactual evidence. Reference is a weak estimator, not a state. Unknown is a valid, first-class answer.
Observed≠Inferred
A selection is not a preference. A model, router, policy, workflow, or user can make it. AgentMeasure records what happened before inferring why.
Measurement model
From reach to value.
Metric families, not a universal KPI. The same semantics map onto search, booking, and compute.
Effect confirmation: planned
Defined does not mean observable in every runtime. Semantics are spec status; observation depends on the harness. See the observability matrix ↓
Provider-side measurement.
No agent-side install.
No agent-side install
Third-party agents never need AgentMeasure. Measurement happens where the capability is delivered.
Provider-side measurement
The SDK wraps your tool handlers and emits canonical observations using one schema and six payload classes.
Not on the critical path
Non-blocking, durable spool with loss accounting. If the SDK fails, your server doesn’t.
Your data can stay local
JSONL on your machine, analytics generated locally. Hosted ingestion is explicit opt-in, later.
MCP is the first reference surface, not a requirement. Closed-source software is fine.
Telemetry records events. AgentMeasure defines what those events mean.
What works today.
Shipped, with tests, conformance vectors, and published profiles.
Numbers track releases and the roadmap.
Start with one scoped review.
Check the effective terms, export fields and upcoming decision before agreeing the review. The open-source tools and AMS-1 standard remain free.
- Local tools, conformance pack and CI action
- AMS-1 statements and public rule library
- Reproducible checks with missing evidence visible
- Run in your environment
- One vendor and one complete billing period
- Data and rule feasibility checked first
- Finance one-pager, findings and evidence bundle
- 30-minute walkthrough; processing scope agreed with you
Design partners are selected by data feasibility and feedback. Public testimonials require separate consent.
- Monthly, quarterly or before a renewal
- Effective rule versions and period comparison
- Vendor count, record volume and manual review agreed in advance
- Repeat need established after the first review
Billing-rule differences, stricter buyer standards and actual credits stay separate. Every delivery includes coverage and unresolved evidence. Findings do not establish a vendor-confirmed refund; negotiation stays with the parties. The verification rules stay public and free; paying does not change what the evidence says.
See the model in two minutes.
What the demo proves
- ● Canonical observations - one portable schema
- ● Operation / attempt correlation, fail-closed
- ● Caller attribution: correlated / declared / unknown
- ● Local analytics without any cloud
- ● Deterministic replay - reproducible numbers
Conformance in CI
Turn measurement assumptions into checks.
One workflow step reads your telemetry fixture and reports PASS / FAIL / UNPROVABLE per invariant. The evidence contract is explicit; what the data cannot decide is a finding, not a zero.
External provider trial
Find out what your agent traffic actually is.
Start with one sample. Scale to 7 days when it pays. Your data stays on your machine.
Billed per AI “resolution” (Intercom Fin, Zendesk AI, Salesforce)? — get an evidence-based check of that bill. Send a sanitized conversations/tickets export, or run our free CLI locally — data never leaves your machine. You get back a two-line settlement statement (AMS-1): Tier 1, everything the vendor’s own published rules would bill — including resolutions only its own system attested; Tier 2, only resolutions an affected party actually confirmed. Every line traces to a named rule, and anything the export cannot prove is listed as unprovable, never guessed. No charge — we are validating the method on real bills.
Not on per-resolution billing? Start smaller — a zero-install measurement check. Send 20–100 anonymized trace or log rows (or point to a public export). We map them locally and send back a short check: what safely counts as attempts vs operations, where retries may inflate usage, and what the telemetry cannot prove. No SDK, no integration, no 7-day commitment. Raw data stays local by default — a sanitized sample is only shared if you explicitly authorize it. A sample check is not a full audit and claims no causal effect. Production samples and synthetic fixtures enter through separate intakes.
Then, if the check pays for itself — the full 7-day audit. One capability, one week of real traffic, a Measurement Report and a 45-minute review call, every business conclusion carrying its evidence class:
How many calls are actually one logical operation?
What does one successful operation really cost?
How much “production” traffic is CI, synthetic, or unknown?
You get: a Measurement Report plus a 45-minute review call. Every business conclusion includes its evidence class.
Three providers, one capability each. If the method breaks on your data, we publish that. We already applied this discipline to six public claims in Benchmark Run #001.
One measurement language
across agent runtimes.
A public record of observation blind spots: what each runtime can and cannot observe.
| Runtime | Operation | Attempt | Consumption | Delegation |
|---|---|---|---|---|
| Claude Code | ● | ● | ◐ | ◐ |
| Codex | ● | ● | ○ | ○ |
| DeepSeek Harness | ● | ● | ○ | ● |
| Pydantic AI | ● | ● | ○ | ○ |
| OTel GenAI | ◐ | ● | ○ | ○ |
Harness profiles, Draft 0.4.4: a fixed questionnaire per runtime, with the mapping rules and blind spots in writing. Read the profiles →
Measure usage. Then measure causality.
What happened?
Observation / correlation / evidence-graded facts
Did the change actually improve it?
Preregistered experiments / effect sizes / honest nulls
Software is becoming callable.
Before capabilities can be priced, compared, or purchased by agents,
they need a shared measurement layer.
That layer is the foundation for Capability as a Service. It is a direction, not a product sold today. The standard defines economic facts; money movement stays with payment rails.
The referee doesn’t open a store.
Rules we wrote down before anyone asked us to bend them.
The spec, the engine, the SDK, conformance, and the local dashboard stay open. The free list is written into governance — not revocable later.
Adoption is the productRaw data never leaves your machine by default. Only aggregated, anonymized evidence returns to the network — only if you opt in.
Your traces stay yoursNo paid placement, no recommendation slots, no custody of funds. The moment a referee sells shelf space, the measurements stop meaning anything.
Neutrality is the product