Early access

Understand the economics of your AI products.

Connect AI and infrastructure costs to customers, workflows, and revenue. Make informed pricing decisions, understand margins, and plan for growth.

Example · accounts with revenue mapped, 30 days
Revenue $157,320
Cost to serve $33,030
Gross margin 79%
Northwind Ltd −13%
Kestrel Health 66%
Brightline 72%
138 other accounts 85%
Vantage Group
Gross margin by account · zero marked
A blended 79% hides an account served at a loss. Vantage carries $2,410 of cost with no revenue mapped, so it is excluded from the blend rather than counted as zero.

The premise

Know what your AI costs to deliver. Decide how to price, build, and grow.

You cannot price what you cannot cost. Everything below is drawn from one dataset, so the three answers agree with each other.

What does our product cost to deliver?

Across workflows, features, and customers, with every layer metered rather than inferred from the model bill.

How do those costs relate to revenue?

Across accounts, usage patterns, and plans, joined to your own revenue events.

What should we change or invest in next?

Model comparisons, measured outcomes, and spending trends, each with the evidence behind it.

Cost to deliver

Delivery cost is six layers. Your invoice shows one.

Tokens are the part you can already see, and they are less than half of it. ai-tally meters vector search, tool calls, embeddings, GPU hours, and egress from their own ingest paths, then attributes the total to the workflow that spent it.

Example · trailing 30 days
What the providers billed you
$16,390
OpenAI $10,480 Anthropic $5,180 Google $730
Total delivery cost · all workflows
$35,440
LLM Compute Vector DB Tool calls Egress Embeddings
Example delivery cost per workflow, trailing 30 days
WorkflowRunsCost / run30-day costShare
Support copilot212,400$0.0710$15,08043%
Doc search124,000$0.0800$9,92028%
Report writer31$139.03$4,31012%
Onboarding agent8,640$0.4086$3,53010%
Internal jobs46,200$0.0563$2,6007%
All workflows391,271$35,440100%
The model bill was 46% of the AI bill. One API key served all five workflows, so the invoice could not have split them.

Cost and revenue

Margins vary more inside your book than between your plans.

Direct spend lands on an account by hashed account id. Shared layers cannot be attributed that way, so they are allocated pro-rata on direct spend, and the rule is named on screen. You are told which half of each number was measured and which was derived.

Example · account economics, last 30 days
Example account economics, last 30 days
AccountRevenueDirectAllocatedMargin
Northwind Ltd$5,400$4,132$1,988−13%
Brightline$17,790$3,362$1,61872%
Kestrel Health$9,530$2,188$1,05266%
Vantage Group$1,627$783
138 other accounts$124,600$12,621$6,06985%
142 accounts$157,320$23,930$11,51079%
Allocated is the $11,510 of shared compute and egress spread pro-rata on direct spend. Northwind bills $5,400 a month and costs $6,120 to serve. Nine accounts run below 0% and together carry $11,420 of cost against $9,060 of revenue. The blended 79% covers only accounts with revenue mapped.

Recoverable cost

Some of that delivery cost bought nothing.

The cheapest margin you can gain is the spend that returned nothing. Each detector names where the waste is, says whether its number is spend already incurred or a saving it estimated, and drills through to the runs behind it. Findings are hypotheses with evidence, and two of the five below refuse to put a number on themselves.

Paid for nothing 1,204 billed runs ended failed or abandoned with no later success. The tokens were charged; no result was produced. The figure is what those runs spent, read off the ledger rather than modelled. How much of it you stop spending depends on which failures you fix. $2,840Observed spend
Wrong-sized model On doc search, claude-sonnet-4 ties gpt-4o inside the eval confidence interval at $0.04169 per call against $0.04984. The figure is what the swap would have saved at this window’s volume, so it is a projection, not a measurement. Cheaper candidates that lose the eval are not counted, however large their saving looks. $1,010Estimated saving
Duplicated work 512 failed runs that a later same-shape success replaced, so the same outcome was paid for twice. The figure is what those replaced attempts spent, counted apart from the runs above, which never succeeded, so no spend lands in both. A plain repeat is not claimed, because without message bodies it is indistinguishable from real multi-turn use. $1,190Observed spend
No measured return The onboarding agent spends $3,530 a month with no attributed value. Top-of-funnel work and revenue that simply is not wired yet look identical from telemetry, so no recoverable amount is claimed. Unbounded
Structural inefficiency Context bloat and runaway loops are judged against each workflow's own median, never a global average. The report writer settled 31 runs in this window; the floor is 50. Below floor
Spent on runs that returned nothing · 30 days $4,030
Estimated saving from the model swap · 30 days $1,010

Forecast

Where the month lands, and the day you cross budget.

Planning for growth needs a number before the month is over. A day-of-week-weighted median projection with an 80% confidence cone, held to a 14-day settled-history floor: a volatile number early in the month is worse than no number at all.

Example snapshot · September 8
$40K $20K $10K $0 BUDGET $30,000 SEP 26 · BREACH SEP 1 SEP 8 · SNAPSHOT SEP 30
Settled to date$9,280
Projected month-end$34,800
80% cone$31.2K – $38.9K
Budget breachSep 26

Model choice

Are you on the right model? Decided on your traffic, not a leaderboard.

The last lever is what you build on. Replay is a separate opt-in feature: an admin turns it on for the whole organization, and only then is a sample of request content captured, replayed against candidate providers under a daily budget cap, and scored by a pairwise judge. Ordinary cost telemetry never carries prompts or answers. Win rates carry Wilson 95% intervals, so a tie reads as a tie. What replay stores.

Example · doc search, 1,240 replayed calls, 30 days
gpt-4o In production
Cost per call $0.04984 · $6,180 / mo over 124,000 calls
BASELINE
claude-sonnet-4
Cost per call $0.04169 · $5,170 / mo, saves $1,010
46–56% WIN
gemini-2.5-flash
Cost per call $0.00900 · $1,116 / mo, saves $5,064
39–49% WIN
1% of doc-search traffic, held under the daily replay budget cap. Only claude-sonnet-4 ties: gemini’s interval excludes 50%, so its $5,064 is not claimable as savings. A candidate with no judged eval pass renders for quality, never a placeholder percentage.

Setup

What you connect, and what each connection buys you.

A provider key alone gets you the first view. The three after it each need one more connection, and every page says which numbers are still missing until you make it.

  1. Send your calls Proxy, Python SDK or OpenTelemetry, whichever suits your stack. One call takes about five minutes and fills in cost by workflow, feature and model.
  2. Connect your cloud bill Compute, vector search and egress come from your cloud bill, through a reference to a role or a secret rather than a pasted key. Skip it and delivery cost is the model bill again.
  3. Tag calls with a customer Per-account cost needs a hashed customer id on the call. The Python SDK hashes it inside your app; with the proxy or OpenTelemetry you hash it yourself. Untagged calls still count toward totals, and toward no account.
  4. Upload what customers pay A CSV, one row per customer per month, keyed on the same ids. Until it lands, margin reads as a blank instead of a zero.

Model comparison is a separate decision. It replays a scrubbed sample of your requests against candidate models, an admin turns it on for the whole organization, and it sends that sample to providers you may not have used before. What is captured and how long it is kept is written out in the security docs.

Know your margins before you price.

Send one call and cost by workflow is there in about five minutes. Accounts, margins and forecasts follow the four connections above.