AI Usage Billing Case Study: A Month of Invoices That Match a Recount to the Billionth of a Dollar
An AI product that bills by the token has to count every request exactly once, even when the network does not cooperate. Clients retry, and a retry can arrive through a different gateway. A gateway that loses its link holds its reports and sends them hours later. A queue can stick at the end of the month, after the invoices are out. In the Meter demo the product’s gateways report every request to one meter, a compiled Precomputing policy inside SQLite, and the invoices are one SQL view over totals the meter keeps ready.
The run covers September 2026 for an invented AI writing product: six customers on three plans, three models, and two gateways, eu and us. About 2% of requests are retried. On 15 September the eu gateway loses its link for six hours. On 27 September a Free customer reaches its quota. And on 2 October, after September has closed, a queue stuck in the us gateway since the evening of the 30th is emptied.

The month in 1-hour windows, read from the meter. Late reports land in the hour their request happened, so the outage leaves no gap.
The Set-Up
| AI usage billing | |
|---|---|
| Customers | Sundial Research (Enterprise, 25% off list), Harbor Legal Drafts, Quarry Road Media and Maple Grove Travel (Pro), Juniper Bakery Blog and Northgate Tutors (Free) |
| Models | Small, medium and large, at example list prices from $0.30 to $48 per million tokens. Model cost is half of list |
| Reports | One per request: request id, time, customer, model, gateway, tokens in and out; 373,351 in all, retries included |
| Rules | A repeated request id refused for seven days; up to 48 hours late accepted; a month closes a day after it ends; every request kept until 90 days after the close |
| Ready answers | Requests, tokens in and tokens out per customer, model and month; tokens per customer and month with a quota |
| Billing | Price, plan and customer tables in the same file; invoices as a view, in billionths of a dollar |
The rules take four lines of the policy:
id request_id refuse repeats 7d
late 48h
period month close 24h
raw until closed + 90d
What Happened
- 1 September. Reports start arriving from both gateways, about 500 requests an hour between them, following the working day.
- Every day. About 2% of requests get no answer in time and are sent again with the same request id, sometimes through the other gateway. The meter refuses each repeat and counts it under
repeat. By the end of the month that is 7,219 reports. - 15 September, 08:00. The eu gateway loses its link to the meter. It keeps serving its customers and holds their reports.
- 15 September, 14:00. The link returns and the gateway delivers the 625 reports it held. Each request counts in the hour it happened, so the gap in the eu chart fills in. Of those reports, 29 had already reached the meter as retries through the us gateway, and are refused as repeats.
- 27 September, 08:08. Juniper Bakery Blog reaches its Free quota of 5 million tokens. The eu gateway reads the quota view before serving and turns Juniper’s requests away, 672 of them before October. It went 309 tokens over, because the gateway reads the quota once a minute.
- 1 October, 00:00. A new month. Juniper’s quota starts again. Some September requests reach the meter after midnight through a slow eu gateway, and 25 of them count in September, when they happened.
- 2 October, 00:00. September closes. From here on nothing can change its totals.
- 2 October, 05:00. A queue stuck in the us gateway since 18:00 on 30 September is emptied: 642 reports. September is closed, so the meter refuses all 642 and counts them under
closed. The invoices do not change, and the refusal count says exactly what was missed.
What the File Can Answer
September’s invoices, read from the meter’s precomputes and the price list with one view:
sqlite> SELECT name, plan, requests, printf('%.2f', due_cents / 100.0) AS due_usd,
...> printf('%.2f', cost_nano / 1e9) AS model_cost_usd
...> FROM invoices WHERE period = '2026-09' ORDER BY due_nano DESC;
name plan requests due_usd model_cost_usd
Sundial Research enterprise 150790 2591.16 1727.44
Harbor Legal Drafts pro 39586 2374.67 1187.34
Quarry Road Media pro 93120 683.48 341.74
Maple Grove Travel pro 61126 155.25 77.63
Juniper Bakery Blog free 4264 0.00 3.26
Northgate Tutors free 1453 0.00 0.44
What the meter refused, and why:
sqlite> SELECT reason, sum(n) AS reports FROM usage_refused GROUP BY reason;
reason reports
closed 642
repeat 7219
And the quotas, as each gateway reads them before serving:
sqlite> SELECT customer, CAST(used AS INTEGER) AS used, CAST(lim AS INTEGER) AS lim,
...> CAST(remaining AS INTEGER) AS remaining, reached FROM monthly_tokens WHERE period = '2026-09';
customer used lim remaining reached
harbor 137230248 300000000 162769752 0
juniper 5000309 5000000 -309 1
maple 79723462 300000000 220276538 0
northgate 1565343 5000000 3434657 0
quarry 172082979 300000000 127917021 0
Sundial is on Enterprise, which has no quota, so it has no row.
The Numbers
| Over the month | Result |
|---|---|
| Requests served | 366,118 |
| Reports sent, retries included | 373,351 |
| Retried reports refused as repeats | 7,219 |
| Reports refused after September closed | 642 |
| September invoices | $5,804.56 due, against $3,337.84 of model cost: a margin of 42.5% |
| Invoice lines and amounts against a separate recount | 12 of 12 lines and 6 of 6 amounts equal, to the billionth of a dollar |
| Hourly windows against the recount | 8,215 of 8,215 equal, request for request and token for token |
| Reading every invoice | Under 1 ms from the precomputes, against 0.36 to 0.61 seconds from every request |
| A quota check | 10 to 13 µs |
| The file | 25.3 MB at the end of the run; 1.9 MB after the dispute window, with the same invoices |
| The native Engine on the same reports | 2.3 to 3.3 seconds, 113,000 to 162,000 a second; its tables equal the demo’s, value for value, apart from the price list and quota limits: 453,536 rows, 2,949,652 values |
From a headless run of the demo’s own code with the page’s SQLite build, and from the native binary on a two-core cloud server (Intel Xeon at 2.8 GHz). The recount is separate code that takes the gateways’ reports in the order they were sent and applies the policy’s rules itself. It never reads the meter.
What the Demo Revealed
The first version of the policy refused reports more than a day late, and closed each month a day after it ended. The stuck queue’s reports were about 35 hours late when they arrived, so the late rule, which is checked first, refused them as late, and the count did not say that they belonged to a month already invoiced. The policy now accepts reports up to 48 hours late. A report for a closed month is refused as closed, which is what the people sending invoices need to know. The Engine and the triggers are tested with both settings and agree on every refusal.
Money is kept in billionths of a dollar, as integers, because floating-point cents do not add up exactly over hundreds of thousands of requests. The invoices round to the cent once, at the end.
Next: On Real Traffic
The meter runs wherever SQLite runs, so a pilot can put it next to an existing billing system: the gateways report to the meter, and the billing system reads its totals at the month’s end. The questions for a pilot are how the rules fit a real product’s plans, and how gateways with real clocks behave. An exact stream goes by the times on the events, so one gateway with a clock far ahead could close a month early.
Try It Yourself
Open the Meter demo, press Play and start a retry storm while it runs: for an hour every request is reported three times, and the invoices do not move. Then cut the eu link and watch the gap fill in when it returns. At the end, every invoice is checked against a recount, and the file is yours to query or download.