Ingram Cloud

Documentation

Usage, cost & budgets

Usage, cost & budgets

Watch what this project consumes (tokens, runs, cost) and cap it. This page is about what each project spends; invoicing your own customers is Customers & metering, and topping up the org wallet those projects draw from is Credits & billing.

What's measured

Every run records token usage; every model call lands in a trace with a computed cost (token counts × a per-model price book, in EUR). The same figure is what the run debits from your org's credit wallet (times any platform margin), so usage you see here and credits you spend there move together.

Reading usage

Totals and a per-day series:

# Authorization: tenant-admin token (server-side only)
curl https://api.cloud.ingram.tech/v1/tenant/usage
# → { "totals": { "runs": …, "input_tokens": …, "output_tokens": …, "total_tokens": … },
#     "series": [ { "date": "2026-06-10", "runs": …, "total_tokens": … }, … ] }

Cost rollups, grouped and filterable:

# Authorization: tenant-admin token (server-side only)
curl "https://api.cloud.ingram.tech/v1/usage?group_by=smith&period=2026-06"
# group_by: app | smith | model | customer
# filters:  from / to (dates), period (YYYY-MM), smith_id, customer_id
# → { "totals": { "tokens": …, "input_tokens": …, "cache_read_tokens": …,
#                 "cache_write_tokens": …, "cost": …, "run_count": … },
#     "groups": [ { "smith": "smt_…", "tokens": …, "input_tokens": …,
#                   "cache_read_tokens": …, "cache_write_tokens": …,
#                   "cost": …, "run_count": … } ] }

tokens is the grand total; input_tokens is the input slice, and cache_read_tokens / cache_write_tokens are sub-slices of input — tokens served from the provider's prompt cache and tokens written to it. The cache hit rate is cache_read_tokens / input_tokens: keep it high by holding your system prompt and tool definitions byte-stable (see Prompt caching).

The console's Observe → Usage & cost renders the same numbers: stat cards (including the cache hit rate), budgets, the usage-over-time chart, top consumers, and a per-model Prompt caching breakdown.

BYOK model-provider keys

Smiths run on Ingram-hosted models by default. To run on your own provider account instead, add a key per project — Settings → Models, or

# Authorization: tenant-admin token (server-side only)
curl -X PUT https://api.cloud.ingram.tech/v1/tenant/model_keys/anthropic \
  -H "Authorization: Bearer $IC_TOKEN" \
  -H "IC-Api-Version: 2026-05-01" \
  -H "Content-Type: application/json" \
  -d '{ "api_key": "sk-ant-…" }'        # optional: "base_url" for proxies

Keys are write-only — never returned by reads (GET /v1/tenant/model_keys shows presence only). The model catalog (GET /v1/tenant/models) marks which models are runnable — hosted models plus any provider you've added a key for; a run that resolves to a provider with no key available — yours or ours — fails with model_key_missing.

Budgets

A budget caps monthly spend for a scope and acts when crossed. limit is the cap in your billing currency — currency on the budget names it (EUR by default):

# Authorization: tenant-admin token (server-side only)
curl -X POST https://api.cloud.ingram.tech/v1/budgets \
  -H "Authorization: Bearer $IC_TOKEN" \
  -H "IC-Api-Version: 2026-05-01" \
  -H "Content-Type: application/json" \
  -d '{ "scope": "tenant", "limit": 100, "action": "block" }'
# scope: "tenant" (whole project) | "agent" + scope_id (an agent id — caps every
#        smith of that design) | "smith" + scope_id | "customer" + scope_id
# action: "warn" | "block"

Scopes nest from broad to narrow — tenantagentsmith, plus customer across a smith's billed party. They're independent caps evaluated together: a run is blocked if any block budget that applies to it is at its limit, so the tightest one wins.

  • Crossing 80% and 100% emits a budget.threshold event (once per tier per month). Wire it to a webhook for alerts.
  • A block budget at 100% rejects new runs with 402 budget_exceeded until the month rolls over or the limit is raised. warn only notifies.
  • GET /v1/budgets/{id}/status returns spend-to-date, percent, and whether it's over; the console shows the same as cards at the top of Usage & cost.

Budget checks run at run creation against the current calendar month (UTC).

A smith can read its own caps mid-run with the read-only read_budget hosted tool — limit, spend-to-date, and remaining for each scope that applies to it — so it can pace its work and tell the user when it's near a cap. The cap is still enforced server-side; reading it never changes a limit.

The tenant budget is the project's funding cap

A tenant-scope budget is special: it bounds how much this project may draw from the org's shared credit wallet in the month. Without one, a project draws freely up to the org's overall balance; with a block one, it stops at its own limit even while the wallet still has funds. Two independent gates, both at run creation: the project budget (402 budget_exceeded) and the empty wallet (402 insufficient_credits).