The author introduced unalloc, an open-source cost-reconciliation tool that combines data from OpenCost, LiteLLM, OpenAI, and Anthropic into one exact ledger and shows the share of spending without an owner.

In a constructed multi-pod deployment scenario using one month of synthetic OpenCost allocations, owner labels only on LeaderWorkerSet leader pods left 66% of the GPU bill without an owner. A natural fallback key assigned 61% of the bill to a Helm chart name and reduced the reported unallocated share to 4%.

On an NVIDIA H100 with vLLM, an inference server, a token meter, which allocates costs by the number of tokens, assigned a retrieval-heavy tenant a 12–14 percentage-point larger share of the bill than an equal time-share meter at every tested load. The author notes that neither of the two meters provides a ground-truth allocation.

Claim check:

  • The author introduced unalloc, an open-source tool that combines cost data from OpenCost, LiteLLM, OpenAI, and Anthropic into one exact ledger and reports the share of spending without an owner. (confirmed by the publication itself: evidence; «We present unalloc, an open-source tool that joins OpenCost, LiteLLM, OpenAI and Anthropic cost data into one exact ledger and reports the share of spend with no owner»)
  • In a constructed multi-pod deployment scenario using one month of synthetic OpenCost allocations, owner labels only on LeaderWorkerSet leader pods left 66% of the GPU bill without an owner. (confirmed by the publication itself: evidence; «in a constructed multi-pod deployment scenario – one month of synthetic OpenCost allocations, not observed billing data – owner labels set only on LeaderWorkerSet leader pods leave 66% of that deployment’s GPU bill unowned»)
  • A natural fallback key assigned 61% of the bill to a Helm chart name and reduced the reported unallocated share to 4%. (confirmed by the publication itself: evidence; «the natural fallback key assigns 61% of it to a Helm chart name while the headline unallocated share falls to 4%»)
  • On an NVIDIA H100 with vLLM, a token meter assigned a retrieval-heavy tenant a 12–14 percentage-point larger share of the bill than an equal time-share meter at every tested load. (confirmed by the publication itself: evidence; «on an NVIDIA H100 running vLLM, a token meter assigns a retrieval-heavy tenant 12-14 percentage points more of the bill than an equal time-share meter at every load tested»)
  • The author notes that neither of the two meters provides a ground-truth allocation. (confirmed by the publication itself: evidence; «Neither meter is a ground truth»)

Primary sources:

score 64.4 out of 100 · kind: research