Skip to content

Day 39 — Cost Receipts and Tests

Today Rankforge becomes measurable. “The run cost about a few cents” is not enough when you need to find the expensive stage, compare two prompts, or prove which source and intermediate artifact reached the final package.

Each scripted or live response returns its own usage:

pub struct StepUsage {
pub input_tokens: u64,
pub output_tokens: u64,
pub reasoning_tokens: u64,
pub cached_tokens: u64,
pub web_search_calls: u64,
}

Provider usage is an external observation. Record the provider, model release, pricing-table version, and whether a field was provider-reported or locally estimated when you add a live adapter.

Rankforge stores cost as micro-US dollars. Pricing is also stored as integer microdollars per million tokens. The calculation uses checked_mul and checked_add, then rounds once.

This avoids binary floating-point values in durable accounting and prevents a huge or malicious usage field from silently wrapping. The displayed decimal is presentation only.

Price tables change. The values copied from the experiment are an input to the estimate, not eternal truth. Version or configure them before using the report for billing.

Every receipt contains:

pub struct StageReceipt {
pub stage: Stage,
pub artifact: String,
pub input_sha256: String,
pub output_sha256: String,
pub usage: StepUsage,
pub estimated_cost_microusd: u64,
}

The research input hash starts with the website snapshot. Each later receipt uses the preceding artifact hash as its input identity. This baseline proves sequence identity. A production request should hash a manifest containing every actual input, including prompt version, source snapshot, accepted artifacts, model release, tool observations, and policy.

Terminal window
cargo test -p rankforge
cargo run -p rankforge -- build \
--source fixtures/rankforge/site.md \
--responses fixtures/rankforge/responses.json \
--output target/rankforge

The integration test launches CARGO_BIN_EXE_rankforge and asserts that all five stage artifacts, blog-final.md, blog-cost.json, and run-report.json exist. It also checks stage and web-search totals. Unit tests break stage order and prove a known integer cost calculation.

Add failure tests for empty source, malformed JSON, missing sections, forbidden punctuation, arithmetic overflow, unwritable output, and a package that does not preserve article coverage.

You are finished when you can:

  • explain why live scraping is optional in the required test path;
  • identify the artifact accepted by every stage;
  • make a skipped angle fail before generation artifacts are written;
  • distinguish section validation from factual verification;
  • trace the final package back through hashes and usage receipts;
  • compare quality, latency, and cost for two pipeline versions.

Finish with the Project 8 Revision →.