State of Inference
Pilot01 Inference cost
USD per verified completionFinancial close / Baseten / matched configurations
One observation. No trend yet.
- GLM 5.3 Flash$0.0045
- DeepSeek V4.1 Flash$0.0102
- GLM 5.3$0.0286
- DeepSeek V4 Pro 0813$0.0526
- Kimi K3$0.1197
Includes inference spend on failed attempts. Hosting and tools excluded. Incompatible runtimes start separate series.
02 Workhorse benchmark
PilotZero observed errors is not proof of reliability. Open a model for its report and sources.
Inference USD estimated from usage and captured rates.
03 Monitor
ExaSource discovery / unreviewed
Checking connection
Watchlist
- Model releases
Open weights, tool use and new endpoints
- Inference economics
Pricing, caching and serving changes
- Practical evaluation
Agent reliability, fx and reproducibility
Source discovery, not benchmark evidence.
04 Changelog / releases
Our benchmark and terminal14 September validation: 144 completed attempts, one provider failure with unknown cost and five unstarted slots. The larger batch is paused. The table retains the earlier five-attempt pilot.
- Rubric v2
Cost first. Errors penalized.
5:1 cost-to-turn weighting and a squared error penalty. Existing pilot rescored without new calls.
- Baseten pilot
Five configurations tested
One matched financial-close task per model. JSON, citations and PDF artifacts verified.
- Financial close v1
The workload is the benchmark
Database retrieval, five calculations, source citations and a machine-verifiable report, running on fx.
The workload is the benchmark
A controlled financial-close task. Four tools. Five calculations. One independently verified result.
14 September validation: 144 completed attempts, one provider failure with unknown cost and five unstarted slots. The larger batch is paused. The table retains the earlier five-attempt pilot.
The agent retrieves financial records from a database, calculates five month-end metrics, cites the JSON source records and generates a PDF. A programmatic verifier checks the complete result. Each attempt passes or fails.
A turn is one model generation, including the final reply. Tool calls within the same generation share a turn. Inference cost is measured in USD using reported usage and the provider rates captured with the run. Local database, PDF and hosting costs are excluded.
Error rate is failed attempts divided by all valid attempts. Wrong calculations, invalid citations, incomplete reports and task deadlines are errors. Invalid host-interrupted measurements are archived separately and block release rather than being attributed to the model.
The provisional v2 score is 100 times efficiency times (1 minus error rate) squared. Efficiency is (5 times cost efficiency plus turn efficiency) divided by 6. At equal resource use, a 10% error rate retains 81% of the score, and a 50% error rate retains 25%.
Cost efficiency is 0.10 divided by (0.10 plus mean inference USD per attempt). Turn efficiency is the smaller of 1 and six divided by mean turns per attempt. Means include failures. The scales and penalty strength are design choices, not calibrated production thresholds. Scores run from zero to 100.
Each release preserves its measurement date, scoring date, frozen configuration and error-rate uncertainty interval. Earlier measurements and rubric exports remain archived. New runtime policies start separate comparison series.
This is a dated pilot, not a live market feed or a general model ranking. Each row belongs to a model and provider configuration. A single provider and one synthetic task family cannot establish performance across providers or real workloads.