Kylace

State of Inference

Pilot
Financial close v1/ Pilot / 1 provider

01 Inference cost

USD per verified completion

Financial close / Baseten / matched configurations

0.0000.0300.0600.0900.120GLM 5.3 Flash: $0.004513DeepSeek V4.1 Flash: $0.010156GLM 5.3: $0.028564DeepSeek V4 Pro 0813: $0.052591Kimi K3: $0.1197262026-09-14History starts hereOne observation. No trend yet.

One observation. No trend yet.

  • GLM 5.3 Flash$0.0045
  • DeepSeek V4.1 Flash$0.0102
  • GLM 5.3$0.0286
  • DeepSeek V4 Pro 0813$0.0526
  • Kimi K3$0.1197

Includes inference spend on failed attempts. Hosting and tools excluded. Incompatible runtimes start separate series.

02 Workhorse benchmark

Pilot

Financial close / Baseten / rubric v2

State of Inference
ModelProviderUSD / attemptTurns / attemptError rateScore
GLM 5.3 FlashBaseten$0.00459.00.0% (0/1)90.8
DeepSeek V4.1 FlashBaseten$0.01028.00.0% (0/1)88.2
GLM 5.3Baseten$0.02868.00.0% (0/1)77.3
DeepSeek V4 Pro 0813Baseten$0.05268.00.0% (0/1)67.1
Kimi K3Baseten$0.11977.00.0% (0/1)52.2

Baseten pilot, measured 2026-09-14. Provisional v2 score. Inference USD estimated from reported usage and captured provider rates. Means include failed attempts. Error rate shows failed attempts over total attempts. Zero observed errors does not establish reliability. One provider and one synthetic task family, not a general ranking.

Zero observed errors is not proof of reliability. Open a model for its report and sources.

Inference USD estimated from usage and captured rates.

score = 100 * ((5c + t) / 6) * (1 - error rate)^2

03 Monitor

Exa

Source discovery / unreviewed

Checking connection

Watchlist
  • Model releases

    Open weights, tool use and new endpoints

  • Inference economics

    Pricing, caching and serving changes

  • Practical evaluation

    Agent reliability, fx and reproducibility

Source discovery, not benchmark evidence.

04 Changelog / releases

Our benchmark and terminal

14 September validation: 144 completed attempts, one provider failure with unknown cost and five unstarted slots. The larger batch is paused. The table retains the earlier five-attempt pilot.

  1. Rubric v2

    Cost first. Errors penalized.

    5:1 cost-to-turn weighting and a squared error penalty. Existing pilot rescored without new calls.

  2. Baseten pilot

    Five configurations tested

    One matched financial-close task per model. JSON, citations and PDF artifacts verified.

  3. Financial close v1

    The workload is the benchmark

    Database retrieval, five calculations, source citations and a machine-verifiable report, running on fx.

Export observations >
Method & evidence

The workload is the benchmark

A controlled financial-close task. Four tools. Five calculations. One independently verified result.

14 September validation: 144 completed attempts, one provider failure with unknown cost and five unstarted slots. The larger batch is paused. The table retains the earlier five-attempt pilot.

The agent retrieves financial records from a database, calculates five month-end metrics, cites the JSON source records and generates a PDF. A programmatic verifier checks the complete result. Each attempt passes or fails.

A turn is one model generation, including the final reply. Tool calls within the same generation share a turn. Inference cost is measured in USD using reported usage and the provider rates captured with the run. Local database, PDF and hosting costs are excluded.

Error rate is failed attempts divided by all valid attempts. Wrong calculations, invalid citations, incomplete reports and task deadlines are errors. Invalid host-interrupted measurements are archived separately and block release rather than being attributed to the model.

The provisional v2 score is 100 times efficiency times (1 minus error rate) squared. Efficiency is (5 times cost efficiency plus turn efficiency) divided by 6. At equal resource use, a 10% error rate retains 81% of the score, and a 50% error rate retains 25%.

Cost efficiency is 0.10 divided by (0.10 plus mean inference USD per attempt). Turn efficiency is the smaller of 1 and six divided by mean turns per attempt. Means include failures. The scales and penalty strength are design choices, not calibrated production thresholds. Scores run from zero to 100.

Each release preserves its measurement date, scoring date, frozen configuration and error-rate uncertainty interval. Earlier measurements and rubric exports remain archived. New runtime policies start separate comparison series.

This is a dated pilot, not a live market feed or a general model ranking. Each row belongs to a model and provider configuration. A single provider and one synthetic task family cannot establish performance across providers or real workloads.