AI workload evaluation and improvement

State of Inference

Top 5 in the pilot

ModelProviderUSD / attemptTurns / attemptError rateScore
GLM 5.3 FlashBaseten$0.00459.00.0% (0/1)90.8
DeepSeek V4.1 FlashBaseten$0.01028.00.0% (0/1)88.2
GLM 5.3Baseten$0.02868.00.0% (0/1)77.3
DeepSeek V4 Pro 0813Baseten$0.05268.00.0% (0/1)67.1
Kimi K3Baseten$0.11977.00.0% (0/1)52.2

Baseten pilot, measured 2026-09-14. Provisional v2 score. Inference USD estimated from reported usage and captured provider rates. Means include failed attempts. Error rate shows failed attempts over total attempts. Zero observed errors does not establish reliability. One provider and one synthetic task family, not a general ranking.

14 September validation: 144 completed attempts, one provider failure with unknown cost and five unstarted slots. The larger batch is paused. The table retains the earlier five-attempt pilot.

In practice

86% lower inference cost

A production agent helped users check financial data. We tested alternative models on its actual tasks to find a less expensive configuration.

Inference cost per turn
  • Previous configuration100%

    Claude Sonnet 5

  • Selected configuration14%

    GLM-5.3-Flash, max effort

Costs calculated from recorded token use with cached input. One workload, 9 September 2026.

We built 22 tests from real user tasks, with automated checks for answer accuracy and tool use. The selected configuration cost less across every case.

How we work

Kylace helps teams improve AI workloads that cost too much, fall short or need to scale. You define what better means: cost, quality, speed or reliability. We bring specialist knowledge and take on the practical work of testing alternatives.

You receive a reusable evaluation and a recommendation backed by results, with tested changes where we find a worthwhile improvement. We build or adapt the evaluation around your real cases and requirements, then prepare the models, providers and test environments. We investigate changes to the model and the wider system, including prompts, tools and how the work is organized.

We maintain the evaluation as your workload changes and investigate relevant developments in models and technology. When we find an improvement that meets your priorities, we explain the benefit, the tradeoffs and whether it warrants a change. Implementation and deployment are agreed with your team.

Contact

If you're responsible for AI systems in production, I'd like to hear where your team needs better results, lower costs or room to scale. Tell me about one workload, and we can discuss what needs to improve and whether Kylace is a good fit. – Alf Viktor Williamsen

Research