When the Token Meter Disappears

August 1st, 2026

As open models become capable and token prices approach zero, model choice may move out of sight and into the agent harness.


I keep looking at the numbers from Kimi K3 and DeepSeek V4, and I am not sure the most interesting story is which model won which benchmark.

Kimi K3 is a 2.8 trillion parameter open weight model with a million token context window and benchmark results that put it in the frontier conversation. DeepSeek V4 Flash has the same context length, yet one million cached input tokens now costs $0.0028 and one million output tokens costs $0.28.

I am old enough to remember when going online made noise, took time, and left you vaguely aware that somewhere a meter was running. Flat rate internet did more than make the bill predictable. It changed our behavior because the connection stopped being the decision and simply became part of the room.

I wonder if model intelligence is approaching the same transition.

When a system can ask three models, compare the work, discard two answers, and try again without anyone caring about the bill, why should a person choose a model at all? At some point selecting a model may feel as strange as selecting the network route that carried this page to your screen.

Someone still has to choose, of course, but perhaps that someone is the harness. The harness decides what context to provide, which tools to allow, how much reasoning to request, when to try another model, and whether the answer is actually good enough. Even the benchmark tables hint at this because the scores move with the reasoning settings, tools, and agent harness wrapped around the model.

Maybe the model is becoming a component while the harness becomes the product.

Are we still arguing about which model is best because model choice matters, or because our harnesses have not yet learned to choose for us?

Token cost may go to zero. Judgment will not.