You run a pipeline that sends every support transcript to a model and asks for structured output back: a sentiment label, a few entities, a score. It works. It has worked for months. The bill has crept up a little every month, and you have put that down to volume. Then a post on APIs.io catches your eye, about a measurement Vonage published called The Hidden Token Tax on JSON Schemas, and you read it with this month's invoice open in another tab.
The mechanism is so simple it is almost embarrassing. The schema you send to get structured output is part of the input, and it goes with every single request. Vonage's team measured it properly: five schemas, from one field at 88 characters to a nested one at 1,730, sent with the same short prompt to Gemini models on Vertex AI and to Claude Haiku 4.5 on Bedrock, at temperature zero, with token counts taken from each provider's own usage API. You trust the numbers because they tell you how they got them.
The number that stops you is this one. With a medium schema and a 100-token prompt, Claude Haiku 4.5 counted 415 input tokens against 121 without the schema. An overhead of 71 percent. Gemini 2.5 Flash added only 24 tokens for the same schema, so it depends heavily on the model. Over a million calls the bill ran from $37.20 to $396.00. Your transcripts are short. You do the arithmetic and do not enjoy it.
There is relief in the next line, though. The overhead is fixed per call, so it fades as the payload grows: 37 percent at 500 input tokens, 3 percent at 10,000. If your prompts are long, this matters less. If they are short and frequent, like yours, it matters a lot.
The advice is concrete, and you start a list as you read: strip fields nobody reads, flatten the nesting, take descriptions and examples out of the schema and put them in the prompt, skip the schema entirely when you only need one label, batch where you can. Vonage reports that cutting one schema from 1,730 characters to 555 saved roughly 650 tokens a call on Claude. You can think of two schemas in your own code that look a lot like that 1,730-character one.
Then the APIs.io post turns, and it is the part you keep thinking about afterward. The schema a model reads at inference time and the schema an API publishes in its contract are two different documents, and their economics run in opposite directions. In the structured-output call, every description and example is paid for on every request, so you trim them. In a published OpenAPI contract, those same descriptions and examples are how a developer, or an agent, understands the API before calling it, and they cost nothing per call. Vonage's advice is right for the first document and would be wrong for the second.
The post then reads Vonage's own record. A Kin Score of 62.2, in the strong band, lifted by developer ergonomics and discoverability, held back by contract governance at 31.8. Agent Readiness of 24.8, agent-aware. And in its public contract, the catalog cannot find OpenAPI examples or error semantics. The company that just showed where examples cost money has not put them where they are free.
You do not read that as a gotcha. You read it as a warning about your own habits, because you can feel the temptation already: having learned to trim, you will want to trim everywhere. You will want to tidy the internal OpenAPI your other teams read, too, and take out the long descriptions and the example payloads because they look like waste now.
So your list gets a second column. In the prompt and the inference-time schema: cut hard, measure, cut again. In the contract: leave every description and example where it is, and add the ones that are missing. The tax is real. It just belongs to one document and not the other.
Read the original: https://apis.io/2026/10/06/vonage-measures-the-token-tax-on-json-schemas-and-the-advice-cuts-against-the-contract/