Choosing an engine.
yarnnn runs your work on engines from several providers, and lets you pick per conversation. Here is how to think about the choice — and where to find the current numbers, which we don't keep ourselves.
Capability
How well an engine handles reasoning, long context, and nuance. Matters most on judgment work; matters least on routine extraction.
Cost
What a provider charges per token, in and out. The spread between the cheapest and most capable engines is large — often more than an order of magnitude.
Speed
How fast the first and last token arrive. A faster engine can be worth more than a smarter one on work you are waiting on.
Why we don't rank them here
Any ranking we published would be out of date within weeks, and it would go out of date quietly — a stale comparison table looks exactly like a current one. Providers ship new models and change prices on their own schedule. So we point you at the people who track this properly, and keep this page to the part that doesn't move.
The part that doesn't move
More capable engines cost more per token. That much is stable. What surprises people is that the cheapest engine is not always the cheapest outcome:
- A weaker engine can cost more. If it needs three attempts, or produces work you rewrite by hand, the cheaper per-token rate buys you a more expensive result. Difficulty is what should pick the engine, not the price list.
- Length drives cost more than choice of engine. A long conversation on a cheap engine can outspend a short one on an expensive engine. What you send matters as much as who you send it to.
- Routine work rarely needs the top engine. Extracting fields, reformatting, summarising something short — the gap between engines narrows as the task gets more mechanical, while the price gap stays wide.
Where the current numbers live
This is independent of us. We don't reproduce their figures, because a copy is a snapshot and theirs are maintained.
For rates straight from the source, each provider publishes its own: Anthropic, OpenAI, Google, DeepSeek and xAI.
Your own workspace is the better benchmark
Benchmarks tell you how engines compare in general. They can't tell you what your work costs, which depends on how you write, how long your conversations run, and what you ask for. Your Usage screen breaks spending down by engine over your own history — for deciding what to use tomorrow, that beats any leaderboard.
You can change engine per conversation, so the cost of guessing wrong is one conversation. See pricing for how usage draws your balance.