Switching provider
Migration questions for people arriving from another provider, including the cases where staying is the better call.
74 questions
vs OpenRouter
In context →- Should I use AI Token Router or OpenRouter?
- Across 8 dimensions compared below, AI Token Router comes out ahead on 5 and OpenRouter on 3 -- so the honest answer is that it depends on which dimensions matter to your workload, and both sections naming a winner are on this page rather than only the flattering one.
- When is AI Token Router the better choice?
- If you are specifically working with open-weight models and want the lowest total cost with fully published pricing — no platform fee on top, cached input visible before you commit, and video models treated as a real category rather than something we do not carry — that is exactly what we built this for.
- How does AI Token Router pricing compare to OpenRouter?
- On platform fee: AI Token Router is None. OpenRouter is ~5.5% on credit purchases (non-crypto).
- How was this AI Token Router vs OpenRouter comparison made?
- Where the OpenRouter figures come from: Every OpenRouter cell in the table is taken from OpenRouter's own published material -- their pricing page, their documentation or their rate card -- not from an aggregator, a press mention or our own testing. Where they publish no number for a dimension, the row says so rather than filling the gap with an estimate: no dollar figure appears in their column anywhere on this page. Where our own figures come from: Our column is the catalogue that prices real requests, including the parts of it that do not flatter us: 7 of the 24 catalogued models are callable today, and the rows say so rather than counting the catalogue as the lineup. The rates behind it were last reconciled with upstream on Sep 2, 2026 and are published in full on the pricing page. When this was last checked: September 2026. Competitor pricing and product scope both move without notice, so every figure here is a statement about that date rather than a permanent one. Anything you are about to make a decision on is worth confirming against OpenRouter's own site before you make it. What is deliberately not compared: There are no benchmark scores, no latency measurements and no tokens-per-second figures in our column -- not on this page and nowhere else on this site -- because we have not measured them. Where a speed claim appears in the OpenRouter column it is theirs, repeated as they publish it and not verified by us. An unmeasured quality or speed number would be the easiest thing on this page to invent, which is precisely why there is not one. How the winner column is decided: Each of the 8 rows is marked for one side: 5 to AI Token Router, 3 to OpenRouter. That marking is our judgement and you are free to disagree with it. The figures underneath it are the part you can check without taking our word for anything. Corrections: If a figure here is stale or wrong, tell us and we will change it -- including when the correction runs in OpenRouter's favour, and including when it costs us a row we currently win. A comparison only ever corrected in its own favour is not being corrected.
- What does OpenRouter do better?
- OpenRouter wins on model catalog size, closed frontier models (gpt, claude, gemini), provider redundancy. Those rows are on the comparison table with OpenRouter marked as the winner, because a comparison that never concedes anything is not a comparison.
vs DeepInfra
In context →- Should I use AI Token Router or DeepInfra?
- Across 6 dimensions compared below, AI Token Router comes out ahead on 4 and DeepInfra on 2 -- so the honest answer is that it depends on which dimensions matter to your workload, and both sections naming a winner are on this page rather than only the flattering one.
- When is DeepInfra the better choice?
- If your workload is text-only, high-volume, and you are optimising purely for the lowest per-token rate, check DeepInfra's number against ours model by model before switching — on several popular models they are cheaper, and we would rather you find that out here than feel misled later. They run their own GPU fleet at real scale and it shows in the pricing.
- How does AI Token Router pricing compare to DeepInfra?
- On headline text-model pricing: AI Token Router is Competitive; cheaper on some models, dearer on others. DeepInfra is Frequently the lowest published rate on open text models.
- How was this AI Token Router vs DeepInfra comparison made?
- Where the DeepInfra figures come from: Every DeepInfra cell in the table is taken from DeepInfra's own published material -- their pricing page, their documentation or their rate card -- not from an aggregator, a press mention or our own testing. Where they publish no number for a dimension, the row says so rather than filling the gap with an estimate: no dollar figure appears in their column anywhere on this page. Where our own figures come from: Our column is the catalogue that prices real requests, including the parts of it that do not flatter us: 7 of the 24 catalogued models are callable today, and the rows say so rather than counting the catalogue as the lineup. The rates behind it were last reconciled with upstream on Sep 2, 2026 and are published in full on the pricing page. When this was last checked: September 2026. Competitor pricing and product scope both move without notice, so every figure here is a statement about that date rather than a permanent one. Anything you are about to make a decision on is worth confirming against DeepInfra's own site before you make it. What is deliberately not compared: There are no benchmark scores, no latency measurements and no tokens-per-second figures in our column -- not on this page and nowhere else on this site -- because we have not measured them. Where a speed claim appears in the DeepInfra column it is theirs, repeated as they publish it and not verified by us. An unmeasured quality or speed number would be the easiest thing on this page to invent, which is precisely why there is not one. How the winner column is decided: Each of the 6 rows is marked for one side: 4 to AI Token Router, 2 to DeepInfra. That marking is our judgement and you are free to disagree with it. The figures underneath it are the part you can check without taking our word for anything. Corrections: If a figure here is stale or wrong, tell us and we will change it -- including when the correction runs in DeepInfra's favour, and including when it costs us a row we currently win. A comparison only ever corrected in its own favour is not being corrected.
- What does DeepInfra do better?
- DeepInfra wins on headline text-model pricing, model catalog size. Those rows are on the comparison table with DeepInfra marked as the winner, because a comparison that never concedes anything is not a comparison.
vs Together AI
In context →- Should I use AI Token Router or Together AI?
- Across 6 dimensions compared below, AI Token Router comes out ahead on 3 and Together AI on 3 -- so the honest answer is that it depends on which dimensions matter to your workload, and both sections naming a winner are on this page rather than only the flattering one.
- When is Together AI the better choice?
- If you need to fine-tune a model and serve it from the same platform, or you want dedicated GPU capacity rather than shared inference, Together AI does something we simply do not do. They are also the safer answer if your procurement process weights vendor track record heavily — we are newer, and pretending otherwise would not survive your first reference call.
- How does AI Token Router pricing compare to Together AI?
- On text-model pricing: AI Token Router is 36–42% below official on every model we carry, 65–68% with a cached prefix. Together AI is Competitive, above DeepInfra on most models.
- How was this AI Token Router vs Together AI comparison made?
- Where the Together AI figures come from: Every Together AI cell in the table is taken from Together AI's own published material -- their pricing page, their documentation or their rate card -- not from an aggregator, a press mention or our own testing. Where they publish no number for a dimension, the row says so rather than filling the gap with an estimate: no dollar figure appears in their column anywhere on this page. Where our own figures come from: Our column is the catalogue that prices real requests, including the parts of it that do not flatter us: 7 of the 24 catalogued models are callable today, and the rows say so rather than counting the catalogue as the lineup. The rates behind it were last reconciled with upstream on Sep 2, 2026 and are published in full on the pricing page. When this was last checked: September 2026. Competitor pricing and product scope both move without notice, so every figure here is a statement about that date rather than a permanent one. Anything you are about to make a decision on is worth confirming against Together AI's own site before you make it. What is deliberately not compared: There are no benchmark scores, no latency measurements and no tokens-per-second figures in our column -- not on this page and nowhere else on this site -- because we have not measured them. Where a speed claim appears in the Together AI column it is theirs, repeated as they publish it and not verified by us. An unmeasured quality or speed number would be the easiest thing on this page to invent, which is precisely why there is not one. How the winner column is decided: Each of the 6 rows is marked for one side: 3 to AI Token Router, 3 to Together AI. That marking is our judgement and you are free to disagree with it. The figures underneath it are the part you can check without taking our word for anything. Corrections: If a figure here is stale or wrong, tell us and we will change it -- including when the correction runs in Together AI's favour, and including when it costs us a row we currently win. A comparison only ever corrected in its own favour is not being corrected.
- What does Together AI do better?
- Together AI wins on fine-tuning, dedicated gpu clusters, enterprise track record. Those rows are on the comparison table with Together AI marked as the winner, because a comparison that never concedes anything is not a comparison.
vs Fireworks AI
In context →- Should I use AI Token Router or Fireworks AI?
- Across 9 dimensions compared below, AI Token Router comes out ahead on 4 and Fireworks AI on 4, and 1 are a genuine tie -- so the honest answer is that it depends on which dimensions matter to your workload, and both sections naming a winner are on this page rather than only the flattering one.
- When is Fireworks AI the better choice?
- If tokens-per-second is the number your product lives or dies on — voice agents, interactive coding, anything where the user is watching the response render — Fireworks is built for exactly that, and they let you buy more of it explicitly through Standard, Priority and Fast tiers rather than making you hope. They also carry more Kimi SKUs than anyone else we looked at, including a US-hosted variant that matters if data residency is on your checklist, and they will fine-tune on open weights and serve the result. If you are already on Azure, they are available there and we are not. None of that is something we can match today.
- How does AI Token Router pricing compare to Fireworks AI?
- On platform fee: AI Token Router is None. No minimum spend, no subscription. Fireworks AI is None published; $1 in starter credits for new accounts.
- How was this AI Token Router vs Fireworks AI comparison made?
- Where the Fireworks AI figures come from: Every Fireworks AI cell in the table is taken from Fireworks AI's own published material -- their pricing page, their documentation or their rate card -- not from an aggregator, a press mention or our own testing. Where they publish no number for a dimension, the row says so rather than filling the gap with an estimate: 1 of the 9 rows carry a dollar figure they publish, and the rest turn on what the product does or does not do. Where our own figures come from: Our column is the catalogue that prices real requests, including the parts of it that do not flatter us: 7 of the 24 catalogued models are callable today, and the rows say so rather than counting the catalogue as the lineup. The rates behind it were last reconciled with upstream on Sep 2, 2026 and are published in full on the pricing page. When this was last checked: September 2026. Competitor pricing and product scope both move without notice, so every figure here is a statement about that date rather than a permanent one. Anything you are about to make a decision on is worth confirming against Fireworks AI's own site before you make it. What is deliberately not compared: There are no benchmark scores, no latency measurements and no tokens-per-second figures in our column -- not on this page and nowhere else on this site -- because we have not measured them. Where a speed claim appears in the Fireworks AI column it is theirs, repeated as they publish it and not verified by us. An unmeasured quality or speed number would be the easiest thing on this page to invent, which is precisely why there is not one. How the winner column is decided: Each of the 9 rows is marked for one side: 4 to AI Token Router, 4 to Fireworks AI, and 1 a genuine tie. That marking is our judgement and you are free to disagree with it. The figures underneath it are the part you can check without taking our word for anything. Corrections: If a figure here is stale or wrong, tell us and we will change it -- including when the correction runs in Fireworks AI's favour, and including when it costs us a row we currently win. A comparison only ever corrected in its own favour is not being corrected.
- What does Fireworks AI do better?
- Fireworks AI wins on model catalog size, latency on large open models, fine-tuning and training, enterprise distribution. Those rows are on the comparison table with Fireworks AI marked as the winner, because a comparison that never concedes anything is not a comparison.
vs Novita AI
In context →- Should I use AI Token Router or Novita AI?
- Across 8 dimensions compared below, AI Token Router comes out ahead on 4 and Novita AI on 3, and 1 are a genuine tie -- so the honest answer is that it depends on which dimensions matter to your workload, and both sections naming a winner are on this page rather than only the flattering one.
- When is Novita AI the better choice?
- Novita is the closest thing on this list to what we are building, and on several axes they are simply further along: 200+ models against our 24, real per-second video billing where we still round every clip to five seconds, and small-model pricing at the market floor. If you generate a lot of short video clips, their billing is straightforwardly better for you than ours is right now. They also give you somewhere to go when serverless stops being enough — sandboxes, GPU instances, bare metal — and we have no answer to that at all.
- How does AI Token Router pricing compare to Novita AI?
- On headline text-model pricing: AI Token Router is 36–43% below each model's official rate, published alongside it. Novita AI is At or near the market floor on small models — Llama 3.1 8B at $0.02/$0.05 per 1M is very hard to beat.
- How was this AI Token Router vs Novita AI comparison made?
- Where the Novita AI figures come from: Every Novita AI cell in the table is taken from Novita AI's own published material -- their pricing page, their documentation or their rate card -- not from an aggregator, a press mention or our own testing. Where they publish no number for a dimension, the row says so rather than filling the gap with an estimate: 3 of the 8 rows carry a dollar figure they publish, and the rest turn on what the product does or does not do. Where our own figures come from: Our column is the catalogue that prices real requests, including the parts of it that do not flatter us: 7 of the 24 catalogued models are callable today, and the rows say so rather than counting the catalogue as the lineup. The rates behind it were last reconciled with upstream on Sep 2, 2026 and are published in full on the pricing page. When this was last checked: September 2026. Competitor pricing and product scope both move without notice, so every figure here is a statement about that date rather than a permanent one. Anything you are about to make a decision on is worth confirming against Novita AI's own site before you make it. What is deliberately not compared: There are no benchmark scores, no latency measurements and no tokens-per-second figures in our column -- not on this page and nowhere else on this site -- because we have not measured them. Where a speed claim appears in the Novita AI column it is theirs, repeated as they publish it and not verified by us. An unmeasured quality or speed number would be the easiest thing on this page to invent, which is precisely why there is not one. How the winner column is decided: Each of the 8 rows is marked for one side: 4 to AI Token Router, 3 to Novita AI, and 1 a genuine tie. That marking is our judgement and you are free to disagree with it. The figures underneath it are the part you can check without taking our word for anything. Corrections: If a figure here is stale or wrong, tell us and we will change it -- including when the correction runs in Novita AI's favour, and including when it costs us a row we currently win. A comparison only ever corrected in its own favour is not being corrected.
- What does Novita AI do better?
- Novita AI wins on model catalog size, headline text-model pricing, video billing granularity. Those rows are on the comparison table with Novita AI marked as the winner, because a comparison that never concedes anything is not a comparison.
vs Replicate
In context →- Should I use AI Token Router or Replicate?
- Across 9 dimensions compared below, AI Token Router comes out ahead on 4 and Replicate on 4, and 1 are a genuine tie -- so the honest answer is that it depends on which dimensions matter to your workload, and both sections naming a winner are on this page rather than only the flattering one.
- When is Replicate the better choice?
- If the model you want is obscure, brand new, or yours, Replicate is the right answer and we are not. Their community push model means a checkpoint that landed on GitHub last week is probably already runnable, their image and video long tail is far past anything we carry, and Cog lets you deploy your own weights behind the same API. The developer experience is genuinely excellent. For exploration, prototyping and creative tooling, that breadth is worth more than a cheaper per-token rate on a model you were not going to use.
- How does AI Token Router pricing compare to Replicate?
- On cost predictability: AI Token Router is A published rate per model; you can price a request before you send it. Replicate is Many models bill by GPU-seconds of runtime ($0.000025–$0.0112/sec), so a cold start or a slow prompt costs more.
- How was this AI Token Router vs Replicate comparison made?
- Where the Replicate figures come from: Every Replicate cell in the table is taken from Replicate's own published material -- their pricing page, their documentation or their rate card -- not from an aggregator, a press mention or our own testing. Where they publish no number for a dimension, the row says so rather than filling the gap with an estimate: 3 of the 9 rows carry a dollar figure they publish, and the rest turn on what the product does or does not do. Where our own figures come from: Our column is the catalogue that prices real requests, including the parts of it that do not flatter us: 7 of the 24 catalogued models are callable today, and the rows say so rather than counting the catalogue as the lineup. The rates behind it were last reconciled with upstream on Sep 2, 2026 and are published in full on the pricing page. When this was last checked: September 2026. Competitor pricing and product scope both move without notice, so every figure here is a statement about that date rather than a permanent one. Anything you are about to make a decision on is worth confirming against Replicate's own site before you make it. What is deliberately not compared: There are no benchmark scores, no latency measurements and no tokens-per-second figures in our column -- not on this page and nowhere else on this site -- because we have not measured them. Where a speed claim appears in the Replicate column it is theirs, repeated as they publish it and not verified by us. An unmeasured quality or speed number would be the easiest thing on this page to invent, which is precisely why there is not one. How the winner column is decided: Each of the 9 rows is marked for one side: 4 to AI Token Router, 4 to Replicate, and 1 a genuine tie. That marking is our judgement and you are free to disagree with it. The figures underneath it are the part you can check without taking our word for anything. Corrections: If a figure here is stale or wrong, tell us and we will change it -- including when the correction runs in Replicate's favour, and including when it costs us a row we currently win. A comparison only ever corrected in its own favour is not being corrected.
- What does Replicate do better?
- Replicate wins on model catalog size, bringing your own weights, how fast a brand-new open model appears, video billing unit. Those rows are on the comparison table with Replicate marked as the winner, because a comparison that never concedes anything is not a comparison.
vs Groq
In context →- Should I use AI Token Router or Groq?
- Across 9 dimensions compared below, AI Token Router comes out ahead on 4 and Groq on 4, and 1 are a genuine tie -- so the honest answer is that it depends on which dimensions matter to your workload, and both sections naming a winner are on this page rather than only the flattering one.
- When is Groq the better choice?
- If latency is the product — a voice agent, a live coding assistant, anything where a human is waiting on the first token — Groq wins and it is not close. Their speed comes from custom silicon rather than a scheduling trick, which means it holds up under load in a way GPU-based competitors cannot match by trying harder. They are also well capitalised and building serious capacity, and Compound gives you web search and code execution without wiring up tools yourself. If your workload is one or two big text models and you need them fast, go there.
- How does AI Token Router pricing compare to Groq?
- On pricing transparency: AI Token Router is Every model's rate published, with the official rate beside it — 36–43% below official. Groq is No rate card on groq.com. Rates live in console docs for some models, and flagship Llama models are marked "Enterprise pricing" with no public number at all.
- How was this AI Token Router vs Groq comparison made?
- Where the Groq figures come from: Every Groq cell in the table is taken from Groq's own published material -- their pricing page, their documentation or their rate card -- not from an aggregator, a press mention or our own testing. Where they publish no number for a dimension, the row says so rather than filling the gap with an estimate: no dollar figure appears in their column anywhere on this page. Where our own figures come from: Our column is the catalogue that prices real requests, including the parts of it that do not flatter us: 7 of the 24 catalogued models are callable today, and the rows say so rather than counting the catalogue as the lineup. The rates behind it were last reconciled with upstream on Sep 2, 2026 and are published in full on the pricing page. When this was last checked: September 2026. Competitor pricing and product scope both move without notice, so every figure here is a statement about that date rather than a permanent one. Anything you are about to make a decision on is worth confirming against Groq's own site before you make it. What is deliberately not compared: There are no benchmark scores, no latency measurements and no tokens-per-second figures in our column -- not on this page and nowhere else on this site -- because we have not measured them. Where a speed claim appears in the Groq column it is theirs, repeated as they publish it and not verified by us. An unmeasured quality or speed number would be the easiest thing on this page to invent, which is precisely why there is not one. How the winner column is decided: Each of the 9 rows is marked for one side: 4 to AI Token Router, 4 to Groq, and 1 a genuine tie. That marking is our judgement and you are free to disagree with it. The figures underneath it are the part you can check without taking our word for anything. Corrections: If a figure here is stale or wrong, tell us and we will change it -- including when the correction runs in Groq's favour, and including when it costs us a row we currently win. A comparison only ever corrected in its own favour is not being corrected.
- What does Groq do better?
- Groq wins on raw throughput, hardware, built-in agentic tooling, speech-to-text. Those rows are on the comparison table with Groq marked as the winner, because a comparison that never concedes anything is not a comparison.
vs Hyperbolic
In context →- Should I use AI Token Router or Hyperbolic?
- Across 8 dimensions compared below, AI Token Router comes out ahead on 6 and Hyperbolic on 2 -- so the honest answer is that it depends on which dimensions matter to your workload, and both sections naming a winner are on this page rather than only the flattering one.
- When is Hyperbolic the better choice?
- If what you actually want is compute rather than tokens, Hyperbolic is a real offer and we have nothing comparable. On-demand H100s, H200s and B200s with no quota limits and no contract, a $5 entry point, credits that are 1:1 with dollars, and a route through reserved clusters to private cloud if the workload grows. Anyone running their own fine-tunes, training jobs, or a model we do not carry should be renting GPUs, not buying tokens — and they are a reasonable place to do it.
- How does AI Token Router pricing compare to Hyperbolic?
- On pricing transparency: AI Token Router is Every model's rate published, with the official rate beside it — 36–43% below official. Hyperbolic is The worst on this list: /pricing and /inference both return 404, and no per-token rate card exists anywhere on their domain.
- How was this AI Token Router vs Hyperbolic comparison made?
- Where the Hyperbolic figures come from: Every Hyperbolic cell in the table is taken from Hyperbolic's own published material -- their pricing page, their documentation or their rate card -- not from an aggregator, a press mention or our own testing. Where they publish no number for a dimension, the row says so rather than filling the gap with an estimate: 1 of the 8 rows carry a dollar figure they publish, and the rest turn on what the product does or does not do. Where our own figures come from: Our column is the catalogue that prices real requests, including the parts of it that do not flatter us: 7 of the 24 catalogued models are callable today, and the rows say so rather than counting the catalogue as the lineup. The rates behind it were last reconciled with upstream on Sep 2, 2026 and are published in full on the pricing page. When this was last checked: September 2026. Competitor pricing and product scope both move without notice, so every figure here is a statement about that date rather than a permanent one. Anything you are about to make a decision on is worth confirming against Hyperbolic's own site before you make it. What is deliberately not compared: There are no benchmark scores, no latency measurements and no tokens-per-second figures in our column -- not on this page and nowhere else on this site -- because we have not measured them. Where a speed claim appears in the Hyperbolic column it is theirs, repeated as they publish it and not verified by us. An unmeasured quality or speed number would be the easiest thing on this page to invent, which is precisely why there is not one. How the winner column is decided: Each of the 8 rows is marked for one side: 6 to AI Token Router, 2 to Hyperbolic. That marking is our judgement and you are free to disagree with it. The figures underneath it are the part you can check without taking our word for anything. Corrections: If a figure here is stale or wrong, tell us and we will change it -- including when the correction runs in Hyperbolic's favour, and including when it costs us a row we currently win. A comparison only ever corrected in its own favour is not being corrected.
- What does Hyperbolic do better?
- Hyperbolic wins on raw gpu access, path beyond serverless. Those rows are on the comparison table with Hyperbolic marked as the winner, because a comparison that never concedes anything is not a comparison.
The rest, answered on their own pages
Every question below is answered in full where it belongs, beside the rate table and the specification it refers to.
Check it against your own numbers
$5 in free credits, no card required.
Every rate quoted above is the rate the API bills.