Skip to content

Licensing, procurement and legal review

The licence is the reason this is possible, so we publish it

Every model here is served under an open licence that permits somebody other than the publisher to host the weights. That is not a compliance footnote — it is the mechanism the entire price difference rests on, which is why the licence is named on every model page rather than summarised in a policy.

24 models catalogued, each with a named licence; 7 callable today, and the pricing table says which is which.

  • No platform fee
  • OpenAI SDK compatible
  • Pay as you go
  • 7 models callable today

Send the licence and data-handling summary

The per-model licence list, the retention schedule, and the answers to the vendor-review questions that come up most — in a form you can forward to legal without editing it.

One email, no sequence. Or skip the email and create an account for $5 in credit.

The question legal actually asks

The blocking question in a model procurement review is rarely about price and almost never about benchmarks. It is: under what right is this vendor serving these weights to us, and what happens to our data once it reaches them. Both have documentary answers, and a vendor who cannot produce them quickly is a vendor whose review stalls for a quarter.

So the answers are published rather than available on request. Every model in the catalogue names its licence on its own page. The models we refuse to serve are listed too, with the reason for each refusal, because a list of what a vendor will not do is more informative to a reviewer than a paragraph asserting that the vendor is careful.

The commercial reason for this alignment is worth stating plainly: an open licence is what makes third-party hosting lawful, third-party hosting is what puts the cost base under our control, and control of the cost base is what produces the price difference. A vendor reselling closed-weight access at a discount would have neither the right nor the margin, which is why we do not do it.

The catalogue, with licences attached

Below is every model we can serve, with its publisher, its licence and its current availability. The licence column is the one a reviewer reads first; the availability column is the one an architect reads first.

Pricing is published on the same table as the licence rather than in a separate quote, and every rate sits 36–42% below what the model's own publisher charges. Nothing on this page is contingent on a procurement conversation.

Input rates on the models we are licensed to serve
Input rates on the models we are licensed to serveThe licence and the price are properties of the same decision: open weights are the only category where a third party can lawfully host and therefore control what the inference costs. On input, DeepSeek V4 Flash is the lowest at $0.090 per 1M tokens and Kimi K3 the highest at $1.85 per 1M tokens, a 20.6x spread. Every rate here runs 36–42% below the model publisher's own.Kimi K2.6$0.55Kimi K2.6, official$0.95Kimi K3$1.85Kimi K3, official$3.00GLM-5.2$0.82GLM-5.2, official$1.40DeepSeek V4 Pro$0.28DeepSeek V4 Pro, official$0.43DeepSeek V4 Flash$0.090DeepSeek V4 Flash, official$0.14
The licence and the price are properties of the same decision: open weights are the only category where a third party can lawfully host and therefore control what the inference costs. On input, DeepSeek V4 Flash is the lowest at $0.090 per 1M tokens and Kimi K3 the highest at $1.85 per 1M tokens, a 20.6x spread. Every rate here runs 36–42% below the model publisher's own.Our published rates, read from the same table the API bills from.
ModelPublisherLicenceModalityCallable today
Kimi K2.6Moonshot AIModified MITtextYes
Kimi K3Moonshot AIModified MITtextYes
GLM-5.2Z.aiMITtextYes
GLM-5.2 AirZ.aiMITtextNot yet served
DeepSeek V4 ProDeepSeekMITtextYes
DeepSeek V4 FlashDeepSeekMITtextYes
Qwen3 Max InstructAlibabaApache 2.0textNot yet served
Qwen3 235B A22BAlibabaApache 2.0textNot yet served
MiniMax M2MiniMaxMITtextNot yet served
Llama 4 MaverickMetaLlama 4 CommunitytextNot yet served
Mistral Large 3Mistral AIApache 2.0textNot yet served
GPT-OSS 120BOpenAIApache 2.0textNot yet served
Wan 2.2 T2V A14BAlibabaApache 2.0videoNot yet served
Wan 2.2 I2V A14BAlibabaApache 2.0videoNot yet served
LTX-2.5LightricksLTXV Open WeightsvideoNot yet served
HunyuanVideo 1.5TencentTencent Hunyuan CommunityvideoNot yet served
CogVideoX-5BZhipu AIApache 2.0videoNot yet served
Mochi 1GenmoApache 2.0videoNot yet served
FLUX.2 [schnell]Black Forest LabsApache 2.0imageNot yet served
Qwen-ImageAlibabaApache 2.0imageYes
Stable Diffusion 3.5 LargeStability AIStability CommunityimageNot yet served
HiDream-I1HiDreamMITimageNot yet served
Qwen3 Embedding 8BAlibabaApache 2.0embeddingYes
BGE-M3BAAIMITembeddingNot yet served
Publisher, licence and availability for every model in the catalogue.

Availability is stated in the table rather than discovered at call time. A model marked unavailable has a published price and no upstream serving it yet; requesting one returns an error naming the model rather than a substitution.

The cost side of the business case

A licence review usually runs in parallel with a cost approval, and the second is easier to produce evidence for. The calculator prices your volume across every model we serve against the publisher's own rate.

The result is a shareable URL, which is the format a finance approver can check rather than take on trust. It is also the format that survives being pasted into a procurement ticket.

M
M
0% of input cached

An agent or chat loop resends the same system prompt on every turn. That repeated prefix is billed at the cached rate, and on most models that is a fifth of the standard input rate. Set this to zero if every request you send is different.

Cheapest for this shape of work
$12.60/mo on DeepSeek V4 Flash
Against the official rate
$19.6036%

That is $7.00 a month, or $84.00 a year, on the same tokens through the same model.

An estimate, not a quote — but it is computed from the same rate table the API bills from, so the only assumptions in it are the ones you set above. Third-party figures are those providers’ own published rates; where a provider does not publish one, the cell is blank rather than guessed.

What a review needs from us, in order

Three documents, all of them already published. None of them requires a call, an NDA or a sales conversation to obtain.

  1. 1

    The licence for each model you intend to use

    Named on the model's own page and in the licence reference, with the family explained. A reviewer needs the licence name and the permission it grants for third-party hosting and for commercial use of the outputs.

  2. 2

    The data-handling schedule

    What is stored, for how long, and what is never stored. Chat and completion traffic is relayed and not retained. Asynchronous job inputs and results are held for 30 days so they can be fetched, or deleted immediately on request. Usage metadata is retained because the invoice is built from it.

  3. 3

    The refusal list

    The models we do not serve and why. A reviewer assessing supply-chain risk learns more from a specific refusal than from a general assurance, and this is the list that shows the policy is applied rather than stated.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://router.xark.io/api/v1"   # ← the only line that changes
)

response = client.chat.completions.create(
    model="z-ai/glm-5.2",
    messages=[{"role": "user", "content": "Hello"}]
)
The integration a reviewer is approving is one base URL.

Licence and privacy questions that are not answered on the published pages go to affiliate@xark.io, which is a monitored address rather than a role alias that bounces.

Send the licence and data-handling summary

The per-model licence list, the retention schedule, and the answers to the vendor-review questions that come up most — in a form you can forward to legal without editing it.

One email, no sequence. Or skip the email and create an account for $5 in credit.

The models we refuse, and why

This list exists because it is the most useful thing a vendor can show a reviewer. Each of these is a model customers ask for, and each is one we will not serve.

Sora 2 — OpenAI
Closed weights — no resale license exists
Veo 3.1 — Google DeepMind
Closed weights — no resale license exists
Runway Gen-4 — Runway
Closed weights — no resale license exists
Kling 2.5 — Kuaishou
Closed weights — no resale license exists
Wan 2.5 / 2.6 — Alibaba
API-only — weights never released, despite the open Wan 2.1/2.2 lineage

What we do with your data

We do not train on prompts or completions, ever. Chat and completion traffic is relayed to the model and not stored by us at all. The inputs and results of asynchronous jobs — video and image generation, which cannot return inline — are held for 30 days from completion so you can fetch them, or deleted immediately on request.

Usage metadata, meaning the model, token counts, latency and cost, is retained for 24 months because the dashboard and your invoice are built from it. The privacy policy carries the authoritative schedule, and where this page and that page differ, that page is the one to cite in a review.

Contracting entity
Innoprise Incubator INC, named identically in the terms, the privacy policy and the compliance page so a reviewer is not reconciling three documents.
Output rights
Determined by the licence of the model that generated them, which is why the licence is on the model page rather than abstracted into a single site-wide statement. Licence families differ on this and the difference matters.
No silent substitution
A request for a specific model is served by that model or fails. Substituting a model would change the licence terms of the output you receive, which is a compliance event and not merely an engineering one.

Procurement and licensing questions

The questions a review actually blocks on, answered without hedging.

What licence is each model served under?

Each model's licence is named on its own page and in the licence reference — Kimi K3 under Modified MIT, GLM-5.2 under MIT, and so on for all 24 catalogued models. They are grouped by family so a reviewer can read the terms once and apply the conclusion to every model under it.

Can we use the outputs commercially?

That is determined by the licence of the model that produced them, which is why the licence is attached to the model rather than asserted site-wide. The licence reference sets out what each family permits. Where a family imposes conditions on commercial use, that is stated on the family page rather than buried in a general term.

Why do you not offer GPT, Claude or Gemini?

Because reselling closed-weight model access frequently violates the origin provider's terms, and that is a risk we will not take on or expose a customer to. It is the same reason the refusal list on this page exists. If your architecture requires those models, a router that carries them is the correct vendor and we say so on our comparison pages.

Which models do you explicitly refuse to serve?

Sora 2 (OpenAI), Veo 3.1 (Google DeepMind), Runway Gen-4 (Runway), Kling 2.5 (Kuaishou), Wan 2.5 / 2.6 (Alibaba). Most are closed-weight, so no resale licence exists. One is API-only — the weights were never released despite an open lineage in earlier versions — which is exactly the case a reviewer is most likely to miss.

Do you train on our prompts?

No, and there is no setting that changes this. Prompts and completions from chat and completion requests are relayed and not stored by us at all. There is no opt-out to configure because there is nothing to opt out of.

How long is our data retained?

Chat and completion traffic is not retained. Asynchronous job inputs and results — video and image generation, which cannot return inline — are held for 30 days from completion so they can be fetched, or deleted immediately on request. Usage metadata is kept for 24 months because the dashboard and the invoice are built from it. The privacy policy carries the authoritative schedule.

Who is the contracting entity?

Innoprise Incubator INC. The same entity is named in the terms, the privacy policy and the compliance page, deliberately, because three documents naming three variants of a company name is how a review stalls on a question nobody intended to raise.

Where do the weights actually run?

On inference capacity we buy and operate rather than on a resold upstream API, which is what the open licence permits and what puts the cost base under our control. Requests for a model we do not currently serve fail with an error naming the model rather than being routed somewhere unexpected.

Which models are actually callable today?

7 of 24: Kimi K2.6, Kimi K3, GLM-5.2, DeepSeek V4 Pro, DeepSeek V4 Flash, Qwen-Image, Qwen3 Embedding 8B. The remainder are catalogued with published prices and marked as not yet served. For a procurement exercise this matters more than usual, because approving a model that cannot be called wastes the review rather than the integration.

Do you substitute a different model if one fails?

No. Substituting would change the licence terms of the output as well as its cost and content, which makes it a compliance event rather than a resilience feature. Requests fail with the upstream status and a machine-readable code instead.

How do we verify your pricing claim independently?

Every row prints the publisher's own rate beside ours, so the comparison can be checked against the publisher's public pricing page rather than accepted. The claim is that our rate sits 36–42% below theirs — not that it is the lowest rate available anywhere, which is a claim nobody can substantiate for long.

Who do we contact for a security or legal question?

affiliate@xark.io for licensing and contractual questions and affiliate@xark.io for security disclosure. Both are monitored addresses. We publish only addresses that actually receive mail, because a bouncing contact in a policy document is worse than a missing one — the sender believes they made contact.

Is there a volume agreement for a reviewed vendor relationship?

Yes. Accounts above $1,000 a month can ask for custom rates, and we reply within 24 hours. Nothing on the published pricing page is contingent on that conversation; it is an option, not a gate.

Send the licence list to whoever has to approve it

Everything a first-pass review needs is already published, which is the point. Leave an address and we will send the summarised version in a form you can forward without editing.

We guarantee at least 20% below the model publisher's own rate on every model we serve. The smallest discount in the catalogue today is 36%, so the guarantee has room in it by design.

Send the licence and data-handling summary

The per-model licence list, the retention schedule, and the answers to the vendor-review questions that come up most — in a form you can forward to legal without editing it.

One email, no sequence. Or skip the email and create an account for $5 in credit.

Related: Licence reference · Compliance · Privacy policy