Catalogue

every model we serve, and what it really costs

One OpenAI-compatible endpoint, open weights only, priced per million tokens with no minimum and no subscription. Each model reaches several classes of GPU through different quantizations, so a 6 GB card and a 24 GB card both have work to do — and hosts keep 80% of what their hardware earns.

what your GPU can serve

your cardmodelsfor example
a 6 GB card3 modelsGemma 4 E4B Instruct, Qwen2.5 1.5B Instruct, Qwen2.5 0.5B Instruct
8 GB3 modelsGemma 4 E4B Instruct, Qwen2.5 1.5B Instruct, Qwen2.5 0.5B Instruct
10 GB5 modelsQwen3.8 27B Instruct, Gemma 4 12B Instruct, Gemma 4 E4B Instruct, Qwen2.5 1.5B Instruct
12 GB7 modelsQwen3.8 27B Instruct, GLM-4.7-Flash, Qwen3 Coder 30B-A3B, Gemma 4 12B Instruct
16 GB9 modelsQwen3.8 27B Instruct, Qwen3.6 35B-A3B, GLM-4.7-Flash, Qwen3 Coder 30B-A3B
24 GB10 modelsQwen3.8 27B Instruct, Qwen3.6 35B-A3B, GLM-4.7-Flash, Qwen3 Coder 30B-A3B

Numbers are the minimum VRAM a build needs. Install mahout, run mahout models, and it tells you exactly what your own card can serve.

Qwen3.8 27B Instruct

qwen3.8-27b

2026 flagship open model for a single card: strong reasoning and code, 262K context, and no EU-sovereign listing anywhere else.

chat code reasoning long-context multilingual tool-use vision

size
27B dense
max context
256K tokens
licence
Apache-2.0
price
€0.25 per million input tokens · €0.75 per million output tokens
smallest card
from 9 GB VRAM
9 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
UD-Q5_K_M 18.4 GB 23 GB VRAM 32K reference quality
UD-Q4_K_M 15.3 GB 20 GB VRAM 32K reference quality
UD-IQ4_XS 13.3 GB 18 GB VRAM 32K balanced
UD-Q3_K_XL 12.2 GB 16 GB VRAM 32K balanced
UD-IQ3_S 11.2 GB 15 GB VRAM 32K balanced
UD-IQ3_XXS 10.2 GB 14 GB VRAM 32K compact — fits smaller cards
UD-Q2_K_XL 9.2 GB 12 GB VRAM 32K compact — fits smaller cards
UD-IQ2_S 7.8 GB 11 GB VRAM 16K compact — fits smaller cards
UD-IQ2_XXS 6.8 GB 9 GB VRAM 16K compact — fits smaller cards

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

Qwen3.6 35B-A3B

qwen3.6-35b-a3b

Fast enough to delegate to: 3B active parameters decode at small-model speed while the full 35B answers, which is what makes it the one to hand sub-agent work to. Sees images, 262K context.

chat code reasoning long-context multilingual tool-use vision

size
35B MoE, 3B active
max context
256K tokens
licence
Apache-2.0
price
€0.18 per million input tokens · €0.55 per million output tokens
smallest card
from 13 GB VRAM
8 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
UD-IQ4_NL_XL 18.2 GB 23 GB VRAM 32K reference quality
UD-IQ4_XS 16.5 GB 21 GB VRAM 32K reference quality
UD-Q3_K_XL 15.7 GB 20 GB VRAM 32K balanced
UD-IQ3_S 12.7 GB 17 GB VRAM 32K balanced
UD-IQ3_XXS 12.3 GB 16 GB VRAM 32K balanced
UD-Q2_K_XL 11.4 GB 15 GB VRAM 32K compact — fits smaller cards
UD-IQ2_XXS 10.0 GB 14 GB VRAM 32K compact — fits smaller cards
UD-IQ1_M 9.4 GB 13 GB VRAM 16K compact — fits smaller cards

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

GLM-4.7-Flash

glm-4.7-flash

MIT licensed end to end — the most permissive terms in the catalogue, and the answer for buyers whose legal team reads the licence. Strong general reasoning from a second vendor.

chat code reasoning long-context multilingual tool-use

size
31B MoE
max context
198K tokens
licence
MIT
price
€0.16 per million input tokens · €0.5 per million output tokens
smallest card
from 11 GB VRAM
9 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_1 17.7 GB 23 GB VRAM 32K reference quality
UD-Q4_K_XL 16.3 GB 21 GB VRAM 32K reference quality
IQ4_XS 15.2 GB 20 GB VRAM 32K balanced
UD-Q3_K_XL 12.8 GB 17 GB VRAM 32K balanced
UD-IQ3_XXS 12.0 GB 16 GB VRAM 32K balanced
UD-Q2_K_XL 11.1 GB 15 GB VRAM 32K compact — fits smaller cards
UD-IQ2_XXS 9.8 GB 14 GB VRAM 32K compact — fits smaller cards
UD-IQ1_S 8.6 GB 12 GB VRAM 16K compact — fits smaller cards
UD-TQ1_0 7.8 GB 11 GB VRAM 16K compact — fits smaller cards

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

Qwen3 Coder 30B-A3B

qwen3-coder-30b-a3b

Built for code and agentic tool use: repository-scale context at 262K, and 3B active parameters so it answers at conversational speed.

chat code reasoning long-context tool-use

size
30B MoE, 3B active
max context
256K tokens
licence
Apache-2.0
price
€0.2 per million input tokens · €0.6 per million output tokens
smallest card
from 11 GB VRAM
9 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_1 17.9 GB 23 GB VRAM 32K reference quality
UD-Q4_K_XL 16.5 GB 21 GB VRAM 32K reference quality
IQ4_XS 15.3 GB 20 GB VRAM 32K balanced
UD-Q3_K_XL 12.9 GB 17 GB VRAM 32K balanced
UD-IQ3_XXS 12.0 GB 16 GB VRAM 32K balanced
UD-Q2_K_XL 11.0 GB 15 GB VRAM 32K compact — fits smaller cards
UD-IQ2_XXS 9.6 GB 14 GB VRAM 32K compact — fits smaller cards
UD-IQ1_M 9.0 GB 13 GB VRAM 16K compact — fits smaller cards
UD-TQ1_0 7.5 GB 11 GB VRAM 16K compact — fits smaller cards

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

Gemma 4 26B-A4B Instruct

gemma-4-26b-a4b

Mixture-of-experts: 26B of knowledge at the speed of a 4B model. Quantization-aware weights published by Google, so the 4-bit build loses far less than a post-hoc quantization.

chat reasoning multilingual long-context

size
26B MoE, 4B active
max context
128K tokens
licence
Gemma Terms of Use
price
€0.15 per million input tokens · €0.45 per million output tokens
smallest card
from 17 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_0 (QAT) 13.4 GB 17 GB VRAM 32K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

gpt-oss 20B

gpt-oss-20b

Open-weight 20B tuned for tool use and structured output. Natively MXFP4, so 4-bit costs it almost nothing.

chat code tool-use reasoning

size
20B MoE
max context
128K tokens
licence
Apache-2.0
price
€0.12 per million input tokens · €0.4 per million output tokens
smallest card
from 14 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_K_M 10.8 GB 14 GB VRAM 32K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

Gemma 4 12B Instruct

gemma-4-12b

The 8 GB workhorse: a capable general assistant on cards most gamers already own, with Google's quantization-aware 4-bit weights.

chat reasoning multilingual

size
12B dense
max context
128K tokens
licence
Gemma Terms of Use
price
€0.08 per million input tokens · €0.25 per million output tokens
smallest card
from 10 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_0 (QAT) 6.5 GB 10 GB VRAM 32K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

Gemma 4 E4B Instruct

gemma-4-e4b

Built for small cards and laptops: 6 GB of VRAM is enough to earn.

chat multilingual

size
8B MatFormer, 4B effective
max context
128K tokens
licence
Gemma Terms of Use
price
€0.05 per million input tokens · €0.15 per million output tokens
smallest card
from 6 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_0 (QAT) 4.8 GB 6 GB VRAM 16K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

Qwen2.5 1.5B Instruct

qwen2.5-1.5b-instruct

Tiny, fast, runs anywhere — the enrollment smoke test and the model behind starter work for a freshly linked machine.

chat

size
1.5B dense
max context
32K tokens
licence
Apache-2.0
price
€0.05 per million input tokens · €0.15 per million output tokens
smallest card
from 3 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_K_M 1.0 GB 3 GB VRAM 32K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

Qwen2.5 0.5B Instruct

qwen2.5-0.5b-instruct

Smallest catalog entry: CPU-only fallback and multi-model swap tests.

chat

size
0.5B dense
max context
32K tokens
licence
Apache-2.0
price
€0.02 per million input tokens · €0.06 per million output tokens
smallest card
from 2 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_K_M 0.5 GB 2 GB VRAM 32K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine, if you host, and on the machine serving you, if you build.

use them, or serve them

Point your SDK's base URL at https://app.elephantpool.ai/api/gateway/v1 and pass the model id — or connect a GPU and get paid for every token it produces.