Catalogue

every model we serve, and what it really costs

One OpenAI-compatible endpoint, open weights only, priced per million tokens with no minimum and no subscription. Each model reaches several classes of GPU through different quantizations, so a 6 GB card and a 24 GB card both have work to do — and hosts keep 80% of what their hardware earns.

what your GPU can serve

Your cardModelsFor example
6 GB 3 Gemma 4 E4B Instruct, Qwen2.5 1.5B Instruct, Qwen2.5 0.5B Instruct
8 GB 3 Gemma 4 E4B Instruct, Qwen2.5 1.5B Instruct, Qwen2.5 0.5B Instruct
10 GB 5 Qwen3.8 27B Instruct, Gemma 4 12B Instruct, Gemma 4 E4B Instruct
12 GB 5 Qwen3.8 27B Instruct, Gemma 4 12B Instruct, Gemma 4 E4B Instruct
16 GB 6 Qwen3.8 27B Instruct, gpt-oss 20B, Gemma 4 12B Instruct
24 GB 7 Qwen3.8 27B Instruct, Gemma 4 26B-A4B Instruct, gpt-oss 20B

Minimum VRAM per build. Install mahout and run mahout models — it tells you exactly what your own card can serve.

Qwen3.8 27B Instruct

qwen3.8-27b

2026 flagship open model for a single card: strong reasoning and code, 262K context, and no EU-sovereign listing anywhere else.

chatcodereasoninglong-contextmultilingualtool-use

size
27B dense
max context
256K tokens
licence
Apache-2.0
price
€0.25 in · €0.75 out per million tokens
smallest card
from 9 GB VRAM
5 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
UD-Q5_K_M 18.4 GB 23 GB VRAM 16K reference quality
UD-Q4_K_M 15.3 GB 20 GB VRAM 16K reference quality
UD-Q3_K_XL 12.2 GB 16 GB VRAM 8K balanced
UD-Q2_K_XL 9.2 GB 12 GB VRAM 8K compact — fits smaller cards
UD-IQ2_XXS 6.8 GB 9 GB VRAM 8K compact — fits smaller cards

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine if you host, and on the machine serving you if you build.

Gemma 4 26B-A4B Instruct

gemma-4-26b-a4b

Mixture-of-experts: 26B of knowledge at the speed of a 4B model. Quantization-aware weights published by Google, so the 4-bit build loses far less than a post-hoc quantization.

chatreasoningmultilinguallong-context

size
26B MoE, 4B active
max context
128K tokens
licence
Gemma Terms of Use
price
€0.15 in · €0.45 out per million tokens
smallest card
from 17 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_0 (QAT) 13.4 GB 17 GB VRAM 8K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine if you host, and on the machine serving you if you build.

gpt-oss 20B

gpt-oss-20b

Open-weight 20B tuned for tool use and structured output. Natively MXFP4, so 4-bit costs it almost nothing.

chatcodetool-usereasoning

size
20B MoE
max context
128K tokens
licence
Apache-2.0
price
€0.12 in · €0.4 out per million tokens
smallest card
from 14 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_K_M 10.8 GB 14 GB VRAM 8K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine if you host, and on the machine serving you if you build.

Gemma 4 12B Instruct

gemma-4-12b

The 8 GB workhorse: a capable general assistant on cards most gamers already own, with Google's quantization-aware 4-bit weights.

chatreasoningmultilingual

size
12B dense
max context
128K tokens
licence
Gemma Terms of Use
price
€0.08 in · €0.25 out per million tokens
smallest card
from 10 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_0 (QAT) 6.5 GB 10 GB VRAM 8K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine if you host, and on the machine serving you if you build.

Gemma 4 E4B Instruct

gemma-4-e4b

Built for small cards and laptops: 6 GB of VRAM is enough to earn.

chatmultilingual

size
8B MatFormer, 4B effective
max context
128K tokens
licence
Gemma Terms of Use
price
€0.05 in · €0.15 out per million tokens
smallest card
from 6 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_0 (QAT) 4.8 GB 6 GB VRAM 8K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine if you host, and on the machine serving you if you build.

Qwen2.5 1.5B Instruct

qwen2.5-1.5b-instruct

Tiny, fast, runs anywhere — the enrollment smoke test and the model behind starter work for a freshly linked machine.

chat

size
1.5B dense
max context
32K tokens
licence
Apache-2.0
price
€0.05 in · €0.15 out per million tokens
smallest card
from 3 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_K_M 1 GB 3 GB VRAM 8K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine if you host, and on the machine serving you if you build.

Qwen2.5 0.5B Instruct

qwen2.5-0.5b-instruct

Smallest catalog entry: CPU-only fallback and multi-model swap tests.

chat

size
0.5B dense
max context
32K tokens
licence
Apache-2.0
price
€0.02 in · €0.06 out per million tokens
smallest card
from 2 GB VRAM
1 build(s) — one per VRAM class
quantizationdownloadneedscontextnotes
Q4_K_M 0.5 GB 2 GB VRAM 8K reference quality

Every build is pinned to an exact upstream revision and verified against its SHA-256 before it runs — on your machine if you host, and on the machine serving you if you build.

use them, or serve them

Point your SDK's base URL at https://app.elephantpool.ai/api/gateway/v1 and pass the model id — or connect a GPU and get paid for every token it produces.