One OpenAI-compatible endpoint, open weights only, priced per million tokens with no minimum
and no subscription. Each model reaches several classes of GPU through different
quantizations, so a 6 GB card and a 24 GB card both have work to do — and hosts keep
80% of what their hardware earns.
Minimum VRAM per build. Install mahout and run
mahout models — it tells you exactly what your own card can serve.
Qwen3.8 27B Instruct
qwen3.8-27b
2026 flagship open model for a single card: strong reasoning and code, 262K context, and no EU-sovereign listing anywhere else.
chatcodereasoninglong-contextmultilingualtool-use
size
27B dense
max context
256K tokens
licence
Apache-2.0
price
€0.25 in · €0.75 out per million tokens
smallest card
from 9 GB VRAM
5 build(s) — one per VRAM class
quantization
download
needs
context
notes
UD-Q5_K_M
18.4 GB
23 GB VRAM
16K
reference quality
UD-Q4_K_M
15.3 GB
20 GB VRAM
16K
reference quality
UD-Q3_K_XL
12.2 GB
16 GB VRAM
8K
balanced
UD-Q2_K_XL
9.2 GB
12 GB VRAM
8K
compact — fits smaller cards
UD-IQ2_XXS
6.8 GB
9 GB VRAM
8K
compact — fits smaller cards
Every build is pinned to an exact upstream revision and verified against its SHA-256
before it runs — on your machine if you host, and on the machine serving you if you build.
Gemma 4 26B-A4B Instruct
gemma-4-26b-a4b
Mixture-of-experts: 26B of knowledge at the speed of a 4B model. Quantization-aware weights published by Google, so the 4-bit build loses far less than a post-hoc quantization.
chatreasoningmultilinguallong-context
size
26B MoE, 4B active
max context
128K tokens
licence
Gemma Terms of Use
price
€0.15 in · €0.45 out per million tokens
smallest card
from 17 GB VRAM
1 build(s) — one per VRAM class
quantization
download
needs
context
notes
Q4_0 (QAT)
13.4 GB
17 GB VRAM
8K
reference quality
Every build is pinned to an exact upstream revision and verified against its SHA-256
before it runs — on your machine if you host, and on the machine serving you if you build.
gpt-oss 20B
gpt-oss-20b
Open-weight 20B tuned for tool use and structured output. Natively MXFP4, so 4-bit costs it almost nothing.
chatcodetool-usereasoning
size
20B MoE
max context
128K tokens
licence
Apache-2.0
price
€0.12 in · €0.4 out per million tokens
smallest card
from 14 GB VRAM
1 build(s) — one per VRAM class
quantization
download
needs
context
notes
Q4_K_M
10.8 GB
14 GB VRAM
8K
reference quality
Every build is pinned to an exact upstream revision and verified against its SHA-256
before it runs — on your machine if you host, and on the machine serving you if you build.
Gemma 4 12B Instruct
gemma-4-12b
The 8 GB workhorse: a capable general assistant on cards most gamers already own, with Google's quantization-aware 4-bit weights.
chatreasoningmultilingual
size
12B dense
max context
128K tokens
licence
Gemma Terms of Use
price
€0.08 in · €0.25 out per million tokens
smallest card
from 10 GB VRAM
1 build(s) — one per VRAM class
quantization
download
needs
context
notes
Q4_0 (QAT)
6.5 GB
10 GB VRAM
8K
reference quality
Every build is pinned to an exact upstream revision and verified against its SHA-256
before it runs — on your machine if you host, and on the machine serving you if you build.
Gemma 4 E4B Instruct
gemma-4-e4b
Built for small cards and laptops: 6 GB of VRAM is enough to earn.
chatmultilingual
size
8B MatFormer, 4B effective
max context
128K tokens
licence
Gemma Terms of Use
price
€0.05 in · €0.15 out per million tokens
smallest card
from 6 GB VRAM
1 build(s) — one per VRAM class
quantization
download
needs
context
notes
Q4_0 (QAT)
4.8 GB
6 GB VRAM
8K
reference quality
Every build is pinned to an exact upstream revision and verified against its SHA-256
before it runs — on your machine if you host, and on the machine serving you if you build.
Qwen2.5 1.5B Instruct
qwen2.5-1.5b-instruct
Tiny, fast, runs anywhere — the enrollment smoke test and the model behind starter work for a freshly linked machine.
chat
size
1.5B dense
max context
32K tokens
licence
Apache-2.0
price
€0.05 in · €0.15 out per million tokens
smallest card
from 3 GB VRAM
1 build(s) — one per VRAM class
quantization
download
needs
context
notes
Q4_K_M
1 GB
3 GB VRAM
8K
reference quality
Every build is pinned to an exact upstream revision and verified against its SHA-256
before it runs — on your machine if you host, and on the machine serving you if you build.
Qwen2.5 0.5B Instruct
qwen2.5-0.5b-instruct
Smallest catalog entry: CPU-only fallback and multi-model swap tests.
chat
size
0.5B dense
max context
32K tokens
licence
Apache-2.0
price
€0.02 in · €0.06 out per million tokens
smallest card
from 2 GB VRAM
1 build(s) — one per VRAM class
quantization
download
needs
context
notes
Q4_K_M
0.5 GB
2 GB VRAM
8K
reference quality
Every build is pinned to an exact upstream revision and verified against its SHA-256
before it runs — on your machine if you host, and on the machine serving you if you build.
use them, or serve them
Point your SDK's base URL at https://app.elephantpool.ai/api/gateway/v1 and pass
the model id — or connect a GPU and get paid for every token it produces.