Model catalog

Every model, explained and priced.

Filter by what you’re building, search by capability or tag, and compare cost, speed, and benchmarks — on one OpenAI-compatible key.

Models
113
One API surface
Providers
17
Failover across all
Showing
102
After your filters

What are you building?

Pick a job to narrow the catalog and surface leaders.

Best value
DeepSeek V4
$0.435 in / $0.870 out
Lowest blended rate, weighting output 3× — where the bill usually lands.
Fastest to first token
Gemini 3.6 Flash
480ms TTFT
Quickest to start streaming — what a waiting user actually feels.
Highest benchmarked
o3
96.7% avg
Mean of MMLU / MATH.
102 of 113
Sort
Amazon
Amazon Nova Lite
nova-lite

Fast, low-cost Nova for high-volume text and light multimodal tasks.

In$0.060
Out$0.240
Ctx300K
TTFT350ms
p5090ms
StreamingTool useJSON modeVision
Amazon
Amazon Nova Micro
nova-micro

Smallest Nova — text-only ultra-cheap option for classification and routing.

In$0.035
Out$0.140
Ctx128K
TTFT280ms
p5070ms
StreamingTool useJSON mode
Amazon
Amazon Nova Pro
nova-pro

Bedrock Nova Pro — multimodal workhorse with 300K context.

In$0.800
Out$3.20
Ctx300K
TTFT600ms
p50160ms
StreamingTool useJSON modeVision
Anthropic
New
Claude Fable 5
claude-fable-5

Anthropic's next-generation model for complex knowledge work and coding, built for sustained autonomous operation across multi-day tasks — plans across stages, delegates to sub-agents, and self-verifies its work.

In$10.00
Out$50.00
Ctx1M
TTFT900ms
p50220ms
StreamingTool useJSON modeVision+1
Anthropic
New
Claude Haiku 4.5
claude-haiku-4-5

Fastest and most compact Claude model. Ideal for classification, extraction, and high-volume tasks where cost and latency matter most.

In$0.800
Out$4.00
Ctx200K
TTFT380ms
p5085ms
StreamingTool useJSON modeVision
Anthropic
New
Claude Opus 4.5
claude-opus-4-5

Most capable Claude model. Sets the bar on advanced reasoning, research synthesis, and nuanced long-form content.

In$15.00
Out$75.00
Ctx200K
TTFT1.8s
p50290ms
StreamingTool useJSON modeVision+1
Anthropic
New
Claude Opus 4.8
claude-opus-4-8

The latest Opus 4.x model — advanced reasoning and research synthesis, with fast mode available for premium-priced low-latency responses.

In$5.00
Out$25.00
Ctx200K
TTFT1.6s
p50260ms
StreamingTool useJSON modeVision+1
Anthropic
New
Claude Sonnet 4.5
claude-sonnet-4-5

Balanced intelligence and speed. Handles complex coding, analysis, and multi-step reasoning without the cost of Opus.

In$3.00
Out$15.00
Ctx200K
TTFT720ms
p50175ms
StreamingTool useJSON modeVision
Anthropic
New
Claude Sonnet 5
claude-sonnet-5

Balanced Claude 5.x model with the full 1M-token context window at standard pricing. Handles complex coding, analysis, and multi-step reasoning at Sonnet-tier cost.

In$2.00
Out$10.00
Ctx1M
TTFT700ms
p50170ms
StreamingTool useJSON modeVision
Mistral
Codestral
codestral

Mistral’s code-specialized model — fill-in-the-middle and multi-language coding.

In$0.200
Out$0.600
Ctx32K
TTFT400ms
p50110ms
StreamingTool useJSON mode
Cohere
Command R
command-r

Cohere Command R — efficient RAG and tool use at budget rates.

In$0.150
Out$0.600
Ctx128K
TTFT500ms
p50140ms
StreamingTool useJSON mode
Cohere
Command R+
command-r-plus

Cohere’s RAG-oriented flagship — strong retrieval-augmented generation and tool use.

In$2.50
Out$10.00
Ctx128K
TTFT700ms
p50180ms
StreamingTool useJSON mode
OpenAI
Free
DALL·E 3
dall-e-3

OpenAI image generation — priced per image upstream.

InFree
OutFree
Ctx0K
TTFT2.0s
p502.0s
DeepSeek
DeepSeek Chat
deepseek-chat

DeepSeek’s chat alias — V3-class coding and general chat at low cost.

In$0.270
Out$1.10
Ctx64K
TTFT700ms
p50180ms
StreamingTool useJSON mode
DeepSeek
DeepSeek R1
deepseek-r1

Open-source reasoning model rivalling o1. Publishes its chain of thought — useful for auditability and research workflows.

In$0.550
Out$2.19
Ctx64K
TTFT1.2s
p50267ms
StreamingReasoning
DeepSeek
DeepSeek V3
deepseek-v3

Mixture-of-experts model with strong coding and math benchmarks at a fraction of frontier model prices.

In$0.270
Out$1.10
Ctx64K
TTFT780ms
p50180ms
StreamingTool useJSON mode
DeepSeek
DeepSeek V3.2
deepseek-v3-2

Bedrock-hosted DeepSeek V3.2 — coding and math at competitive MoE rates.

In$0.620
Out$1.85
Ctx128K
TTFT700ms
p50175ms
StreamingTool useJSON mode
DeepSeek
New
DeepSeek V4
deepseek-v4

DeepSeek's flagship (Pro tier) — 1M-token context with strong reasoning and coding at a fraction of frontier-tier cost.

In$0.435
Out$0.870
Ctx1M
TTFT750ms
p50190ms
StreamingTool useJSON modeReasoning
Mistral
Devstral 2 123B
devstral-2-123b

Mistral Devstral 2 — large coding-focused model on Bedrock.

In$0.400
Out$2.00
Ctx128K
TTFT800ms
p50200ms
StreamingTool useJSON mode
Google
Gemini 1.5 Flash
gemini-1.5-flash

Prior Gemini Flash — cheap long-context option.

In$0.075
Out$0.300
Ctx1M
TTFT350ms
p5090ms
StreamingTool useJSON modeVision
Google
Gemini 1.5 Pro
gemini-1.5-pro

Prior Gemini Pro generation — 1M context for long documents.

In$3.50
Out$10.50
Ctx1M
TTFT800ms
p50200ms
StreamingTool useJSON modeVision
Google
Gemini 2.0 Flash
gemini-2.0-flash

Low-latency workhorse with a 1M context window and multimodal inputs.

In$0.100
Out$0.400
Ctx1M
TTFT410ms
p5098ms
StreamingTool useJSON modeVision
Google
Gemini 2.5 Flash
gemini-2.5-flash

Best-value Gemini model. Handles massive context at low cost with strong performance on coding and document analysis.

In$0.300
Out$2.50
Ctx1M
TTFT580ms
p50140ms
StreamingTool useJSON modeVision
Google
Preview
Gemini 2.5 Flash Native Audio
gemini-2.5-flash-preview-native-audio

Gemini 2.5 Flash with native audio — preview long-context multimodal.

In$0.300
Out$2.50
Ctx1M
TTFT500ms
p50130ms
StreamingTool useVision
Google
Gemini 2.5 Pro
gemini-2.5-pro

Google's highest-capability model with native multimodal reasoning across text, image, video, and audio — all within a 1M-token window.

In$1.25
Out$10.00
Ctx1M
TTFT1.4s
p50310ms
StreamingTool useJSON modeVision+1
Google
New
Gemini 3.5 Flash
gemini-3.5-flash

Previous Flash tier — prefer 3.6 Flash for new workloads.

In$1.50
Out$9.00
Ctx1M
TTFT500ms
p50130ms
StreamingTool useJSON modeVision+1
Google
New
Gemini 3.5 Flash-Lite
gemini-3.5-flash-lite

High-throughput Flash-Lite for extraction, classification, and subagent loops.

In$0.300
Out$2.50
Ctx1M
TTFT280ms
p5070ms
StreamingTool useJSON modeVision
Google
New
Gemini 3.6 Flash
gemini-3.6-flash

July 2026 Flash workhorse — stronger coding/agentic loops than 3.5 at lower output cost, 1M context.

In$1.50
Out$7.50
Ctx1M
TTFT480ms
p50120ms
StreamingTool useJSON modeVision+1
Google
Gemma 3 12B
gemma-3-12b-it

Open Gemma 3 12B instruct — compact open-weight generalist on Bedrock.

In$0.090
Out$0.290
Ctx128K
TTFT400ms
p50110ms
StreamingTool useJSON mode
Google
Gemma 3 27B
gemma-3-27b-it

Open Gemma 3 27B instruct — strong open-weight generalist on Bedrock.

In$0.230
Out$0.380
Ctx128K
TTFT550ms
p50150ms
StreamingTool useJSON mode
Google
Gemma 3 4B
gemma-3-4b-it

Smallest Gemma 3 instruct — ultra-cheap classification and light chat.

In$0.040
Out$0.080
Ctx128K
TTFT280ms
p5070ms
StreamingJSON mode
Google
FreePreview
Gemma 4 26B A4B (Community)
community/gemma-4-26b-a4b

Shared community capacity for Google Gemma 4 26B A4B IT (multimodal).

InFree
OutFree
Ctx256K
TTFT850ms
p50360ms
StreamingTool useJSON modeVision+1
Google
Free
Gemma 4 31B (Community)
community/gemma-4-31b

Shared community capacity for Google Gemma 4 31B IT. Rate-limited free pool.

InFree
OutFree
Ctx128K
TTFT800ms
p50350ms
StreamingTool useJSON mode
Zhipu
GLM-4.7
glm-4-7

Prior GLM-4.7 generation — capable bilingual generalist.

In$0.600
Out$2.20
Ctx128K
TTFT650ms
p50170ms
StreamingTool useJSON mode
Zhipu
GLM-4.7 Flash
glm-4-7-flash

Fast GLM-4.7 tier for high-throughput bilingual workloads.

In$0.070
Out$0.400
Ctx128K
TTFT300ms
p5080ms
StreamingTool useJSON mode
Zhipu
New
GLM-5
glm-5

Zhipu GLM-5 — strong Chinese/English bilingual reasoning and coding.

In$1.00
Out$3.20
Ctx128K
TTFT700ms
p50180ms
StreamingTool useJSON modeReasoning
OpenAI
New
GPT-4.1 Mini
gpt-4.1-mini

Lightweight 4.1 variant with the same 1M-token window at a fraction of the cost. Ideal for agentic pipelines that need long context cheaply.

In$0.400
Out$1.60
Ctx1M
TTFT520ms
p50130ms
StreamingTool useJSON modeVision
OpenAI
GPT-4.1 Nano
gpt-4.1-nano

Cheapest 4.1-family model — high-throughput extraction and subagent loops with 1M context.

In$0.100
Out$0.400
Ctx1M
TTFT350ms
p5090ms
StreamingTool useJSON mode
OpenAI
GPT-4o
gpt-4o

Flagship multimodal model optimised for real-time dialogue. Natively processes text, images, and structured data.

In$2.50
Out$10.00
Ctx128K
TTFT900ms
p50210ms
StreamingTool useJSON modeVision
OpenAI
Preview
GPT-4o Audio Preview
gpt-4o-audio-preview

GPT-4o with native audio input/output (preview).

In$2.50
Out$10.00
Ctx128K
TTFT900ms
p50220ms
StreamingTool use
OpenAI
GPT-4o Mini
gpt-4o-mini

Affordable GPT-4o sibling that handles everyday tasks with excellent speed. Strong default for high-volume deployments.

In$0.150
Out$0.600
Ctx128K
TTFT490ms
p50130ms
StreamingTool useJSON modeVision
OpenAI
GPT-5.5
gpt-5.5

OpenAI's most capable model — advanced coding, research, analysis, software operation, document workflows, and long-running agentic tasks with less orchestration.

In$5.00
Out$30.00
Ctx272K
TTFT950ms
p50230ms
StreamingTool useJSON modeVision+1
OpenAI
New
GPT-5.6
gpt-5.6

Successor to GPT-5.5 (Sol tier) — OpenAI's current flagship for advanced coding, research, analysis, and long-running agentic tasks.

In$5.00
Out$30.00
Ctx272K
TTFT900ms
p50220ms
StreamingTool useJSON modeVision+1
Groq
New
GPT-OSS 120B
openai/gpt-oss-120b

Groq-hosted open-weight flagship replacing Llama 3.3 70B. Strong reasoning at low latency.

In$0.150
Out$0.600
Ctx128K
TTFT480ms
p50110ms
StreamingTool useJSON modeReasoning
OpenAI
GPT-OSS 120B (Bedrock)
gpt-oss-120b-bedrock

Open GPT-OSS 120B via Bedrock — large open-weight generalist.

In$0.150
Out$0.600
Ctx128K
TTFT800ms
p50200ms
StreamingTool useJSON mode
Groq
New
GPT-OSS 20B
openai/gpt-oss-20b

Groq-hosted open-weight model replacing Llama 3.1 8B. Fast, cheap, and ideal for high-volume prototyping.

In$0.075
Out$0.300
Ctx128K
TTFT320ms
p5065ms
StreamingTool useJSON mode
OpenAI
GPT-OSS 20B (Bedrock)
gpt-oss-20b-bedrock

Open GPT-OSS 20B via Bedrock — compact open-weight option.

In$0.070
Out$0.300
Ctx128K
TTFT500ms
p50140ms
StreamingTool useJSON mode
OpenAI
Free
GPT-OSS 20B (Community)
community/gpt-oss-20b

Shared community capacity for GPT-OSS 20B. Strong default for light prototyping.

InFree
OutFree
Ctx128K
TTFT700ms
p50300ms
StreamingTool useJSON mode
OpenAI
GPT-OSS Safeguard 120B
gpt-oss-safeguard-120b

Safety-tuned GPT-OSS 120B on Bedrock for moderation-style workloads.

In$0.150
Out$0.600
Ctx128K
TTFT800ms
p50200ms
StreamingJSON mode
OpenAI
GPT-OSS Safeguard 20B
gpt-oss-safeguard-20b

Safety-tuned GPT-OSS 20B on Bedrock — compact moderation helper.

In$0.070
Out$0.200
Ctx128K
TTFT500ms
p50140ms
StreamingJSON mode
xAI
Grok 3
grok-3

Prior Grok flagship — strong reasoning and coding with a 131K context window.

In$3.00
Out$15.00
Ctx131K
TTFT900ms
p50220ms
StreamingTool useJSON modeReasoning
xAI
Grok 3 Mini
grok-3-mini

Lightweight Grok 3 — fast and cheap for lighter agent loops.

In$0.300
Out$0.500
Ctx131K
TTFT400ms
p50100ms
StreamingTool useJSON mode
xAI
New
Grok 4.3
grok-4.3

Long-context Grok — 1M tokens with strong reasoning at a lower rate than Grok 4.5.

In$1.25
Out$2.50
Ctx1M
TTFT900ms
p50220ms
StreamingTool useJSON modeVision+1
xAI
New
Grok 4.5
grok-4.5

xAI's current frontier model — strong reasoning and agentic coding with a 500K-token context window.

In$2.00
Out$6.00
Ctx500K
TTFT850ms
p50210ms
StreamingTool useJSON modeVision+1
Moonshot
Kimi K2 Thinking
kimi-k2-thinking

Thinking-mode Kimi for multi-step reasoning and research synthesis.

In$0.600
Out$2.50
Ctx256K
TTFT1.2s
p50280ms
StreamingTool useJSON modeReasoning
Moonshot
Kimi K2.5
kimi-k2-5

General Kimi K2.5 — long-context reasoning and coding at Bedrock rates.

In$0.600
Out$3.00
Ctx256K
TTFT750ms
p50190ms
StreamingTool useJSON modeReasoning
Moonshot
New
Kimi K2.7 Code
kimi-k2.7-code

Moonshot coding-focused Kimi — 256K context for large repos and agent loops.

In$0.950
Out$4.00
Ctx256K
TTFT800ms
p50200ms
StreamingTool useJSON modeReasoning
Poolside
FreePreview
Laguna S 2.1 (Community)
community/laguna-s-2.1

Shared community capacity for Poolside Laguna S 2.1.

InFree
OutFree
Ctx256K
TTFT800ms
p50350ms
StreamingTool useReasoning
Poolside
FreePreview
Laguna XS 2.1 (Community)
community/laguna-xs-2.1

Shared community capacity for Poolside Laguna XS 2.1 — compact coding model.

InFree
OutFree
Ctx256K
TTFT700ms
p50300ms
StreamingTool useReasoning
InclusionAI
FreePreview
Ling 3.0 Flash (Community)
community/ling-3.0-flash

Shared community capacity for InclusionAI Ling 3.0 Flash (MoE).

InFree
OutFree
Ctx128K
TTFT750ms
p50320ms
StreamingTool useReasoning
Meta
Llama 3 70B Instruct
llama3-70b-instruct

Meta Llama 3 70B instruct on Bedrock — classic open-weight workhorse.

In$2.65
Out$3.50
Ctx8K
TTFT700ms
p50180ms
StreamingTool useJSON mode
Meta
Llama 3 8B Instruct
llama3-8b-instruct

Meta Llama 3 8B instruct — small, fast open-weight option on Bedrock.

In$0.300
Out$0.600
Ctx8K
TTFT300ms
p5080ms
StreamingTool useJSON mode
Meta
Llama 3.1 405B (Fireworks)
accounts/fireworks/models/llama-v3p1-405b-instruct

Meta Llama 3.1 405B instruct via Fireworks.

In$3.00
Out$3.00
Ctx128K
TTFT900ms
p50220ms
StreamingTool useJSON mode
Meta
Llama 3.1 70B (Fireworks)
accounts/fireworks/models/llama-v3p1-70b-instruct

Meta Llama 3.1 70B instruct via Fireworks.

In$0.900
Out$0.900
Ctx128K
TTFT600ms
p50160ms
StreamingTool useJSON mode
Meta
Llama 3.3 70B SpecDec
llama-3.3-70b-specdec

Groq speculative-decoding Llama 3.3 70B — faster throughput variant.

In$0.590
Out$0.990
Ctx128K
TTFT400ms
p50100ms
StreamingTool useJSON mode
Mistral
Magistral Small
magistral-small-2509

Mistral Magistral Small — reasoning-oriented small model on Bedrock.

In$0.500
Out$1.50
Ctx128K
TTFT700ms
p50180ms
StreamingTool useJSON modeReasoning
MiniMax
MiniMax M2
minimax-m2

MiniMax M2 — long-context agentic model via Bedrock.

In$0.300
Out$1.20
Ctx200K
TTFT700ms
p50180ms
StreamingTool useJSON mode
MiniMax
MiniMax M2.1
minimax-m2-1

MiniMax M2.1 — long-context agentic model via Bedrock.

In$0.300
Out$1.20
Ctx200K
TTFT700ms
p50180ms
StreamingTool useJSON mode
MiniMax
New
MiniMax M2.5
minimax-m2-5

MiniMax M2.5 — long-context agentic model via Bedrock.

In$0.300
Out$1.20
Ctx200K
TTFT700ms
p50180ms
StreamingTool useJSON mode
Mistral
Ministral 3 14B
ministral-3-14b-instruct

Ministral 3 14B — mid-size open instruct model on Bedrock.

In$0.200
Out$0.200
Ctx128K
TTFT450ms
p50120ms
StreamingTool useJSON mode
Mistral
Ministral 3 3B
ministral-3-3b-instruct

Tiny Ministral 3 — ultra-cheap classification and light chat on Bedrock.

In$0.100
Out$0.100
Ctx128K
TTFT250ms
p5060ms
StreamingJSON mode
Mistral
Ministral 3 8B
ministral-3-8b-instruct

Compact Ministral 3 8B — small open instruct model on Bedrock.

In$0.150
Out$0.150
Ctx128K
TTFT350ms
p5090ms
StreamingTool useJSON mode
Mistral
Mistral 7B Instruct v0.2
mistral-7b-instruct-v0-2

Classic Mistral 7B instruct on Bedrock.

In$0.150
Out$0.200
Ctx32K
TTFT350ms
p5090ms
Streaming
Mistral
Mistral Large
mistral-large

Flagship Mistral model with top-tier coding, reasoning, and multilingual performance. Native function calling support.

In$3.00
Out$9.00
Ctx128K
TTFT820ms
p50198ms
StreamingTool useJSON mode
Mistral
New
Mistral Large 3
mistral-large-3-675b-instruct

Mistral's current flagship (675B) — top-tier coding, reasoning, and multilingual performance, routed via Bedrock.

In$0.500
Out$1.50
Ctx128K
TTFT780ms
p50190ms
StreamingTool useJSON modeReasoning
Mistral
Mistral Small
mistral-small

Efficient European model for classification, summarisation, and structured extraction. Strong on French and other EU languages.

In$0.200
Out$0.600
Ctx32K
TTFT490ms
p50120ms
StreamingTool useJSON mode
Mistral
Mixtral 8x7B Instruct
mixtral-8x7b-instruct-v0-1

Mixtral MoE instruct on Bedrock — classic open MoE generalist.

In$0.450
Out$0.700
Ctx32K
TTFT600ms
p50160ms
Streaming
Nvidia
FreePreview
Nemotron 3 Nano 30B (Community)
community/nemotron-3-nano-30b

Shared community capacity for NVIDIA Nemotron 3 Nano 30B A3B.

InFree
OutFree
Ctx250K
TTFT800ms
p50340ms
StreamingTool useReasoning
Nvidia
FreePreview
Nemotron 3 Nano Omni 30B (Community)
community/nemotron-3-nano-omni-30b

Shared community capacity for NVIDIA Nemotron 3 Nano Omni — multimodal reasoning.

InFree
OutFree
Ctx250K
TTFT900ms
p50380ms
StreamingTool useVisionReasoning
Nvidia
FreePreview
Nemotron 3 Super 120B (Community)
community/nemotron-3-super-120b

Shared community capacity for NVIDIA Nemotron 3 Super.

InFree
OutFree
Ctx256K
TTFT1.0s
p50420ms
StreamingJSON modeReasoning
Nvidia
FreePreview
Nemotron 3 Ultra 550B (Community)
community/nemotron-3-ultra-550b

Shared community capacity for NVIDIA Nemotron 3 Ultra (MoE). Large free context window.

InFree
OutFree
Ctx977K
TTFT1.2s
p50500ms
StreamingTool useReasoning
Nvidia
FreePreview
Nemotron 3.5 Content Safety (Community)
community/nemotron-3.5-content-safety

Shared community capacity for NVIDIA Nemotron 3.5 content-safety classifier.

InFree
OutFree
Ctx125K
TTFT600ms
p50250ms
StreamingVisionReasoning
Nvidia
Nemotron Nano 12B v2
nemotron-nano-12b-v2

NVIDIA Nemotron Nano 12B — slightly larger nano tier for general chat.

In$0.200
Out$0.600
Ctx128K
TTFT450ms
p50120ms
StreamingTool useJSON mode
Nvidia
FreePreview
Nemotron Nano 12B VL (Community)
community/nemotron-nano-12b-vl

Shared community capacity for NVIDIA Nemotron Nano 12B V2 vision-language.

InFree
OutFree
Ctx125K
TTFT750ms
p50320ms
StreamingTool useVisionReasoning
Nvidia
Nemotron Nano 3 30B
nemotron-nano-3-30b

NVIDIA Nemotron Nano 3 30B — open reasoning-capable mid-size model.

In$0.060
Out$0.240
Ctx128K
TTFT600ms
p50160ms
StreamingTool useJSON modeReasoning
Nvidia
Nemotron Nano 9B v2
nemotron-nano-9b-v2

NVIDIA Nemotron Nano 9B — compact open model for light workloads.

In$0.060
Out$0.230
Ctx128K
TTFT350ms
p5090ms
StreamingJSON mode
Nvidia
Free
Nemotron Nano 9B v2 (Community)
community/nemotron-nano-9b

Shared community capacity for NVIDIA Nemotron Nano 9B v2. Good for low-stakes tasks.

InFree
OutFree
Ctx128K
TTFT750ms
p50320ms
StreamingJSON mode
Nvidia
New
Nemotron Super 3 120B
nemotron-super-3-120b

NVIDIA Nemotron Super 3 120B — open reasoning and coding at Bedrock rates.

In$0.150
Out$0.650
Ctx128K
TTFT800ms
p50200ms
StreamingTool useJSON modeReasoning
Cohere
FreePreview
North Mini Code (Community)
community/north-mini-code

Shared community capacity for Cohere North Mini Code.

InFree
OutFree
Ctx250K
TTFT750ms
p50320ms
StreamingReasoning
OpenAI
New
o3
o3

OpenAI's most powerful reasoning model. Spends more compute thinking before answering — ideal for math, science, and strategic planning.

In$10.00
Out$40.00
Ctx200K
TTFT5.0s
p50850ms
StreamingTool useJSON modeReasoning
OpenAI
o3-mini
o3-mini

Compact reasoning model tuned for STEM tasks. Faster and cheaper than o3 while retaining strong chain-of-thought capabilities.

In$1.10
Out$4.40
Ctx128K
TTFT2.8s
p50420ms
StreamingTool useJSON modeReasoning
OpenAI
New
o4-mini
o4-mini

Next-generation compact reasoning model. Faster than o3-mini with improved accuracy on code, math, and agentic workflows.

In$1.10
Out$4.40
Ctx200K
TTFT2.2s
p50380ms
StreamingTool useJSON modeVision+1
Qwen
Qwen3 32B
qwen3-32b

Dense Qwen3 32B — solid generalist at low cost on Bedrock.

In$0.150
Out$0.600
Ctx128K
TTFT500ms
p50140ms
StreamingTool useJSON mode
Qwen
Qwen3 Coder 30B
qwen3-coder-30b-a3b

Qwen3 MoE coder — strong coding at low cost with 256K context.

In$0.150
Out$0.600
Ctx256K
TTFT550ms
p50150ms
StreamingTool useJSON mode
Qwen
New
Qwen3 Coder Next
qwen3-coder-next

Latest Qwen3 coding model — strong agentic coding with 256K context.

In$0.500
Out$1.20
Ctx256K
TTFT650ms
p50170ms
StreamingTool useJSON mode
Qwen
Qwen3 Next 80B
qwen3-next-80b-a3b

Qwen3 Next MoE — long-context generalist for analysis and agents.

In$0.140
Out$1.20
Ctx256K
TTFT700ms
p50180ms
StreamingTool useJSON mode
Qwen
Qwen3 VL 235B
qwen3-vl-235b-a22b

Large vision-language Qwen3 — document and image understanding at MoE scale.

In$0.530
Out$2.66
Ctx256K
TTFT900ms
p50220ms
StreamingTool useJSON modeVision
OpenAI
text-embedding-3-large
text-embedding-3-large

OpenAI embedding model — higher-dimensional vectors.

In$0.130
OutFree
Ctx8K
TTFT250ms
p5060ms
OpenAI
text-embedding-3-small
text-embedding-3-small

OpenAI embedding model — small, cheap vectors for retrieval.

In$0.020
OutFree
Ctx8K
TTFT200ms
p5050ms
Mistral
Voxtral Mini 3B
voxtral-mini-3b-2507

Mistral Voxtral Mini — compact audio-aware model on Bedrock.

In$0.040
Out$0.040
Ctx32K
TTFT400ms
p50100ms
Streaming
Mistral
Voxtral Small 24B
voxtral-small-24b-2507

Mistral Voxtral Small — larger audio-aware model on Bedrock.

In$0.100
Out$0.300
Ctx32K
TTFT600ms
p50160ms
Streaming
OpenAI
Free
Whisper
whisper-1

OpenAI speech-to-text — priced per minute upstream.

InFree
OutFree
Ctx0K
TTFT1.5s
p501.5s

Three steps to your first call

Getting started
01

Create a free account

Email and a password. No card, no sales call. Community models are free from the first minute.

02

Create an inference key

Name it, scope it to the models you want, and set a spend ceiling. The cap is enforced before each request, not reconciled after.

03

Change one line

Point your existing OpenAI-compatible client at the Relixr base URL. Same SDK, same request shape, same response shape.