Skip to main content

Provider guides

Text API coverage​

The primary coverage focus is DeepSeek, Qwen, GLM/Z.ai, Kimi, MiniMax, Tencent HY3/HY4 and Xiaomi MiMo. Existing providers remain supported. The native JSON interface described in the function reference preserves reasoning and tools across protocol round trips. All model IDs are passed through, including account-specific deployments and dated versions. Models do not need a built-in pricing/catalog entry to be called.

The table is a documentation snapshot reviewed on 2026-09-14, not a live account availability check. Native transport is tested with mock HTTP for all seven providers. Real inference, model limits, regional availability and pricing are not live-verified. No new hard-coded prices or limits are inferred from model names. Check each linked catalog for current lifecycle and account terms.

ProviderChat selectionOther documented text APIs and endpointsCredentials
DeepSeekdeepseek; explicit V4 Flash/Pro model IDsThinking/tools, JSON output; FIM uses a separately selected beta endpoint/modelDEEPSEEK_API_KEY
Qwenqwen / dashscope; explicit Qwen model IDChat/Responses/embeddings; Messages at https://dashscope-intl.aliyuncs.com/apps/anthropic/v1/messages; Responses at https://dashscope-intl.aliyuncs.com/api/v2/apps/protocols/compatible-mode/v1/responsesDASHSCOPE_API_KEY / QWEN_API_KEY
GLMzai; explicit GLM model IDChat, embeddings; Messages at https://api.z.ai/api/anthropic/v1/messagesZAI_API_KEY
Kimikimi / moonshot; explicit Kimi IDChat/Responses/Messages; Messages at https://api.moonshot.ai/anthropic/v1/messagesMOONSHOT_API_KEY / KIMI_API_KEY
MiniMaxminimax; explicit MiniMax IDMessages at https://api.minimax.io/anthropic/v1/messagesMINIMAX_API_KEY
Tencenthunyuan / tencent; hy3, hy4-previewChat, Responses, Messages; append the respective path to the selected regional hostHUNYUAN_API_KEY / TOKENHUB_API_KEY
MiMomimo, xiaomi, xiaomi_mimo; mimo-v2.5-pro default, explicit IDs supportedChat at https://api.xiaomimimo.com/v1/chat/completions; Messages at https://api.xiaomimimo.com/anthropic/v1/messagesMIMO_API_KEY via api-key header

Tencent's documented regional bases are Singapore https://tokenhub-intl.tencentcloudmaas.com/v1, Guangzhou https://tokenhub.tencentcloudmaas.com/v1, and US https://tokenhub-us.tencentcloudmaas.tech/v1. Select the region matching your key explicitly. The existing Hunyuan default remains for compatibility. HY4 is a preview; activation and model availability are controlled by Tencent.

Qwen native reranking has model-dependent paths and payloads. Use qwen3-rerank with the full workspace /compatible-mode/v1/reranks endpoint, api := 'rerank', and a body containing model, query, documents, and optional top_n. GTE reranking uses nested input/parameters and a different endpoint; send its documented body unchanged. The existing ai_rerank is completion-based scoring, not a claim that every provider has a native reranking model. No native embeddings or reranking are claimed for providers without a verified public reference.

Reasoning options vary: DeepSeek, GLM and MiMo use thinking; Qwen uses enable_thinking and supports preserve_thinking on selected models. Preserve returned reasoning state when replying to tool calls. Some Qwen models require streaming; use native JSON with stream:true and read the buffered event array. Use the provider's JSON-object response format when JSON Schema is unsupported, then validate the result locally. Coding/subscription keys may require different endpoints from general API keys. No automatic account, region or protocol fallback is performed. Image, audio, video generation and asynchronous media jobs are out of scope for this rollout.

SELECT ai_complete('Explain columnar storage.', provider := 'mimo',
request_options := '{"thinking":{"type":"disabled"}}', max_tokens := 256);

This page gives one simple end-to-end example for each supported provider. Examples assume the extension is installed and loaded:

INSTALL ai FROM community;
LOAD ai;

(Or build from source with GEN=ninja make release and run ./build/release/duckdb.)

Each guide uses a DuckDB secret for provider, model, and base URL settings. The API key is read from the matching environment variable when the provider call runs. If you prefer a fully DuckDB-managed secret, add API_KEY '...' to the CREATE SECRET statement.

After any provider call, inspect the local usage buffer:

SELECT function_name, provider, protocol, model, http_status, elapsed_ms
FROM ai_usage()
ORDER BY event_id DESC
LIMIT 1;

Provider matrix​

ProviderProtocolDefault modelCredentialsBase URL behavior
ollamaOllama chat and embeddingsllama3.2; embeddings use nomic-embed-textOptional OLLAMA_API_KEYDefaults to http://localhost:11434; OLLAMA_HOST overrides it.
openaiOpenAI-compatible chat and embeddingsgpt-5.6-luna; embeddings use text-embedding-3-smallOPENAI_API_KEYDefaults to https://api.openai.com/v1.
azureOpenAI-compatible chat and embeddingsgpt-4o; embeddings use text-embedding-3-smallAZURE_OPENAI_API_KEYAppends /openai/v1 to AZURE_OPENAI_BASE_URL, AZURE_OPENAI_ENDPOINT, or secret BASE_URL when needed.
anthropic / claudeAnthropic Messagesclaude-haiku-4-5ANTHROPIC_API_KEY or CLAUDE_API_KEYDefaults to https://api.anthropic.com/v1.
bedrockOpenAI-compatible chatopenai.gpt-oss-120bAWS_BEDROCK_API_KEY, AWS_BEARER_TOKEN_BEDROCK, or BEDROCK_API_KEYSet AWS_REGION, AWS_BEDROCK_REGION, AWS_BEDROCK_BASE_URL, or secret BASE_URL.
cerebrasOpenAI-compatible chatgpt-oss-120bCEREBRAS_API_KEYDefaults to https://api.cerebras.ai/v1.
cloudflare / workers_aiOpenAI-compatible chat and embeddings@cf/zai-org/glm-4.7-flash; embeddings use @cf/baai/bge-base-en-v1.5CLOUDFLARE_API_KEY, CLOUDFLARE_API_TOKEN, or CLOUDFLARE_AUTH_TOKENDerives the endpoint from CLOUDFLARE_ACCOUNT_ID, or accepts CLOUDFLARE_WORKERS_AI_BASE_URL, CLOUDFLARE_AI_BASE_URL, or secret BASE_URL.
cohereOpenAI-compatible chat and embeddingscommand-a-plus-05-2026; embeddings use embed-v4.0COHERE_API_KEYDefaults to https://api.cohere.ai/compatibility/v1.
dashscope / qwenOpenAI-compatible chat and embeddingsqwen-plus; embeddings use text-embedding-v4DASHSCOPE_API_KEY, ALIBABA_API_KEY, or QWEN_API_KEYDefaults to https://dashscope-intl.aliyuncs.com/compatible-mode/v1; override with a workspace-specific base URL when needed.
deepinfraOpenAI-compatible chat and embeddingsmeta-llama/Meta-Llama-3.1-8B-Instruct-Turbo; embeddings use BAAI/bge-large-en-v1.5DEEPINFRA_API_KEYDefaults to https://api.deepinfra.com/v1/openai.
fireworksOpenAI-compatible chat and embeddingsaccounts/fireworks/models/gpt-oss-20b; embeddings use nomic-ai/nomic-embed-text-v1.5FIREWORKS_API_KEYDefaults to https://api.fireworks.ai/inference/v1.
gemini / gcp / googleOpenAI-compatible chat and embeddingsgemini-3.7-flash; embeddings use gemini-embedding-2GEMINI_API_KEYDefaults to Google's OpenAI-compatible endpoint.
groqOpenAI-compatible chatopenai/gpt-oss-20bGROQ_API_KEYDefaults to https://api.groq.com/openai/v1.
huggingface / hfOpenAI-compatible chatopenai/gpt-oss-120bHF_TOKEN, HUGGINGFACE_API_KEY, or HUGGING_FACE_HUB_TOKENDefaults to https://router.huggingface.co/v1.
hunyuan / tencent_hunyuanOpenAI-compatible chat through Tencent TokenHubhy3HUNYUAN_API_KEY, TOKENHUB_API_KEY, or TENCENT_TOKENHUB_API_KEYDefaults to https://tokenhub.tencentmaas.com/v1; TOKENHUB_BASE_URL selects another TokenHub region.
minimaxOpenAI-compatible chatMiniMax-M2.7MINIMAX_API_KEY or MINI_MAX_API_KEYDefaults to https://api.minimax.io/v1.
mistralOpenAI-compatible chat and embeddingsmistral-small-latest; embeddings use mistral-embedMISTRAL_API_KEYDefaults to https://api.mistral.ai/v1.
moonshot / kimiOpenAI-compatible chatkimi-k3MOONSHOT_API_KEY or KIMI_API_KEYDefaults to https://api.moonshot.ai/v1.
nebius / nebius_token_factoryOpenAI-compatible chatmeta-llama/Meta-Llama-3.1-70B-InstructNEBIUS_API_KEY or TOKEN_FACTORY_API_KEYDefaults to https://api.tokenfactory.nebius.com/v1.
nvidia / nvidia_nimOpenAI-compatible chatnvidia/nemotron-3-super-120b-a12bNVIDIA_API_KEYDefaults to https://integrate.api.nvidia.com/v1.
zai / zhipuOpenAI-compatible chat and embeddingsglm-4.7-flash; embeddings use embedding-3ZAI_API_KEYDefaults to https://api.z.ai/api/paas/v4.
deepseekOpenAI-compatible chatdeepseek-v4-flashDEEPSEEK_API_KEYDefaults to https://api.deepseek.com.
openrouterOpenAI-compatible chat and embeddingsopenai/gpt-4o-mini; embeddings use openai/text-embedding-3-smallOPENROUTER_API_KEYDefaults to https://openrouter.ai/api/v1.
databricksOpenAI-compatible chatdatabricks-gpt-oss-120bDATABRICKS_TOKENDerives /serving-endpoints from DATABRICKS_HOST, or accepts full Model Serving, AI Gateway, or chat-completions URLs.
snowflakeOpenAI-compatible chatclaude-sonnet-4-5SNOWFLAKE_PAT or SNOWFLAKE_TOKENDerives /api/v2/cortex/v1 from Snowflake account URL, host, or account id.
perplexityOpenAI-compatible chatsonarPERPLEXITY_API_KEYDefaults to https://api.perplexity.ai.
poeOpenAI-compatible chatGPT-5.4POE_API_KEYDefaults to https://api.poe.com/v1.
qianfan / ernieOpenAI-compatible chaternie-4.5-turbo-128kQIANFAN_API_KEY, BAIDU_QIANFAN_API_KEY, or BAIDU_API_KEYDefaults to https://qianfan.baidubce.com/v2.
sambanovaOpenAI-compatible chatMeta-Llama-3.3-70B-InstructSAMBANOVA_API_KEYDefaults to https://api.sambanova.ai/v1.
siliconflowOpenAI-compatible chatQwen/Qwen2.5-72B-InstructSILICONFLOW_API_KEYDefaults to https://api.siliconflow.com/v1.
togetherOpenAI-compatible chat and embeddingsmeta-llama/Llama-3.3-70B-Instruct-Turbo; embeddings use intfloat/multilingual-e5-large-instructTOGETHER_API_KEYDefaults to https://api.together.xyz/v1.
stepfun / stepOpenAI-compatible chatstep-3.5-flashSTEPFUN_API_KEY or STEP_API_KEYDefaults to https://api.stepfun.com/v1.
vercel / vercel_ai_gatewayOpenAI-compatible chat and embeddingsopenai/gpt-4o-mini; embeddings use openai/text-embedding-3-smallAI_GATEWAY_API_KEY, VERCEL_AI_GATEWAY_API_KEY, or VERCEL_OIDC_TOKENDefaults to https://ai-gateway.vercel.sh/v1.
vertex / google_vertexOpenAI-compatible chatgoogle/gemini-2.5-flashVERTEX_AI_ACCESS_TOKEN, GOOGLE_CLOUD_ACCESS_TOKEN, or VERTEX_API_KEYDerives the endpoint from GOOGLE_CLOUD_PROJECT, or accepts VERTEX_AI_BASE_URL, GOOGLE_VERTEX_BASE_URL, or secret BASE_URL.
volcengine / doubaoOpenAI-compatible chatdoubao-seed-2-1-pro-260628VOLCENGINE_API_KEY, ARK_API_KEY, or DOUBAO_API_KEYDefaults to https://ark.cn-beijing.volces.com/api/v3.
xai / grokOpenAI-compatible chatgrok-4.6XAI_API_KEYDefaults to https://api.x.ai/v1.
typesafe / jevTypeSafe System One evaluationjev-latestTYPESAFE_API_KEYDefaults to https://api.typesafe.ai/v1 and calls /systemone.
openai_privacy_filterDedicated redaction endpointopenai/privacy-filterOptional OPENAI_PRIVACY_FILTER_API_KEYDefaults to http://localhost:8080 and calls POST /redact.
openai_compatible / localOpenAI-compatible chat and embeddingsgpt-4o-mini; embeddings use text-embedding-3-smallOptional OPENAI_COMPATIBLE_API_KEYRequires BASE_URL or OPENAI_COMPATIBLE_BASE_URL.
llamacpp / llama.cppOpenAI-compatible chat and embeddingsdefault (llama-server answers with its loaded model)Optional LLAMACPP_API_KEY (llama-server --api-key)Defaults to http://localhost:8080/v1. Embeddings need llama-server --embeddings.

For guidance on choosing providers, credentials, logging, cost, throughput, and PII workflows, see Best practices.

TypeSafe Jev​

Jev evaluates state against typed questions and returns probabilities, choices, and rubric scores. It does not generate text or embeddings. Use typesafe (or its alias jev) for the direct TypeSafe API. The supported entry points are ai_jev, ai_provider_call, ai_classify, and ai_filter. Use ai_jev for typed choices, rubric scores and probabilities with automatic row batching. Other AI task functions, including ai_score, require a completion provider. Generation options such as temperature, max_tokens, and system_prompt are rejected. Put task guidance in classification instructions or in native question instructions.

export TYPESAFE_API_KEY='...'
./build/release/duckdb
LOAD ai;
CREATE OR REPLACE SECRET typesafe_ai (
TYPE duckdb_ai,
AI_PROVIDER 'typesafe',
MODEL 'jev-latest'
);

SELECT ai_classify(
'I was charged twice.', ['billing', 'technical', 'other'],
secret := 'typesafe_ai'
) AS department;

SELECT ai_filter(
'Our production imports are blocked.',
'Does the text describe blocked production work?',
secret := 'typesafe_ai'
) AS production_blocked;

Classification sends a native Choice question. Filtering sends a Noul question and returns true when its probability is at least 0.5. Both keep one request per row. For typed multi-question results and batches of up to 32 rows, use:

SELECT ai_jev('I was charged twice. Please refund me.', {
department: MAP {'billing': 'Payments and refunds', 'other': 'Other requests'},
refund_requested: MAP {'true': 'Asks for money back', 'false': 'Does not ask for money back'}
}, secret := 'typesafe_ai') AS decision;

Use decision.department and decision.refund_requested downstream. Follow the typed Jev cookbook to save results before filtering or exporting them. ai_jev does not require DuckDB's JSON extension.

For full probability distributions or custom instructions and shared state, use ai_provider_call:

SELECT ai_provider_call(
'{
"model": "jev-1.13.0",
"state": "I was charged twice. Please refund the duplicate.",
"questions": {
"refund_requested": {
"type": "noul",
"instructions": "Does the text explicitly request a refund?"
},
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments and refunds",
"other": "Other requests"
}
}
}
}',
secret := 'typesafe_ai'
);

Native requests specify model in the body, even when a secret contains a model. The result preserves the complete response, including answers, per-option probabilities, confidence, the resolved model version, and usage. The extension appends /systemone to the base URL automatically.

As checked on September 18, 2026, TypeSafe lists jev-1.13.0 with aliases jev-latest and jev-preview, at $0.042 per million input tokens and free output tokens. Pin the version when tuning decision thresholds. Choice supports up to 255 options, and Score accepts 2 to 10 ordered levels. A Score uses the level indices, so a three-level rubric produces values from 0 to 2.

The extension's mock tests check HTTP contracts, not Jev's live latency or prediction quality. Follow the typed Jev cookbook to reduce repeated state and calls. See TypeSafe's API reference, models and limits, and known limitations.

Ollama​

Use Ollama for local models without a hosted API key.

ollama serve
ollama pull qwen3.8:27b
ollama pull nomic-embed-text
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET ollama_ai (
TYPE duckdb_ai,
AI_PROVIDER 'ollama',
MODEL 'qwen3.8:27b'
);

SELECT ai_complete(
'Write one sentence about why DuckDB is useful for analytics.',
secret := 'ollama_ai',
timeout_seconds := 120
) AS answer;

SELECT ai_embed(
'DuckDB local analytics',
secret := 'ollama_ai',
model := 'nomic-embed-text'
)[1] AS first_embedding_value;

If Ollama runs on a non-default host, set OLLAMA_HOST before starting DuckDB or add BASE_URL 'http://host:11434' to the secret.

OpenAI​

export OPENAI_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET openai_ai (
TYPE duckdb_ai,
AI_PROVIDER 'openai',
MODEL 'gpt-5.6-luna'
);

SELECT ai_complete(
'Summarize DuckDB in one short sentence.',
secret := 'openai_ai'
) AS answer;

SELECT ai_complete_json(
'Return a JSON object with keys "name" and "kind" for DuckDB.',
secret := 'openai_ai',
response_schema := '{
"type": "object",
"properties": {
"name": {"type": "string"},
"kind": {"type": "string"}
},
"required": ["name", "kind"]
}'
) AS profile_json;

SELECT ai_embed(
'DuckDB vector search',
secret := 'openai_ai',
model := 'text-embedding-3-small'
)[1] AS first_embedding_value;

Azure OpenAI​

Use your Azure OpenAI resource URL and deployment name. The extension appends /openai/v1 to the base URL when needed.

export AZURE_OPENAI_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET azure_ai (
TYPE duckdb_ai,
AI_PROVIDER 'azure',
BASE_URL 'https://my-resource.openai.azure.com',
MODEL 'gpt-4o'
);

SELECT ai_complete(
'Explain DuckDB to a data analyst in one sentence.',
secret := 'azure_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'azure_ai',
model := 'text-embedding-3-small'
)[1] AS first_embedding_value;

If your Azure deployment names differ from the model names above, use the deployment name in MODEL or model := ....

Claude / Anthropic​

export ANTHROPIC_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET claude_ai (
TYPE duckdb_ai,
AI_PROVIDER 'anthropic',
MODEL 'claude-haiku-4-5'
);

SELECT ai_complete(
'Explain DuckDB in one concise paragraph.',
secret := 'claude_ai',
max_tokens := 160
) AS answer;

SELECT ai_summarize(
'DuckDB is an in-process analytical database built for fast local queries.',
secret := 'claude_ai',
max_tokens := 80
) AS summary;

SELECT ai_complete_json(
'Extract DuckDB as a database profile.',
secret := 'claude_ai',
response_schema := '{
"type": "object",
"properties": {"name": {"type": "string"}},
"required": ["name"],
"additionalProperties": false
}'
) AS profile;

For response_schema := ..., the extension sends Anthropic's output_config.format JSON Schema request shape. Claude is configured for completion calls; embeddings are not configured for this provider.

Gemini​

Gemini uses Google's OpenAI-compatible endpoint.

export GEMINI_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET gemini_ai (
TYPE duckdb_ai,
AI_PROVIDER 'gemini',
MODEL 'gemini-3.7-flash'
);

SELECT ai_complete(
'Write one sentence about DuckDB extensions.',
secret := 'gemini_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'gemini_ai',
model := 'gemini-embedding-2'
)[1] AS first_embedding_value;

Provider aliases gcp, google, and google_gemini also resolve to Gemini. Gemini 3.7 Flash and 3.6 Flash deprecate sampling parameters, so duckdb_ai omits temperature for these models even when the generic SQL option is provided. Built-in cost estimates use Google's introductory 3.7/3.6 Flash pricing through December 31, 2026 and roll over to the published standard rate on January 1, 2027.

Mistral​

export MISTRAL_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET mistral_ai (
TYPE duckdb_ai,
AI_PROVIDER 'mistral',
MODEL 'mistral-small-latest'
);

SELECT ai_complete(
'Give one practical use case for DuckDB and AI functions.',
secret := 'mistral_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'mistral_ai',
model := 'mistral-embed'
)[1] AS first_embedding_value;

Z.ai​

export ZAI_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET zai_ai (
TYPE duckdb_ai,
AI_PROVIDER 'zai',
MODEL 'glm-4.7-flash'
);

SELECT ai_complete(
'Summarize DuckDB in one sentence.',
secret := 'zai_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'zai_ai',
model := 'embedding-3'
)[1] AS first_embedding_value;

Provider aliases zai and zhipu resolve to the same provider.

DeepSeek​

export DEEPSEEK_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET deepseek_ai (
TYPE duckdb_ai,
AI_PROVIDER 'deepseek',
MODEL 'deepseek-v4-flash'
);

SELECT ai_complete(
'Write one sentence about using SQL with language models.',
secret := 'deepseek_ai'
) AS answer;

SELECT ai_classify(
'Customer says the invoice is overdue.',
'billing, support, sales, other',
secret := 'deepseek_ai'
) AS category;

DeepSeek is configured for completion calls. Embeddings are not configured for this provider. Built-in cost estimates use the published peak-hour cache-miss rates; DeepSeek charges 50% less during its documented off-peak windows.

OpenRouter​

export OPENROUTER_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET openrouter_ai (
TYPE duckdb_ai,
AI_PROVIDER 'openrouter',
MODEL 'openai/gpt-4o-mini'
);

SELECT ai_complete(
'Explain DuckDB extensions in one sentence.',
secret := 'openrouter_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'openrouter_ai',
model := 'openai/text-embedding-3-small'
)[1] AS first_embedding_value;

OpenRouter model names include the upstream provider prefix, for example openai/gpt-4o-mini.

Set OPENROUTER_HTTP_REFERER and OPENROUTER_X_TITLE to send OpenRouter's optional attribution headers.

Groq​

export GROQ_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET groq_ai (
TYPE duckdb_ai,
AI_PROVIDER 'groq',
MODEL 'openai/gpt-oss-20b'
);

SELECT ai_complete(
'Explain vectorized SQL execution in one sentence.',
secret := 'groq_ai'
) AS answer;

Groq is configured for completion calls. Embeddings are not configured for this provider.

Together AI​

export TOGETHER_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET together_ai (
TYPE duckdb_ai,
AI_PROVIDER 'together',
MODEL 'meta-llama/Llama-3.3-70B-Instruct-Turbo'
);

SELECT ai_complete(
'Write one sentence about local-first analytics.',
secret := 'together_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'together_ai',
model := 'intfloat/multilingual-e5-large-instruct'
)[1] AS first_embedding_value;

Fireworks AI​

export FIREWORKS_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET fireworks_ai (
TYPE duckdb_ai,
AI_PROVIDER 'fireworks',
MODEL 'accounts/fireworks/models/gpt-oss-20b'
);

SELECT ai_complete(
'Give one reason to enrich data inside SQL.',
secret := 'fireworks_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'fireworks_ai',
model := 'nomic-ai/nomic-embed-text-v1.5'
)[1] AS first_embedding_value;

Fireworks model IDs are passed through unchanged. This supports serverless base models such as accounts/fireworks/models/gpt-oss-20b, fast routers such as accounts/fireworks/routers/kimi-k2p6-turbo, account deployments, and embedding models such as fireworks/qwen3-embedding-8b. Use a model ID available to the configured Fireworks account.

DeepInfra​

export DEEPINFRA_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET deepinfra_ai (
TYPE duckdb_ai,
AI_PROVIDER 'deepinfra',
MODEL 'meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo'
);

SELECT ai_complete(
'Summarize why open-weight hosted inference is useful.',
secret := 'deepinfra_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'deepinfra_ai',
model := 'BAAI/bge-large-en-v1.5'
)[1] AS first_embedding_value;

Cerebras​

export CEREBRAS_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET cerebras_ai (
TYPE duckdb_ai,
AI_PROVIDER 'cerebras',
MODEL 'gpt-oss-120b'
);

SELECT ai_complete(
'Explain fast inference for analytical workflows in one sentence.',
secret := 'cerebras_ai'
) AS answer;

Cerebras is configured for completion calls. Embeddings are not configured for this provider.

Cohere​

export COHERE_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET cohere_ai (
TYPE duckdb_ai,
AI_PROVIDER 'cohere',
MODEL 'command-a-plus-05-2026'
);

SELECT ai_complete(
'Classify this support note as billing, support, sales, or other.',
secret := 'cohere_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'cohere_ai',
model := 'embed-v4.0'
)[1] AS first_embedding_value;

Cloudflare Workers AI​

Cloudflare Workers AI's OpenAI-compatible endpoint is account scoped. Set CLOUDFLARE_ACCOUNT_ID, or provide a full BASE_URL.

export CLOUDFLARE_API_KEY='...'
export CLOUDFLARE_ACCOUNT_ID='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET cloudflare_ai (
TYPE duckdb_ai,
AI_PROVIDER 'workers_ai',
MODEL '@cf/zai-org/glm-4.7-flash'
);

SELECT ai_complete(
'Explain edge-hosted inference in one sentence.',
secret := 'cloudflare_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'cloudflare_ai',
model := '@cf/baai/bge-base-en-v1.5'
)[1] AS first_embedding_value;

Aliases cloudflare, workers_ai, cloudflare_workers_ai, and cloudflare_ai resolve to the same provider.

Alibaba Cloud Model Studio / DashScope​

DashScope exposes Qwen models through an OpenAI-compatible endpoint. The default base URL uses the existing international compatible-mode endpoint; use a secret BASE_URL, DASHSCOPE_BASE_URL, or ALIBABA_MODEL_STUDIO_BASE_URL for a workspace-specific endpoint.

export DASHSCOPE_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET dashscope_ai (
TYPE duckdb_ai,
AI_PROVIDER 'qwen',
MODEL 'qwen-plus'
);

SELECT ai_complete(
'Write one sentence about analytics agents.',
secret := 'dashscope_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'dashscope_ai',
model := 'text-embedding-v4'
)[1] AS first_embedding_value;

Aliases dashscope, qwen, alibaba, alibaba_model_studio, and model_studio resolve to the same provider.

Nebius Token Factory​

export NEBIUS_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET nebius_ai (
TYPE duckdb_ai,
AI_PROVIDER 'nebius_token_factory',
MODEL 'meta-llama/Meta-Llama-3.1-70B-Instruct'
);

SELECT ai_complete(
'Summarize why managed open-weight inference matters.',
secret := 'nebius_ai'
) AS answer;

Nebius is configured for completion calls. Embeddings are not configured for this provider.

SambaNova Cloud​

export SAMBANOVA_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET sambanova_ai (
TYPE duckdb_ai,
AI_PROVIDER 'sambanova',
MODEL 'Meta-Llama-3.3-70B-Instruct'
);

SELECT ai_complete(
'Explain fast open-weight inference in one sentence.',
secret := 'sambanova_ai'
) AS answer;

Aliases sambanova, sambanova_ai, samba_nova, and sambacloud resolve to the same provider. Embeddings are not configured for this provider.

SiliconFlow​

export SILICONFLOW_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET siliconflow_ai (
TYPE duckdb_ai,
AI_PROVIDER 'siliconflow',
MODEL 'Qwen/Qwen2.5-72B-Instruct'
);

SELECT ai_complete(
'Write one sentence about open model hosting.',
secret := 'siliconflow_ai'
) AS answer;

Aliases siliconflow and silicon_flow resolve to the same provider. Embeddings are not configured for this provider.

Vercel AI Gateway​

export AI_GATEWAY_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET vercel_ai (
TYPE duckdb_ai,
AI_PROVIDER 'vercel_ai_gateway',
MODEL 'openai/gpt-4o-mini'
);

SELECT ai_complete(
'Explain model gateway routing in one sentence.',
secret := 'vercel_ai'
) AS answer;

SELECT ai_embed(
'DuckDB vector search',
secret := 'vercel_ai',
model := 'openai/text-embedding-3-small'
)[1] AS first_embedding_value;

Aliases vercel, vercel_ai_gateway, vercel_gateway, and ai_gateway resolve to the same provider.

Moonshot AI / Kimi​

export MOONSHOT_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET kimi_ai (
TYPE duckdb_ai,
AI_PROVIDER 'kimi',
MODEL 'kimi-k3'
);

SELECT ai_complete(
'Explain code-oriented model inference in one sentence.',
secret := 'kimi_ai'
) AS answer;

Aliases moonshot, kimi, moonshot_ai, and kimi_api resolve to the same provider. The current Kimi chat model IDs are passed through unchanged, including kimi-k3, kimi-k2.7-code, kimi-k2.7-code-highspeed, kimi-k2.6, and kimi-k2.5. Moonshot V1 model IDs are also accepted while they remain available to the configured account. Embeddings are not configured for this provider.

Baidu Qianfan / ERNIE​

export QIANFAN_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET qianfan_ai (
TYPE duckdb_ai,
AI_PROVIDER 'ernie',
MODEL 'ernie-4.5-turbo-128k'
);

SELECT ai_complete(
'Write one sentence about enterprise AI platforms.',
secret := 'qianfan_ai'
) AS answer;

Aliases qianfan, baidu, baidu_qianfan, ernie, and wenxin resolve to the same provider. Embeddings are not configured for this provider.

Tencent TokenHub / Hunyuan​

export TOKENHUB_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET hunyuan_ai (
TYPE duckdb_ai,
AI_PROVIDER 'tencent_hunyuan',
MODEL 'hy3'
);

SELECT ai_complete(
'Summarize governed model APIs in one sentence.',
secret := 'hunyuan_ai'
) AS answer;

Aliases hunyuan, tencent, and tencent_hunyuan resolve to the same provider. HUNYUAN_API_KEY and TENCENT_HUNYUAN_API_KEY remain accepted for existing configurations. Set TOKENHUB_BASE_URL for the international or a custom regional TokenHub endpoint. Embeddings are not configured for this provider.

StepFun​

export STEPFUN_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET stepfun_ai (
TYPE duckdb_ai,
AI_PROVIDER 'step',
MODEL 'step-3.5-flash'
);

SELECT ai_complete(
'Write one sentence about fast chat models.',
secret := 'stepfun_ai'
) AS answer;

Aliases stepfun, step, and step_fun resolve to the same provider. Embeddings are not configured for this provider.

MiniMax​

export MINIMAX_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET minimax_ai (
TYPE duckdb_ai,
AI_PROVIDER 'minimax',
MODEL 'MiniMax-M2.7'
);

SELECT ai_complete(
'Explain model choice for SQL enrichment in one sentence.',
secret := 'minimax_ai'
) AS answer;

Aliases minimax and mini_max resolve to the same provider. Embeddings are not configured for this provider.

Poe​

export POE_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET poe_ai (
TYPE duckdb_ai,
AI_PROVIDER 'poe',
MODEL 'GPT-5.4'
);

SELECT ai_complete(
'Summarize model routing in one sentence.',
secret := 'poe_ai'
) AS answer;

Aliases poe and poe_api resolve to the same provider. Embeddings are not configured for this provider. Poe's Chat Completions compatibility layer ignores response_format, so duckdb_ai rejects non-text response_format values and response_schema for this provider instead of promising unenforced structured output.

Volcengine Ark / Doubao​

export VOLCENGINE_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET doubao_ai (
TYPE duckdb_ai,
AI_PROVIDER 'doubao',
MODEL 'doubao-seed-2-1-pro-260628'
);

SELECT ai_complete(
'Explain hosted model APIs in one sentence.',
secret := 'doubao_ai'
) AS answer;

Aliases volcengine, volcano_engine, volcengine_ark, doubao, and ark resolve to the same provider. Embeddings are not configured for this provider.

Hugging Face​

export HF_TOKEN='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET huggingface_ai (
TYPE duckdb_ai,
AI_PROVIDER 'huggingface',
MODEL 'openai/gpt-oss-120b'
);

SELECT ai_complete(
'Write one sentence about open-source model routing.',
secret := 'huggingface_ai'
) AS answer;

Aliases hf, hugging_face, huggingface_hub, and hf_inference resolve to huggingface. Embeddings are not configured by default; use openai_compatible with an embedding-capable router endpoint if needed.

xAI / SpaceXAI​

export XAI_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET xai_ai (
TYPE duckdb_ai,
AI_PROVIDER 'xai',
MODEL 'grok-4.6'
);

SELECT ai_complete(
'Summarize this SQL migration risk in one sentence.',
secret := 'xai_ai'
) AS answer;

Aliases xai, x.ai, x-ai, and grok resolve to the same provider. Embeddings are not configured for this provider.

Perplexity​

export PERPLEXITY_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET perplexity_ai (
TYPE duckdb_ai,
AI_PROVIDER 'perplexity',
MODEL 'sonar'
);

SELECT ai_complete(
'Give one current consideration for managed model APIs.',
secret := 'perplexity_ai'
) AS answer;

Perplexity is configured for completion calls. Embeddings are not configured for this provider.

NVIDIA NIM​

export NVIDIA_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET nvidia_ai (
TYPE duckdb_ai,
AI_PROVIDER 'nvidia_nim',
MODEL 'nvidia/nemotron-3-super-120b-a12b'
);

SELECT ai_complete(
'Explain GPU-hosted inference in one sentence.',
secret := 'nvidia_ai'
) AS answer;

Aliases nvidia, nvidia_nim, and nim resolve to the same provider. Embeddings are not configured for this provider.

Amazon Bedrock​

Bedrock's OpenAI-compatible endpoint is regional. Set either a full base URL or a region; the extension derives https://bedrock-mantle.<region>.api.aws/v1 from the region.

export AWS_BEARER_TOKEN_BEDROCK='...'
export AWS_REGION='us-east-1'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET bedrock_ai (
TYPE duckdb_ai,
AI_PROVIDER 'bedrock',
MODEL 'openai.gpt-oss-120b'
);

SELECT ai_complete(
'Summarize why governed enterprise inference matters.',
secret := 'bedrock_ai'
) AS answer;

Aliases bedrock, aws_bedrock, amazon_bedrock, and bedrock_mantle resolve to the same provider. Embeddings are not configured by default.

Google Vertex AI​

Vertex AI's OpenAI-compatible endpoint is project and location scoped. Set a full base URL, or set GOOGLE_CLOUD_PROJECT and optionally GOOGLE_CLOUD_LOCATION.

export GOOGLE_CLOUD_PROJECT='my-project'
export GOOGLE_CLOUD_LOCATION='global'
export VERTEX_AI_ACCESS_TOKEN="$(gcloud auth print-access-token)"
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET vertex_ai (
TYPE duckdb_ai,
AI_PROVIDER 'vertex',
MODEL 'google/gemini-2.5-flash'
);

SELECT ai_complete(
'Explain BigQuery and DuckDB together in one sentence.',
secret := 'vertex_ai'
) AS answer;

Aliases vertex, google_vertex, vertex_ai, and gcp_vertex resolve to the same provider. Embeddings are not configured for this provider.

Databricks​

Databricks Model Serving exposes chat endpoints through an OpenAI-compatible API. Use a Databricks personal access token or service-principal token, and set the model to the serving endpoint name.

export DATABRICKS_TOKEN='...'
export DATABRICKS_HOST='https://<workspace>.cloud.databricks.com'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET databricks_ai (
TYPE duckdb_ai,
AI_PROVIDER 'databricks',
MODEL 'databricks-gpt-oss-120b'
);

SELECT ai_complete(
'Explain Delta Lake in one sentence.',
secret := 'databricks_ai'
) AS answer;

If BASE_URL is omitted, set DATABRICKS_HOST; the extension derives https://<workspace>/serving-endpoints. You can also set a secret BASE_URL, DATABRICKS_BASE_URL, or per-call base_url := ... to use a full /serving-endpoints, /ai-gateway/mlflow/v1, or /chat/completions endpoint. Aliases mosaic, mosaic_ai, and databricks_ai resolve to databricks.

Databricks reasoning models can return message.content as typed reasoning and text blocks. duckdb_ai returns the text blocks and ignores reasoning summaries. Databricks Claude Sonnet 5 and Claude Opus 5 model services reject sampling temperature, so the extension omits temperature for those model IDs even when the generic SQL option is set. Older Claude and other Databricks models retain explicit temperature support.

Snowflake Cortex REST​

Snowflake Cortex REST exposes a Chat Completions API compatible with the OpenAI request shape. Use a Snowflake Programmatic Access Token, OAuth token, or JWT with a role that can call Cortex REST.

export SNOWFLAKE_PAT='...'
export SNOWFLAKE_ACCOUNT_URL='https://<account-identifier>.snowflakecomputing.com'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET snowflake_ai (
TYPE duckdb_ai,
AI_PROVIDER 'snowflake',
MODEL 'claude-sonnet-4-5'
);

SELECT ai_complete(
'Summarize why governed model inference matters.',
secret := 'snowflake_ai'
) AS answer;

If BASE_URL is omitted, set SNOWFLAKE_ACCOUNT_URL, SNOWFLAKE_HOST, or SNOWFLAKE_ACCOUNT; the extension derives https://<account>.snowflakecomputing.com/api/v2/cortex/v1. You can also set a secret BASE_URL, SNOWFLAKE_BASE_URL, or per-call base_url := ... to use a full /api/v2/cortex/v1 or /chat/completions endpoint. Snowflake model IDs include values such as claude-sonnet-4-5, llama4-maverick, and llama3.3-70b, depending on region and model access.

OpenAI Privacy Filter​

OpenAI Privacy Filter is an open-weight PII detection and masking model. The extension calls it through a small HTTP service so the model can run either on the same machine as DuckDB or in your own cloud deployment.

Use local hosting when unredacted PII should not leave the machine running DuckDB:

# Host a wrapper around the openai/privacy-filter Python package.
# The wrapper should expose POST /redact.
export OPF_CHECKPOINT="$HOME/.opf/privacy_filter"
./serve-privacy-filter --host 127.0.0.1 --port 8080
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET privacy_filter_local (
TYPE duckdb_ai,
AI_PROVIDER 'openai_privacy_filter',
BASE_URL 'http://localhost:8080'
);

SELECT ai_redact(
'email alice@example.com token fake-token',
secret := 'privacy_filter_local'
) AS redacted_text;

Use cloud hosting when multiple DuckDB clients should share one managed Privacy Filter deployment:

CREATE OR REPLACE SECRET privacy_filter_cloud (
TYPE duckdb_ai,
AI_PROVIDER 'openai_privacy_filter',
BASE_URL 'https://privacy-filter.example.com',
API_KEY '...'
);

SELECT ai_redact(
internal_note,
secret := 'privacy_filter_cloud'
) AS redacted_note
FROM support_tickets;

The service contract is intentionally small:

POST /redact
Content-Type: application/json
Authorization: Bearer ... # optional

{"text":"email alice@example.com","model":"openai/privacy-filter"}

The response should include one of redacted_text, masked_text, text, or output:

{"redacted_text":"email [PRIVATE_EMAIL]"}

Aliases privacy_filter, pii_filter, and opf resolve to openai_privacy_filter. The default base URL is http://localhost:8080; set OPENAI_PRIVACY_FILTER_BASE_URL, DUCKDB_AI_BASE_URL, a secret BASE_URL, or a per-call base_url := ... to point at a different local or cloud service. For cloud auth, use a secret API_KEY, OPENAI_PRIVACY_FILTER_API_KEY, or DUCKDB_AI_API_KEY.

OpenAI-compatible / Local gateway​

Use this path for vLLM, LM Studio, LiteLLM, Ollama's /v1 endpoint, or any gateway that exposes an OpenAI-compatible chat API.

For local Ollama's OpenAI-compatible endpoint:

ollama serve
ollama pull qwen3.8:27b
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET local_openai_ai (
TYPE duckdb_ai,
AI_PROVIDER 'local',
BASE_URL 'http://localhost:11434/v1',
MODEL 'qwen3.8:27b'
);

SELECT ai_complete(
'Write one sentence about local AI models in DuckDB.',
secret := 'local_openai_ai'
) AS answer;

For a hosted OpenAI-compatible gateway, add an API key:

export OPENAI_COMPATIBLE_API_KEY='...'
./build/release/duckdb
LOAD ai;

CREATE OR REPLACE SECRET gateway_ai (
TYPE duckdb_ai,
AI_PROVIDER 'openai_compatible',
BASE_URL 'https://gateway.example/v1',
MODEL 'provider/model-name'
);

SELECT ai_complete(
'Write one sentence about DuckDB.',
secret := 'gateway_ai'
) AS answer;

Aliases local, openai_compatible, openai-compatible, local_openai, local-models, and local_models all use the OpenAI-compatible protocol.

llama.cpp​

Use the llamacpp provider (aliases llama.cpp, llama-cpp, llama_cpp, llama-server, llama_server) for a local llama.cpp llama-server. It uses the OpenAI-compatible protocol and defaults to http://localhost:8080/v1, so no base URL setup is needed for a default server:

llama-server -m model.gguf --embeddings
./build/release/duckdb
LOAD ai;
SET duckdb_ai_provider = 'llama.cpp';

SELECT ai_complete('Write one sentence about DuckDB.') AS answer;
SELECT ai_embed('DuckDB');

The request model defaults to default; llama-server ignores it and answers with its loaded model, so no model configuration is needed either. Set LLAMACPP_BASE_URL for a non-default host or port, and LLAMACPP_API_KEY when the server runs with --api-key. Embedding calls require starting llama-server with --embeddings.

response_schema := ... uses llama.cpp's direct response_format.schema request shape so llama-server can enforce the supplied JSON Schema.

For throughput, start llama-server with parallel slots (for example --parallel 4) and match max_concurrent_requests so DuckDB keeps every slot busy; llama.cpp reuses cached prompt prefixes automatically, so shared system prompts stay cheap across rows.