Remote models v7

Remote models run outside your Postgres host and are called over an API. You register one with aidb.create_model(), giving it a name, a provider, a configuration object, and credentials. After that, it's used by name in any AIDB SQL function or pipeline step, the same way as a local model.

Model providers

A provider tells aidb.create_model() which API to call and which configuration to expect. Pick the row that matches the service you're connecting to and the task you need:

ProviderServiceTaskConfig helperDetails
openai_embeddingsOpenAI, or any server with an OpenAI-compatible embeddings API (for example, Ollama)Text embeddingsaidb.embeddings_config()OpenAI-compatible endpoints
openai_completionsOpenAI, or any server with an OpenAI-compatible chat completions APIText generationaidb.completions_config()OpenAI-compatible endpoints
openai_responsesOpenAI Responses APIText generation, agents with native tool callingaidb.openai_responses_config()OpenAI Responses API
openai_responses_azureAzure AI Foundry, Responses APIText generation, agents with native tool callingaidb.openai_responses_config()OpenAI Responses API
anthropic_messagesAnthropic Messages APIText generation, agents with native tool callingaidb.anthropic_messages_config()Anthropic Messages API
anthropic_messages_azureAzure AI Foundry, Anthropic Messages APIText generation, agents with native tool callingaidb.anthropic_messages_config()Anthropic Messages API
anthropic_messages_bedrockAWS Bedrock, Anthropic modelsText generation, agents with native tool callingaidb.anthropic_messages_config()Anthropic Messages API
nim_completionsNVIDIA NIMText generationaidb.completions_config()NVIDIA NIM
nim_embeddingsNVIDIA NIMText embeddingsaidb.embeddings_config()NVIDIA NIM
nim_clipNVIDIA NIMText and image embeddingsaidb.nim_clip_config()NVIDIA NIM
nim_paddle_ocrNVIDIA NIMOCR (text extraction from images)aidb.nim_ocr_config()NVIDIA NIM
nim_rerankingNVIDIA NIMRerankingaidb.nim_reranking_config()NVIDIA NIM
geminiGoogle GeminiText generationaidb.gemini_config()Google Gemini
openrouter_chatOpenRouterText generationaidb.openrouter_chat_config()OpenRouter
openrouter_embeddingsOpenRouterText embeddingsaidb.openrouter_embeddings_config()OpenRouter
embeddingsLegacy. Same as openai_embeddings.Text embeddingsaidb.embeddings_config()Legacy providers
completionsLegacy. Same as openai_completions.Text generationaidb.completions_config()Legacy providers

For OpenAI models, openai_responses and openai_completions accept the same model identifiers. Use openai_responses for agents and tool calling. Use openai_completions for plain text generation. For Claude models, use anthropic_messages.

To list the providers registered in your database:

SELECT server_name, server_description FROM aidb.model_providers ORDER BY server_name;

Registering a remote model

Every remote model is registered the same way:

SELECT aidb.create_model(
    name        => 'my_openai_embedder',
    provider    => 'openai_embeddings',
    config      => aidb.embeddings_config(model => 'text-embedding-3-small'),
    credentials => '{"api_key": "sk-..."}'::JSONB
);
  • config holds the model identifier and provider settings. Build it with the provider's config helper from the table above.
  • credentials holds the API key or basic-auth pair. config must not contain api_key or basic_auth; aidb.create_model() rejects it. To keep the secret out of the database, pass credentials_env or credentials_k8s_secret instead. See aidb.create_model.

Validation at creation

By default, aidb.create_model() sends a small test request to the provider before registering the model. If the request fails, for example because of a wrong API key or URL, the error is reported and nothing is registered. The request needs network access from the Postgres host and can incur a small usage cost.

Pass validate => false to skip the test request:

SELECT aidb.create_model(
    name        => 'my_openai_embedder',
    provider    => 'openai_embeddings',
    config      => aidb.embeddings_config(model => 'text-embedding-3-small'),
    credentials => '{"api_key": "sk-..."}'::JSONB,
    validate    => false
);

To test a registered model later, call aidb.validate_model().

OpenAI-compatible endpoints

These two providers work with OpenAI and with any server that implements the same API, for example Ollama or vLLM. The url parameter selects the server. Omit it for OpenAI.

openai_embeddings

Text embeddings through the /v1/embeddings API. Configure it with aidb.embeddings_config().

Models: any embedding model the server offers. For OpenAI, for example, text-embedding-3-small and text-embedding-3-large. For Ollama, any embedding model you've pulled, for example, nomic-embed-text.

-- OpenAI
SELECT aidb.create_model(
    'my_openai_embedder',
    'openai_embeddings',
    config      => aidb.embeddings_config(model => 'text-embedding-3-small'),
    credentials => '{"api_key": "sk-..."}'::JSONB
);

-- Ollama on your own server
SELECT aidb.create_model(
    'my_ollama_embedder',
    'openai_embeddings',
    config => aidb.embeddings_config(
        model => 'nomic-embed-text',
        url   => 'http://llama.local:11434/v1/embeddings'
    )
);

aidb.embeddings_config() parameters:

ParameterTypeDefaultDescription
modelTEXTRequiredModel identifier as expected by the API.
urlTEXTNULLAPI endpoint URL. Defaults to OpenAI's endpoint.
max_concurrent_requestsINTEGERNULLMaximum concurrent requests to the endpoint.
max_batch_sizeINTEGERNULLMaximum number of inputs per batch request.
input_typeTEXTNULLInput type sent when embedding documents. Provider-specific.
input_type_queryTEXTNULLInput type sent when embedding queries. Provider-specific.
is_hcp_modelBOOLEANNULLSet to true if the model runs on Hybrid Manager.
Note

pgvector can't index vectors with more than 2000 dimensions. If your model produces more, use aidb.vector_index_disabled_config() in the pipeline step and manage the index yourself. See the pgvector documentation.

openai_completions

Text generation through the /v1/chat/completions API. Configure it with aidb.completions_config().

Models: any chat model the server offers. For OpenAI, for example, gpt-4o or gpt-5.4-mini. For Ollama, any chat model you've pulled, for example, llama3.2.

aidb.generate_text() rejects tools, tool_choice, and response_format for this provider. When an agent uses this provider, AIDB describes the tools in the prompt text and parses tool calls out of the model's text reply. For native tool calling, use openai_responses.

SELECT aidb.create_model(
    'my_openai_llm',
    'openai_completions',
    config      => aidb.completions_config(
        model       => 'gpt-4o',
        temperature => 0.2
    ),
    credentials => '{"api_key": "sk-..."}'::JSONB
);

aidb.completions_config() parameters:

ParameterTypeDefaultDescription
modelTEXTRequiredModel identifier as expected by the API.
urlTEXTNULLAPI endpoint URL. Defaults to OpenAI's endpoint.
temperatureDOUBLE PRECISIONNULLSampling temperature.
top_pDOUBLE PRECISIONNULLNucleus sampling threshold.
seedBIGINTNULLRandom seed for reproducible outputs.
system_promptTEXTNULLSystem prompt prepended to every request.
max_tokensJSONBNULLMax tokens config, from aidb.max_tokens_config().
thinkingBOOLEANNULLfalse turns model reasoning off; true turns it on; NULL leaves the model's default.
max_concurrent_requestsINTEGERNULLMaximum concurrent requests to the endpoint.
extra_argsJSONBNULLAdditional provider-specific request fields.
is_hcp_modelBOOLEANNULLSet to true if the model runs on Hybrid Manager.

OpenAI Responses API

openai_responses

Text generation and agents through OpenAI's /v1/responses API. Configure it with aidb.openai_responses_config(). Tool calling is native: tools and tool_choice are sent as real request fields and tool calls are parsed from the API's tool-call responses. response_format gives structured output.

Models: the same identifiers as openai_completions. Commonly used families:

Model identifierDescription
gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-lunaThe GPT-5.6 family, in three tiers of capability, latency, and cost.
gpt-5.4-miniSmaller, lower-cost tier of the GPT-5.4 generation with good tool-calling accuracy.
gpt-4oLong-established multimodal model.

This isn't a full catalog. Check OpenAI's model documentation for the current lineup.

SELECT aidb.create_model(
    'my_gpt',
    'openai_responses',
    config      => aidb.openai_responses_config(model => 'gpt-5.1'),
    credentials => '{"api_key": "sk-..."}'::JSONB
);

See aidb.openai_responses_config for all parameters, and Tool calling and structured output for tools, tool_choice, and response_format.

openai_responses_azure

The same API and config helper as openai_responses, for models hosted on Azure AI Foundry. url is required; set it to your resource's Responses endpoint.

Models: the OpenAI models your Azure resource has deployed. Availability on Azure can lag behind OpenAI's own API for new models.

SELECT aidb.create_model(
    'my_gpt_azure',
    'openai_responses_azure',
    config      => aidb.openai_responses_config(
        model => 'gpt-5.1',
        url   => 'https://<resource>.openai.azure.com/openai/v1/responses'
    ),
    credentials => '{"api_key": "..."}'::JSONB
);

Anthropic Messages API

anthropic_messages

Text generation and agents through Anthropic's /v1/messages API. Configure it with aidb.anthropic_messages_config(). Tool calling is native, as with openai_responses.

Models: commonly used Claude models:

Model identifierDescription
claude-opus-4-6Flagship model for demanding reasoning and agentic tool-calling tasks.
claude-sonnet-5Mid-tier model for general-purpose agents and everyday text generation.
claude-haiku-4-5Fastest, lowest-cost current tier, for high-volume or latency-sensitive calls.
claude-3-5-haikuLong-established, widely used variant.

This isn't a full catalog. Check Anthropic's model documentation for the current lineup.

SELECT aidb.create_model(
    'my_claude',
    'anthropic_messages',
    config      => aidb.anthropic_messages_config(model => 'claude-opus-4-6'),
    credentials => '{"api_key": "sk-ant-..."}'::JSONB
);
Note

The Messages API has no structured-output field. When you pass response_format, AIDB sends it as a forced call to an internal tool. Because of this, response_format can't be combined with tools or tool_choice in the same call.

See aidb.anthropic_messages_config for all parameters.

anthropic_messages_azure and anthropic_messages_bedrock

The same API body and config helper as anthropic_messages, for Claude models hosted on Azure AI Foundry or AWS Bedrock. url is required for both.

  • anthropic_messages_azure: set url to your resource's Messages endpoint.
  • anthropic_messages_bedrock: set url to the regional bedrock-runtime base endpoint. AIDB appends the model ID to the path. Authentication uses a Bedrock API key as a bearer token, not AWS SigV4.

Models: the Claude models your Azure resource or Bedrock region has enabled. On Bedrock, use Bedrock's model IDs, for example, anthropic.claude-opus-4-6-v1:0. Availability on Azure and Bedrock can lag behind Anthropic's own API for new models.

-- Azure AI Foundry
SELECT aidb.create_model(
    'my_claude_azure',
    'anthropic_messages_azure',
    config      => aidb.anthropic_messages_config(
        model => 'claude-opus-4-6',
        url   => 'https://<resource>.services.ai.azure.com/anthropic/v1/messages'
    ),
    credentials => '{"api_key": "..."}'::JSONB
);

-- AWS Bedrock
SELECT aidb.create_model(
    'my_claude_bedrock',
    'anthropic_messages_bedrock',
    config      => aidb.anthropic_messages_config(
        model => 'anthropic.claude-opus-4-6-v1:0',
        url   => 'https://bedrock-runtime.us-east-1.amazonaws.com'
    ),
    credentials => '{"api_key": "..."}'::JSONB
);

NVIDIA NIM

Five providers connect to NVIDIA NIM microservices, either hosted on build.nvidia.com or running in your own environment. Omit url for build.nvidia.com. Set url to your own NIM endpoint otherwise.

To get an API key for build.nvidia.com, create an account there, select a model, and generate a key from the model's page.

nim_completions

Text generation. Configure it with aidb.completions_config(), the same helper as openai_completions. See openai_completions for the parameters.

Models: any NIM chat model, for example, meta/llama-3.3-70b-instruct or nvidia/llama-3.3-nemotron-super-49b-v1.

SELECT aidb.create_model(
    'my_nim_llm',
    'nim_completions',
    config      => aidb.completions_config(model => 'meta/llama-3.3-70b-instruct'),
    credentials => '{"api_key": "nvapi-..."}'::JSONB
);

SELECT aidb.generate_text('my_nim_llm', 'Tell me a short, one sentence story');

nim_embeddings

Text embeddings. Configure it with aidb.embeddings_config(), the same helper as openai_embeddings. See openai_embeddings for the parameters.

Models: any NIM text embedding model, for example, nvidia/llama-3.2-nv-embedqa-1b-v2 or nvidia/llama-3.2-nemoretriever-300m-embed-v1.

NIM embedding models take an input type with each request. input_type is sent when embedding documents and defaults to passage; allowed values are passage, query, and value. input_type_query is sent when embedding queries and defaults to query; allowed values are query and value.

SELECT aidb.create_model(
    'my_nim_embedder',
    'nim_embeddings',
    config      => aidb.embeddings_config(model => 'nvidia/llama-3.2-nv-embedqa-1b-v2'),
    credentials => '{"api_key": "nvapi-..."}'::JSONB
);

nim_clip

Text and image embeddings in one vector space. Configure it with aidb.nim_clip_config().

Models: nvidia/nvclip.

SELECT aidb.create_model(
    'my_nim_clip',
    'nim_clip',
    config      => aidb.nim_clip_config(model => 'nvidia/nvclip'),
    credentials => '{"api_key": "nvapi-..."}'::JSONB
);

aidb.nim_clip_config() parameters:

ParameterTypeDefaultDescription
modelTEXTNULLNIM CLIP model identifier.
urlTEXTNULLNIM endpoint URL. Defaults to build.nvidia.com.
is_hcp_modelBOOLEANNULLSet to true if the model runs on Hybrid Manager.

nim_paddle_ocr

OCR: extracts text from images. Configure it with aidb.nim_ocr_config(). The model identifier is optional.

SELECT aidb.create_model(
    'my_nim_ocr',
    'nim_paddle_ocr',
    config      => aidb.nim_ocr_config(),
    credentials => '{"api_key": "nvapi-..."}'::JSONB
);

aidb.nim_ocr_config() parameters:

ParameterTypeDefaultDescription
modelTEXTNULLNIM OCR model identifier.
urlTEXTNULLNIM endpoint URL. Defaults to build.nvidia.com.
is_hcp_modelBOOLEANNULLSet to true if the model runs on Hybrid Manager.

nim_reranking

Reranking: scores candidate texts against a query. Configure it with aidb.nim_reranking_config(). Use the model with aidb.rerank_text().

Models: any NIM reranking model, for example, nvidia/nv-rerankqa-mistral-4b-v3 or nvidia/llama-3.2-nv-rerankqa-1b-v2.

SELECT aidb.create_model(
    'my_reranker',
    'nim_reranking',
    config      => aidb.nim_reranking_config(model => 'nvidia/nv-rerankqa-mistral-4b-v3'),
    credentials => '{"api_key": "nvapi-..."}'::JSONB
);

aidb.nim_reranking_config() parameters:

ParameterTypeDefaultDescription
modelTEXTNULLNIM reranking model identifier.
urlTEXTNULLNIM endpoint URL. Defaults to build.nvidia.com.
is_hcp_modelBOOLEANNULLSet to true if the model runs on Hybrid Manager.

Google Gemini

gemini

Text generation with Google's Gemini models. Configure it with aidb.gemini_config().

Models: any Gemini model, for example, gemini-2.0-flash.

api_key is the helper's first parameter and has no default. Pass NULL for it and supply the key through credentials.

SELECT aidb.create_model(
    'my_gemini',
    'gemini',
    config      => aidb.gemini_config(
        api_key => NULL,
        model   => 'gemini-2.0-flash'
    ),
    credentials => '{"api_key": "AIza..."}'::JSONB
);

aidb.gemini_config() parameters:

ParameterTypeDefaultDescription
api_keyTEXTRequiredPass NULL. Supply the key through credentials instead.
modelTEXTNULLGemini model identifier.
urlTEXTNULLAPI endpoint URL override.
max_concurrent_requestsINTEGERNULLMaximum concurrent requests to the API.
thinking_budgetINTEGERNULLToken budget for extended thinking. Gemini 2.x models only.

OpenRouter

OpenRouter is a gateway to models from many vendors under one API and one API key. Model identifiers use OpenRouter's own slugs, for example, openai/gpt-4o or anthropic/claude-opus-4-6.

openrouter_chat

Text generation. Configure it with aidb.openrouter_chat_config(). Tool calling works the same way as with openai_completions: aidb.generate_text() rejects tools, and agents get tool calling through the prompt text.

Models: any OpenRouter chat model, for example, anthropic/claude-3-5-haiku or openai/gpt-4o.

SELECT aidb.create_model(
    'my_or_chat',
    'openrouter_chat',
    config      => aidb.openrouter_chat_config('anthropic/claude-3-5-haiku'),
    credentials => '{"api_key": "sk-or-..."}'::JSONB
);

aidb.openrouter_chat_config() parameters:

ParameterTypeDefaultDescription
modelTEXTRequiredOpenRouter model identifier.
urlTEXTNULLAPI endpoint URL override.
max_concurrent_requestsINTEGERNULLMaximum concurrent requests.
max_tokensJSONBNULLMax tokens config, from aidb.max_tokens_config().

openrouter_embeddings

Text embeddings. Configure it with aidb.openrouter_embeddings_config().

Models: any OpenRouter embedding model, for example, mistral/mistral-embed.

SELECT aidb.create_model(
    'my_or_embedder',
    'openrouter_embeddings',
    config      => aidb.openrouter_embeddings_config('mistral/mistral-embed'),
    credentials => '{"api_key": "sk-or-..."}'::JSONB
);

aidb.openrouter_embeddings_config() parameters:

ParameterTypeDefaultDescription
modelTEXTRequiredOpenRouter embeddings model identifier.
urlTEXTNULLAPI endpoint URL override.
max_concurrent_requestsINTEGERNULLMaximum concurrent requests.
max_batch_sizeINTEGERNULLMaximum inputs per batch request.

Legacy providers

aidb.model_providers also lists two providers from earlier AIDB versions:

ProviderUse instead
embeddingsopenai_embeddings
completionsopenai_completions

They behave the same as their openai_* counterparts and take the same config helpers. Models registered with them keep working. Use the openai_* names for new models.

See the Models reference for full details on every config helper.