Local models v7

Local models run on your Postgres host. No external API is involved. AIDB registers a set of default models at install time, and you can register more with aidb.create_model(). A local model is used by name in any AIDB SQL function or pipeline step, the same way as a remote model.

Local models run on one of two inference runtimes, Candle or llama.cpp. Both run on the CPU. Usage is identical for both; the provider decides which runtime loads the model.

Model providers

A provider tells aidb.create_model() which runtime to use and which configuration to expect. Pick the row that matches your model's file format and the task you need:

ProviderRuntimeModel filesTaskConfigDetails
bert_localCandleHugging Face safetensorsText embeddingsaidb.bert_config()bert_local
clip_localCandleHugging Face safetensorsText and image embeddingsaidb.clip_config()clip_local
t5_localCandleHugging Face safetensorsText generation and text embeddingsaidb.t5_config()t5_local
llama_instruct_localCandleHugging Face safetensorsText generationaidb.llama_config()llama_instruct_local
llamacpp_generatellama.cppGGUFText generationjsonb_build_object()llamacpp_generate
llamacpp_embeddingsllama.cppGGUFText embeddingsjsonb_build_object()llamacpp_embeddings
llamacpp_rerankingllama.cppGGUFRerankingjsonb_build_object()llamacpp_reranking
llamacpp_ocrllama.cppGGUFOCR (text extraction from images)jsonb_build_object()llamacpp_ocr
dummynonenoneTesting: returns fake datanoneDefault models

For every local provider, the config key that names the Hugging Face repository is hf_model. The old key model still works but logs a deprecation warning. The Candle config helpers write the key as model; to avoid the warning, build the config with jsonb_build_object('hf_model', ...) instead.

To list the providers registered in your database:

SELECT server_name, server_description FROM aidb.model_providers ORDER BY server_name;

Default models

These models are registered in every AIDB installation. Use them by name without any setup. The model files download from Hugging Face the first time a model is used; see Local model downloads.

Model nameProviderDescription
bertbert_localGeneral-purpose English text embeddings. Good default for text search and RAG.
clipclip_localJoint text and image embeddings for multimodal search.
t5t5_localLightweight text-to-text generation: translation, summarization, question answering.
llamallama_instruct_localInstruction-following chat and completion model.
bge-small-en-v1.5-f16llamacpp_embeddingsCompact, fast English text embeddings.
nomic-embed-text-v1.5-Q8_0llamacpp_embeddingsGeneral-purpose English embeddings with a 2048-token context window.
bge-m3-f16llamacpp_embeddingsMultilingual embeddings with an 8192-token context window, for long or non-English documents.
qwen3-embedding-0.6b-Q8_0llamacpp_embeddingsLarger embedding model with a 16384-token context window.
qwen3-embedding-4b-Q8_0llamacpp_embeddingsLargest default embedding model, 20480-token context window. Higher quality, more compute.
qwen3.5-0.8b-Q8_0llamacpp_generateSmall, fast chat and completion model with a 262144-token context window.
llama-3.2-1b-instruct-Q8_0llamacpp_generateCompact Meta Llama 3.2 instruction model for lightweight chat and completion.
lightonocr-2-1b-Q8_0llamacpp_ocrCompact vision-language model that extracts text from images.
dummydummyReturns fake data, for testing pipelines without running inference.
-- Text embedding with the default BERT model
SELECT aidb.encode_text('bert', 'The quick brown fox');

-- Image embedding with the default CLIP model
SELECT aidb.encode_image('clip', pg_read_binary_file('/tmp/photo.jpg')::BYTEA);

-- Text generation with the default T5 model
SELECT aidb.generate_text('t5', 'Translate to French: Hello, world.');

-- Text embedding with a default llama.cpp embedding model
SELECT aidb.encode_text('bge-small-en-v1.5-f16', 'The quick brown fox');

-- Text generation with a default llama.cpp model
SELECT aidb.generate_text('qwen3.5-0.8b-Q8_0', 'Translate to French: Hello, world.');

Custom models

Register any other model with aidb.create_model(), giving it a name, a provider, and a config:

SELECT aidb.create_model(
    name     => 'my_model',
    provider => '<provider>',
    config   => <config>
);

The provider sections below list the models EDB has tested with each provider. You aren't limited to them: any Hugging Face model in a format the provider supports can be registered the same way.

Validation at creation

By default, aidb.create_model() downloads and loads the model files and runs a small test inference before registering the model. If the model identifier doesn't exist, a local path is wrong, or the weights can't be loaded, the error is reported and nothing is registered.

Pass validate => false to register the model without downloading or loading it. The download then happens on first use.

SELECT aidb.create_model('my_bert', 'bert_local',
    config   => aidb.bert_config('sentence-transformers/all-MiniLM-L6-v2'),
    validate => false);

To test a registered model later, call aidb.validate_model().

Candle providers

The four Candle providers load Hugging Face models in safetensors format. Each has a config helper that takes the Hugging Face model identifier as its first argument. The Postgres process needs network access to download the model files and write access to the cache directory. For air-gapped environments, download the files in advance and set cache_dir to their location.

bert_local

Text embeddings with aidb.encode_text(). Configure it with aidb.bert_config().

Tested models:

Model identifierDefault modelDescription
sentence-transformers/all-MiniLM-L6-v2bertGeneral-purpose English sentence embeddings. Fast and small.
sentence-transformers/paraphrase-multilingual-mpnet-base-v2Multilingual embeddings.
sentence-transformers/multi-qa-MiniLM-L6-cos-v1Tuned for question answering and retrieval.
sentence-transformers/paraphrase-TinyBERT-L6-v2Smaller and faster than all-MiniLM-L6-v2, at some quality cost.
SELECT aidb.create_model(
    'my_bert',
    'bert_local',
    config => aidb.bert_config('sentence-transformers/paraphrase-multilingual-mpnet-base-v2')
);

aidb.bert_config() parameters:

ParameterTypeDefaultDescription
modelTEXTRequiredHugging Face model identifier.
revisionTEXTNULLModel revision or branch.
cache_dirTEXTNULLLocal directory for caching model files.

clip_local

Text embeddings with aidb.encode_text() and image embeddings with aidb.encode_image(), in one vector space. Configure it with aidb.clip_config().

Tested models:

Model identifierDefault modelDescription
openai/clip-vit-base-patch32clipJoint text and image embeddings.
SELECT aidb.create_model(
    'my_clip',
    'clip_local',
    config => aidb.clip_config('openai/clip-vit-base-patch32')
);

aidb.clip_config() parameters:

ParameterTypeDefaultDescription
modelTEXTRequiredHugging Face model identifier.
revisionTEXTNULLModel revision or branch.
cache_dirTEXTNULLLocal directory for caching model files.
image_sizeINTEGERNULLInput image size in pixels.

t5_local

Text generation with aidb.generate_text() and text embeddings with aidb.encode_text(). Configure it with aidb.t5_config().

Tested models:

Model identifierDefault modelDescription
t5-smallt5Fastest, smallest T5 variant.
t5-baseLarger: better quality, slower.
t5-largeLargest: highest quality, most resource use.
SELECT aidb.create_model(
    'my_t5',
    't5_local',
    config => aidb.t5_config('google/flan-t5-base')
);

aidb.t5_config() parameters:

ParameterTypeDefaultDescription
modelTEXTRequiredHugging Face model identifier.
revisionTEXTNULLModel revision or branch.
temperatureDOUBLE PRECISIONNULLSampling temperature.
seedBIGINTNULLRandom seed for reproducible outputs.
max_tokensINTEGERNULLMaximum tokens to generate.
repeat_penaltyREALNULLPenalty applied to repeated tokens.
repeat_last_nINTEGERNULLNumber of trailing tokens considered for the repeat penalty.
model_pathTEXTNULLLocal path to model weights. Overrides model.
cache_dirTEXTNULLLocal directory for caching model files.
top_pDOUBLE PRECISIONNULLNucleus sampling threshold.

llama_instruct_local

Text generation with aidb.generate_text(). Configure it with aidb.llama_config().

Tested models:

Model identifierDefault modelDescription
TinyLlama/TinyLlama-1.1B-Chat-v1.0llamaSmall, fast instruction-following chat model.
HuggingFaceTB/SmolLM2-135M-InstructSmallest and fastest of these variants; lowest quality.
HuggingFaceTB/SmolLM2-360M-InstructSmall; better quality than the 135M variant.
HuggingFaceTB/SmolLM2-1.7B-InstructLargest of these variants; better quality, more resource use.
SELECT aidb.create_model(
    'my_llama',
    'llama_instruct_local',
    config => aidb.llama_config(
        'meta-llama/Llama-3.2-3B-Instruct',
        temperature => 0.5
    )
);

aidb.llama_config() parameters:

ParameterTypeDefaultDescription
modelTEXTRequiredHugging Face model identifier.
revisionTEXTNULLModel revision or branch.
cache_dirTEXTNULLLocal directory for caching model files.
system_promptTEXTNULLDefault system prompt.
use_flash_attentionBOOLEANNULLEnable flash attention.
model_pathTEXTNULLLocal path to model weights. Overrides model.
seedBIGINTNULLRandom seed for reproducible outputs.
temperatureDOUBLE PRECISIONNULLSampling temperature.
top_pDOUBLE PRECISIONNULLNucleus sampling threshold.
sample_lenINTEGERNULLMaximum tokens to generate.
use_kv_cacheBOOLEANNULLEnable KV cache.
repeat_penaltyREALNULLPenalty applied to repeated tokens.
repeat_last_nINTEGERNULLNumber of trailing tokens considered for the repeat penalty.

llama.cpp providers

The four llama.cpp providers load GGUF model files, a common format for quantized open-weight models. They have no config helper; build config with jsonb_build_object(). Name the model either by Hugging Face repository (hf_model plus model_file) or by a file on disk (local_path). Use aidb.max_threads to control how many threads llama.cpp uses.

llamacpp_generate

Text generation with aidb.generate_text(). Also supports tools, tool_choice, and response_format for tool calling and structured output; see Tool calling and structured output.

Tested models:

hf_modelmodel_fileDefault modelDescription
unsloth/Qwen3.5-0.8B-GGUFQwen3.5-0.8B-Q8_0.ggufqwen3.5-0.8b-Q8_0Small, fast chat model with a 262144-token context.
unsloth/Llama-3.2-1B-Instruct-GGUFLlama-3.2-1B-Instruct-Q8_0.ggufllama-3.2-1b-instruct-Q8_0Compact Meta Llama 3.2 instruction model.
SELECT aidb.create_model(
    name => 'my_llamacpp_chat',
    provider => 'llamacpp_generate',
    config => jsonb_build_object(
        'local_path', '/var/lib/aidb/models/qwen2.5-3b-instruct-q4_k_m.gguf',
        'n_ctx', 4096,
        'temperature', 0.2
    )
);

Config keys:

KeyTypeDefaultDescription
hf_modelTEXTRequired unless local_path is setHugging Face repo id, for example, Qwen/Qwen2-0.5B-Instruct-GGUF. model is a deprecated alias that logs a warning.
model_fileTEXTRequired unless local_path is setFilename of the .gguf inside the repo. For a split GGUF, name any one shard; every shard downloads and llama.cpp reassembles them.
revisionTEXTmainHugging Face revision.
local_pathTEXTNULLAbsolute path to a local .gguf file. Overrides hf_model/model_file/revision.
cache_dirTEXTNULLLocal directory for caching downloaded model files.
n_ctxINTEGER4096Context window, in tokens. Capped at the value the GGUF was trained with.
chat_templateTEXTNULLOverride the chat template baked into the GGUF.
temperatureDOUBLE PRECISION0.7Sampling temperature.
top_pDOUBLE PRECISION0.9Nucleus sampling threshold.
max_tokensINTEGER512Maximum tokens to generate.
seedBIGINTNULLRandom seed for reproducible outputs.
system_promptTEXTNULLSystem prompt prepended to every request.
repeat_penaltyREAL1.1Penalty applied to repeated tokens.
repeat_last_nINTEGER64Number of trailing tokens considered for the repeat penalty.
thinkingBOOLEANNULLfalse turns reasoning off and removes reasoning blocks from the output. true turns reasoning on for models that need a prompt to reason. Unset leaves the model's default.

llamacpp_embeddings

Text embeddings with aidb.encode_text().

Tested models:

hf_modelmodel_fileDefault modelDescription
unsloth/bge-small-en-v1.5-GGUFbge-small-en-v1.5-f16.ggufbge-small-en-v1.5-f16Compact, fast English embeddings.
nomic-ai/nomic-embed-text-v1.5-GGUFnomic-embed-text-v1.5.Q8_0.ggufnomic-embed-text-v1.5-Q8_0General-purpose English embeddings, 2048-token context.
CompendiumLabs/bge-m3-ggufbge-m3-f16.ggufbge-m3-f16Multilingual embeddings, 8192-token context.
Qwen/Qwen3-Embedding-0.6B-GGUFQwen3-Embedding-0.6B-Q8_0.ggufqwen3-embedding-0.6b-Q8_0Larger embedding model, 16384-token context.
Qwen/Qwen3-Embedding-4B-GGUFQwen3-Embedding-4B-Q8_0.ggufqwen3-embedding-4b-Q8_0Largest default embedding model, 20480-token context.
SELECT aidb.create_model(
    name => 'my_llamacpp_embed',
    provider => 'llamacpp_embeddings',
    config => jsonb_build_object(
        'hf_model', 'nomic-ai/nomic-embed-text-v1.5-GGUF',
        'model_file', 'nomic-embed-text-v1.5.Q4_K_M.gguf',
        'revision', '0188c9bf409793f810680a5a431e7b899c46104c',
        'n_ctx', 2048,
        'normalize', true
    )
);

Config keys:

KeyTypeDefaultDescription
hf_modelTEXTRequired unless local_path is setHugging Face repo id, for example, nomic-ai/nomic-embed-text-v1.5-GGUF. model is a deprecated alias that logs a warning.
model_fileTEXTRequired unless local_path is setFilename of the .gguf inside the repo. For a split GGUF, name any one shard; every shard downloads and llama.cpp reassembles them.
revisionTEXTmainHugging Face revision.
local_pathTEXTNULLAbsolute path to a local .gguf file. Overrides hf_model/model_file/revision.
cache_dirTEXTNULLLocal directory for caching downloaded model files.
n_ctxINTEGER2048Context window, in tokens. Capped at the value the GGUF was trained with.
poolingTEXTunspecifiedPooling strategy: unspecified (use the GGUF's own metadata), none, mean, cls, or last.
normalizeBOOLEANtrueL2-normalize output embedding vectors.
query_prefixTEXTNULLText prepended to every input embedded as a query, for example, search_query: for nomic-embed-text.
document_prefixTEXTNULLText prepended to every input embedded as a document, for example, search_document: for nomic-embed-text.

llamacpp_reranking

Reranking with aidb.rerank_text(). Supports two reranker architectures. The architecture is read from the GGUF's general.architecture metadata: BERT and XLM-RoBERTa family models score as a cross-encoder, and the qwen3 architecture scores as decoder-only using the Qwen3-Reranker yes/no verdict token. Other architectures are rejected when the model loads.

Tested models:

hf_modelmodel_fileArchitectureDescription
gpustack/bge-reranker-v2-m3-GGUFbge-reranker-v2-m3-Q4_K_M.ggufCross-encoderMultilingual reranker. Scores each query and document pair with a classification head.
ggml-org/Qwen3-Reranker-0.6B-Q8_0-GGUFqwen3-reranker-0.6b-q8_0.ggufDecoder-onlyQwen3-Reranker family. Scores each pair from the model's yes/no verdict token.

Neither is registered as a default model.

-- Cross-encoder
SELECT aidb.create_model(
    name => 'my_llamacpp_reranker',
    provider => 'llamacpp_reranking',
    config => jsonb_build_object(
        'hf_model', 'gpustack/bge-reranker-v2-m3-GGUF',
        'model_file', 'bge-reranker-v2-m3-Q4_K_M.gguf',
        'revision', '3093af03b1a635e67b084b1d8c03c5f5e020fd05',
        'n_ctx', 2048
    )
);

-- Decoder-only, with a custom task instruction
SELECT aidb.create_model(
    name => 'my_llamacpp_reranker_decoder',
    provider => 'llamacpp_reranking',
    config => jsonb_build_object(
        'hf_model', 'ggml-org/Qwen3-Reranker-0.6B-Q8_0-GGUF',
        'model_file', 'qwen3-reranker-0.6b-q8_0.gguf',
        'revision', 'a02f48bb4f057028298c21fa033da2b30d7742d5',
        'n_ctx', 2048,
        'instruction', 'Retrieve legal passages relevant to the query'
    )
);

Config keys:

KeyTypeDefaultDescription
hf_modelTEXTRequired unless local_path is setHugging Face repo id, for example, gpustack/bge-reranker-v2-m3-GGUF. model is a deprecated alias that logs a warning.
model_fileTEXTRequired unless local_path is setFilename of the .gguf inside the repo. For a split GGUF, name any one shard; every shard downloads and llama.cpp reassembles them.
revisionTEXTmainHugging Face revision.
local_pathTEXTNULLAbsolute path to a local .gguf file. Overrides hf_model/model_file/revision.
cache_dirTEXTNULLLocal directory for caching downloaded model files.
n_ctxINTEGER2048Context window, in tokens. A query and document pair that doesn't fit is rejected at call time.
reranker_typeTEXTNULL (read from the GGUF)Force the scoring mechanism: cross_encoder or decoder_only. Set it only if the detected architecture is wrong.
instructionTEXTBuilt-in generic retrieval instructionTask instruction for decoder-only rerankers. Ignored by cross-encoders.
Note

llamacpp_reranking returns a raw model logit as logit_score, not a normalized 0 to 1 relevance score. Compare scores within one aidb.rerank_text() call, not across providers or calls.

llamacpp_ocr

Extracts text from images with aidb.perform_ocr() and in PerformOcr pipeline steps. Needs a vision-language GGUF model plus its multimodal projector (mmproj) file.

Tested models:

hf_modelmodel_filemmproj_fileDefault modelDescription
ggml-org/LightOnOCR-2-1B-GGUFLightOnOCR-2-1B-Q8_0.ggufmmproj-LightOnOCR-2-1B-Q8_0.gguflightonocr-2-1b-Q8_0Compact vision-language model for OCR.
SELECT aidb.create_model(
    name => 'my_llamacpp_ocr',
    provider => 'llamacpp_ocr',
    config => jsonb_build_object(
        'hf_model', 'ggml-org/LightOnOCR-2-1B-GGUF',
        'model_file', 'LightOnOCR-2-1B-Q8_0.gguf',
        'mmproj_file', 'mmproj-LightOnOCR-2-1B-Q8_0.gguf',
        'revision', '7d4daf70c2856a155b1413def00f4ae59843a9cf',
        'n_ctx', 8192
    )
);

Config keys:

KeyTypeDefaultDescription
hf_modelTEXTRequired unless local_path is setHugging Face repo id, for example, ggml-org/LightOnOCR-2-1B-GGUF. model is a deprecated alias that logs a warning.
model_fileTEXTRequired unless local_path is setFilename of the base model .gguf inside the repo.
mmproj_fileTEXTRequired unless mmproj_local_path is setFilename of the multimodal projector .gguf, in the same repo and revision as the base model.
mmproj_local_pathTEXTNULLAbsolute path to a local mmproj .gguf file. Overrides mmproj_file.
revisionTEXTmainHugging Face revision.
local_pathTEXTNULLAbsolute path to a local base model .gguf file. Overrides hf_model/model_file/revision.
cache_dirTEXTNULLLocal directory for caching downloaded model files.
n_ctxINTEGER8192Context window, in tokens.
promptTEXTBuilt-in transcription instructionOCR instruction sent with each image. How much it influences the output depends on the model.
temperatureDOUBLE PRECISION0.0Sampling temperature.

Local model downloads

When a local model is used for the first time, or its cached files are missing or corrupt, AIDB downloads the model files from Hugging Face into the configured cache_dir.

Progress messages

For each file, AIDB emits a downloading line when the request starts and a complete line when it finishes. For large files, a progress line reports bytes and percent complete about every 30 seconds.

SELECT aidb.generate_text('t5', 'test');
Output
NOTICE:  [t5_local] loading t5-small (rev refs/pr/15) -> /var/.../snapshots/refs/pr/15
NOTICE:  [hf] downloading https://huggingface.co/t5-small/resolve/refs%2Fpr%2F15/config.json (expected: 1.2 KB)
NOTICE:  [hf] complete: https://huggingface.co/t5-small/resolve/refs%2Fpr%2F15/config.json (1.2 KB) in 0.9s
NOTICE:  [hf] downloading https://huggingface.co/t5-small/resolve/refs%2Fpr%2F15/model.safetensors (size unknown)
NOTICE:  [hf] downloading t5-small/model.safetensors: 115.4 MB / 230.8 MB (50.0%)
NOTICE:  [hf] downloading t5-small/model.safetensors: 216.4 MB / 230.8 MB (93.7%)
NOTICE:  [hf] complete: https://huggingface.co/t5-small/resolve/refs%2Fpr%2F15/model.safetensors (230.8 MB) in 83.5s

These lines go to the client at NOTICE level. Use aidb.download_log_level to send them to the server log instead, or to turn them off.

Retries and resume

If a download is interrupted, AIDB retries it. Each retry resumes from where the previous attempt stopped. Only attempts that make no progress count toward aidb.download_max_attempts (default 30), with exponential backoff between attempts. A Ctrl-C from psql takes effect between retries.

Verification and self-heal

AIDB checks each downloaded file against the size and, when available, the SHA-256 hash published by Hugging Face, and records them in a sidecar file next to the download. On every later use, it rechecks the file size against the sidecar. If a cached file is found corrupt when a model loads, AIDB deletes that snapshot and downloads it again once.

Using a registered model

A registered model is referenced by name in any AIDB function. Custom and default models work the same way:

-- Use a custom BERT model for embedding
SELECT aidb.encode_text('my_bert', 'Hello world');

-- Use a custom BERT model in a pipeline step
SELECT aidb.create_pipeline(
    name => 'my_pipeline',
    source => 'my_source_table',
    source_key_column => 'id',
    source_data_column => 'content',
    step_1 => 'KnowledgeBase',
    step_1_options => aidb.knowledge_base_config(model => 'my_bert', data_format => 'Text')
);

Some functions, such as aidb.generate_text() and aidb.summarize_text(), also accept a per-call configuration. See Configuration model for creation-time and call-time settings, and the Models reference for every config helper.

Troubleshooting

llama.cpp model loading and inference issues

If a llama.cpp model fails to load, times out, or behaves unexpectedly, turn on llama.cpp's own diagnostic logging. It writes GGUF metadata, tensor shapes, and backend and model initialization details to the server log:

aidb.enable_llamacpp_logs = true

The parameter is read once when the extension initializes, so restart Postgres after changing it. The output is verbose, hundreds of lines per model load, so turn it on only while investigating and off again afterwards. It applies only to llama.cpp models. See aidb.enable_llamacpp_logs.