Skip to content

AI Model Selection ​

Last Updated: September 2026

Choose the AI models that best fit your needs - from powerful cloud models to privacy-focused local options. Scrapalot supports OpenAI, Anthropic, Google Gemini, DeepSeek, Groq, OpenRouter, Z.ai, local runtimes (Ollama, LM Studio, vLLM, GGUF) and any OpenAI-compatible provider.

Your Model, Your Choice ​

Scrapalot is model-agnostic. Use any combination of:

  • Cloud models for best quality and speed
  • Local models for privacy and cost savings
  • Self-hosted models for complete control
  • Mix and match different models for different tasks

Cloud Model Providers ​

Model names are not listed here on purpose

Providers rename and retire models every few months, so any list written into a docs page is wrong within a release or two. Settings → AI Providers shows the models your configured providers actually offer right now, fetched live — that is the authoritative list. This page describes what each provider is good at, which changes far more slowly.

Scrapalot talks to providers, not to a fixed set of models. Add a provider with your own key and every model it exposes becomes selectable.

ProviderReach for it when
OpenAIGeneral-purpose quality, the widest tooling ecosystem, and embeddings
AnthropicVery long documents, careful instruction-following, detailed explanation
GoogleMultimodal input and large context windows
DeepSeekStrong quality per unit cost — a common default for high-volume work
GroqLatency-critical work; very fast inference
OpenRouterOne key, many vendors — useful for trying models before committing
Z.aiGLM-family models through their own API
Any OpenAI-compatible endpointSelf-hosted or hosted providers that speak the OpenAI API (for example NVIDIA NIM)
Scrapalot AIThe bundled system provider on hosted plans — no key of your own needed

Pricing is per-provider and per-model, billed by them, not by Scrapalot — except on paid hosted plans, where the bundled Scrapalot AI models are included.

Picking a model without guessing ​

Rather than copying a recommendation, sort by what the task needs:

  • Hard reasoning, high stakes — the provider's current flagship. Slower, pricier, best answers.
  • Everyday questions over documents — a mid-tier model. This is where most work belongs.
  • High volume or latency-sensitive — the smallest model that still answers correctly, or a fast-inference provider.
  • Long documents — check the context window before anything else; it decides how much of a document reaches the model in one pass.

Local & Self-Hosted Options ​

Local GGUF Models ​

What: Run AI models directly on your server

Privacy benefits:

  • Your data never leaves your infrastructure
  • No API costs
  • No rate limits
  • Complete control

Choosing one: open-weight families move fast, so pick by size rather than by name. Roughly, a 7-8B model runs on a modern CPU or a small GPU and handles summarizing and straightforward questions; a 13-14B model needs a GPU and reasons noticeably better; a 70B-class model needs serious VRAM and is where local quality approaches the cloud flagships. Ollama's and LM Studio's own catalogues list what is current.

Requirements:

  • GPU recommended (NVIDIA preferred)
  • CPU-only possible but slower
  • 16GB+ RAM recommended
  • Storage for model files (5-50GB per model)

Ollama ​

What: Local model server with easy management

Benefits:

  • Simple model download and switching
  • Automatic model management
  • Good performance out of box
  • Active community

Best for:

  • Development and testing
  • Privacy-sensitive deployments
  • Cost control
  • Learning and experimentation

Setup: Install Ollama, run models locally

LM Studio ​

What: Desktop app for running local models

Benefits:

  • User-friendly GUI
  • GPU acceleration built-in
  • Easy model browsing
  • Great for getting started

Best for:

  • Personal use
  • Small teams
  • Development
  • Model experimentation

Desktop Local AI ​

What: The Scrapalot desktop app's own local runtime (llama.cpp, or MLX on Apple Silicon)

How it works: the desktop app exposes the model on your machine and chat turns that select it are relayed to your computer, so prompts and answers for that model stay local.

vLLM ​

What: Production-grade local inference server

Benefits:

  • Optimized for throughput
  • Advanced GPU utilization
  • Production-ready
  • Scales well

Best for:

  • Large deployments
  • High concurrent users
  • Maximum performance
  • Enterprise use

Choosing Models ​

For Chat (Answer Generation) ​

PriorityCloudLocalWhen
QualityYour provider's current flagship70B-classComplex questions, critical accuracy
BalanceMid-tier13-14BMost use cases
Speed / volumeSmallest model that still answers correctly7-8BHigh volume, fast responses

The exact model names live in Settings → AI Providers, populated from what your providers actually offer today.

Scrapalot's default is sentence-transformers/all-MiniLM-L6-v2 at 384 dimensions — local, no key, no external calls, and fast enough to embed a large library without a GPU. Most people never need to change it.

The embedding model is a deployment-wide choice made by an administrator, not a per-collection or per-user setting: every collection is searched in the same vector space, which is what lets one question search several collections at once. The vector store is sized for 384-dimension embeddings.

Embedding models can come from local sentence-transformers / Hugging Face models, Ollama or vLLM, or cloud providers such as OpenAI and Google. The admin view lists every embedding model in the catalogue and flags which ones can actually run.

Changing the embedding model means re-embedding

Embeddings from different models are not comparable, even at the same dimension count. Switching the model invalidates every stored vector, so the whole library has to be re-processed under the new model.

For Reranking (result ordering) ​

After the initial search, a reranker re-orders results so the most relevant passages come first. Reranking is optional and comes in two tiers:

  • Local (CPU) — the default. Runs on the server with no key and no external calls, using a compact cross-encoder. Best for privacy and zero extra cost.
  • External API (Jina / Cohere) — a GPU-backed reranking service. Faster on large result sets and often more accurate, but needs an API key and sends passages to a third party.

An administrator can switch the mode at any time (no redeploy) from Settings → Providers → System keys, where the matching API key is also entered. Until a key is set, the local CPU reranker is used.

Hardware Requirements ​

For Local Models ​

Minimum (CPU only):

  • 16GB RAM
  • Any modern CPU
  • 50GB disk space
  • Models: 7B parameters or smaller

Recommended (GPU):

  • 16GB+ RAM
  • NVIDIA GPU with 8GB+ VRAM
  • 100GB disk space
  • Models: up to 13B parameters

High Performance (Multi-GPU):

  • 32GB+ RAM
  • Multiple NVIDIA GPUs (16GB+ each)
  • 200GB+ disk space
  • Models: 70B+ parameters

Performance Guide ​

Model SizeCPU RAMGPU VRAMSpeedQuality
3B8GB4GBFastGood
7B16GB8GBMediumBetter
13B32GB16GBSlowerGreat
70B128GB40GB+SlowBest

Model Caching ​

What it does: Local models (embeddings, rerankers, GGUF chat models) are loaded lazily on first use and then kept in memory, so the service starts quickly and only the first request that needs a given model pays its load time.

How it helps:

  • Fast startup — no model is loaded until something needs it
  • Repeated queries reuse the loaded model
  • Automatic management, no configuration needed

Cost Considerations ​

Cloud Models ​

Pricing structure:

  • Pay per token (input + output), billed by the provider
  • Flagship models cost several times what mid-tier models do, for the same question
  • Embeddings are cheap by comparison — a whole library costs less than a day of chat

Typical costs:

  • 1000 documents embedded: $1-5
  • 1000 chat queries: $10-100 (varies by model)
  • Monthly for active use: $50-500 (depends on volume)

Local Models ​

One-time costs:

  • GPU hardware: $500-5000
  • Setup time: Few hours
  • Model downloads: Free

Ongoing costs:

  • Electricity only
  • No per-query charges
  • No rate limits
  • Scales with hardware

Break-even: Typically 3-6 months for moderate use

Switching Models ​

Easy Migration ​

Change chat model:

  • Select new model in settings
  • Existing conversations unaffected
  • Takes effect immediately

Change embedding model (administrators):

  • Select the new model for the deployment
  • Re-process the library with the new embeddings
  • Search quality may improve or change

Mix and match:

  • Different chat models per conversation
  • Chat and embeddings can use different providers
  • System tasks (routing, extraction, research planning) run on the system provider, independent of your chat model

Best Practices ​

Model Selection ​

Start simple:

  1. Begin with a mid-tier cloud model
  2. Evaluate quality and cost on your own documents
  3. Try local models if privacy or cost matters
  4. Move up a tier only where answers are actually falling short

Optimize for use case:

  • Customer support: fast, small models — latency matters more than depth
  • Research: the highest-quality model you can justify — depth beats speed here
  • Internal docs: Local models for privacy
  • Public content: Cloud models for scale

Performance Tuning ​

For local models:

  • Start with smaller models (7B)
  • Increase size if quality insufficient
  • Use GPU acceleration when available
  • Monitor memory usage

For cloud models:

  • Use cheaper models for embeddings
  • Reserve expensive models for complex queries
  • Monitor API costs
  • Set reasonable rate limits

Privacy & Compliance ​

When privacy matters:

  • Use local models exclusively
  • Self-host embeddings and chat
  • Data never leaves your infrastructure
  • Full audit trail

When cloud is acceptable:

  • Use encrypted connections
  • Review provider privacy policies
  • Consider data residency requirements
  • Understand data retention

Troubleshooting ​

Model Loading Slow ​

For local models:

  • Check available RAM/VRAM
  • Reduce model size
  • Enable model caching
  • Use faster storage (SSD)

Poor Answer Quality ​

Try:

  • Switch to larger/better model
  • Adjust temperature settings
  • Improve document chunking
  • Use better embedding model

Out of Memory ​

Solutions:

  • Use smaller model
  • Reduce context window
  • Add more RAM/VRAM
  • Use quantized models (Q4, Q5)

API Errors ​

Check:

  • API key is valid
  • Rate limits not exceeded
  • Sufficient API credits
  • Service status

Cross-Service Model Sync ​

Model providers and their models are owned by the Python AI service and synced to the Kotlin backend via Redis Streams SAGA (P→K direction). This ensures the Kotlin backend always has up-to-date model resolution data without directly querying Python's database.


Scrapalot works with any model. Start with what's convenient, optimize as you learn your usage patterns. You're never locked in.

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.