AI Model Selection
Last Updated: September 2026
Choose the AI models that best fit your needs - from powerful cloud models to privacy-focused local options. Scrapalot supports OpenAI, Anthropic, Google Gemini, DeepSeek, Groq, OpenRouter, Z.ai, local runtimes (Ollama, LM Studio, vLLM, GGUF) and any OpenAI-compatible provider.
Your Model, Your Choice
Scrapalot is model-agnostic. Use any combination of:
- Cloud models for best quality and speed
- Local models for privacy and cost savings
- Self-hosted models for complete control
- Mix and match different models for different tasks
Cloud Model Providers
Model names are not listed here on purpose
Providers rename and retire models every few months, so any list written into a docs page is wrong within a release or two. Settings → AI Providers shows the models your configured providers actually offer right now, fetched live — that is the authoritative list. This page describes what each provider is good at, which changes far more slowly.
Scrapalot talks to providers, not to a fixed set of models. Add a provider with your own key and every model it exposes becomes selectable.
| Provider | Reach for it when |
|---|---|
| OpenAI | General-purpose quality, the widest tooling ecosystem, and embeddings |
| Anthropic | Very long documents, careful instruction-following, detailed explanation |
| Multimodal input and large context windows | |
| DeepSeek | Strong quality per unit cost — a common default for high-volume work |
| Groq | Latency-critical work; very fast inference |
| OpenRouter | One key, many vendors — useful for trying models before committing |
| Z.ai | GLM-family models through their own API |
| Any OpenAI-compatible endpoint | Self-hosted or hosted providers that speak the OpenAI API (for example NVIDIA NIM) |
| Scrapalot AI | The bundled system provider on hosted plans — no key of your own needed |
Pricing is per-provider and per-model, billed by them, not by Scrapalot — except on paid hosted plans, where the bundled Scrapalot AI models are included.
Picking a model without guessing
Rather than copying a recommendation, sort by what the task needs:
- Hard reasoning, high stakes — the provider's current flagship. Slower, pricier, best answers.
- Everyday questions over documents — a mid-tier model. This is where most work belongs.
- High volume or latency-sensitive — the smallest model that still answers correctly, or a fast-inference provider.
- Long documents — check the context window before anything else; it decides how much of a document reaches the model in one pass.
Local & Self-Hosted Options
Local GGUF Models
What: Run AI models directly on your server
Privacy benefits:
- Your data never leaves your infrastructure
- No API costs
- No rate limits
- Complete control
Choosing one: open-weight families move fast, so pick by size rather than by name. Roughly, a 7-8B model runs on a modern CPU or a small GPU and handles summarizing and straightforward questions; a 13-14B model needs a GPU and reasons noticeably better; a 70B-class model needs serious VRAM and is where local quality approaches the cloud flagships. Ollama's and LM Studio's own catalogues list what is current.
Requirements:
- GPU recommended (NVIDIA preferred)
- CPU-only possible but slower
- 16GB+ RAM recommended
- Storage for model files (5-50GB per model)
Ollama
What: Local model server with easy management
Benefits:
- Simple model download and switching
- Automatic model management
- Good performance out of box
- Active community
Best for:
- Development and testing
- Privacy-sensitive deployments
- Cost control
- Learning and experimentation
Setup: Install Ollama, run models locally
LM Studio
What: Desktop app for running local models
Benefits:
- User-friendly GUI
- GPU acceleration built-in
- Easy model browsing
- Great for getting started
Best for:
- Personal use
- Small teams
- Development
- Model experimentation
Desktop Local AI
What: The Scrapalot desktop app's own local runtime (llama.cpp, or MLX on Apple Silicon)
How it works: the desktop app exposes the model on your machine and chat turns that select it are relayed to your computer, so prompts and answers for that model stay local.
vLLM
What: Production-grade local inference server
Benefits:
- Optimized for throughput
- Advanced GPU utilization
- Production-ready
- Scales well
Best for:
- Large deployments
- High concurrent users
- Maximum performance
- Enterprise use
Choosing Models
For Chat (Answer Generation)
| Priority | Cloud | Local | When |
|---|---|---|---|
| Quality | Your provider's current flagship | 70B-class | Complex questions, critical accuracy |
| Balance | Mid-tier | 13-14B | Most use cases |
| Speed / volume | Smallest model that still answers correctly | 7-8B | High volume, fast responses |
The exact model names live in Settings → AI Providers, populated from what your providers actually offer today.
For Embeddings (Search)
Scrapalot's default is sentence-transformers/all-MiniLM-L6-v2 at 384 dimensions — local, no key, no external calls, and fast enough to embed a large library without a GPU. Most people never need to change it.
The embedding model is a deployment-wide choice made by an administrator, not a per-collection or per-user setting: every collection is searched in the same vector space, which is what lets one question search several collections at once. The vector store is sized for 384-dimension embeddings.
Embedding models can come from local sentence-transformers / Hugging Face models, Ollama or vLLM, or cloud providers such as OpenAI and Google. The admin view lists every embedding model in the catalogue and flags which ones can actually run.
Changing the embedding model means re-embedding
Embeddings from different models are not comparable, even at the same dimension count. Switching the model invalidates every stored vector, so the whole library has to be re-processed under the new model.
For Reranking (result ordering)
After the initial search, a reranker re-orders results so the most relevant passages come first. Reranking is optional and comes in two tiers:
- Local (CPU) — the default. Runs on the server with no key and no external calls, using a compact cross-encoder. Best for privacy and zero extra cost.
- External API (Jina / Cohere) — a GPU-backed reranking service. Faster on large result sets and often more accurate, but needs an API key and sends passages to a third party.
An administrator can switch the mode at any time (no redeploy) from Settings → Providers → System keys, where the matching API key is also entered. Until a key is set, the local CPU reranker is used.
Hardware Requirements
For Local Models
Minimum (CPU only):
- 16GB RAM
- Any modern CPU
- 50GB disk space
- Models: 7B parameters or smaller
Recommended (GPU):
- 16GB+ RAM
- NVIDIA GPU with 8GB+ VRAM
- 100GB disk space
- Models: up to 13B parameters
High Performance (Multi-GPU):
- 32GB+ RAM
- Multiple NVIDIA GPUs (16GB+ each)
- 200GB+ disk space
- Models: 70B+ parameters
Performance Guide
| Model Size | CPU RAM | GPU VRAM | Speed | Quality |
|---|---|---|---|---|
| 3B | 8GB | 4GB | Fast | Good |
| 7B | 16GB | 8GB | Medium | Better |
| 13B | 32GB | 16GB | Slower | Great |
| 70B | 128GB | 40GB+ | Slow | Best |
Model Caching
What it does: Local models (embeddings, rerankers, GGUF chat models) are loaded lazily on first use and then kept in memory, so the service starts quickly and only the first request that needs a given model pays its load time.
How it helps:
- Fast startup — no model is loaded until something needs it
- Repeated queries reuse the loaded model
- Automatic management, no configuration needed
Cost Considerations
Cloud Models
Pricing structure:
- Pay per token (input + output), billed by the provider
- Flagship models cost several times what mid-tier models do, for the same question
- Embeddings are cheap by comparison — a whole library costs less than a day of chat
Typical costs:
- 1000 documents embedded: $1-5
- 1000 chat queries: $10-100 (varies by model)
- Monthly for active use: $50-500 (depends on volume)
Local Models
One-time costs:
- GPU hardware: $500-5000
- Setup time: Few hours
- Model downloads: Free
Ongoing costs:
- Electricity only
- No per-query charges
- No rate limits
- Scales with hardware
Break-even: Typically 3-6 months for moderate use
Switching Models
Easy Migration
Change chat model:
- Select new model in settings
- Existing conversations unaffected
- Takes effect immediately
Change embedding model (administrators):
- Select the new model for the deployment
- Re-process the library with the new embeddings
- Search quality may improve or change
Mix and match:
- Different chat models per conversation
- Chat and embeddings can use different providers
- System tasks (routing, extraction, research planning) run on the system provider, independent of your chat model
Best Practices
Model Selection
Start simple:
- Begin with a mid-tier cloud model
- Evaluate quality and cost on your own documents
- Try local models if privacy or cost matters
- Move up a tier only where answers are actually falling short
Optimize for use case:
- Customer support: fast, small models — latency matters more than depth
- Research: the highest-quality model you can justify — depth beats speed here
- Internal docs: Local models for privacy
- Public content: Cloud models for scale
Performance Tuning
For local models:
- Start with smaller models (7B)
- Increase size if quality insufficient
- Use GPU acceleration when available
- Monitor memory usage
For cloud models:
- Use cheaper models for embeddings
- Reserve expensive models for complex queries
- Monitor API costs
- Set reasonable rate limits
Privacy & Compliance
When privacy matters:
- Use local models exclusively
- Self-host embeddings and chat
- Data never leaves your infrastructure
- Full audit trail
When cloud is acceptable:
- Use encrypted connections
- Review provider privacy policies
- Consider data residency requirements
- Understand data retention
Troubleshooting
Model Loading Slow
For local models:
- Check available RAM/VRAM
- Reduce model size
- Enable model caching
- Use faster storage (SSD)
Poor Answer Quality
Try:
- Switch to larger/better model
- Adjust temperature settings
- Improve document chunking
- Use better embedding model
Out of Memory
Solutions:
- Use smaller model
- Reduce context window
- Add more RAM/VRAM
- Use quantized models (Q4, Q5)
API Errors
Check:
- API key is valid
- Rate limits not exceeded
- Sufficient API credits
- Service status
Cross-Service Model Sync
Model providers and their models are owned by the Python AI service and synced to the Kotlin backend via Redis Streams SAGA (P→K direction). This ensures the Kotlin backend always has up-to-date model resolution data without directly querying Python's database.
Related Documentation
- RAG Strategy - How models power retrieval
- Document Processing - Embedding usage
- Database Design - Model configuration storage
- Deployment Guide - Production model setup
Scrapalot works with any model. Start with what's convenient, optimize as you learn your usage patterns. You're never locked in.