GPU Acceleration Setup
Last Updated: September 2026
How to use GPUs with a self-hosted Scrapalot: NVIDIA GPU images for the AI service and workers, dedicated GPU worker servers, and GPU-backed local model engines.
Overview
Scrapalot runs fine on CPU only — the reference deployment is CPU-only and uses cloud LLM providers. A GPU helps in two independent places:
| Where | What gets faster | How |
|---|---|---|
| AI service and workers | Embeddings, cross-encoder reranking, OCR / layout models | Build the CUDA image and apply the GPU Compose override (NVIDIA only) |
| LLM inference | Chat, deep research and agents on local models | Run a GPU-backed engine — Ollama, LM Studio, vLLM or any OpenAI-compatible server — and add it as a provider |
For AMD, Intel or Apple GPUs, run the LLM engine natively on that hardware (Ollama and LM Studio support them) and connect it as a provider; the containerised CUDA image is for NVIDIA GPUs.
NVIDIA GPU in Docker
Prerequisites
- NVIDIA GPU with a current driver (
nvidia-smiworks on the host) - NVIDIA Container Toolkit installed and the
nvidiaruntime registered with Docker (docker infolists it under Runtimes) - The CPU deployment from the Deployment Guide working first
What the GPU image changes
The default scrapalot-chat image ships CPU PyTorch on purpose. The Dockerfile takes a build argument to switch:
TORCH_VARIANT | PyTorch | Use |
|---|---|---|
cpu (default) | CPU wheels | CPU-only hosts |
gpu | torch 2.7.0 from the CUDA 12.8 (cu128) wheel index | NVIDIA hosts, including Blackwell GPUs (sm_120) |
The same image runs scrapalot-chat, scrapalot-workers and scrapalot-workers-graph, so one build covers all three.
Build and run
The GPU override file docker-compose.gpu.yaml sets TORCH_VARIANT=gpu for the build and gives the three containers access to all GPUs:
cd scrapalot-chat/docker-scrapalot
# Build the CUDA image
docker compose -f docker-compose.yaml -f docker-compose.gpu.yaml build scrapalot-chat
# Recreate the AI containers with GPU access
docker compose -f docker-compose.yaml -f docker-compose.gpu.yaml up -d \
scrapalot-chat scrapalot-workers scrapalot-workers-graphAlways pass both files for these services on a GPU host. The base docker-compose.yaml deliberately contains no GPU reservation, because that would make docker compose up fail on hosts without the NVIDIA runtime.
With CI/CD, the AI-service deploy workflow has a gpu input that builds the CUDA image and layers the same override.
Verify
docker exec scrapalot-chat python -c \
"import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.device_count())"
# expected: 2.7.0+cu128 True <number of GPUs>
docker exec scrapalot-chat nvidia-smiIf the version ends in +cpu, the container is still running the CPU image — rebuild with the override file and recreate the containers.
Keep bulk PDF parsing at low concurrency on GPU
Many document-processing jobs each initialising their own CUDA context at the same time can make PDF text extraction return empty results, which then shows up as documents failing with low extraction yield. Bulk parsing is CPU-bound anyway: keep document-processing concurrency low on the GPU image and use the GPU for OCR and embeddings.
Dedicated GPU Worker Server
To keep the main server CPU-only and offload heavy document work (GPU OCR, embeddings) to a separate NVIDIA machine, run only scrapalot-workers there with docker-compose.gpu-workers.yml. It uses a CUDA image built from docker-scrapalot/Dockerfile.workers (CUDA 12.1 runtime base):
# On the GPU server, from the scrapalot-chat checkout
docker build -f docker-scrapalot/Dockerfile.workers -t scrapalot-workers-cuda:latest .
cd docker-scrapalot
docker compose -f docker-compose.yaml -f docker-compose.gpu-workers.yml up -d scrapalot-workersThe remote worker connects back to the main server's services, so set in its .env:
| Variable | Points to |
|---|---|
CELERY_BROKER_URL, CELERY_RESULT_BACKEND | The main server's Redis |
POSTGRES_HOST, POSTGRES_PORT, POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DB | The main server's pgvector |
NEO4J_URI, NEO4J_USER, NEO4J_PASSWORD | The main server's Neo4j |
Connect the two machines over a private network or VPN (for example WireGuard or Tailscale). Never expose Redis, PostgreSQL or Neo4j to the internet to make this work.
GPU-Backed Local Models
The recommended way to run local LLMs is a dedicated inference engine on the GPU machine, registered as a provider:
- Install Ollama, LM Studio or vLLM on the GPU machine and load a model.
- Make its API reachable from the Scrapalot server — the same host, the Docker network, or a private network/VPN to another machine.
- In Scrapalot, open Settings → AI Providers, add the provider (Ollama, LM Studio, vLLM or OpenAI-compatible) with its base URL, and pick the model.
# Check reachability from the Scrapalot server first
curl http://<gpu-host>:11434/api/tags # Ollama
curl http://<gpu-host>:1234/v1/models # LM Studio
curl http://<gpu-host>:8000/v1/models # vLLMExample: a GPU machine on another network
When the GPU machine sits elsewhere (for example at an office while Scrapalot runs in the cloud), join both machines to a VPN such as Tailscale and use the GPU machine's VPN address as the provider base URL:
# On both machines
curl -fsSL https://tailscale.com/install.sh | sh
sudo tailscale up
tailscale status # note the GPU machine's 100.x.y.z address
# From the Scrapalot server
curl http://100.x.y.z:1234/v1/modelsBind the engine to the VPN interface (or use VPN ACLs) so it is not reachable from the public internet, and enable the engine's API key if it has one.
Desktop users' own GPUs
Users of the Scrapalot desktop app can use a model running on their own computer (Ollama, LM Studio, vLLM) in cloud chats: the desktop app relays the requests over its existing connection, so the full pipeline — RAG, collections, citations, history — works with a local model and no port has to be opened on the user's machine.
Built-in llama.cpp (Local AI)
The AI image also contains llama-cpp-python for running GGUF models inside the container, switched on by llm.local_ai.enabled in configs/config.yaml (LOCAL_AI_ENABLED, off by default) and tuned in the llm.advanced block:
llm:
advanced:
gpu_layers: ${LLM_GPU_LAYERS:-auto} # layers to offload (0 = CPU only)
context_size: ${LLM_CONTEXT_SIZE:-32768}
batch_size: ${LLM_BATCH_SIZE:-1024}
threads: ${LLM_THREADS:-4}
vulkan:
enabled: ${LLM_VULKAN_ENABLED:-true}
prefer_vulkan: ${LLM_VULKAN_PREFER:-false}
device_selection: ${LLM_VULKAN_DEVICE_SELECTION:-auto}The image's llama-cpp-python is a CPU build. GPU offload through it requires rebuilding llama-cpp-python with a GPU backend yourself; for GPU inference an external engine (above) is simpler and faster.
Local development (outside Docker)
For development on a workstation with an NVIDIA GPU:
conda activate scrapalot-chat
# CUDA PyTorch — same pins as the GPU image (requirements-gpu.txt)
pip install torch==2.7.0 torchvision==0.22.0 torchaudio==2.7.0 \
--index-url https://download.pytorch.org/whl/cu128 --force-reinstall
# Optional: llama-cpp-python with CUDA (build from source)
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --no-cache-dir --force-reinstall
# or with Vulkan (NVIDIA, AMD, Intel)
CMAKE_ARGS="-DGGML_VULKAN=on" pip install llama-cpp-python --no-cache-dir --force-reinstall
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"Use --force-reinstall for PyTorch: installing requirements.txt afterwards would otherwise put the CPU build back.
Blackwell GPUs
NVIDIA Blackwell cards (RTX 50-series, RTX PRO Blackwell) need CUDA 12.8 kernels. The GPU image's torch 2.7.0+cu128 includes Blackwell (sm_120) support, so embeddings, reranking and OCR run on the GPU with the standard GPU build — no nightly PyTorch is required. Use a driver recent enough for CUDA 12.8.
Troubleshooting
could not select device driver "nvidia" The NVIDIA Container Toolkit is missing or Docker was not restarted after installing it. Check docker info | grep -i runtimes.
torch.cuda.is_available() is False in the container
- The image is the CPU variant (
+cpuin the version) — rebuild withdocker-compose.gpu.yaml - The container was started without the override file — recreate it with both
-ffiles nvidia-smifails inside the container — driver/toolkit problem on the host
GPU out of memory Several containers share the GPUs (the reservation is not exclusive). Lower worker concurrency, use smaller embedding/reranker models, or give the LLM engine its own GPU.
Cannot reach a remote GPU engine Test with curl from the Scrapalot server (see above), check the VPN status on both machines and the GPU machine's firewall, and make sure the engine listens on an address other than 127.0.0.1.
Monitoring
nvidia-smi -l 1
nvidia-smi --query-gpu=name,temperature.gpu,utilization.gpu,memory.used,memory.total --format=csvNext Steps
- Model Management - Providers and local models
- Deployment Guide - Services and configuration
- Cloud Infrastructure - TLS, proxy and CI/CD