Skip to content

GPU Acceleration Setup ​

Last Updated: September 2026

How to use GPUs with a self-hosted Scrapalot: NVIDIA GPU images for the AI service and workers, dedicated GPU worker servers, and GPU-backed local model engines.

Overview ​

Scrapalot runs fine on CPU only — the reference deployment is CPU-only and uses cloud LLM providers. A GPU helps in two independent places:

WhereWhat gets fasterHow
AI service and workersEmbeddings, cross-encoder reranking, OCR / layout modelsBuild the CUDA image and apply the GPU Compose override (NVIDIA only)
LLM inferenceChat, deep research and agents on local modelsRun a GPU-backed engine — Ollama, LM Studio, vLLM or any OpenAI-compatible server — and add it as a provider

For AMD, Intel or Apple GPUs, run the LLM engine natively on that hardware (Ollama and LM Studio support them) and connect it as a provider; the containerised CUDA image is for NVIDIA GPUs.

NVIDIA GPU in Docker ​

Prerequisites ​

  • NVIDIA GPU with a current driver (nvidia-smi works on the host)
  • NVIDIA Container Toolkit installed and the nvidia runtime registered with Docker (docker info lists it under Runtimes)
  • The CPU deployment from the Deployment Guide working first

What the GPU image changes ​

The default scrapalot-chat image ships CPU PyTorch on purpose. The Dockerfile takes a build argument to switch:

TORCH_VARIANTPyTorchUse
cpu (default)CPU wheelsCPU-only hosts
gputorch 2.7.0 from the CUDA 12.8 (cu128) wheel indexNVIDIA hosts, including Blackwell GPUs (sm_120)

The same image runs scrapalot-chat, scrapalot-workers and scrapalot-workers-graph, so one build covers all three.

Build and run ​

The GPU override file docker-compose.gpu.yaml sets TORCH_VARIANT=gpu for the build and gives the three containers access to all GPUs:

bash
cd scrapalot-chat/docker-scrapalot

# Build the CUDA image
docker compose -f docker-compose.yaml -f docker-compose.gpu.yaml build scrapalot-chat

# Recreate the AI containers with GPU access
docker compose -f docker-compose.yaml -f docker-compose.gpu.yaml up -d \
  scrapalot-chat scrapalot-workers scrapalot-workers-graph

Always pass both files for these services on a GPU host. The base docker-compose.yaml deliberately contains no GPU reservation, because that would make docker compose up fail on hosts without the NVIDIA runtime.

With CI/CD, the AI-service deploy workflow has a gpu input that builds the CUDA image and layers the same override.

Verify ​

bash
docker exec scrapalot-chat python -c \
  "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.device_count())"
# expected: 2.7.0+cu128 True <number of GPUs>

docker exec scrapalot-chat nvidia-smi

If the version ends in +cpu, the container is still running the CPU image — rebuild with the override file and recreate the containers.

Keep bulk PDF parsing at low concurrency on GPU

Many document-processing jobs each initialising their own CUDA context at the same time can make PDF text extraction return empty results, which then shows up as documents failing with low extraction yield. Bulk parsing is CPU-bound anyway: keep document-processing concurrency low on the GPU image and use the GPU for OCR and embeddings.

Dedicated GPU Worker Server ​

To keep the main server CPU-only and offload heavy document work (GPU OCR, embeddings) to a separate NVIDIA machine, run only scrapalot-workers there with docker-compose.gpu-workers.yml. It uses a CUDA image built from docker-scrapalot/Dockerfile.workers (CUDA 12.1 runtime base):

bash
# On the GPU server, from the scrapalot-chat checkout
docker build -f docker-scrapalot/Dockerfile.workers -t scrapalot-workers-cuda:latest .

cd docker-scrapalot
docker compose -f docker-compose.yaml -f docker-compose.gpu-workers.yml up -d scrapalot-workers

The remote worker connects back to the main server's services, so set in its .env:

VariablePoints to
CELERY_BROKER_URL, CELERY_RESULT_BACKENDThe main server's Redis
POSTGRES_HOST, POSTGRES_PORT, POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DBThe main server's pgvector
NEO4J_URI, NEO4J_USER, NEO4J_PASSWORDThe main server's Neo4j

Connect the two machines over a private network or VPN (for example WireGuard or Tailscale). Never expose Redis, PostgreSQL or Neo4j to the internet to make this work.

GPU-Backed Local Models ​

The recommended way to run local LLMs is a dedicated inference engine on the GPU machine, registered as a provider:

  1. Install Ollama, LM Studio or vLLM on the GPU machine and load a model.
  2. Make its API reachable from the Scrapalot server — the same host, the Docker network, or a private network/VPN to another machine.
  3. In Scrapalot, open Settings → AI Providers, add the provider (Ollama, LM Studio, vLLM or OpenAI-compatible) with its base URL, and pick the model.
bash
# Check reachability from the Scrapalot server first
curl http://<gpu-host>:11434/api/tags     # Ollama
curl http://<gpu-host>:1234/v1/models     # LM Studio
curl http://<gpu-host>:8000/v1/models     # vLLM

Example: a GPU machine on another network ​

When the GPU machine sits elsewhere (for example at an office while Scrapalot runs in the cloud), join both machines to a VPN such as Tailscale and use the GPU machine's VPN address as the provider base URL:

bash
# On both machines
curl -fsSL https://tailscale.com/install.sh | sh
sudo tailscale up
tailscale status          # note the GPU machine's 100.x.y.z address

# From the Scrapalot server
curl http://100.x.y.z:1234/v1/models

Bind the engine to the VPN interface (or use VPN ACLs) so it is not reachable from the public internet, and enable the engine's API key if it has one.

Desktop users' own GPUs ​

Users of the Scrapalot desktop app can use a model running on their own computer (Ollama, LM Studio, vLLM) in cloud chats: the desktop app relays the requests over its existing connection, so the full pipeline — RAG, collections, citations, history — works with a local model and no port has to be opened on the user's machine.

Built-in llama.cpp (Local AI) ​

The AI image also contains llama-cpp-python for running GGUF models inside the container, switched on by llm.local_ai.enabled in configs/config.yaml (LOCAL_AI_ENABLED, off by default) and tuned in the llm.advanced block:

yaml
llm:
  advanced:
    gpu_layers: ${LLM_GPU_LAYERS:-auto}      # layers to offload (0 = CPU only)
    context_size: ${LLM_CONTEXT_SIZE:-32768}
    batch_size: ${LLM_BATCH_SIZE:-1024}
    threads: ${LLM_THREADS:-4}
    vulkan:
      enabled: ${LLM_VULKAN_ENABLED:-true}
      prefer_vulkan: ${LLM_VULKAN_PREFER:-false}
      device_selection: ${LLM_VULKAN_DEVICE_SELECTION:-auto}

The image's llama-cpp-python is a CPU build. GPU offload through it requires rebuilding llama-cpp-python with a GPU backend yourself; for GPU inference an external engine (above) is simpler and faster.

Local development (outside Docker) ​

For development on a workstation with an NVIDIA GPU:

bash
conda activate scrapalot-chat

# CUDA PyTorch — same pins as the GPU image (requirements-gpu.txt)
pip install torch==2.7.0 torchvision==0.22.0 torchaudio==2.7.0 \
  --index-url https://download.pytorch.org/whl/cu128 --force-reinstall

# Optional: llama-cpp-python with CUDA (build from source)
CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --no-cache-dir --force-reinstall
# or with Vulkan (NVIDIA, AMD, Intel)
CMAKE_ARGS="-DGGML_VULKAN=on" pip install llama-cpp-python --no-cache-dir --force-reinstall

python -c "import torch; print(torch.__version__, torch.cuda.is_available())"

Use --force-reinstall for PyTorch: installing requirements.txt afterwards would otherwise put the CPU build back.

Blackwell GPUs ​

NVIDIA Blackwell cards (RTX 50-series, RTX PRO Blackwell) need CUDA 12.8 kernels. The GPU image's torch 2.7.0+cu128 includes Blackwell (sm_120) support, so embeddings, reranking and OCR run on the GPU with the standard GPU build — no nightly PyTorch is required. Use a driver recent enough for CUDA 12.8.

Troubleshooting ​

could not select device driver "nvidia" The NVIDIA Container Toolkit is missing or Docker was not restarted after installing it. Check docker info | grep -i runtimes.

torch.cuda.is_available() is False in the container

  • The image is the CPU variant (+cpu in the version) — rebuild with docker-compose.gpu.yaml
  • The container was started without the override file — recreate it with both -f files
  • nvidia-smi fails inside the container — driver/toolkit problem on the host

GPU out of memory Several containers share the GPUs (the reservation is not exclusive). Lower worker concurrency, use smaller embedding/reranker models, or give the LLM engine its own GPU.

Cannot reach a remote GPU engine Test with curl from the Scrapalot server (see above), check the VPN status on both machines and the GPU machine's firewall, and make sure the engine listens on an address other than 127.0.0.1.

Monitoring ​

bash
nvidia-smi -l 1
nvidia-smi --query-gpu=name,temperature.gpu,utilization.gpu,memory.used,memory.total --format=csv

Next Steps ​

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.