Production Deployment Guide
Last Updated: September 2026
How to run Scrapalot on your own infrastructure with Docker Compose: the services, the configuration an operator must provide, and day-2 operations.
Who this guide is for
The Scrapalot cloud (SaaS) service is the standard way to use Scrapalot. These deployment guides are intended for Enterprise on-premise deployments, where an operator or admin runs a licensed Scrapalot instance on their own infrastructure. On-premise deployment is offered on-demand as part of an Enterprise arrangement — contact sales to discuss it.
Deployment Overview
Docker Compose
Single-server deployment — the supported topology
Cloud / VPS
The same stack on a cloud server, with TLS and CI/CD
GPU acceleration
Optional NVIDIA GPU for embeddings, OCR and local models
Architecture Components
Scrapalot is a set of containers defined in one Compose file (docker-scrapalot/docker-compose.yaml in the scrapalot-chat repository).
| Service | Image | Role |
|---|---|---|
nginx-proxy-manager | jc21/nginx-proxy-manager | Reverse proxy, Let's Encrypt TLS |
scrapalot-ui | scrapalot-ui (built from scrapalot-ui) | Web app |
scrapalot-gw | scrapalot-gw (built from scrapalot-gw) | Single API entry point: auth, routing, rate limits |
scrapalot-backend | scrapalot-backend (built from scrapalot-backend) | Users, auth, workspaces, collections, notes, settings, subscriptions |
scrapalot-chat | scrapalot-chat (built from scrapalot-chat) | AI service: RAG, deep research, document pipeline, MCP server, streaming |
scrapalot-workers | same image as chat | Celery workers (documents, fast, research queues) + the scheduler |
scrapalot-workers-graph | same image as chat | Celery worker for knowledge-graph extraction |
pgvector | pgvector/pgvector:pg18 | PostgreSQL 18 with pgvector — holds both databases |
redis | redis:8-alpine | Cache, Celery broker, cross-service event streams |
neo4j | neo4j:2025-community (APOC) | Knowledge graph and personal-memory graph |
coturn | coturn/coturn:4.6-alpine | STUN/TURN relay for workspace voice and calls |
autoheal | willfarrell/autoheal | Restarts containers whose health check fails |
scrapalot-discord | same image as chat | Discord community bot — not needed on-premise |
postgres-backend | postgres:16-alpine | Legacy separate backend database, only with the local / microservices profiles — not used by the default layout |
Databases. A single pgvector container holds two databases, created on first start by init-db.sh: scrapalot (AI data: documents, chunks, embeddings — public schema) and scrapalot_backend (users, workspaces, notes, settings — scrapalot schema). Schema migrations run automatically when the services start (Alembic for the AI service, Liquibase for the backend).
Request path. Everything public goes through the reverse proxy: the web app to scrapalot-ui, and the API domain (REST, SSE streaming, WebSockets, /mcp) to scrapalot-gw. The gateway forwards to the backend and AI service on the private Docker network. Databases and internal service ports are bound to localhost or the Docker network only.
Docker Compose Deployment
Requirements
- Linux host with Docker Engine and the Compose plugin
- 8 vCPU / 16 GB RAM as the reference size (the default per-container memory limits add up to more than 16 GB; they are ceilings, not reservations). More RAM helps large libraries and graph extraction.
- 100 GB+ disk for databases, uploads and models
- A domain with DNS records for the app and the API (e.g.
app.example.com,api.example.com) - The
scrapalot-chat,scrapalot-backend,scrapalot-gwandscrapalot-uisources checked out side by side (the Compose file builds the UI, backend and gateway from sibling directories), or pre-built images supplied with your Enterprise arrangement
Step 1: Create the environment file
cd scrapalot-chat/docker-scrapalot
cp example.env .env
# edit .envRequired settings — set all of these before the first start. Never keep the sample or fallback values in production.
| Variable | Used by | Notes |
|---|---|---|
JWT_SECRET | gateway, backend, AI service | Random, ≥ 32 characters, identical for all three (they read the same .env). The Compose file passes it to the gateway and backend with no fallback, so authentication does not work until it is set |
TURN_STATIC_AUTH_SECRET | coturn, backend | Compose refuses to start without it. Generate with openssl rand -hex 32 |
POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DB | pgvector, AI service, workers | Database superuser for the pgvector container |
POSTGRES_BACKEND_USER, POSTGRES_BACKEND_PASSWORD | backend | Same credentials as above in the default single-container layout |
REDIS_PASSWORD | all services | Redis requirepass |
NEO4J_USER, NEO4J_PASSWORD | neo4j, AI service, workers | Set explicitly — every service must see the same value |
MCP_ENCRYPTION_KEY | backend, AI service, workers | Base64-encoded 32-byte key (openssl rand -base64 32). Encrypts provider API keys, connector credentials and MCP tokens at rest (AES-256-GCM). Without it these secrets are stored unencrypted |
FRONTEND_URL | backend, AI service | Public URL of the web app |
BACKEND_BASE_URL | AI service | Public URL of the API |
VITE_API_BASE_URL, VITE_LLM_INFERENCE_ENDPOINT | UI build | https://api.example.com/api/v1 and …/api/v1/llm-inference — baked into the UI image at build time |
CORS_ALLOWED_ORIGINS | gateway, backend | Your web app origin(s) |
Optional settings
| Variable | Purpose |
|---|---|
GOOGLE_OAUTH_CLIENT_ID, GOOGLE_OAUTH_CLIENT_SECRET, GOOGLE_OAUTH_REDIRECT_URI | "Sign in with Google" and the Google Drive connector (…/api/v1/auth/google/callback) |
EMAIL_ENABLED, SMTP_HOST, SMTP_PORT, SMTP_USERNAME, SMTP_PASSWORD, EMAIL_FROM | Outgoing mail (invitations, share notifications). Off unless EMAIL_ENABLED=true |
HUGGINGFACE_TOKEN | Downloading gated embedding / reranker models |
FIRECRAWL_API_KEY, TAVILY_API_KEY, SERPAPI_KEY | Web search providers for web search and deep research |
GRAPH_ENABLED | Knowledge-graph features on/off — set it explicitly (true/false) so the AI service and the workers agree |
TURN_HOST | Hostname clients use for STUN/TURN |
GATEWAY_RATE_LIMIT_FREE, …_RESEARCHER, …_PRO | Per-tier AI request limits per second (defaults 3 / 10 / 50) |
JWT_EXPIRATION_MS, JWT_REFRESH_EXPIRATION_MS | Token lifetimes (defaults 4 h / 90 days) |
USER_CPUSET, BATCH_CPUSET | CPU pinning for user-facing vs. batch containers. The defaults (0-4 / 5-7) assume 8 cores — set them to match your host (e.g. 0-2 / 3 on 4 cores) |
CHAT_MEMORY, WORKERS_MEMORY, WORKERS_GRAPH_MEMORY, PGVECTOR_MEMORY, NEO4J_MEMORY, NEO4J_HEAP_SIZE, NEO4J_PAGECACHE_SIZE, BACKEND_MEMORY, GW_MEMORY | Container memory limits and Neo4j memory. Keep the Neo4j page cache at least as large as the graph store |
LLM provider API keys (OpenAI, Anthropic, OpenRouter, DeepSeek, …) are not environment variables: an administrator adds them in the app under Settings → AI Providers, where the system provider used for Scrapalot's own agents is also chosen.
Step 2: Adjust host-specific files
coturn/turnserver.conf— set your public IP (external-ip),realm/server-name, and the certificate path. Skip if you do not use voice.- In
docker-compose.yaml, the Neo4j Bolt advertised address is set for the hosted service — point it at your host or remove it.
Step 3: Start the stack
# The compose file uses an external network; create it once
docker network create docker-scrapalot_scrapalot-network
# Build (or pull your supplied images and set SCRAPALOT_*_IMAGE in .env)
docker compose build
# Infrastructure first
docker compose up -d pgvector redis neo4j nginx-proxy-manager coturn
# Then the application
docker compose up -d scrapalot-chat scrapalot-workers scrapalot-workers-graph \
scrapalot-backend scrapalot-gw scrapalot-ui autohealStart order matters: databases → AI service and backend → gateway last (it waits for the backend to be healthy). Listing the services explicitly keeps the Discord bot and the legacy postgres-backend out.
The first start of scrapalot-chat takes a few minutes (model downloads and migrations); its health check allows for that.
Step 4: Verify
docker compose ps # all services "healthy"
curl -f http://localhost:8080/actuator/health # gateway
docker compose logs -f scrapalot-chat # AI service start-upThen configure TLS and the proxy hosts (Cloud / VPS guide), open the web app, register the first account and configure an AI provider.
Published ports
| Port | Service | Expose publicly? |
|---|---|---|
| 80, 443 | nginx-proxy-manager | Yes |
| 81 | nginx-proxy-manager admin UI | No — restrict to your admin network |
| 3000 | scrapalot-ui | No — reach it through the proxy |
| 8080 | scrapalot-gw | No — reach it through the proxy |
| 3478, 5349 + UDP relay range | coturn (host network) | Yes, if you use voice |
Everything else (databases, backend, AI service) listens on localhost or the internal Docker network only.
Configuration Management
Besides .env, the AI service reads configs/config.yaml (mounted into the container). Most keys accept environment overrides with ${VAR:-default} syntax. Examples:
documents:
max_concurrent_jobs_per_user: 3
max_file_size_mb: 50
llm:
advanced:
gpu_layers: ${LLM_GPU_LAYERS:-auto}
context_size: ${LLM_CONTEXT_SIZE:-32768}
batch_size: ${LLM_BATCH_SIZE:-1024}
threads: ${LLM_THREADS:-4}Most runtime behaviour (models, prompts, feature toggles, per-user settings) is managed in the app by administrators and stored in the database, not in files. Changes to configs/ or .env take effect after the affected containers are recreated (docker compose up -d <service>).
CI/CD
Each repository ships a GitHub Actions deploy workflow. They run on a runner on the deployment host, deploy on every push to main (path-filtered) and can be started manually with an environment of dev, rc or prod. Secrets are stored per GitHub environment.
| Repository | Workflow | What it does |
|---|---|---|
scrapalot-chat | CICD-Infrastructure (manual only) | Deploys infrastructure services: all, essential, redis, neo4j, nginx-proxy-manager or workers |
scrapalot-chat | CICD-Backend | Builds the AI image when needed and redeploys scrapalot-chat and the workers; optional gpu input builds the CUDA image and layers the GPU override |
scrapalot-backend | CICD-Backend-Kotlin | Builds and redeploys the Kotlin backend |
scrapalot-gw | CICD-Gateway | Builds and redeploys the gateway |
scrapalot-ui | CICD-UI | Builds and redeploys the web app |
Database migrations are applied by the services themselves on start-up, so a deploy needs no separate migration step. Code changes are live only after the container restarts — there is no hot reload.
Without CI/CD, the manual equivalent is git pull in each repository followed by docker compose build <service> and docker compose up -d <service>.
Monitoring
- Health checks: every application container has a Docker health check;
autohealrestarts containers that turn unhealthy. - Gateway / backend metrics: Spring Boot Actuator exposes
health,info,metricsandprometheusendpoints — scrape them from inside the Docker network. - Logs:
docker compose logs -f <service>; container logs are rotated (20 MB × 3 files per container). - Resource usage:
docker stats.
Backup and Recovery
Back up everything that is not reproducible from the images:
| What | How | Notes |
|---|---|---|
scrapalot database | pg_dump -Fc inside the pgvector container | Documents, chunks, embeddings — the largest item |
scrapalot_backend database | pg_dump -Fc | Users, workspaces, notes, settings |
| Postgres roles | pg_dumpall -g | |
| Neo4j graph | neo4j-admin database dump | Needs a short Neo4j stop. The graph can be re-extracted from documents, but that costs hours of LLM calls |
scrapalot_data volume | tar, excluding models/, cache/, logs/, tmp/ | Uploaded files, thumbnails, generated content |
| Redis volume | BGSAVE + tar | Queue state and caches — optional |
npm_data, npm_letsencrypt volumes | tar | Proxy hosts and certificates |
.env and configs/ | copy to a secret store | Contains your secrets |
Keep at least one copy off the host and test a restore periodically. To restore: recreate the host, restore .env/configs/, start pgvector, redis and neo4j, load the dumps and volumes, then start the application services.
Local Models
Scrapalot can use local LLM engines instead of (or alongside) cloud APIs:
- Ollama, LM Studio, vLLM or any OpenAI-compatible server — run the engine on the Scrapalot host or another machine on your network and add it as a provider in Settings → AI Providers. This is the recommended way to use GPU-backed models.
- Desktop app relay — a desktop user's local engine can serve their chats through the full cloud pipeline over the desktop app's connection, without opening any port on their machine.
See GPU Acceleration and Model Management.
Next Steps
- Cloud / VPS Deployment - TLS, proxy hosts and CI/CD on a cloud server
- GPU Setup Guide - NVIDIA GPU images and local model engines
- Architecture Overview - Understand the system architecture
- Security - Security features and operator checklist