Skip to content

Production Deployment Guide ​

Last Updated: September 2026

How to run Scrapalot on your own infrastructure with Docker Compose: the services, the configuration an operator must provide, and day-2 operations.

Who this guide is for

The Scrapalot cloud (SaaS) service is the standard way to use Scrapalot. These deployment guides are intended for Enterprise on-premise deployments, where an operator or admin runs a licensed Scrapalot instance on their own infrastructure. On-premise deployment is offered on-demand as part of an Enterprise arrangement — contact sales to discuss it.

Deployment Overview ​

Docker

Docker Compose

Single-server deployment — the supported topology

Cloud

Cloud / VPS

The same stack on a cloud server, with TLS and CI/CD

GPU

GPU acceleration

Optional NVIDIA GPU for embeddings, OCR and local models

Architecture Components ​

Scrapalot is a set of containers defined in one Compose file (docker-scrapalot/docker-compose.yaml in the scrapalot-chat repository).

ServiceImageRole
nginx-proxy-managerjc21/nginx-proxy-managerReverse proxy, Let's Encrypt TLS
scrapalot-uiscrapalot-ui (built from scrapalot-ui)Web app
scrapalot-gwscrapalot-gw (built from scrapalot-gw)Single API entry point: auth, routing, rate limits
scrapalot-backendscrapalot-backend (built from scrapalot-backend)Users, auth, workspaces, collections, notes, settings, subscriptions
scrapalot-chatscrapalot-chat (built from scrapalot-chat)AI service: RAG, deep research, document pipeline, MCP server, streaming
scrapalot-workerssame image as chatCelery workers (documents, fast, research queues) + the scheduler
scrapalot-workers-graphsame image as chatCelery worker for knowledge-graph extraction
pgvectorpgvector/pgvector:pg18PostgreSQL 18 with pgvector — holds both databases
redisredis:8-alpineCache, Celery broker, cross-service event streams
neo4jneo4j:2025-community (APOC)Knowledge graph and personal-memory graph
coturncoturn/coturn:4.6-alpineSTUN/TURN relay for workspace voice and calls
autohealwillfarrell/autohealRestarts containers whose health check fails
scrapalot-discordsame image as chatDiscord community bot — not needed on-premise
postgres-backendpostgres:16-alpineLegacy separate backend database, only with the local / microservices profiles — not used by the default layout

Databases. A single pgvector container holds two databases, created on first start by init-db.sh: scrapalot (AI data: documents, chunks, embeddings — public schema) and scrapalot_backend (users, workspaces, notes, settings — scrapalot schema). Schema migrations run automatically when the services start (Alembic for the AI service, Liquibase for the backend).

Request path. Everything public goes through the reverse proxy: the web app to scrapalot-ui, and the API domain (REST, SSE streaming, WebSockets, /mcp) to scrapalot-gw. The gateway forwards to the backend and AI service on the private Docker network. Databases and internal service ports are bound to localhost or the Docker network only.

Docker Compose Deployment ​

Requirements ​

  • Linux host with Docker Engine and the Compose plugin
  • 8 vCPU / 16 GB RAM as the reference size (the default per-container memory limits add up to more than 16 GB; they are ceilings, not reservations). More RAM helps large libraries and graph extraction.
  • 100 GB+ disk for databases, uploads and models
  • A domain with DNS records for the app and the API (e.g. app.example.com, api.example.com)
  • The scrapalot-chat, scrapalot-backend, scrapalot-gw and scrapalot-ui sources checked out side by side (the Compose file builds the UI, backend and gateway from sibling directories), or pre-built images supplied with your Enterprise arrangement

Step 1: Create the environment file ​

bash
cd scrapalot-chat/docker-scrapalot
cp example.env .env
# edit .env

Required settings — set all of these before the first start. Never keep the sample or fallback values in production.

VariableUsed byNotes
JWT_SECRETgateway, backend, AI serviceRandom, ≥ 32 characters, identical for all three (they read the same .env). The Compose file passes it to the gateway and backend with no fallback, so authentication does not work until it is set
TURN_STATIC_AUTH_SECRETcoturn, backendCompose refuses to start without it. Generate with openssl rand -hex 32
POSTGRES_USER, POSTGRES_PASSWORD, POSTGRES_DBpgvector, AI service, workersDatabase superuser for the pgvector container
POSTGRES_BACKEND_USER, POSTGRES_BACKEND_PASSWORDbackendSame credentials as above in the default single-container layout
REDIS_PASSWORDall servicesRedis requirepass
NEO4J_USER, NEO4J_PASSWORDneo4j, AI service, workersSet explicitly — every service must see the same value
MCP_ENCRYPTION_KEYbackend, AI service, workersBase64-encoded 32-byte key (openssl rand -base64 32). Encrypts provider API keys, connector credentials and MCP tokens at rest (AES-256-GCM). Without it these secrets are stored unencrypted
FRONTEND_URLbackend, AI servicePublic URL of the web app
BACKEND_BASE_URLAI servicePublic URL of the API
VITE_API_BASE_URL, VITE_LLM_INFERENCE_ENDPOINTUI buildhttps://api.example.com/api/v1 and …/api/v1/llm-inference — baked into the UI image at build time
CORS_ALLOWED_ORIGINSgateway, backendYour web app origin(s)

Optional settings

VariablePurpose
GOOGLE_OAUTH_CLIENT_ID, GOOGLE_OAUTH_CLIENT_SECRET, GOOGLE_OAUTH_REDIRECT_URI"Sign in with Google" and the Google Drive connector (…/api/v1/auth/google/callback)
EMAIL_ENABLED, SMTP_HOST, SMTP_PORT, SMTP_USERNAME, SMTP_PASSWORD, EMAIL_FROMOutgoing mail (invitations, share notifications). Off unless EMAIL_ENABLED=true
HUGGINGFACE_TOKENDownloading gated embedding / reranker models
FIRECRAWL_API_KEY, TAVILY_API_KEY, SERPAPI_KEYWeb search providers for web search and deep research
GRAPH_ENABLEDKnowledge-graph features on/off — set it explicitly (true/false) so the AI service and the workers agree
TURN_HOSTHostname clients use for STUN/TURN
GATEWAY_RATE_LIMIT_FREE, …_RESEARCHER, …_PROPer-tier AI request limits per second (defaults 3 / 10 / 50)
JWT_EXPIRATION_MS, JWT_REFRESH_EXPIRATION_MSToken lifetimes (defaults 4 h / 90 days)
USER_CPUSET, BATCH_CPUSETCPU pinning for user-facing vs. batch containers. The defaults (0-4 / 5-7) assume 8 cores — set them to match your host (e.g. 0-2 / 3 on 4 cores)
CHAT_MEMORY, WORKERS_MEMORY, WORKERS_GRAPH_MEMORY, PGVECTOR_MEMORY, NEO4J_MEMORY, NEO4J_HEAP_SIZE, NEO4J_PAGECACHE_SIZE, BACKEND_MEMORY, GW_MEMORYContainer memory limits and Neo4j memory. Keep the Neo4j page cache at least as large as the graph store

LLM provider API keys (OpenAI, Anthropic, OpenRouter, DeepSeek, …) are not environment variables: an administrator adds them in the app under Settings → AI Providers, where the system provider used for Scrapalot's own agents is also chosen.

Step 2: Adjust host-specific files ​

  • coturn/turnserver.conf — set your public IP (external-ip), realm / server-name, and the certificate path. Skip if you do not use voice.
  • In docker-compose.yaml, the Neo4j Bolt advertised address is set for the hosted service — point it at your host or remove it.

Step 3: Start the stack ​

bash
# The compose file uses an external network; create it once
docker network create docker-scrapalot_scrapalot-network

# Build (or pull your supplied images and set SCRAPALOT_*_IMAGE in .env)
docker compose build

# Infrastructure first
docker compose up -d pgvector redis neo4j nginx-proxy-manager coturn

# Then the application
docker compose up -d scrapalot-chat scrapalot-workers scrapalot-workers-graph \
  scrapalot-backend scrapalot-gw scrapalot-ui autoheal

Start order matters: databases → AI service and backend → gateway last (it waits for the backend to be healthy). Listing the services explicitly keeps the Discord bot and the legacy postgres-backend out.

The first start of scrapalot-chat takes a few minutes (model downloads and migrations); its health check allows for that.

Step 4: Verify ​

bash
docker compose ps                               # all services "healthy"
curl -f http://localhost:8080/actuator/health   # gateway
docker compose logs -f scrapalot-chat           # AI service start-up

Then configure TLS and the proxy hosts (Cloud / VPS guide), open the web app, register the first account and configure an AI provider.

Published ports ​

PortServiceExpose publicly?
80, 443nginx-proxy-managerYes
81nginx-proxy-manager admin UINo — restrict to your admin network
3000scrapalot-uiNo — reach it through the proxy
8080scrapalot-gwNo — reach it through the proxy
3478, 5349 + UDP relay rangecoturn (host network)Yes, if you use voice

Everything else (databases, backend, AI service) listens on localhost or the internal Docker network only.

Configuration Management ​

Besides .env, the AI service reads configs/config.yaml (mounted into the container). Most keys accept environment overrides with ${VAR:-default} syntax. Examples:

yaml
documents:
  max_concurrent_jobs_per_user: 3
  max_file_size_mb: 50

llm:
  advanced:
    gpu_layers: ${LLM_GPU_LAYERS:-auto}
    context_size: ${LLM_CONTEXT_SIZE:-32768}
    batch_size: ${LLM_BATCH_SIZE:-1024}
    threads: ${LLM_THREADS:-4}

Most runtime behaviour (models, prompts, feature toggles, per-user settings) is managed in the app by administrators and stored in the database, not in files. Changes to configs/ or .env take effect after the affected containers are recreated (docker compose up -d <service>).

CI/CD ​

Each repository ships a GitHub Actions deploy workflow. They run on a runner on the deployment host, deploy on every push to main (path-filtered) and can be started manually with an environment of dev, rc or prod. Secrets are stored per GitHub environment.

RepositoryWorkflowWhat it does
scrapalot-chatCICD-Infrastructure (manual only)Deploys infrastructure services: all, essential, redis, neo4j, nginx-proxy-manager or workers
scrapalot-chatCICD-BackendBuilds the AI image when needed and redeploys scrapalot-chat and the workers; optional gpu input builds the CUDA image and layers the GPU override
scrapalot-backendCICD-Backend-KotlinBuilds and redeploys the Kotlin backend
scrapalot-gwCICD-GatewayBuilds and redeploys the gateway
scrapalot-uiCICD-UIBuilds and redeploys the web app

Database migrations are applied by the services themselves on start-up, so a deploy needs no separate migration step. Code changes are live only after the container restarts — there is no hot reload.

Without CI/CD, the manual equivalent is git pull in each repository followed by docker compose build <service> and docker compose up -d <service>.

Monitoring ​

  • Health checks: every application container has a Docker health check; autoheal restarts containers that turn unhealthy.
  • Gateway / backend metrics: Spring Boot Actuator exposes health, info, metrics and prometheus endpoints — scrape them from inside the Docker network.
  • Logs: docker compose logs -f <service>; container logs are rotated (20 MB × 3 files per container).
  • Resource usage: docker stats.

Backup and Recovery ​

Back up everything that is not reproducible from the images:

WhatHowNotes
scrapalot databasepg_dump -Fc inside the pgvector containerDocuments, chunks, embeddings — the largest item
scrapalot_backend databasepg_dump -FcUsers, workspaces, notes, settings
Postgres rolespg_dumpall -g
Neo4j graphneo4j-admin database dumpNeeds a short Neo4j stop. The graph can be re-extracted from documents, but that costs hours of LLM calls
scrapalot_data volumetar, excluding models/, cache/, logs/, tmp/Uploaded files, thumbnails, generated content
Redis volumeBGSAVE + tarQueue state and caches — optional
npm_data, npm_letsencrypt volumestarProxy hosts and certificates
.env and configs/copy to a secret storeContains your secrets

Keep at least one copy off the host and test a restore periodically. To restore: recreate the host, restore .env/configs/, start pgvector, redis and neo4j, load the dumps and volumes, then start the application services.

Local Models ​

Scrapalot can use local LLM engines instead of (or alongside) cloud APIs:

  • Ollama, LM Studio, vLLM or any OpenAI-compatible server — run the engine on the Scrapalot host or another machine on your network and add it as a provider in Settings → AI Providers. This is the recommended way to use GPU-backed models.
  • Desktop app relay — a desktop user's local engine can serve their chats through the full cloud pipeline over the desktop app's connection, without opening any port on their machine.

See GPU Acceleration and Model Management.

Next Steps ​

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.