Skip to content

Cloud Infrastructure Deployment ​

Last Updated: September 2026

Running the Scrapalot Docker Compose stack on a cloud server or VPS, with TLS, a reverse proxy and automated deployments.

Overview ​

The cloud topology is the same single-server Compose stack described in the Deployment Guide, plus:

  • Nginx Proxy Manager terminating TLS (Let's Encrypt) in front of the web app and the API gateway
  • GitHub Actions deploy workflows that rebuild and restart services on every push to main
  • A host firewall that exposes only SSH, HTTP and HTTPS (plus the TURN ports if you use voice)

Prerequisites ​

Server

  • Ubuntu 24.04 LTS (any recent Linux with Docker works)
  • Reference size: 8 vCPU / 16 GB RAM, 100 GB+ disk (a separate data volume for Docker is recommended)
  • Public IP address and root or sudo access

Domain

  • DNS A records for the web app (e.g. app.example.com) and the API (e.g. api.example.com)

Accounts and keys

  • An LLM provider key (configured in the app after start-up) or a local model engine
  • Optional: Google OAuth client (Google sign-in and Drive connector), SMTP credentials, web search API keys, HuggingFace token

Step 1: Prepare the Server ​

bash
# Docker Engine + Compose plugin
curl -fsSL https://get.docker.com | sh

# Firewall: SSH, HTTP, HTTPS only
ufw allow OpenSSH && ufw allow 80/tcp && ufw allow 443/tcp && ufw --force enable
# If you use workspace voice (coturn), also open 3478/udp+tcp, 5349/tcp and the relay UDP range from coturn/turnserver.conf

# Swap, so a large document batch cannot OOM the box
fallocate -l 8G /swapfile && chmod 600 /swapfile && mkswap /swapfile && swapon /swapfile
echo '/swapfile none swap sw 0 0' >> /etc/fstab

Check out (or receive) the scrapalot-chat, scrapalot-backend, scrapalot-gw and scrapalot-ui sources side by side in one deployment directory — the Compose file in scrapalot-chat/docker-scrapalot/ builds the other images from their sibling directories.

Step 2: Configure and Start ​

Follow Steps 1–4 of the Deployment Guide: create .env with the required secrets (JWT_SECRET, TURN_STATIC_AUTH_SECRET, database / Redis / Neo4j passwords, MCP_ENCRYPTION_KEY, public URLs), create the external Docker network, then start the infrastructure services and the application services.

Use your real HTTPS URLs from the start:

VariableExample
FRONTEND_URLhttps://app.example.com
BACKEND_BASE_URLhttps://api.example.com
VITE_API_BASE_URLhttps://api.example.com/api/v1
VITE_LLM_INFERENCE_ENDPOINThttps://api.example.com/api/v1/llm-inference
CORS_ALLOWED_ORIGINShttps://app.example.com
GOOGLE_OAUTH_REDIRECT_URIhttps://api.example.com/api/v1/auth/google/callback

The VITE_* values are compiled into the UI image — rebuild scrapalot-ui after changing them.

Step 3: TLS and Reverse Proxy ​

Open the Nginx Proxy Manager admin UI on port 81 — preferably through an SSH tunnel (ssh -L 8181:localhost:81 <host>) rather than opening the port — and replace the default administrator account immediately.

Create two proxy hosts:

1. Web app

  • Domain: app.example.com
  • Forward to: http://scrapalot-ui:3000
  • Enable: Block Common Exploits, Websockets Support
  • SSL: request a Let's Encrypt certificate; enable Force SSL, HTTP/2, HSTS

2. API

  • Domain: api.example.com
  • Forward to: http://scrapalot-gw:8080 — the gateway, not the backend or the AI service
  • Enable: Block Common Exploits, Websockets Support
  • SSL: as above
  • Advanced tab:
nginx
location / {
    proxy_pass http://scrapalot-gw:8080;
    proxy_http_version 1.1;

    # WebSockets (job progress, notifications, notes collaboration)
    proxy_set_header Upgrade $http_upgrade;
    proxy_set_header Connection 'upgrade';
    proxy_cache_bypass $http_upgrade;

    proxy_set_header Host $host;
    proxy_set_header X-Real-IP $remote_addr;
    proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header X-Forwarded-Proto $scheme;

    # Streaming answers: no buffering, long reads (deep research can stream for many minutes)
    proxy_buffering off;
    proxy_connect_timeout 75s;
    proxy_send_timeout 1800s;
    proxy_read_timeout 1800s;

    # Document uploads
    client_max_body_size 500m;
}

Do not add CORS headers in Nginx — the gateway handles CORS, and duplicate headers make browsers reject responses. Set the allowed origins with CORS_ALLOWED_ORIGINS instead.

After saving, verify HTTP/2 is actually served:

bash
curl -sI -o /dev/null -w "%{http_version}\n" https://api.example.com/actuator/health   # want: 2

Public surface ​

HostnameTargetAuth
app.example.comscrapalot-uiApp login
api.example.comscrapalot-gwJWT / API key

That is the whole public surface. The proxy admin UI, Neo4j Browser and the databases stay on localhost; reach them over an SSH tunnel when needed:

bash
ssh -N -L 7474:localhost:7474 -L 7687:localhost:7687 <host>
# Neo4j Browser at http://localhost:7474, Bolt at bolt://localhost:7687

Architecture ​

Background workers run in two containers built from the AI image:

ContainerQueuesPurpose
scrapalot-workersdocuments, fast, research + schedulerDocument parsing, chunking and embedding; summaries and light tasks; background research runs
scrapalot-workers-graphgraph_extractionKnowledge-graph entity and relationship extraction, kept apart so long graph jobs never block uploads

Data Persistence ​

VolumeContents
pgvector_dataPostgreSQL data (both databases)
neo4j_data, neo4j_logs, neo4j_pluginsKnowledge graph
redis_dataRedis append-only file
scrapalot_dataUploaded files, thumbnails, model cache, logs
npm_data, npm_letsencryptProxy configuration and certificates
backend_logsBackend logs

See Backup and Recovery for what to back up.

CI/CD Pipeline ​

Each repository has a deploy workflow that runs on a GitHub Actions runner registered on the deployment host:

RepositoryWorkflowTrigger
scrapalot-chatCICD-Infrastructure — Redis, Neo4j, Nginx Proxy Manager, workersManual only (initial setup and infrastructure changes)
scrapalot-chatCICD-Backend — AI service and workersPush to main (path-filtered) or manual
scrapalot-backendCICD-Backend-KotlinPush to main or manual
scrapalot-gwCICD-GatewayPush to main or manual
scrapalot-uiCICD-UIPush to main or manual

Manual runs take an environment (dev, rc, prod); secrets and variables are defined per GitHub environment, not per repository. The AI-service workflow only rebuilds the image when dependencies or Docker files changed and only restarts the workers when their code changed; its gpu input builds the CUDA image for NVIDIA hosts (see GPU Setup).

The services apply their own database migrations on start-up.

Manual deployment ​

Without GitHub Actions:

bash
cd scrapalot-chat && git pull
cd docker-scrapalot
docker compose build scrapalot-chat
docker compose up -d scrapalot-chat scrapalot-workers scrapalot-workers-graph

Do the same in the other repositories for scrapalot-backend, scrapalot-gw and scrapalot-ui.

Common Tasks ​

bash
docker compose ps                                  # status and health
docker compose logs -f --tail=100 scrapalot-chat   # follow one service
docker stats --no-stream                           # CPU / memory per container
docker compose up -d --force-recreate scrapalot-chat   # recreate after .env / config changes
docker exec scrapalot-workers supervisorctl status     # worker processes

Troubleshooting ​

Service won't start ​

bash
docker compose logs scrapalot-chat
docker compose config >/dev/null   # validates .env interpolation (e.g. a missing required secret)
docker compose ps                  # are pgvector / redis / neo4j healthy?

A container pinned to CPUs the host does not have fails to start — set USER_CPUSET / BATCH_CPUSET for your core count.

502 / 504 from the API ​

  • The API proxy host must point to scrapalot-gw:8080
  • 504 on long answers: raise proxy_read_timeout and turn off proxy_buffering
  • Check docker compose ps scrapalot-gw scrapalot-backend — the gateway waits for a healthy backend

Everything returns 401 ​

JWT_SECRET differs between services, or was changed without recreating all of them. Set it once in .env and recreate the gateway, backend and AI service.

WebSocket keeps reconnecting ​

Enable Websockets Support on the API proxy host and keep the Upgrade / Connection headers.

SSL certificate issues ​

Let's Encrypt certificates renew automatically while Nginx Proxy Manager is running. DNS must point at the server and port 80 must be reachable for the HTTP challenge.

Out of memory ​

bash
free -h
docker stats --no-stream

Lower concurrency (GRAPH_WORKER_CONCURRENCY), adjust the *_MEMORY limits in .env, or add RAM. Keep the Neo4j page cache (NEO4J_PAGECACHE_SIZE) at least as large as the graph store.

Security Best Practices ​

  • Expose only 22, 80 and 443 (plus TURN ports if used); keep port 81 and all database ports closed
  • Use strong, unique values for every secret in .env; never commit .env
  • Restrict SSH to keys; enable 2FA on the accounts that can push to main, since pushes deploy
  • Use Nginx Proxy Manager access lists for any administrative hostnames you do publish
  • Keep the host and images updated; review logs and back up regularly

See also Security.

Next Steps ​

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.