Skip to content

Quick Start Guide ​

The fastest way to start is the web app — sign up at scrapalot.app and you are asking questions about your own documents in minutes, with nothing to install.

Prefer an app? There is a native desktop app for Windows, macOS and Linux, and a native Android app.

Want to run it yourself? The Community Edition is free and open source (AGPL-3.0) and starts with one docker compose up.

Platforms at a glance

Web app (any modern browser), desktop app (Windows, macOS, Linux), and Android app. All three share one account. See Apps & Platforms.

Choose Your Path ​

PathWho it's forWhat you needTime
Hosted cloudAlmost everyoneAn email address~5 min
Community EditionSelf-hosters, developersDocker + an LLM API key~20 min
Enterprise on-premiseCompanies with data-residency rulesAn Enterprise arrangementOn request

What You'll Have ​

By the end of this guide:

Cloud

A running Scrapalot

On our cloud, or on your own machine

Documents

Upload & chat with documents

PDF, EPUB, Office, audio and video

AI

AI-powered answers

With source citations

Web

Web, desktop and Android

One account across all three

Models

Cloud or local AI models

Your choice

Privacy

Your data stays yours

Export and migrate anytime

No prerequisites — you need an email address and a browser.

1. Create your account ​

  1. Go to scrapalot.app/sign-up
  2. Enter your email and password
  3. Click Create Account — the free Researcher plan needs no credit card
  4. You are logged in and land on the dashboard

2. Install an app (optional) ​

The web app is fully featured, so this step is optional:

See Apps & Platforms for install details, including the extra Gatekeeper step on macOS.

Path B: Self-host the Community Edition ​

The Community Edition is the open-source core: document ingestion, hybrid RAG chat with citations, collaborative notes, read-aloud, and bring-your-own-key model support. It runs as one Docker Compose stack.

Prerequisites ​

RequirementVersionCheck commandWhy
Docker24+docker --versionRuns the whole stack
Docker Composev2docker compose versionOrchestrates the services
An LLM key——OpenAI, Anthropic or DeepSeek — or a local GGUF model

You do not need Node, Python or PostgreSQL

Everything runs inside containers, including PostgreSQL 18 with pgvector. You only install those toolchains if you intend to develop on Scrapalot itself — see Contributing.

1. Clone and configure ​

bash
git clone https://github.com/sime2408/scrapalot.git
cd scrapalot

cp .env.example .env

Edit .env and set at minimum:

VariableValue
POSTGRES_PASSWORDA strong database password
JWT_SECRETA random 32+ character secret
OPENAI_API_KEYYour LLM key (or ANTHROPIC_API_KEY / DEEPSEEK_API_KEY)

2. Launch the stack ​

bash
docker compose up -d --build
docker compose logs -f

The first build takes a while. Database migrations run automatically on first boot — Liquibase for the Kotlin backend, Alembic for the Python backend, against the freshly created scrapalot and scrapalot_backend databases.

3. Open the app ​

Once the containers report healthy:

ServiceURL
Web apphttp://localhost:3000
API gatewayhttp://localhost:8080
Kotlin backendhttp://localhost:8091
Python AI backendhttp://localhost:8090
PostgreSQLlocalhost:15432
Redislocalhost:6379

Open http://localhost:3000 and you get the Scrapalot sign-up page. The first account you create is yours to keep — there is no external sign-up service involved.

To stop everything: docker compose down (add -v to also wipe the data volumes).

Path C: Enterprise on-premise ​

Enterprise customers can have the full product — deep research, knowledge graph, AI Scientist, voice, integrations — deployed to their own infrastructure, with support and SLAs. See the Deployment guide or contact sales.

First Steps in Scrapalot ​

The rest of this guide applies to every path.

1. Upload your first document ​

  1. Open Knowledge Stacks in the sidebar and select a collection (or create one)
  2. Drag files onto the Upload tab, or click Browse Files
  3. Click Compose to start processing
  4. Watch live progress as the document is parsed, chunked and embedded — it is ready to query when processing completes

Supported formats:

KindExtensions
Documents.pdf, .epub, .docx, .rtf, .txt, .md
Spreadsheets.xlsx, .xls, .csv, .tsv
Audio (transcribed).mp3, .wav, .m4a, .ogg, .flac, .aac, .opus, and more
Video (transcribed).mp4, .mov, .mkv, .avi, .webm, and more

Files up to 200 MB each. Processing time depends on size and whether OCR is needed — a short text PDF takes seconds, a scanned book takes minutes. See Uploading Documents for details.

2. Ask your first question ​

  1. Select a Knowledge Stack in the chat toolbar, or type @ to pick a single document (@/ for a collection)
  2. Type a question in the chat box
  3. Press Enter
  4. Read the answer — every claim carries a citation back to the source passage

Example questions:

  • "What is this document about?"
  • "Summarize the main points"
  • "What does it say about [specific topic]?"
  • "List all the key findings"

Click any citation to jump to the exact passage it came from. See Asking Questions for search modes and Deep Research.

Configuration Options ​

Using local models (Ollama) ​

Install Ollama:

bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Pull a model — a small one first if you are unsure about your hardware
ollama pull nemotron-3-nano:4b     # or nemotron-3-nano:30b on a bigger machine

# Start Ollama (runs on http://localhost:11434 by default)
ollama serve

Configure in the Scrapalot UI:

  1. Open Settings → AI Providers tab
  2. Click Add
  3. Select Ollama from the provider list
  4. The endpoint auto-fills with http://localhost:11434
  5. Click Fetch Models to list the models you have pulled
  6. Select the ones you want and save the provider

AI SettingsAI Providers tab showing Ollama, vLLM, and LM Studio configuration

No API keys needed — the models run on your own machine.

Ollama must be reachable from Scrapalot

http://localhost:11434 means "localhost as seen by Scrapalot". On the hosted cloud, our servers cannot reach your laptop — local providers work when you self-host, or when Scrapalot runs on the same machine or network as Ollama. Expose it with a tunnel (for example Tailscale) if you want to use it from the cloud.

Other local model providers

vLLM and LM Studio are configured the same way:

  • vLLM: high-performance inference server (requires a custom endpoint URL)
  • LM Studio: desktop app (default: http://localhost:1234/v1)

All three live in the AI Providers tab. The Local AI tab is different: in the desktop app it downloads and runs models on your own computer; on the web it is an admin view of the server's GPU models.

Advanced settings ​

Other configuration lives in the Settings UI:

  • General — theme, accent color, language, font preferences
  • Voice — voice assistant, wake word, your own Whisper key
  • Documents — citation style, chunking strategy
  • Prompts — custom prompt templates
  • Workspaces — workspaces and team members
  • Integrations — MCP servers for your assistant, and API keys to use Scrapalot from your own agent
  • Account — profile, plan and usage

General SettingsGeneral settings for theme and appearance customization

Troubleshooting ​

Hosted cloud ​

Can't sign in — use Forgot password on the login page. If the reset email does not arrive within a few minutes, check spam, then email support.

A document is stuck in "Processing" — large scanned PDFs need OCR and can take several minutes. If it has not moved in 30 minutes, report it with the document name and size.

A feature is greyed out or returns "upgrade required" — it belongs to a higher plan. See Pricing for what each tier includes.

Self-hosted (Community Edition) ​

A container will not start

bash
docker compose ps          # which service is unhealthy
docker compose logs chat   # or: backend, gw, ui, pgvector, redis

"Connection to database failed"

bash
# Is Postgres healthy?
docker compose ps pgvector
docker compose logs pgvector

# Most common cause: POSTGRES_PASSWORD in .env was changed after the
# volume was created. Either restore the old password, or start clean:
docker compose down -v && docker compose up -d --build

Port already in use

bash
# Find what holds the port (3000, 8080, 8090, 8091, 15432, 6379)
sudo lsof -i :3000

# Then change the host-side port mapping in docker-compose.yml

The UI loads but every request fails — the UI talks to the gateway, not to the backends directly. Check the gateway first: docker compose logs gw.

Answers come back empty or without citations — no LLM key is configured, or the key is rejected. Check docker compose logs chat for authentication errors from your provider, and confirm the key in .env and in Settings → AI Providers.

Next Steps ​

Learn more ​

Advanced features ​

Deployment ​

Tips for Success ​

Performance ​

  1. Use SSD storage for faster document processing
  2. Allocate 8 GB+ RAM to Docker when self-hosting the full stack
  3. Use local models for privacy and cost savings
  4. Split large corpora into collections so retrieval stays focused

Best practices ​

  1. Organize documents into collections by topic
  2. Use descriptive names for easy searching
  3. Ask specific questions for better answers
  4. Review sources to verify information
  5. Share collections with your team

Security (self-hosted) ​

  1. Set a strong POSTGRES_PASSWORD and JWT_SECRET before first boot
  2. Use strong passwords for user accounts
  3. Put a TLS-terminating reverse proxy in front of the gateway before exposing it
  4. Keep API keys in .env and never commit that file
  5. Back up the Postgres volume regularly

Getting Help ​

Community support ​

Documentation ​

  • User Guide — every feature, explained
  • Editions — what the Community Edition includes versus the hosted product
  • Architecture — how retrieval, the knowledge graph and deep research fit together

You're Ready! ​

Start uploading documents and asking questions.


Try Scrapalot Now

Web app (free Researcher plan): start at scrapalot.app — no install needed.

Desktop app: Windows, macOS and Linux.

Android app: download the APK and sideload it.

Self-hosting: Community Edition on GitHub — free, AGPL-3.0.

Need the full research suite or team features? See Pricing for Pro, Team and Enterprise.

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.