Skip to content

Architecture Overview ​

Last Updated: September 2026

Scrapalot is a proprietary, enterprise-grade RAG platform built with modern, scalable microservices architecture for intelligent document processing and AI-powered knowledge retrieval.

Cloud-First Platform

Scrapalot is delivered as a managed cloud (SaaS) service - the standard way to use it. Enterprise customers can request on-premise deployment of this architecture to their own infrastructure on demand.

System Overview ​

Scrapalot uses a modern microservices architecture with clear separation of concerns:

Request Flow ​

User asks a question:

  1. React UI sends request to Gateway (port 8080)
  2. Gateway routes to Kotlin Backend (port 8091)
  3. Kotlin authenticates, validates, retrieves user context
  4. Kotlin calls Python Chat via gRPC (port 9091) with user IDs
  5. Python performs AI operations (RAG, LLM, embeddings)
  6. Results stream back through the chain to user as a sequence of typed packets (status, reasoning, answer text, citations, charts, research progress, and so on)

Almost all traffic takes this path. A small set of AI-only REST endpoints (research export and templates, hypotheses and what-if analysis, the Research Council, citation stance classification, DOI import, podcasts, voice, "explain" tools and the MCP server endpoint) plus the real-time WebSockets are routed by the Gateway straight to the Python service. The Python service also calls back into the Kotlin backend over gRPC when it needs data Kotlin owns, such as notes and annotations.

Core Components ​

Web Interface ​

Modern, responsive React application

Features:

  • Real-time streaming answers
  • Document management
  • Multi-language support (English, Spanish, French, German, Italian, Croatian)
  • Dark/light themes
  • Mobile-friendly

Technology:

  • React 18.3.1 with TypeScript 5.9.3
  • Tailwind CSS + Shadcn/ui
  • Real-time STOMP WebSocket updates
  • Vite 5 for fast builds
  • i18next (en, es, fr, de, it, hr translations)

API Gateway ​

Spring Cloud Gateway routing layer

Capabilities:

  • Single entry point for all UI requests
  • Routes to appropriate backend services
  • JWT / API-key validation
  • Per-subscription-tier rate limiting on AI requests
  • Circuit breaking
  • WebSocket passthrough

Technology:

  • Spring Cloud Gateway
  • Port 8080
  • Routes: /api/v1/** → Kotlin Backend by default; a short list of AI-only paths and the WebSocket endpoints → Python Chat

Kotlin Backend ​

User data and business logic service

Capabilities:

  • User authentication (JWT, OAuth, Sessions)
  • All user data management (users, workspaces, collections, notes, projects)
  • Chat sessions and messages, workspace chat, direct messages and huddles
  • Settings (source of truth), subscriptions and API key management
  • Lightweight AI tasks (describe, translate, summarize) via Spring AI
  • Calls Python AI service via gRPC for anything that needs documents, embeddings or the knowledge graph

Technology:

  • Kotlin 2.1.0 + Spring Boot 3.4.1, Spring AI 1.0
  • PostgreSQL (scrapalot_backend database)
  • gRPC client (calls Python on port 9091) and gRPC server (called by Python and the Gateway)
  • STOMP WebSocket broker for workspace chat, DMs and huddles
  • Liquibase migrations
  • MapStruct for DTOs

Python Chat Service ​

Pure AI/ML operations service

Capabilities:

  • Advanced RAG strategies (21 strategies + 9 orchestrators)
  • Multi-provider AI model support
  • Document processing and embeddings (22 chunking strategies)
  • Vector search with pgvector
  • Knowledge graph (Neo4j) and the opt-in Personal Brain (per-user memory)
  • Deep Research (5-phase architecture)
  • Real-time job and research progress over STOMP, collaborative notes over Y.js

Technology:

  • Python 3.12.8 + FastAPI (health, WebSockets and a few AI-only REST endpoints)
  • PostgreSQL with pgvector
  • LangChain 1.3 (RAG framework)
  • Pydantic AI 2.9 (agent framework)
  • Edge-TTS (text-to-speech)
  • gRPC server (port 9091)
  • Celery workers for background jobs

Data Storage ​

Flexible, powerful databases

PostgreSQL with pgvector (port 5432):

  • Two databases: scrapalot_backend (Kotlin), scrapalot (Python)
  • Vector search capabilities for RAG
  • Local Docker container deployment
  • Workspace-scoped access enforced by the application

Redis (port 6379):

  • DB 0: Python cache
  • DB 1: Kotlin
  • DB 2: Gateway (rate limiting)
  • DB 3: Celery broker (background jobs)
  • Redis Streams for cross-service data sync (see below)

Neo4j (port 7687):

  • Knowledge graph storage
  • Entity and relationship extraction
  • Cross-document connection discovery
  • Personal Brain memory subgraph (opt-in per user)

Benefits:

  • Scalable to millions of documents
  • Fast semantic search with pgvector
  • Secure multi-tenant isolation
  • Real-time event propagation

Cross-Service Sync (Redis Streams) ​

Each service owns its own tables; changes are replicated through Redis Streams (scrapalot:stream:<name>) with consumer groups. Settings and model-provider changes use a SAGA handshake (the remote side commits and acknowledges on saga_ack before the local commit lands).

DirectionStreams
Kotlin → Pythonworkspaces, collections, users, connectors, mcp_servers, annotations, message_feedback, user_settings (SAGA)
Python → Kotlinmodel_providers (SAGA), token_usage, collection_summary, session_summary, job_push
Bothsaga_ack

Background Workers (Celery) ​

QueueContainerWork
documentsscrapalot-workersUpload, reprocess and batch document processing, connector imports
fastscrapalot-workersSummaries, hierarchy rebuilds, connector discovery and backup, podcasts, paper generation, maintenance
researchscrapalot-workersDeep Research runs started with "Run in background"
graph_extractionscrapalot-workers-graphEntity extraction, graph housekeeping, Personal Brain memory extraction

A Celery Beat scheduler runs in the main worker container for periodic maintenance.

AI Models ​

Your choice of providers

Cloud Options:

  • OpenAI, Anthropic, Google Gemini
  • DeepSeek, Groq, OpenRouter, Z.ai
  • Any OpenAI-compatible endpoint
  • Scrapalot AI (bundled models on hosted plans)

Local Options:

  • Ollama (easy local server)
  • LM Studio (GPU-accelerated)
  • Direct GGUF models (llama.cpp)
  • vLLM (production inference)
  • The desktop app's built-in local runtime

Flexibility:

  • Mix and match providers
  • Change models anytime
  • Export and migrate your data anytime

RAG Engine: How Search Works ​

Scrapalot uses Tri-Modal Fusion - three search methods working together:

Why Three Search Methods? ​

Semantic Search (Vector embeddings):

  • Understands conceptual similarity
  • "renewable energy" finds "solar power"
  • Best for: Conceptual questions

Keyword Search (BM25):

  • Exact term matching
  • "error 221" finds exactly that
  • Best for: Technical terms, codes, specific phrases

Graph Search (Neo4j):

  • Understands relationships
  • "How are X and Y connected?"
  • Best for: Relationship questions

Intelligent Routing: The system automatically chooses the best method(s) for each question. You don't need to think about it - it just works.

How Documents Become Knowledge ​

Processing steps:

  1. Upload - You upload document
  2. Extract - Text extracted from PDF/Word/etc
  3. Chunk - Intelligently split into searchable segments
  4. Embed - Generate vector embeddings
  5. Index - Store in database with metadata
  6. Ready - Available for search immediately

Processing time:

  • Small doc (10 pages): ~30 seconds
  • Medium doc (100 pages): ~2 minutes
  • Large doc (500 pages): ~10 minutes

Security & Privacy ​

Data Protection ​

Multi-tenant isolation:

  • Your data completely separate from others
  • Workspace-based access control checked on every request
  • The AI service only receives the user, workspace and collection IDs the backend has already authorized

Encryption:

  • All connections encrypted (TLS/SSL)
  • API keys encrypted at rest
  • Secure credential storage

Access control:

  • JWT authentication
  • Role-based permissions
  • Workspace sharing controls

Privacy Options ​

Cloud deployment:

  • Use managed services or self-hosted Docker
  • Your data isolated in your account
  • Encrypted in transit and at rest

On-premise (Enterprise, on-demand):

  • Complete data sovereignty
  • Never leaves your infrastructure
  • Full control over everything
  • Local AI models for zero external calls

Deployment Options ​

Perfect for getting started:

  • Single server deployment
  • Self-hosted PostgreSQL with pgvector (Docker)
  • Cloud AI models (OpenAI, Claude)
  • 10 minutes to running

Requirements:

  • 4GB RAM minimum
  • Docker installed
  • Internet connection

Production Deployment ​

For serious use:

  • Load-balanced API servers
  • Database with replicas
  • Redis caching layer
  • Background worker pool
  • Monitoring and logging

Scaling:

  • Horizontal API scaling
  • Worker pool sizing
  • Read replicas for database
  • CDN for frontend

Privacy-First Deployment ​

For maximum data control (Enterprise on-premise, on-demand):

  • On-premise infrastructure
  • Local AI models only
  • Air-gapped if needed
  • Complete audit trail

Technology Stack ​

Frontend (scrapalot-ui) ​

  • React 18.3.1 + TypeScript 5.9.3
  • Vite 5 (fast builds)
  • Tailwind CSS + Shadcn/ui (Radix primitives)
  • STOMP WebSocket for real-time updates
  • i18next (English, Spanish, French, German, Italian, Croatian)
  • TipTap (core 2.10.3, newer extensions 2.27.1) collaborative editor
  • Framer Motion 12.23.24 (animations)

Gateway (scrapalot-gw) ​

  • Spring Cloud Gateway
  • Kotlin 2.1.0
  • Port 8080 (single entry point)
  • Redis-backed per-tier rate limiting

Kotlin Backend (scrapalot-backend) ​

  • Kotlin 2.1.0 + Spring Boot 3.4.1
  • PostgreSQL (local Docker)
  • Liquibase (migrations)
  • MapStruct (DTO mapping)
  • gRPC client and server
  • Spring AI (lightweight generation tasks)
  • Redis Streams (cross-service sync)

Python Chat (scrapalot-chat) ​

  • Python 3.12.8 + FastAPI
  • SQLAlchemy / SQLModel ORM, Alembic migrations
  • LangChain 1.3 (RAG framework)
  • Pydantic AI 2.9 (agent framework)
  • llama-cpp-python (local models)
  • Edge-TTS (text-to-speech)
  • gRPC server (port 9091)
  • Celery (background workers)

Databases ​

  • PostgreSQL 18 with pgvector (two databases)
  • Redis 8 (Streams sync, Celery broker, rate limiting, caching)
  • Neo4j (knowledge graph and Personal Brain)

AI Providers (all optional) ​

  • OpenAI, Anthropic, Google Gemini
  • Ollama, vLLM, LM Studio, LlamaCPP
  • DeepSeek, Groq, OpenRouter, Z.ai
  • Any OpenAI-compatible endpoint

What Makes Scrapalot Different ​

Most RAG systems use only vector search. Scrapalot combines three methods for superior accuracy.

Intelligent Routing ​

AI automatically selects the best search strategy for each question. No manual configuration needed.

Context Expansion ​

Understands document structure to provide complete context, not just isolated chunks.

Model Flexibility ​

Use any AI model - cloud, local, or mixed. Switch anytime without data migration.

Data Portability ​

Export and migrate your research data anytime. Enterprise on-premise deployment available on demand.

Production Ready ​

Enterprise security, multi-tenancy, real-time streaming, comprehensive monitoring.

Next Steps ​

Explore More

Get Started:

Learn the Features:

Advanced Topics:


Scrapalot makes advanced RAG accessible. Complex technology, simple experience.

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.