Resources
Community resources, contributing guidelines, and additional information about Scrapalot.
Community
GitHub Repository
Community Edition: github.com/sime2408/scrapalot (AGPL-3.0)
- Source code for the self-hostable stack
- Issue tracking
- Pull requests
- Discussions — questions, ideas, show and tell
Getting Help
Documentation:
- Getting Started - Installation and quick start
- User Guide - Comprehensive usage guide
- Architecture - Technical architecture details
Support Channels:
- GitHub Discussions - Ask questions and get help from the community
- Discord - Real-time chat, self-hosting help, and roadmap discussion
- GitHub Issues - Bug reports and feature requests
- Email - Account and billing questions; priority support on Enterprise
Contributing
How to Contribute
We welcome contributions from the community! Here's how you can help:
1. Report Bugs
- Check existing issues first
- Provide detailed reproduction steps
- Include environment details
- Add relevant logs or screenshots
2. Suggest Features
- Start an Ideas discussion, or open an issue
- Explain the use case
- Describe expected behavior
- Consider implementation approach
3. Submit Pull Requests
- Fork the repository
- Create a feature branch
- Write tests for new features
- Follow coding standards
- Submit PR with clear description
Development Setup
The Community Edition is a single repository containing all four services (ui/, gw/, backend/, chat/) and a Docker Compose file that runs the whole stack.
# Clone the Community Edition
git clone https://github.com/sime2408/scrapalot.git
cd scrapalot
# Configure: set POSTGRES_PASSWORD, JWT_SECRET, and an LLM key
cp .env.example .env
# Build and launch the whole stack
docker compose up -d --build
docker compose logs -fThe web app comes up on http://localhost:3000. Database migrations (Liquibase for the Kotlin backend, Alembic for the Python backend) run automatically on first boot.
Working on a single service while the rest of the stack runs in Docker:
# Frontend
cd ui && npm install && npm run dev
# Python AI backend tests
docker compose exec chat python -m pytest tests/
# Kotlin backend tests
cd backend && ./gradlew testCoding Standards
Python (Backend):
- Follow PEP 8 style guide
- Use type hints
- Write docstrings for functions
- Add tests for new features
TypeScript (Frontend):
- Follow ESLint configuration
- Use TypeScript types
- Write component tests
- Document complex logic
Documentation:
- Update relevant docs
- Add examples
- Include configuration details
- No emojis in documentation
Pull Request Process
Fork and Branch
- Fork the repository
- Create descriptive branch name
- Keep changes focused
Development
- Write code following standards
- Add tests
- Update documentation
Testing
- Run all tests locally
- Verify changes work
- Check for regressions
Submit PR
- Clear title and description
- Reference related issues
- Request review
Review Process
- Address review comments
- Update based on feedback
- Maintain clean commit history
License
Scrapalot follows an open-core model — see Editions for the full breakdown.
- Community Edition — the free, self-hostable core, published under the AGPL-3.0 open-source license. Run it on your own infrastructure with Docker Compose.
- Hosted product (free Researcher tier, paid Pro / Team / Enterprise) — a proprietary, managed cloud service that adds the advanced AI surfaces (deep research, knowledge graph, AI Scientist, voice, integrations), team collaboration, and the native desktop and Android apps.
Copyright (c) 2024-2026 Scrapalot. The hosted product and its proprietary modules — all rights reserved; the Community Edition is licensed under AGPL-3.0.
Changelog
Version 1.0.0 (Latest)
Features:
- 21 RAG strategies with 10 orchestrators
- Multi-database architecture (PostgreSQL + pgvector, Redis, Neo4j)
- 19 selectable chunking strategies
- Cloud and local model support
- Real-time WebSocket communication
- Background workers for document processing
Improvements:
- Optimized vector search performance
- Enhanced GPU acceleration support
- Improved error handling
- Better documentation
Bug Fixes:
- Fixed authentication edge cases
- Resolved memory leaks in workers
- Corrected workspace permission edge cases
Roadmap
Short-term (Next 3 months)
Features:
- Additional RAG strategies
- More cloud provider integrations
- Enhanced model management UI
- Improved monitoring dashboard
Improvements:
- Performance optimizations
- Better error messages
- Enhanced documentation
- More examples
Long-term (6-12 months)
Features:
- Multi-tenant architecture
- Advanced analytics
- Custom RAG strategy builder
- Enterprise SSO integration
Improvements:
- Scalability enhancements
- Advanced caching
- Automated testing framework
- Comprehensive benchmarks
Acknowledgments
Technologies
Scrapalot is built with:
Backend:
- Kotlin + Spring Boot - Accounts, workspaces, notes, settings
- Spring Cloud Gateway - API gateway
- Python + gRPC - AI service (RAG, deep research, document processing)
- PostgreSQL + pgvector - Vector database
- Redis - Caching and event streams
- Neo4j - Knowledge graph
- Celery - Background workers
Frontend:
- React - UI framework
- TypeScript - Type safety
- Vite - Build tool
- STOMP WebSocket - Real-time communication
AI/ML:
- OpenAI, Anthropic, Google and other providers - Cloud models
- llama.cpp - Local model inference
- sentence-transformers - Embeddings
- Pydantic AI and LangChain - Agents and RAG
- Whisper - Speech-to-text
Apps:
- Electron - Desktop app
- TipTap + Y.js - Collaborative notes editor
Contributors
Thank you to all contributors who have helped make Scrapalot better!
FAQ
General Questions
Q: Is Scrapalot free to use? A: Yes, in two ways. The Community Edition is free and open source (AGPL-3.0) — self-host it with Docker Compose, no quotas. On the hosted cloud, the Researcher tier is free; paid Pro, Team and Enterprise tiers add the advanced surfaces. See Editions and Pricing.
Q: Can I use my own AI models? A: Yes, Scrapalot supports local models (Ollama, LM Studio, vLLM, and GGUF models run by the desktop app) and cloud providers.
Q: What document formats are supported? A: PDF, EPUB, DOCX, RTF, TXT, MD, XLSX, XLS, CSV, TSV, and audio/video (transcribed). See Uploading Documents for the full list.
Q: How do I deploy to production? A: See Deployment Guide for detailed instructions.
Technical Questions
Q: How does RAG work in Scrapalot? A: See RAG Architecture for comprehensive explanation.
Q: Can I customize chunking strategies? A: Yes, 19 chunking strategies can be selected under Settings → Documents. See Document Processing for details.
Q: Is my data secure? A: Yes — workspace isolation, role-based access, encrypted connections, and Google sign-in. See Security.
Related Links
- GitHub Repository - Source code
- Documentation - Full documentation
- API Reference - REST API docs
- Architecture - Technical details
Have questions? Ask on GitHub Discussions, chat on Discord, or open an issue.