Uploading Documents
Last Updated: September 2026
Learn how to add documents to Scrapalot, what happens while they are processed, and how to manage them afterwards.
Upload Methods
Supported File Formats
| Kind | Extensions |
|---|---|
| Documents | .pdf, .epub, .docx, .rtf, .txt, .md |
| Spreadsheets & data | .xlsx, .xls, .csv, .tsv |
| Audio (transcribed) | .mp3, .wav, .m4a, .ogg, .webm, .flac, .aac, .opus, .wma, .aiff, .amr |
| Video (transcribed) | .mp4, .mov, .mkv, .avi, .m4v, .mpeg, .mpg, .flv, .wmv |
Audio and video are transcribed to text automatically, then indexed like any other document.
Not supported: the legacy Word format .doc (save it as .docx first), PowerPoint, HTML and JSON files, and DRM-protected e-books.
File Size Limits
- Maximum file size: 200 MB per file
- Files larger than 20 MB, PDFs with more than 1,000 pages, and scanned PDFs that need OCR are held back and wait for your confirmation before processing, because they take significant time and memory (see Deferred documents)
- Your plan's document and storage quotas also apply — see Pricing
Document Processing Pipeline
Processing Stages Explained
Stage 1: Text extraction
What happens:
- The file is read and its text extracted. PDFs are parsed with PyMuPDF4LLM by default, or with Docling for layout-aware extraction; pages without a text layer fall back to OCR
- EPUBs are split into their chapters
- Audio and video are transcribed with Whisper
- Title, authors and other metadata are extracted
Common errors:
- "DRM-protected file — cannot be processed"
- "Empty document — no extractable text"
- "Scanned PDF — enable OCR in settings and retry"
Stage 2: Structure and chunking
What happens:
- A document hierarchy is built: book, chapters, sections and paragraphs
- The text is split into chunks with the selected splitting strategy (default: Enhanced Markdown)
- The hierarchy is what later lets answers expand from a matched snippet to its full section
You can choose the strategy under Settings → Documents → Advanced Chunking Options. 19 strategies are available, including Semantic, Recursive, Proposition, Hierarchical, Topic-Based, Sliding Window, Agentic, Concept-Aware, Narrative Structure, Token-Based, Document Structure, HTML Structure and Code Structure.
See: Document Processing Architecture
Stage 3: Embeddings
What happens:
- Each chunk is converted to a vector embedding and stored in PostgreSQL with pgvector
- The default embedding model is
all-MiniLM-L6-v2(384 dimensions)
Stage 4: Knowledge graph (optional)
What happens:
- Entities (people, concepts, places, works…) and their relationships are extracted into the knowledge graph
- This links your documents to each other and powers graph-aware retrieval
The knowledge graph is part of the Pro plan and above.
Processing options
Before you click Compose, Processing options on the Upload tab let you choose which steps run:
- Titles and content — read the file and pull out its text
- Embeddings — index the text so search can find it
- Knowledge graph — extract entities and link them
A step you skip can be run later from the document's status in the library (Build embeddings, Build graph, Create summary).
Document States
The Upload tab groups documents into All, Processing, Pending, Failed and Ready.
How to Upload Documents
Upload tab
Open Knowledge Stacks
- From the sidebar, then select the collection to upload into (or create one)
Add your files
- Drag and drop files onto the Upload tab, or click Browse Files
- Add as many files as you like at once
Click Compose
- Staged files show "Waiting to compose" until you do
Monitor progress
- Each file shows live progress and its current step
- Processing continues on the server if you close the dialog
Document ready
- It shows Ready for search and can be queried immediately
Only a limited number of processing jobs run at the same time. If you see Processing Limit Reached or Too many collections processing, wait for the current jobs to finish (or abort one) before starting more.
Deferred documents
Files over 20 MB, PDFs over 1,000 pages and scanned PDFs are marked Deferred instead of starting straight away. A banner shows how many are waiting; click Process all to start them.
Import papers by identifier
Use Import papers to paste DOI, arXiv or PubMed identifiers (ids or full links). Scrapalot fetches the title, authors, year, venue and abstract from Semantic Scholar and creates a document per paper — no PDF upload needed. For papers that have one, Find open-access PDF looks up a free copy via Unpaywall.
External Connectors
Import from cloud storage and other apps in the Connectors tab of Knowledge Stacks:
- Google Drive, Dropbox, OneDrive
- Notion, Confluence, Slack
- Zotero
Connectors are part of the Pro plan and above. See the Integrations Guide.
Mounted folders (desktop app)
In the desktop app, Settings → Workspaces → Mount a folder points Scrapalot at books you already have on disk. Subfolders become collections, your files stay where they are, and new books can be imported automatically. See Apps & Platforms.
Attach to a chat
You can also attach a file directly to a conversation with Attach Files in the chat toolbar. Attached documents count towards your plan's document quota.
Best Practices
Organizing Documents
Create topic-based collections:
Research Papers/
├── AI Safety/
├── Climate Science/
└── Medical Research/
Work Documents/
├── Q1 Reports/
├── Product Specs/
└── Meeting Notes/Benefits:
- Faster, more accurate searches
- Better automatic collection routing
- Easier management and sharing
See Managing Collections.
Optimize for Processing
Before uploading:
Remove password protection and DRM
- Encrypted PDFs and DRM-protected e-books cannot be processed
Prefer text PDFs over scans
- A native PDF is parsed in seconds; a scan needs OCR page by page
- OCR quality is the ceiling on answer quality — a clean copy beats any setting
Check file integrity
- Verify the file opens correctly
Mind the size
- Split or compress files over 200 MB
Troubleshooting
Upload Fails
"File too large (max 200MB)"
- Compress the PDF or split the file
"Storage Quota Exceeded"
- Delete documents you no longer need, or move up a plan — see Pricing
"Already exists"
- The same file is already in this collection. Delete the existing copy first, or rename the file
Unsupported file
- Check the formats above; convert
.docto.docx
Processing Stuck or Failed
Failed documents show the reason next to them. Most can be retried with Retry processing:
- "Worker stalled" / "Task never picked up by worker" — retry
- "Processing exceeded the time limit" — retry; very large documents may need splitting
- "LLM quota exhausted" — you have run out of AI tokens or provider credit; add credit, switch provider, or wait for the monthly reset
- "Source file no longer on disk" — upload the file again
- "likely scanned-image or DRM-protected" — the file has little extractable text
If a document has not moved in half an hour, see Troubleshooting.
Document Not Searchable
Possible causes:
- Document still processing (check its status)
- Embeddings were skipped in Processing options — run Build embeddings
- The wrong collection is selected in chat
- The document contains images only and OCR produced little text
Managing Documents
The Library tab of Knowledge Stacks lists every document in your workspace.
Find and organize
- Search by title, author or collection, filter and sort
- Tags and priority (high, normal, low) per document
- Smart Collections — saved searches that act as filtered views
- Move to… or add a document to another collection (drag and drop works too)
- Review duplicates — merge or keep copies
Read and summarize
Click a document to open it in the built-in PDF, EPUB, DOCX or spreadsheet viewer, with annotations. From a document you can generate a Book Summary, per-chapter summaries, and read the summary aloud.
Export
- Export Citations — export document metadata in citation formats
- Download — save the original file to your computer
Delete
Select one or more documents and choose what to remove:
- Delete all — the document, its embeddings, graph and file
- Delete embeddings, Delete graph or Delete file — remove only that part
Deletion is Permanent
Deleted documents cannot be recovered. Download a copy first if needed.
Storage & Quotas
Cloud plans
| Plan | Documents | Storage |
|---|---|---|
| Researcher (Free) | 100 | 1 GB |
| Pro | 5,000 | 5 GB |
| Team | 50,000 | 30 GB |
| Enterprise | Unlimited | Unlimited |
Your current usage is shown in Settings → Account.
Desktop app
In the desktop app you can keep original files on This computer instead of the Scrapalot cloud (Settings → Account → Where your books are kept). Files kept locally have no storage limit, but are not available in the browser or on your other devices.
Related Topics
- Asking Questions - Search your uploaded documents
- Collections Management - Organize your documents
- Document Processing Architecture - Technical details
- Integrations - External connectors
Next: Once your documents are uploaded, learn how to ask effective questions to get the best answers.