Skip to content

Uploading Documents ​

Last Updated: September 2026

Learn how to add documents to Scrapalot, what happens while they are processed, and how to manage them afterwards.

Upload Methods ​

Supported File Formats ​

KindExtensions
Documents.pdf, .epub, .docx, .rtf, .txt, .md
Spreadsheets & data.xlsx, .xls, .csv, .tsv
Audio (transcribed).mp3, .wav, .m4a, .ogg, .webm, .flac, .aac, .opus, .wma, .aiff, .amr
Video (transcribed).mp4, .mov, .mkv, .avi, .m4v, .mpeg, .mpg, .flv, .wmv

Audio and video are transcribed to text automatically, then indexed like any other document.

Not supported: the legacy Word format .doc (save it as .docx first), PowerPoint, HTML and JSON files, and DRM-protected e-books.

File Size Limits

  • Maximum file size: 200 MB per file
  • Files larger than 20 MB, PDFs with more than 1,000 pages, and scanned PDFs that need OCR are held back and wait for your confirmation before processing, because they take significant time and memory (see Deferred documents)
  • Your plan's document and storage quotas also apply — see Pricing

Document Processing Pipeline ​

Processing Stages Explained ​

Stage 1: Text extraction ​

What happens:

  • The file is read and its text extracted. PDFs are parsed with PyMuPDF4LLM by default, or with Docling for layout-aware extraction; pages without a text layer fall back to OCR
  • EPUBs are split into their chapters
  • Audio and video are transcribed with Whisper
  • Title, authors and other metadata are extracted

Common errors:

  • "DRM-protected file — cannot be processed"
  • "Empty document — no extractable text"
  • "Scanned PDF — enable OCR in settings and retry"

Stage 2: Structure and chunking ​

What happens:

  • A document hierarchy is built: book, chapters, sections and paragraphs
  • The text is split into chunks with the selected splitting strategy (default: Enhanced Markdown)
  • The hierarchy is what later lets answers expand from a matched snippet to its full section

You can choose the strategy under Settings → Documents → Advanced Chunking Options. 19 strategies are available, including Semantic, Recursive, Proposition, Hierarchical, Topic-Based, Sliding Window, Agentic, Concept-Aware, Narrative Structure, Token-Based, Document Structure, HTML Structure and Code Structure.

See: Document Processing Architecture

Stage 3: Embeddings ​

What happens:

  • Each chunk is converted to a vector embedding and stored in PostgreSQL with pgvector
  • The default embedding model is all-MiniLM-L6-v2 (384 dimensions)

Stage 4: Knowledge graph (optional) ​

What happens:

  • Entities (people, concepts, places, works…) and their relationships are extracted into the knowledge graph
  • This links your documents to each other and powers graph-aware retrieval

The knowledge graph is part of the Pro plan and above.

Processing options ​

Before you click Compose, Processing options on the Upload tab let you choose which steps run:

  • Titles and content — read the file and pull out its text
  • Embeddings — index the text so search can find it
  • Knowledge graph — extract entities and link them

A step you skip can be run later from the document's status in the library (Build embeddings, Build graph, Create summary).

Document States ​

The Upload tab groups documents into All, Processing, Pending, Failed and Ready.

How to Upload Documents ​

Upload tab ​

  1. Open Knowledge Stacks

    • From the sidebar, then select the collection to upload into (or create one)
  2. Add your files

    • Drag and drop files onto the Upload tab, or click Browse Files
    • Add as many files as you like at once
  3. Click Compose

    • Staged files show "Waiting to compose" until you do
  4. Monitor progress

    • Each file shows live progress and its current step
    • Processing continues on the server if you close the dialog
  5. Document ready

    • It shows Ready for search and can be queried immediately

Only a limited number of processing jobs run at the same time. If you see Processing Limit Reached or Too many collections processing, wait for the current jobs to finish (or abort one) before starting more.

Deferred documents ​

Files over 20 MB, PDFs over 1,000 pages and scanned PDFs are marked Deferred instead of starting straight away. A banner shows how many are waiting; click Process all to start them.

Import papers by identifier ​

Use Import papers to paste DOI, arXiv or PubMed identifiers (ids or full links). Scrapalot fetches the title, authors, year, venue and abstract from Semantic Scholar and creates a document per paper — no PDF upload needed. For papers that have one, Find open-access PDF looks up a free copy via Unpaywall.

External Connectors ​

Import from cloud storage and other apps in the Connectors tab of Knowledge Stacks:

  • Google Drive, Dropbox, OneDrive
  • Notion, Confluence, Slack
  • Zotero

Connectors are part of the Pro plan and above. See the Integrations Guide.

Mounted folders (desktop app) ​

In the desktop app, Settings → Workspaces → Mount a folder points Scrapalot at books you already have on disk. Subfolders become collections, your files stay where they are, and new books can be imported automatically. See Apps & Platforms.

Attach to a chat ​

You can also attach a file directly to a conversation with Attach Files in the chat toolbar. Attached documents count towards your plan's document quota.

Best Practices ​

Organizing Documents ​

Create topic-based collections:

Research Papers/
├── AI Safety/
├── Climate Science/
└── Medical Research/

Work Documents/
├── Q1 Reports/
├── Product Specs/
└── Meeting Notes/

Benefits:

  • Faster, more accurate searches
  • Better automatic collection routing
  • Easier management and sharing

See Managing Collections.

Optimize for Processing ​

Before uploading:

  1. Remove password protection and DRM

    • Encrypted PDFs and DRM-protected e-books cannot be processed
  2. Prefer text PDFs over scans

    • A native PDF is parsed in seconds; a scan needs OCR page by page
    • OCR quality is the ceiling on answer quality — a clean copy beats any setting
  3. Check file integrity

    • Verify the file opens correctly
  4. Mind the size

    • Split or compress files over 200 MB

Troubleshooting ​

Upload Fails ​

"File too large (max 200MB)"

  • Compress the PDF or split the file

"Storage Quota Exceeded"

  • Delete documents you no longer need, or move up a plan — see Pricing

"Already exists"

  • The same file is already in this collection. Delete the existing copy first, or rename the file

Unsupported file

  • Check the formats above; convert .doc to .docx

Processing Stuck or Failed ​

Failed documents show the reason next to them. Most can be retried with Retry processing:

  • "Worker stalled" / "Task never picked up by worker" — retry
  • "Processing exceeded the time limit" — retry; very large documents may need splitting
  • "LLM quota exhausted" — you have run out of AI tokens or provider credit; add credit, switch provider, or wait for the monthly reset
  • "Source file no longer on disk" — upload the file again
  • "likely scanned-image or DRM-protected" — the file has little extractable text

If a document has not moved in half an hour, see Troubleshooting.

Document Not Searchable ​

Possible causes:

  1. Document still processing (check its status)
  2. Embeddings were skipped in Processing options — run Build embeddings
  3. The wrong collection is selected in chat
  4. The document contains images only and OCR produced little text

Managing Documents ​

The Library tab of Knowledge Stacks lists every document in your workspace.

Find and organize ​

  • Search by title, author or collection, filter and sort
  • Tags and priority (high, normal, low) per document
  • Smart Collections — saved searches that act as filtered views
  • Move to… or add a document to another collection (drag and drop works too)
  • Review duplicates — merge or keep copies

Read and summarize ​

Click a document to open it in the built-in PDF, EPUB, DOCX or spreadsheet viewer, with annotations. From a document you can generate a Book Summary, per-chapter summaries, and read the summary aloud.

Export ​

  • Export Citations — export document metadata in citation formats
  • Download — save the original file to your computer

Delete ​

Select one or more documents and choose what to remove:

  • Delete all — the document, its embeddings, graph and file
  • Delete embeddings, Delete graph or Delete file — remove only that part

Deletion is Permanent

Deleted documents cannot be recovered. Download a copy first if needed.

Storage & Quotas ​

Cloud plans ​

PlanDocumentsStorage
Researcher (Free)1001 GB
Pro5,0005 GB
Team50,00030 GB
EnterpriseUnlimitedUnlimited

Your current usage is shown in Settings → Account.

Desktop app ​

In the desktop app you can keep original files on This computer instead of the Scrapalot cloud (Settings → Account → Where your books are kept). Files kept locally have no storage limit, but are not available in the browser or on your other devices.


Next: Once your documents are uploaded, learn how to ask effective questions to get the best answers.

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.