Skip to content

External Connectors ​

Last Updated: September 2026

Automatically fetch and sync documents from external sources. Keep your knowledge base up-to-date without manual uploads.

What Are Connectors? ​

Connectors integrate Scrapalot with external services to automatically:

  • Fetch documents from cloud storage and web sources
  • Sync on schedule to keep content current
  • Handle authentication securely
  • Monitor for updates and fetch new content
  • Respect rate limits to avoid service issues

Connector Architecture ​

Connectors use Redis Streams K→P sync to propagate connector configuration from the Kotlin backend (owner) to the Python AI service (consumer). This ensures both services stay in sync when connectors are created, updated, or deleted. Credentials are stored encrypted. The actual listing, downloading and processing of files runs on the background workers (see Background Workers).

Discover, then import ​

A connected folder can be far larger than what you want on the server. Instead of importing everything on connect, a cloud-storage connector first discovers what the source holds — it records each remote file's name and metadata without downloading anything. Your library shows those discovered files next to the documents already imported, and you import the ones you want (one at a time or in bulk). Only imported files are downloaded, chunked and embedded.

If a collection already contains documents you uploaded by hand, pointing a connector at the same folder links them to their remote originals (matched by name and size) instead of importing duplicates; whatever is still missing is reported for import.

Backup to cloud storage ​

The Google Drive, Dropbox and OneDrive connectors also work in the other direction: you can push a document stored in Scrapalot out to your cloud storage as a backup copy.

Supported Sources ​

Workspace Connectors (7) ​

Google Drive ​

Automatically sync folders from Google Drive

Use cases:

  • Team documentation stored in shared folders
  • Project files that update regularly
  • Policies and procedures that change

Features:

  • Pick files through the Google Picker after signing in with Google (OAuth 2.0)
  • Read folders shared as "anyone with the link", including subfolders, without granting account access
  • Google Docs, Sheets and Slides are exported to PDF on import
  • Backup copies can be written back to Drive

Setup:

  1. Add the Google Drive connector to a collection
  2. Either sign in with Google and pick files, or paste the link of a publicly shared folder
  3. Choose a sync schedule
  4. Import the discovered files you want

Dropbox ​

Sync files from Dropbox cloud storage

Use cases:

  • Team documents stored in Dropbox
  • Shared folders and project files
  • Automatic sync when files change

Features:

  • OAuth 2.0 authentication (using a Dropbox app you register)
  • Folder sync with subfolders
  • Backup copies can be written back to Dropbox

Notion ​

Import pages and databases from Notion

Use cases:

  • Team wikis and documentation
  • Project management databases
  • Knowledge bases stored in Notion

Features:

  • Authenticates with a Notion internal-integration token; only pages you share with the integration are visible
  • Page and database import
  • Automatic sync on schedule

Confluence ​

Sync documentation from Atlassian Confluence

Use cases:

  • Enterprise documentation
  • Team knowledge bases
  • Technical specifications

Features:

  • Space and page import (optionally limited to one space or page tree)
  • Hierarchical content sync
  • Cloud (email + API token) and Server / Data Center (username + password)

Slack ​

Import messages and files from Slack channels

Use cases:

  • Team conversations and decisions
  • Shared files and documents
  • Knowledge scattered across channels

Features:

  • Channel history import
  • Authenticates with a Slack app's bot token (read scopes for channels, groups and users)

OneDrive / SharePoint ​

Connect to Microsoft OneDrive and SharePoint document libraries

One connector covers both — they share the same Microsoft Graph backend.

Use cases:

  • Enterprise document management
  • Team collaboration files
  • Compliance documentation

Features:

  • Drive and document library sync
  • Folder filtering
  • Microsoft Entra ID (Azure AD) app registration with read permissions
  • Backup copies can be written back to OneDrive

Zotero ​

Import references and PDFs from Zotero libraries

Use cases:

  • Academic reference management
  • Research paper collections
  • Bibliography management

Features:

  • Library and collection sync (authenticates with a Zotero API key)
  • PDF attachment import
  • Metadata extraction
  • Citation data preservation

A YouTube connector (video transcripts and metadata) is listed in the app as coming soon.

Academic Search Sources ​

These are not workspace connectors you configure and sync — they are literature sources that Deep Research and the notes assistant query live while working.

arXiv ​

Access preprints from arXiv repository

Use cases:

  • Latest research in physics, math, CS, and more
  • Pre-publication papers
  • Technical research

Features:

  • Search by topic, author, or arXiv ID
  • Download full paper PDFs
  • Filter by category and date
  • Automatic metadata extraction

Semantic Scholar ​

AI-powered academic paper search

Use cases:

  • Comprehensive literature search
  • Finding influential papers
  • Research trend analysis

Features:

  • Semantic search across academic papers
  • Citation and reference tracking
  • Author and topic analysis
  • Relevance-based ranking

PubMed ​

Biomedical and life-sciences literature

Use cases:

  • Clinical and biomedical research
  • Systematic reviews
  • Drug and trial literature

Features:

  • Search across MEDLINE and PubMed Central
  • Abstract and metadata retrieval
  • MeSH-aware querying

OpenAlex ​

Open catalogue of scholarly works

Use cases:

  • Broad cross-disciplinary coverage
  • Citation graphs and institutional data
  • Open-access link resolution

Features:

  • Very large index of works, authors and venues
  • Citation counts and references
  • Open-access status per work

Sync Scheduling ​

Schedule Options ​

Manual:

  • Fetch only when you trigger it
  • Good for one-time imports
  • Full control over timing

Hourly:

  • Keep content very current
  • Good for rapidly changing content
  • Higher API usage

Daily:

  • Balance between freshness and efficiency
  • Recommended for most use cases
  • Runs during low-activity hours

Weekly:

  • Light API usage
  • Good for stable content
  • Minimal resource impact

Automatic Updates ​

What happens during sync:

  1. Connector checks source for new/updated documents
  2. Downloads only changed files
  3. Queues documents for processing
  4. Updates existing documents if modified
  5. Sends notification when complete

Smart syncing:

  • Only fetches what changed
  • Deduplicates identical content
  • Preserves existing document metadata
  • Maintains citation links

Authentication & Security ​

OAuth 2.0 (Google Drive, Dropbox, OneDrive) ​

Secure, standard authentication:

  • Authorize once, tokens refresh automatically
  • Revoke access anytime from the provider's account settings
  • No password storage

Permission scope (Google Drive):

  • Scrapalot requests Google's narrow drive.file scope: it can see only the files you pick and the files it created itself (backup copies)
  • It cannot browse the rest of your Drive
  • Publicly shared folders are read without any account access at all

API Keys ​

Simple key-based authentication:

  • Store keys securely encrypted
  • Never exposed in logs
  • Easy to rotate
  • Revoke anytime

Security:

  • Keys encrypted at rest
  • Transmitted over TLS
  • Access controlled per user

Error Handling ​

Automatic Retry ​

If fetching fails:

  • Automatic retry with exponential backoff
  • Skip problematic documents, continue with others
  • Detailed error logging
  • User notification of issues

Common failures handled:

  • Temporary network issues
  • Rate limit exceeded (waits and retries)
  • Document temporarily unavailable
  • Authentication token expired (auto-refresh)

Notifications ​

You're informed when:

  • Sync completes successfully
  • Documents fail to fetch
  • Authentication expires
  • Rate limits approached
  • Service unavailable

Monitoring & Management ​

Connector Status ​

Track connector health:

  • Last successful sync time
  • Next scheduled sync
  • Documents fetched
  • Success/failure counts
  • Current status (active, paused, error)

Available actions:

  • Trigger manual sync
  • Pause/resume syncing
  • Edit configuration
  • View sync history
  • Delete connector

Sync History ​

View past activity:

  • Sync timestamps
  • Success/failure status
  • Documents processed
  • Error messages
  • Processing time

Use for:

  • Troubleshooting issues
  • Verifying sync schedule
  • Monitoring API usage
  • Audit trail

Rate Limiting & Quotas ​

Automatic Rate Management ​

Respects API limits:

  • Configurable delays between requests
  • Automatic backoff on limit warnings
  • Queue management to spread load
  • Pause and resume on quota exhaustion

Google Drive:

  • Discovery costs one API request per folder, so a folder walked within the last hour is not walked again
  • Automatic throttling built-in

Best Practices ​

Connector Setup ​

Optimize your connectors:

  • Use specific folders/URLs, not entire drives
  • Filter by relevant file types
  • Set appropriate sync frequency
  • Group related content in same connector

Performance ​

Efficient syncing:

  • Schedule during low-usage hours
  • Avoid hourly sync unless necessary
  • Use manual sync for one-time imports
  • Monitor document count growth

Organization ​

Keep it maintainable:

  • Name connectors descriptively
  • Document what each connector fetches
  • Review and clean unused connectors
  • Archive completed syncs

Security ​

Protect your data:

  • Use minimum necessary permissions
  • Review connector access regularly
  • Rotate API keys periodically
  • Remove unused connectors

Troubleshooting ​

Connector Won't Authenticate ​

Check:

  • Credentials are correct
  • OAuth consent not expired
  • API key is valid
  • Service is accessible

Solutions:

  • Re-authorize OAuth
  • Generate new API key
  • Check firewall/network
  • Verify service status

No Documents Fetched ​

Common causes:

  • Empty folder/source
  • File type filters too restrictive
  • Permission issues
  • Rate limit reached

Solutions:

  • Verify source has content
  • Adjust file type filters
  • Check permissions
  • Review quota usage

Sync Failing Repeatedly ​

Investigate:

  • Error messages in history
  • Service health status
  • Authentication validity
  • Network connectivity

Fix:

  • Address specific error
  • Re-authenticate if needed
  • Check source availability
  • Contact support if persistent

Use Case Examples ​

Team Documentation ​

Scenario: Engineering team stores docs in Google Drive

Setup:

  • Connect to the shared Drive folder
  • Daily sync schedule
  • Import the PDFs and Markdown files you need

Benefits:

  • Always current documentation
  • No manual uploads
  • Automatic processing
  • Team stays informed

Team Wiki ​

Scenario: Engineering knowledge lives in Confluence and Notion

Setup:

  • Confluence connector limited to the engineering space
  • Notion connector with the relevant pages shared to the integration
  • Weekly sync

Benefits:

  • One place to ask questions across both wikis
  • Updated automatically
  • Citation to the original page

Research Library ​

Scenario: A researcher keeps papers in Zotero

Setup:

  • Zotero connector for the relevant collection
  • Daily sync
  • PDF attachments imported

Benefits:

  • Papers searchable alongside your other documents
  • Bibliographic metadata preserved for citations

Connectors automate document management so you never have to manually upload updates. Set it up once and forget it.

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.