Model Context Protocol (MCP)
Last Updated: September 2026
Enable AI assistants like Claude to query your Scrapalot knowledge base directly. MCP integration lets you ask questions about your documents during AI conversations.
What is MCP?
Model Context Protocol allows AI assistants to access external context sources, including your Scrapalot documents.
Benefits:
- Ask Claude to search your documents during conversations
- Get answers with citations from your knowledge base
- Seamless integration with AI development tools
- Real-time document access
Use Cases
Development with Claude
Scenario: Working with Claude Desktop on a coding project
With MCP:
- Ask Claude: "Search my Scrapalot docs for the authentication setup"
- Claude queries your knowledge base
- Gets relevant context and citations
- Provides answer based on your documents
Without MCP:
- Manually search your documents
- Copy relevant sections
- Paste into Claude conversation
- Ask question again
Documentation Research
Scenario: Researching across multiple document sets
With MCP:
- "Find all mentions of deployment in my technical docs"
- "What does my security policy say about data retention?"
- "Search my meeting notes for decisions about the new feature"
Setup Guide
Prerequisites
- A Scrapalot account with API access (the
api_accessfeature — Pro plan and above) and anscp-API key (see below) - Documents uploaded to at least one collection
- An MCP-compatible client — Claude Code, Claude Desktop, Gemini CLI, or any client that speaks MCP over Streamable HTTP
Note: Scrapalot's MCP endpoint is
https://api.scrapalot.app/mcp(cloud) orhttp://localhost:8080/mcp(self-hosted, through the Gateway). It is a remote Streamable-HTTP server — there is no binary to install. Every request must carry your API key asAuthorization: Bearer scp-.... Don't confuse this with the MCP servers you add under Settings → Integrations, which is the opposite direction: those are external MCP servers that Scrapalot's own agent calls out to.
Getting Your API Key
The easy way is Settings → Integrations → Connect your agent: it mints the key, then hands you a finished config snippet for Claude Code, Claude Desktop, Gemini CLI or the OpenAI SDK with the key already substituted. Existing keys are listed there too, and can be disabled or deleted.
The same thing over the REST API, if you would rather script it:
curl -X POST https://api.scrapalot.app/api/v1/auth/api-keys \
-H "Authorization: Bearer <your-JWT-from-login>" \
-H "Content-Type: application/json" \
-d '{"name": "claude-code"}'The response contains plain_text_key (scp-...) — shown exactly once, store it safely. Requires a Pro-or-higher plan. Keys can be listed, disabled and deleted via the same /api/v1/auth/api-keys endpoints.
Treat the key like your password. An
scp-key authenticates as your whole account (documents, memory, notes — not just MCP). Rotate it if it leaks.
Configuration
Claude Code (one command):
claude mcp add --transport http scrapalot https://api.scrapalot.app/mcp \
--header "Authorization: Bearer scp-..."or per-project via .mcp.json:
{
"mcpServers": {
"scrapalot": {
"type": "http",
"url": "https://api.scrapalot.app/mcp",
"headers": { "Authorization": "Bearer scp-..." }
}
}
}Claude Desktop — remote HTTP servers with a Bearer header are easiest via the mcp-remote bridge. Edit the config file (macOS: ~/Library/Application Support/Claude/claude_desktop_config.json, Windows: %APPDATA%\Claude\claude_desktop_config.json):
{
"mcpServers": {
"scrapalot": {
"command": "npx",
"args": [
"-y", "mcp-remote",
"https://api.scrapalot.app/mcp",
"--header", "Authorization: Bearer scp-..."
]
}
}
}Gemini CLI (~/.gemini/settings.json):
{
"mcpServers": {
"scrapalot": {
"httpUrl": "https://api.scrapalot.app/mcp",
"headers": { "Authorization": "Bearer scp-..." }
}
}
}OpenAI Responses API — register it as a remote MCP server:
tools=[{
"type": "mcp",
"server_url": "https://api.scrapalot.app/mcp",
"headers": {"Authorization": "Bearer scp-..."},
}]Using MCP with Claude
Basic Queries
Ask Claude to search your docs:
Can you search my Scrapalot documents for information about deployment?Specify a collection:
Search the "Technical Docs" collection for authentication setupGet specific information:
Find the deployment checklist in my documentsAdvanced Queries
Multi-collection search:
Search both "API Docs" and "User Guides" for rate limiting informationSemantic search:
Find documents related to security best practicesFollow-up questions:
Based on what you found, what's the recommended approach?Available Tools
Remote callers get a default-deny allowlist: only the tools below are listed and callable over the network.
Discovery & corpus:
get_current_user— who the key belongs tolist_workspaces,get_default_workspace— your workspaceslist_collections— collections you can querylist_documents— documents inside a collectionlist_providers,list_models— available LLM providers/models
Asking questions:
chat_with_documents(query, collection_id, model_name?, incognito?)— a full RAG answer with citations from one collection. Passincognito: trueif the turn should neither read nor teach your personal memory.create_session,list_sessions— conversation continuity
Personal memory (your second brain):
memory_query(question)— recall what Scrapalot knows about youmemory_status()— per-tier usage vs quotamemory_add(text, context?)— deliberately save something (only when you explicitly ask to be remembered)memory_forget(memory_id)— forget one memory (query first to get its id)
Background jobs:
list_active_jobs,get_job_status— watch document processingcancel_job(job_id)— stop one of your own jobs
Building the corpus:
create_workspace(name, …),create_collection(name, workspace_id, …)— new containersupload_document_content(content, filename, collection_id)— upload a document by passing its content. Processing then runs as a normal background job, so watch it withget_job_status.
The rest of the REST API:
list_scrapalot_api(search?, method?)— search the REST operations the typed tools don't cover (notes, sharing, subscriptions, connectors, research templates, …)describe_scrapalot_api(method, path)— parameters and body shape of one operationcall_scrapalot_api(method, path, query?, body?, confirm_destructive?)— call it as you.DELETErequiresconfirm_destructive: true; authentication and API-key management routes are not reachable this way.
Resources and prompts: read-only resources scrapalot://workspaces, scrapalot://collections and scrapalot://providers, plus prompt templates prompt_rag_query, prompt_deep_research and prompt_document_summary.
Everything runs under your identity, your workspace scoping, your plan and your storage quota — the same limits the web app applies.
Not available remotely, by design: uploading by server-side file path (send the content with upload_document_content instead), signing in with a username and password (use your API key), and internal test tooling.
Troubleshooting
MCP Server Not Responding
Verify the endpoint is up (a 401 without a key means the server is alive and auth is working):
curl -s -o /dev/null -w "%{http_code}\n" -X POST https://api.scrapalot.app/mcp \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","method":"tools/list","id":1}'
# expected: 401Beware the wrong host: https://scrapalot.app/mcp (no api.) is the web app and returns the SPA page with a 200 — the MCP endpoint lives on api.scrapalot.app.
Review Claude Desktop logs:
- macOS:
~/Library/Logs/Claude/mcp.log - Windows:
%APPDATA%\Claude\Logs\mcp.log
Authentication Errors
Verify API key:
- Check key in config is correct (starts with
scp-) - Ensure the key is active and hasn't expired
- Test the key with a direct API call:
curl -H "Authorization: Bearer scp-..." \
https://api.scrapalot.app/api/v1/collectionsIf key invalid:
- Mint a new key via
POST /api/v1/auth/api-keys - Update the client config
- Restart the client
No Results Returned
Common causes:
- Collection has no documents, or they are still processing
- Wrong
collection_id(list them withlist_collections) - Query doesn't match content
Solutions:
- Verify documents in the collection and that processing finished (
list_active_jobs) - Try broader search terms
- Make sure the collection is in a workspace you own or that was shared with you
Connection Refused (self-hosted)
- The public surface is the Gateway (
:8080) — port 8090 is the internal Python service and should not be exposed - Verify the gateway container is up:
docker psandcurl -s -o /dev/null -w "%{http_code}" -X POST http://localhost:8080/mcp(expect401)
Best Practices
Query Formulation
Be specific:
- Good: "Find the API rate limit configuration"
- Poor: "Tell me about limits"
Use collection names:
- "Search the Security Policies collection for..."
- "In the API Documentation, find..."
Iterate:
- Start broad, then narrow down
- Ask follow-up questions
- Request clarification
Organization
Organize collections meaningfully:
- Separate by topic or project
- Use descriptive collection names
- Keep related documents together
Maintain document quality:
- Clear, descriptive filenames
- Well-structured content
- Regular updates
Security
Protect your API key:
- Store in config only
- Don't commit to version control
- Rotate periodically
- Use separate named keys per client (
claude-code,gemini-cli, …) so one can be revoked without breaking the others
Know the blast radius:
- An
scp-key currently authenticates as your whole account over the REST API, not just the MCP allowlist — scoped/read-only keys are not available yet - Disable or delete a key the moment a machine holding it is retired
Integration Examples
Research Workflow
Research new technology:
- Upload relevant documentation to Scrapalot
- During Claude conversation, search docs for specific topics
- Get contextual answers with citations
- Dive deeper into specific areas
Documentation Writing
Create documentation with Claude:
- Store existing docs in Scrapalot
- Ask Claude to search for related information
- Get consistent terminology and approach
- Reference existing content
Code Review
Review code with context:
- Store architectural docs in Scrapalot
- During code review, ask Claude to check against standards
- Get design pattern references
- Ensure consistency with existing code
Advanced Configuration
Custom MCP Client
Any MCP SDK that supports Streamable HTTP works. Python example (FastMCP):
from fastmcp import Client
from fastmcp.client.transports import StreamableHttpTransport
transport = StreamableHttpTransport(
"https://api.scrapalot.app/mcp/",
headers={"Authorization": "Bearer scp-..."},
)
async with Client(transport) as client:
tools = await client.list_tools()
answer = await client.call_tool(
"chat_with_documents",
{"query": "…", "collection_id": "<uuid>"},
)Prefer plain REST? The same key also works against the OpenAI-compatible endpoint (POST /api/v1/chat/completions with model: "scrapalot:<workspace>:<collection>") — see the API Reference.
Scrapalot Claude Code Plugin
Separately from the MCP endpoint, Scrapalot publishes a Claude Code plugin marketplace:
claude plugin marketplace add https://api.scrapalot.app/marketplace.json
claude plugin install scrapalot@scrapalotThe plugin ("Scrapalot AI Toolkit") is tooling for people who develop and operate a Scrapalot stack — corpus audits, RAG answer grading, DevOps and code-quality commands, skills, agents and hooks. It does not configure the MCP connection; to query your library from Claude Code, add the MCP server as shown above.
Multiple Scrapalot Instances
Connect to multiple Scrapalot servers by registering the endpoint twice under different names:
{
"mcpServers": {
"scrapalot-cloud": {
"type": "http",
"url": "https://api.scrapalot.app/mcp",
"headers": { "Authorization": "Bearer scp-cloud-key" }
},
"scrapalot-selfhosted": {
"type": "http",
"url": "http://localhost:8080/mcp",
"headers": { "Authorization": "Bearer scp-dev-key" }
}
}
}Related Documentation
- API Reference - REST API endpoints
- Security - Authentication and access control
- RAG Strategy - How search works
- User Guide - Using Scrapalot features
MCP integration brings your Scrapalot knowledge base directly into your AI conversations. Set it up once and search your documents naturally during any Claude conversation.