Graph RAG: Relationship-Aware Search
Last Updated: September 2026
Graph RAG enhances traditional search by understanding relationships between concepts in your documents. Perfect for when you need to understand how ideas connect.
What is Graph RAG?
Traditional search finds documents based on keywords or semantic similarity. Graph RAG goes further by:
- Understanding relationships between entities (people, places, concepts, events)
- Traversing connections to find related information
- Answering relationship questions like "How are X and Y connected?"
- Discovering insights through multi-hop reasoning
How It Works
When to Use Graph RAG
Perfect For
Relationship Questions:
- "How are Person A and Person B connected?"
- "What companies does Organization X work with?"
- "Which projects involve both Technology A and Technology B?"
Multi-Hop Reasoning:
- "What does Company X's partner do?"
- "Who worked on projects related to the same technology?"
- "Trace the connection between Concept A and Concept B"
Entity-Centric Queries:
- "Show me everything about Product X"
- "What are all the locations mentioned for Company Y?"
- "List all people who worked at Organization Z"
Not Needed For
Simple keyword search:
- "Find documents about topic X"
- "Show me the latest updates"
- Standard semantic similarity works better
Single-concept queries:
- "What is RAG?"
- "Explain how authentication works"
- Regular vector search is sufficient
Entity Types Detected
Extraction combines a fast statistical named-entity pass with an LLM pass (running on the system model) that understands context. Six entity types are extracted from text:
- Concept — ideas, theories, technologies, methods
- Person — people mentioned in the text
- Place — cities, countries, regions, locations
- Event — historical events and dated happenings
- Term — domain-specific terminology
- Quote — notable quotations
Images, tables and equations pulled out of PDFs during multimodal extraction are also represented as graph nodes, so they can be connected to the passages that discuss them.
Relationship Types
Graph Schema (actual relationship types)
Structural hierarchy:
- CONTAINS: Collection → Book, Section → Chunk
- HAS_CHAPTER: Book → Chapter
- HAS_SECTION: Chapter → Section
- NEXT: Chapter → Chapter (sequential order)
Entity relationships:
- MENTIONS: Book / Chunk → Entity
- REFERENCES / DESCRIBES / DISCUSSES: Chunk → Entity (typed)
- CO_OCCURS_WITH: Entity → Entity (weighted co-occurrence)
- SHARED_ENTITY: Book → Book (cross-book linking)
- IN_COMMUNITY: Entity → Community (detected clusters); HAS_PARENT_COMMUNITY links clusters into a hierarchy
Communities: related entities are grouped into clusters, and each cluster gets a short generated report. These summaries help answer broad "what are the main themes?" questions across a collection.
Real-World Examples
Example 1: Finding Connections
Question: "How are Carl Jung and alchemy connected?"
Graph RAG Process:
- Identifies "Carl Jung" (Person) and "alchemy" (Concept)
- Searches for paths between them via shared chunks/entities
- Finds: Carl Jung ←[:MENTIONS]– Chunk –[:MENTIONS]→ alchemy (or via CO_OCCURS_WITH)
- Retrieves the connecting document chunks
- Generates answer with full context
Answer: An explanation grounded in the passages where both are discussed together, with citations.
Example 2: Multi-Hop Discovery
Question: "What products use the same technology as Project Alpha?"
Graph RAG Process:
- Finds Project Alpha entity
- Identifies technology it uses
- Traverses to other projects using same technology
- Collects information about those projects
- Assembles comprehensive answer
Answer: Lists all related projects with context about each
Example 3: Temporal Queries
Question: "What happened in the company in 2024?"
Graph RAG Process:
- Identifies "2024" date entity
- Finds all events linked to that timeframe
- Gathers related people, projects, decisions
- Organizes chronologically
- Provides timeline
Answer: Timeline of 2024 events with context
Setup & Configuration
Neo4j Integration
Graph RAG is backed by a Neo4j database. It is enabled on the hosted service; on-premise deployments can switch graph features off, and the rest of Scrapalot (vector and keyword search, chat, Deep Research) keeps working without them.
Automatic Processing:
- The document structure (book → chapters → sections → chunks) is written to the graph right after processing
- Entities and relationships are extracted afterwards on a dedicated background worker, so a large book is searchable before its graph is finished
- Documents that still lack a graph are picked up automatically during off-peak hours
- No manual configuration needed
Automatic housekeeping: a nightly maintenance run merges duplicate entities and spelling variants, resolves conflicting entity types, removes orphaned entities, recomputes co-occurrence weights and importance scores (PageRank), and rebuilds communities.
Performance Considerations
Graph Size
Small graphs (1000s of entities):
- Very fast queries
- Minimal resource usage
- Works on modest hardware
Medium graphs (10,000s of entities):
- Fast with proper indexing
- Moderate resource usage
- Recommended for most deployments
Large graphs (100,000+ entities):
- Requires optimization
- Higher resource needs
Query Performance
Fast queries:
- Direct entity lookups
- 1-2 hop relationships
- Indexed properties
Slower queries:
- Deep traversals (3+ hops)
- Complex pattern matching
- Full graph scans
Optimization:
- Automatic indexing on common properties
- Query depth limits
- Result set size limits
- Smart caching
Combining with Vector Search
Tri-Modal Fusion
Graph RAG works alongside:
- Dense semantic search (vector embeddings)
- Sparse keyword search (BM25)
- Graph-based search (entity relationships)
Intelligent routing:
- System chooses best search method(s)
- Combines results when beneficial
- Balances precision and recall
Example:
- Question mentions specific entities → Use graph search
- Question is conceptual → Use vector search
- Question has exact terms → Use keyword search
- Complex question → Use all three, fuse results
Privacy & Data Sovereignty
Data Storage
What's stored in Neo4j:
- Entity names and types
- Relationship types
- Document references
- Confidence scores
What's NOT stored:
- Full document content (in PostgreSQL)
- Vector embeddings (in pgvector)
- User data (in PostgreSQL)
Self-Hosting
Complete control:
- Run Neo4j on your infrastructure
- Data never leaves your network
- Full audit trail
- Custom backup strategy
Personal Brain (per-user memory)
The same Neo4j database also holds the Personal Brain (shown in the app as Second brain): an opt-in memory that learns durable facts and preferences about you from your conversations and feeds them back into the assistant, so you don't have to repeat yourself.
- Off by default. You turn it on for yourself; it never activates for anyone by default.
- Bounded. Memories are kept in three small, capped groups — who you are, how you prefer to work, and what you know. When a group fills up, similar entries are consolidated rather than piling up.
- Extracted once per conversation, in the background after you leave it — not on every message.
- Linked to your documents' graph, so a memory about a topic can connect to the entities in your library.
- You stay in control. In the app it is called Second brain (Settings → Account, "Learn about me"). There you can view, edit, forget or clear memories and export the whole brain. An incognito chat mode skips memory entirely.
- Sharing is deliberate. A memory can be shared with your team; a sensitivity check blocks personal or sensitive content, and every share is recorded.
Monitoring
Graph Health
Track graph metrics:
- Total entities
- Total relationships
- Entity type distribution
- Relationship type distribution
- Query performance
Access via:
- Admin dashboard
- Neo4j browser
- Query logs
Usage Patterns
Understand how Graph RAG helps:
- Questions using graph search
- Average traversal depth
- Most common entity types
- Popular relationship queries
Best Practices
Document Preparation
Maximize graph value:
- Use clear entity names
- Maintain consistent terminology
- Include context about relationships
- Structure content logically
Query Formulation
Get better results:
- Name specific entities
- Ask about relationships explicitly
- Use "how" and "why" questions
- Request connections and paths
Graph Maintenance
Keep graph healthy:
- Monitor entity quality
- Review relationship accuracy
- Clean up duplicates
- Update deprecated entities
Troubleshooting
No Relationship Found
Common causes:
- Entities not in same document context
- Relationship type not detected
- Traversal depth limit reached
- Indirect connection too distant
Solutions:
- Check entity names are correct
- Review document content
- Increase traversal depth
- Try semantic search instead
Slow Graph Queries
Optimize:
- Reduce traversal depth
- Limit result set size
- Use more specific entity names
- Check graph size
Poor Entity Detection
Improve:
- Use clearer entity names in documents
- Add context around entities
- Review detection confidence
- Consider manual entity tagging
Related Documentation
- RAG Strategy - How Graph RAG fits in
- Context Expansion - Enhanced understanding
- Model Management - Entity extraction models
- Deployment Guide - Neo4j setup
Graph RAG is powerful but optional. Start with standard RAG, add Graph RAG when you need relationship understanding.