Skip to content

Graph RAG: Relationship-Aware Search ​

Last Updated: September 2026

Graph RAG enhances traditional search by understanding relationships between concepts in your documents. Perfect for when you need to understand how ideas connect.

What is Graph RAG? ​

Traditional search finds documents based on keywords or semantic similarity. Graph RAG goes further by:

  • Understanding relationships between entities (people, places, concepts, events)
  • Traversing connections to find related information
  • Answering relationship questions like "How are X and Y connected?"
  • Discovering insights through multi-hop reasoning

How It Works ​

When to Use Graph RAG ​

Perfect For ​

Relationship Questions:

  • "How are Person A and Person B connected?"
  • "What companies does Organization X work with?"
  • "Which projects involve both Technology A and Technology B?"

Multi-Hop Reasoning:

  • "What does Company X's partner do?"
  • "Who worked on projects related to the same technology?"
  • "Trace the connection between Concept A and Concept B"

Entity-Centric Queries:

  • "Show me everything about Product X"
  • "What are all the locations mentioned for Company Y?"
  • "List all people who worked at Organization Z"

Not Needed For ​

Simple keyword search:

  • "Find documents about topic X"
  • "Show me the latest updates"
  • Standard semantic similarity works better

Single-concept queries:

  • "What is RAG?"
  • "Explain how authentication works"
  • Regular vector search is sufficient

Entity Types Detected ​

Extraction combines a fast statistical named-entity pass with an LLM pass (running on the system model) that understands context. Six entity types are extracted from text:

  • Concept — ideas, theories, technologies, methods
  • Person — people mentioned in the text
  • Place — cities, countries, regions, locations
  • Event — historical events and dated happenings
  • Term — domain-specific terminology
  • Quote — notable quotations

Images, tables and equations pulled out of PDFs during multimodal extraction are also represented as graph nodes, so they can be connected to the passages that discuss them.

Relationship Types ​

Graph Schema (actual relationship types) ​

Structural hierarchy:

  • CONTAINS: Collection → Book, Section → Chunk
  • HAS_CHAPTER: Book → Chapter
  • HAS_SECTION: Chapter → Section
  • NEXT: Chapter → Chapter (sequential order)

Entity relationships:

  • MENTIONS: Book / Chunk → Entity
  • REFERENCES / DESCRIBES / DISCUSSES: Chunk → Entity (typed)
  • CO_OCCURS_WITH: Entity → Entity (weighted co-occurrence)
  • SHARED_ENTITY: Book → Book (cross-book linking)
  • IN_COMMUNITY: Entity → Community (detected clusters); HAS_PARENT_COMMUNITY links clusters into a hierarchy

Communities: related entities are grouped into clusters, and each cluster gets a short generated report. These summaries help answer broad "what are the main themes?" questions across a collection.

Real-World Examples ​

Example 1: Finding Connections ​

Question: "How are Carl Jung and alchemy connected?"

Graph RAG Process:

  1. Identifies "Carl Jung" (Person) and "alchemy" (Concept)
  2. Searches for paths between them via shared chunks/entities
  3. Finds: Carl Jung ←[:MENTIONS]– Chunk –[:MENTIONS]→ alchemy (or via CO_OCCURS_WITH)
  4. Retrieves the connecting document chunks
  5. Generates answer with full context

Answer: An explanation grounded in the passages where both are discussed together, with citations.

Example 2: Multi-Hop Discovery ​

Question: "What products use the same technology as Project Alpha?"

Graph RAG Process:

  1. Finds Project Alpha entity
  2. Identifies technology it uses
  3. Traverses to other projects using same technology
  4. Collects information about those projects
  5. Assembles comprehensive answer

Answer: Lists all related projects with context about each

Example 3: Temporal Queries ​

Question: "What happened in the company in 2024?"

Graph RAG Process:

  1. Identifies "2024" date entity
  2. Finds all events linked to that timeframe
  3. Gathers related people, projects, decisions
  4. Organizes chronologically
  5. Provides timeline

Answer: Timeline of 2024 events with context

Setup & Configuration ​

Neo4j Integration ​

Graph RAG is backed by a Neo4j database. It is enabled on the hosted service; on-premise deployments can switch graph features off, and the rest of Scrapalot (vector and keyword search, chat, Deep Research) keeps working without them.

Automatic Processing:

  • The document structure (book → chapters → sections → chunks) is written to the graph right after processing
  • Entities and relationships are extracted afterwards on a dedicated background worker, so a large book is searchable before its graph is finished
  • Documents that still lack a graph are picked up automatically during off-peak hours
  • No manual configuration needed

Automatic housekeeping: a nightly maintenance run merges duplicate entities and spelling variants, resolves conflicting entity types, removes orphaned entities, recomputes co-occurrence weights and importance scores (PageRank), and rebuilds communities.

Performance Considerations ​

Graph Size ​

Small graphs (1000s of entities):

  • Very fast queries
  • Minimal resource usage
  • Works on modest hardware

Medium graphs (10,000s of entities):

  • Fast with proper indexing
  • Moderate resource usage
  • Recommended for most deployments

Large graphs (100,000+ entities):

  • Requires optimization
  • Higher resource needs

Query Performance ​

Fast queries:

  • Direct entity lookups
  • 1-2 hop relationships
  • Indexed properties

Slower queries:

  • Deep traversals (3+ hops)
  • Complex pattern matching
  • Full graph scans

Optimization:

  • Automatic indexing on common properties
  • Query depth limits
  • Result set size limits
  • Smart caching

Tri-Modal Fusion ​

Graph RAG works alongside:

  • Dense semantic search (vector embeddings)
  • Sparse keyword search (BM25)
  • Graph-based search (entity relationships)

Intelligent routing:

  • System chooses best search method(s)
  • Combines results when beneficial
  • Balances precision and recall

Example:

  • Question mentions specific entities → Use graph search
  • Question is conceptual → Use vector search
  • Question has exact terms → Use keyword search
  • Complex question → Use all three, fuse results

Privacy & Data Sovereignty ​

Data Storage ​

What's stored in Neo4j:

  • Entity names and types
  • Relationship types
  • Document references
  • Confidence scores

What's NOT stored:

  • Full document content (in PostgreSQL)
  • Vector embeddings (in pgvector)
  • User data (in PostgreSQL)

Self-Hosting ​

Complete control:

  • Run Neo4j on your infrastructure
  • Data never leaves your network
  • Full audit trail
  • Custom backup strategy

Personal Brain (per-user memory) ​

The same Neo4j database also holds the Personal Brain (shown in the app as Second brain): an opt-in memory that learns durable facts and preferences about you from your conversations and feeds them back into the assistant, so you don't have to repeat yourself.

  • Off by default. You turn it on for yourself; it never activates for anyone by default.
  • Bounded. Memories are kept in three small, capped groups — who you are, how you prefer to work, and what you know. When a group fills up, similar entries are consolidated rather than piling up.
  • Extracted once per conversation, in the background after you leave it — not on every message.
  • Linked to your documents' graph, so a memory about a topic can connect to the entities in your library.
  • You stay in control. In the app it is called Second brain (Settings → Account, "Learn about me"). There you can view, edit, forget or clear memories and export the whole brain. An incognito chat mode skips memory entirely.
  • Sharing is deliberate. A memory can be shared with your team; a sensitivity check blocks personal or sensitive content, and every share is recorded.

Monitoring ​

Graph Health ​

Track graph metrics:

  • Total entities
  • Total relationships
  • Entity type distribution
  • Relationship type distribution
  • Query performance

Access via:

  • Admin dashboard
  • Neo4j browser
  • Query logs

Usage Patterns ​

Understand how Graph RAG helps:

  • Questions using graph search
  • Average traversal depth
  • Most common entity types
  • Popular relationship queries

Best Practices ​

Document Preparation ​

Maximize graph value:

  • Use clear entity names
  • Maintain consistent terminology
  • Include context about relationships
  • Structure content logically

Query Formulation ​

Get better results:

  • Name specific entities
  • Ask about relationships explicitly
  • Use "how" and "why" questions
  • Request connections and paths

Graph Maintenance ​

Keep graph healthy:

  • Monitor entity quality
  • Review relationship accuracy
  • Clean up duplicates
  • Update deprecated entities

Troubleshooting ​

No Relationship Found ​

Common causes:

  • Entities not in same document context
  • Relationship type not detected
  • Traversal depth limit reached
  • Indirect connection too distant

Solutions:

  • Check entity names are correct
  • Review document content
  • Increase traversal depth
  • Try semantic search instead

Slow Graph Queries ​

Optimize:

  • Reduce traversal depth
  • Limit result set size
  • Use more specific entity names
  • Check graph size

Poor Entity Detection ​

Improve:

  • Use clearer entity names in documents
  • Add context around entities
  • Review detection confidence
  • Consider manual entity tagging

Graph RAG is powerful but optional. Start with standard RAG, add Graph RAG when you need relationship understanding.

Open-core — Community Edition under AGPL-3.0 · Hosted product is proprietary.