All concepts

Knowledge Graphs (Neo4j)

Entities as nodes, relations as typed edges; query with Cypher, power GraphRAG

RAG & Retrieval · Intermediate · ~5 min

In plain English

Facts stored as connections — this person works at that company, which owns this product — so you can walk the links to an answer instead of searching text.

Why it's worth your time

It's the structure behind entity-heavy domains, and the thing that makes multi-hop questions answerable exactly rather than approximately.

If you remember three things

  • Nodes are entities, edges are relationships
  • Queries traverse; they don't rank
  • Entity resolution is the hard part, not the storage

Overview

A knowledge graph stores information as entities (nodes) connected by typed, directed relationships (edges), with key-value properties on both — a property graph. Neo4j is the leading native graph database for this, queried in Cypher, whose MATCH clause draws the pattern you want like ASCII-art of the graph. Because edges are first-class, multi-hop questions are cheap traversals, and the same structure powers GraphRAG, retrieving a connected subgraph as grounded context for an LLM.

In an interview

A knowledge graph models data as nodes for entities and typed, directed edges for their relationships, with properties on both — a property graph. It's stored in a native graph database like Neo4j and queried with Cypher, where you MATCH a pattern instead of writing JOINs. Multi-hop questions — friend-of-a-friend, supply-chain paths — are cheap traversals. For RAG this enables GraphRAG: retrieve a relevant subgraph as structured, grounded context, in contrast to vector search, which fetches semantically similar text but knows nothing about how facts relate.

Production defaults

Model first
decide your entity and relationship types before ingesting anything. Retrofitting a schema is painful
Resolution
budget real effort for deduplicating entities. It's where these projects succeed or fail
Combine
graph for exact relationships, vectors for fuzzy recall. Use each for what it's good at

What breaks

  • Traversals return nothing — Entities that should be one node are several. Resolve duplicates before blaming the query.
  • The graph is stale — No update path from the source systems. A graph that isn't maintained decays faster than an index.

Watch it explained

What is a Knowledge Graph? — IBM Technology, 5:36

Related