Build Knowledge Graphs Your AI Can Trust in 90 Days
TL;DR
If you are a consultant or educator, your method lives in your head, a few decks and a lot of call recordings. A knowledge graph is one way to make that usable by AI. But check first that a graph is what you need. Many founder-led businesses need a structured wiki before they need a graph database.
If a graph is the right call, build it in this order:
- Design the schema.
- Prepare the data and tag every fact with its source.
- Extract entities and relations.
- Merge the duplicates.
- Load the graph.
- Test it before anyone relies on it.
When to build knowledge graphs, and when a wiki will do
An AI model does not get smarter on its own. The weights are frozen. What compounds is the memory and context you give it. Structure beats raw intelligence.
That structure does not have to be a graph.
In our own stack, the memory layer is a wiki: a living, linked map of how the business connects. Agents read that map before they act. Alongside it sits a vector index, so the AI can search by meaning and not only by keyword. That covers most of what a founder-led business asks of its knowledge.
| Your questions look like | Start with |
|---|---|
| "Where is the document on X?" | A knowledge base |
| "What is our method for X, and why?" | A structured wiki |
| "Find me things similar to X" | A vector index |
| "Which clients share this problem, and what worked?" | A knowledge graph |
A graph is worth the work when the answer depends on relationships across many records. If that is not you yet, read our guide to knowledge architecture for AI and our walkthrough on building an LLM wiki. DevWiz also has a plain explainer on what RAG is, which is the simpler retrieval layer.
If a graph is the right call, the rest of this post is the pipeline.
What you need before you start
- Skills. Python, basic SQL, and a working sense of how prompts behave.
- Tools. A graph database, access to a model for extraction, and a vector index.
- Data. Start with a few hundred documents, not the whole archive.
- One question. The single business question the graph has to answer.
Neo4j is the common choice for the database. Claude Code can write and run the pipeline scripts with you.
Before you build anything, run one sanity test. Take one sentence. Ask the model for one fact, shaped as subject, relation, object. Load it. If that does not produce a clean node, nothing after it will work. The freeCodeCamp handbook on building a knowledge graph with Python and Neo4j walks through the setup.
Design the schema first
Get the schema wrong and you rebuild. One rule does most of the work. For every property, ask: does this belong to the thing itself, or does it describe a link to something else? A date of birth belongs to the person. "Works at" links a person to a company. The first becomes an attribute. The second becomes an edge.
That is the core idea in OntoKG. The authors call it intrinsic-relational routing, and they applied it to the January 2026 Wikidata dump. Deciding it up front saves you from making schema calls under deadline pressure.
Then pick your graph type. Property graphs are the kind Neo4j uses. RDF is a W3C standard that plugs into public knowledge bases more easily. Choose once. Converting later is painful.
Prepare the data and tag the source
Prep decides how good the graph can be.
- Split your text. Test two or three chunk sizes on a sample before you commit. Small chunks give the model focus and lose context. Big chunks keep context and muddle the extraction.
- Clean your tables. Map every column to one name before you load. "cust_name" and "client" should become the same field.
- Tag every fact with its source. Record the document and where in it the fact came from. When someone asks, you need an answer in seconds.
Store the source as its own node, linked to every fact. Adding it later means touching every fact you have already loaded.
Extract entities and relations
This is where the model does most of the work. It needs guardrails, or it invents things.
- Force a fixed format. Ask for JSON that matches your schema. Reject and retry anything that does not.
- Limit the choices. Give the model a short list of allowed entity and relation types.
- Score and tag. Store a confidence score and a source with every fact.
- Check by hand. Sample a batch of facts each run and mark them right or wrong.
How you structure the task matters as much as which model you pick. The RAKG framework works sentence by sentence to find entities. Then it uses retrieval to build the relations between them. Then it uses a model as a judge to filter out invented facts. On the MINE dataset it reached 95.91% accuracy, against 89.71% for GraphRAG.
Merge the duplicates
"Jon Smith" and "Jonathan Smith" as two records will quietly wreck your graph.
- Group first. Use cheap rules, such as the same email domain, to find likely matches.
- Compare second. Use embeddings to find near matches among those groups.
- Give each entity a permanent ID as soon as you are sure it is unique.
- Send the hard cases up. Rules handle the easy matches. A model or a person decides the rest.
- Log every merge. If one is wrong, you need to split it without a rebuild.
Load the graph
Three rules for loading.
- Use MERGE, not CREATE. CREATE always makes a new node. Run the script twice and you have duplicates. MERGE only creates one if nothing matches.
- Batch your writes. The freeCodeCamp handbook timed 1,000 rows at 2,758 ms written one statement at a time, and 64 ms as one batch. That is 43 times faster.
- Add constraints first. Put uniqueness rules and indexes on your ID fields before you load, not after.
Query it for AI retrieval
A graph nobody can query is an expensive spreadsheet.
Sort the question first. A simple lookup needs one hop. A relationship question needs two or three. Anything longer needs a cap, or the query never returns.
Then pull only the part of the graph that bears on the question. This step is easy to get wrong. Two facts being connected does not mean they support the answer.
PAGE-RAG, a 2026 retrieval method, names that problem the connectivity-support gap. It builds a small graph over the retrieved candidates and scores each path on relevance and source. Then it passes only the supporting facts to the model. On a weighted average across three multi-hop benchmarks, it improved support F1 by 10.4 points and answer F1 by 3.3 points over a strong retriever. We think the same idea applies to a graph you own.
Always cap how deep a query can go in production.
Test it two ways
You cannot eyeball a graph and call it done. You need two kinds of test.
| Test | Question | How |
|---|---|---|
| Inside the graph | Are the facts right? | Sample facts by hand. Check coverage. Look for conflicts |
| Outside the graph | Does it help? | Run a real task with the graph and without it |
A 2025 review in Applied Sciences looks at knowledge graph construction across extraction, learning and evaluation, and is a useful map of the methods published since 2022.
A graph that scores well on paper and does not improve a real task is not worth the storage.
Keep it alive
Left alone, a graph fills with stale facts.
- Update on a rhythm that matches how fast the data changes.
- Date every fact. Record when it was last checked, not just that it exists.
- Sweep on a schedule for orphan nodes, conflicts and missed merges.
- Name an owner for corrections.
- Archive, do not delete. Mark old facts as replaced and keep the history.
- Version the schema like a code release.
Privacy belongs here too. A graph connects facts in ways no single document did. Leave out personal data you do not need. Decide how long you keep it. Know what your licence allows before you load someone else's material. Our post on PII in AI covers the basics.
What a 90-day pilot looks like
| Weeks | What happens |
|---|---|
| 1 to 2 | Pick the one question. Set the schema. Gather a small dataset |
| 3 to 6 | Build extraction and merging. Run the sanity test. Load the first batch |
| 7 to 10 | Test both ways. Fix what is broken. Tighten the source tagging |
| 11 to 13 | Write it up. Train the team. Agree the update rhythm |
Set the pass mark before you start. A precision score, a task that got better, or a query that returns the right answer every time.
The harder problem
The graph technology is the easy part. The hard part is getting what is in your head onto the page.
For a consultant or an educator, the IP is not in the documents. It is in how you decide. When a client says this, you do that. What good looks like. What not good enough looks like. No extraction pipeline finds that in a PDF, because you never wrote it down.
That is the work we start with in the 90-day program. Pull out the calls only you make. Write them as structured knowledge your AI employees can read. Build the system around that with Claude Code. Then the graph, if you need one, has something worth storing.
James Killick
Find out what is still locked in your head
Before you choose a wiki or a graph, find out how much of the business depends on knowledge only you hold.
Take the Founder Bottleneck Assessment. It scores five dimensions in six minutes, then names your next move.
Sources
- RAKG: Document-level Retrieval Augmented Knowledge Graph Construction, arXiv
- Knowledge Graph Construction: Extraction, Learning, and Evaluation, Applied Sciences
- OntoKG: Ontology-Oriented Knowledge Graph Construction with Intrinsic-Relational Routing, arXiv
- PAGE-RAG: Provenance-Aware Graph Evidence Promotion, arXiv
- How to Build a Knowledge Graph with Python and Neo4j, freeCodeCamp
Frequently Asked Questions
James Killick
Founder
The AI Orchestrator. 10+ years building digital products and 200+ apps shipped, now helping $1M+ educators and consultants turn their IP into AI-powered delivery systems.
James Killick founded and runs The AI Orchestrators.
More from James Killick