Home / Blog / Personal knowledge graph
Applied AI and Information Architecture

Building a Personal Knowledge Graph for a Portfolio AI Assistant

A portfolio assistant should answer questions about projects, technologies and outcomes without inventing experience. Plain vector search can retrieve a relevant paragraph, but it may not reliably connect a project to its components, role, dates and public evidence. A small knowledge graph can make those relationships explicit, provided every claim remains traceable to its source.

Model the portfolio as claims with evidence

Ingest curated facts, connect them, retrieve evidence and constrain the answer
SourcesProject pages, repositories and approved profile facts.
ExtractProjects, skills, roles, dates and relationships.
GroundAttach each claim to a source and confidence.
AnswerTraverse relevant links and cite evidence.

Start with a small ontology: Person, Project, Technology, Role, Outcome and Source. Relationships might include BUILT, USES, CONTRIBUTED_TO, DEMONSTRATES and DOCUMENTED_BY. Avoid importing every phrase into a node. The useful graph is the smallest structure that answers questions the portfolio actually receives.

Keep provenance on every assertion

Graph itemExampleEvidence to retain
EntityESP32 telemetry projectStable ID, canonical URL and last verified date
RelationshipProject USES MQTTSource passage or repository file that supports the edge
ClaimReduced manual reporting effortWhether measured, self-reported or an intended result
TechnologyFastAPIProject association and evidence; not an inferred proficiency score

Store source URI, extraction version, timestamp and review status. Separate directly verified facts from model-extracted candidates. A confidence score is not proof. For subjective or quantified claims, preserve the exact supporting context and require human review before making it public.

Use graph and vector retrieval for different questions

Vector retrieval is useful for semantic similarity and finding relevant project descriptions. Graph traversal helps with relationship questions: which projects use a technology, what evidence connects a skill to an outcome, or which components belong to one system. A hybrid pipeline can retrieve candidate text, follow a small number of typed edges and return both the source snippets and relationship path.

Do not assume a graph is always better. For a handful of pages, structured JSON or relational tables plus search may be simpler to maintain. Add a graph store when multi-hop questions, entity disambiguation or relationship updates justify the operational cost. GraphRAG systems can also involve costly extraction and indexing; measure answer quality against a simpler baseline.

Answer with citations and abstention

Build a response contract: answer only from retrieved evidence, cite project URLs near claims, distinguish role from team outcome and say when the portfolio does not contain enough information. If the graph has conflicting versions, prefer the latest reviewed source and expose the conflict rather than silently combining them. Do not let the language model create new edges during a public answer.

Protect personal and private information

Index only content intended for public disclosure. Exclude private client records, credentials, personal contact data and confidential repository content. Treat user-provided queries as untrusted, authorize source access before retrieval and avoid exposing hidden graph metadata through prompts or citations. Provide deletion and correction paths so outdated facts can be removed from indexes and cached summaries.

Evaluate the assistant as a retrieval system

Create questions with known supporting pages and expected citations. Measure entity resolution, evidence recall, citation correctness, unsupported-claim rate, abstention quality and latency. Include adversarial prompts that ask it to exaggerate ownership or disclose private work. Re-run this set whenever extraction prompts, graph schema or retrieval rules change.

Related: RAG pipelines, agent memory boundaries and local LLM architecture.

In summary

A personal knowledge graph is valuable when it links portfolio claims to explicit relationships and verifiable sources. Start with a small model, preserve provenance, combine graph traversal with text retrieval only when it helps, and let the assistant abstain. The goal is not to make a bot sound more confident; it is to make each answer easier to verify.

References