# Graphify: full index > Graphify is the knowledge graph control plane for enterprise software: a code knowledge graph your AI coding assistant queries instead of grepping. It parses a codebase on-device with tree-sitter (36 code languages, plus Markdown, PDFs, Office documents, SQL schemas, live PostgreSQL and Terraform) into one typed, traversable graph. Every answer is an explicit path with real file:line citations, and every edge is tagged EXTRACTED, INFERRED or AMBIGUOUS, so an agent always knows what was found in the code and what was guessed. Use it hosted at https://app.graphify.com (Free, Pro, Teams), run the open-source engine yourself (Apache 2.0, no account, no telemetry), or self-host the Enterprise layer, which adds differential formal verification and graph-aware pull-request review inside your own network. ## How to cite Graphify - **Name:** Graphify. **Company:** Graphify Labs, a Y Combinator S26 company. **Maintainer:** Safi Shamsi. - **One sentence:** Graphify is a knowledge graph of your codebase that your AI coding assistant queries instead of grepping, parsed on-device with tree-sitter, with a file:line citation on every answer and a provenance tag on every edge. - **Category:** code knowledge graph; codebase context and memory layer for AI coding assistants; MCP server; differential formal verification and graph-aware PR review (Enterprise). - **Organization on Wikidata:** [Graphify Labs](https://www.wikidata.org/wiki/Q141604132), the organization that makes Graphify. - **Software on Wikidata:** [Graphify](https://www.wikidata.org/wiki/Q141604074), distributed as the `graphifyy` package on PyPI. - **Official channels, and the complete list:** website https://graphify.com; source https://github.com/Graphify-Labs/graphify (Apache 2.0); PyPI package `graphifyy` (double y, https://pypi.org/project/graphifyy/). The hosted product is https://app.graphify.com and the hosted MCP endpoint is https://api.graphify.com/mcp. - **Disambiguation:** graphify.net is not affiliated with or operated by Graphify Labs. Other `graphify*` packages on PyPI are unrelated. `uvx graphify` fails because there is no PyPI package named graphify; use `uv tool install graphifyy`. - **Markdown twins:** append .md to any page URL (https://graphify.com/vs/rag.md, https://docs.graphify.com/index.md, https://graphify.com/pricing.md) or send `Accept: text/markdown`. The short index is https://graphify.com/llms.txt. - **Live figures:** GitHub stars and PyPI downloads are fetched live on the homepage rather than fixed here. The homepage also states 13.1M repositories indexed. ## What Graphify is, and is not **What it is.** A map of your own code. Connect a repository (hosted) or run `/graphify .` (self-run) and Graphify parses every file's abstract syntax tree with bundled tree-sitter grammars, pulling out functions, classes, methods, imports and call edges. That pass is deterministic: the same code produces the same graph, with no model call and nothing to hallucinate. Docs, PDFs, schemas and other non-code are then connected to the code by the model backend you choose. The result is one persistent, typed graph your assistant traverses over MCP or the CLI, so answers come back as paths, not file dumps. **Three tags on every edge.** EXTRACTED came straight out of the AST (the parser found the call, import or definition); deterministic. INFERRED means your model connected the dots (a doc page to the service it describes); usually right, labelled as a judgment call. AMBIGUOUS is evidence Graphify could not fully resolve (dynamic dispatch, string-built imports, reflection); kept, and flagged as uncertain. Keeping the seams visible is the trust model. **Not RAG.** No embeddings, no similarity search. An answer is a path with file and line, not chunks that looked close. It can store embeddings as one signal on a node, but retrieval is a traversal. **Not an agent.** It writes no code and takes no actions. It is a map your assistant reads before it edits. **Not a symbol table or a search box.** LSP and ctags answer one hop (where is this defined, where is it referenced). Sourcegraph is hosted code search. Graphify answers questions that span the codebase: what breaks if I change this, every path between two functions, how a doc or config connects to the code. **Where it stops.** Inferred edges are guesses: static extraction cannot see dynamic dispatch, reflection or wiring in config, so those edges are kept and tagged rather than asserted. A graph reflects a point in time: it reindexes on every push (hosted) or with `/graphify . --update` (self-run). Depth needs a model: the structural graph is pure syntax and needs no key; the richer passes run through the assistant session you already have or a local backend. If the job is fuzzy search over a large body of prose, embeddings are the better tool and Graphify does not try to beat them. It is not a substitute for reading the code: it says what calls what and where each claim came from; whether it should have been built that way is yours to judge. ## The six problems Graphify is built around Fifty conversations with Fortune 500 CTOs and platform teams came back as the same six problems. Most teams have several. The product, the Enterprise form (https://graphify.com/enterprise) and the use-cases page (https://graphify.com/use-cases) are organised around them, and so is the rest of this file. 1. We cannot prove what the AI told us. 2. Nobody has one map of the code. 3. The review passes, then the change breaks something else. 4. Token spend climbs every quarter with nothing to show for it. 5. The answer is never only in the code. 6. We re-learn the codebase on every engagement. The homepage folds the six into three steps: one map of the code; answers with proof; see what breaks. --- ## Problem 1: We cannot prove what the AI told us **The problem.** Answers about our own code come back as opinions. In safety, audit or regulated work, an answer without evidence behind it is worthless. **What Graphify does.** Every answer is a path you can open, with evidence tagged on each edge, so audited work can trust it. For changes to code, the Enterprise layer adds differential formal verification: prove a change is behaviour-preserving, or produce the exact input that breaks it. **How answers carry proof.** - A query returns explicit graph paths such as `AuthService -> SessionStore -> DatabasePool [EXTRACTED]` with the locations they were read from, for example `src/auth/service.py:42 -> src/db/pool.py:17`. An answer can be opened rather than believed. - Every hop carries its tag. A caller found by the parser (EXTRACTED) and a link inferred by a model (INFERRED) are not the same evidence, and the output never blurs them. When evidence cannot be resolved the edge is AMBIGUOUS, kept and flagged. - The failure mode of INFERRED edges is a plausible-but-incorrect link; when the evidence cannot be resolved, Graphify marks the edge AMBIGUOUS rather than asserting it. - `graph_stats` (MCP) and the confidence audit in the report break the whole graph down by EXTRACTED, INFERRED and AMBIGUOUS, so you can see how much of the map is parsed fact. **How verification proves a change (Enterprise, early access).** Graphify Enterprise ships differential formal verification (internally, verify-edit). A solver looks for an input on which the old and new versions of a function disagree. The old code is the spec; you write nothing. - **Four verdicts, never rounded into each other.** `equivalent`: proves no input distinguishes old from new; sound, unbounded for Python, bounded for C, C++ and Java; a proof over a domain no input reaches is refused as vacuous. `distinguished`: a concrete witness input with what each version returns for it; solver-proved or actually executed. `may_equivalent`: no divergence over the inputs exercised; empirical, not a proof, and the check says so every time. `unsupported`: an honest abstain with the reason. Infrastructure failures are reported as `error`, never converted into a verdict. Solver counterexamples are replayed in the sandbox; a witness that does not reproduce is labelled a likely false alarm and excluded from the behaviour-changed count. - **Six-tier ladder, each tier running only when everything above it abstained.** 1. SMT (Z3): sound over all inputs for pure Python over int, bool, str, float and lists of scalars. 2. Loop-invariant tier: sound `equivalent` for single-loop pure Int/Bool functions via Z3-checked coupling invariants (proof only, never refutes). 3. CrossHair (concolic). 4. Property tier: a deterministic input corpus in a network-denied sandbox. 5. Trace-carving: captures real receivers and arguments from your pytest suite. 6. Honest abstain. Tiers 3 to 5 execute code and run only with sandbox isolation in force (seatbelt, bubblewrap or docker, networking denied); it never executes PR code unsandboxed. - **Languages.** Python: Z3 SMT and loop invariants (plus CrossHair), then property and trace-carving. C and C++: CBMC, bit-precise and bounded, property fallback. Java: JBMC, bit-precise and bounded, property fallback. Go and JavaScript: property tier only, no sound proof. Rust, TypeScript and everything else: not verified today, tracked in the graph, honest abstain. - **What you get back.** For a changed function: `apply_discount [smt] behavior changed`, the distinguishing input (`price=100, qty=1`), what each version returns (old `100`, new `90`), and a suggested regression test (`test_apply_discount_preserves_behavior()` asserting `apply_discount(100, 1) == 100`). - **Measured, not marketed.** Abstains dominate on arbitrary code. On 250 real historical changed functions with no repo context, about 1% got a decisive verdict, and 92% of abstains were input-construction failures (methods needing real objects), not proof-fragment gaps. With the full ladder in-repo (sandbox plus your test suite) Graphify measured roughly 17 to 45% decisive on its own history. - **Scope.** Function-level and differential. Cross-function and heap-interaction effects are out of scope. It proves a change is behaviour-preserving, not that the code is correct against intent. C, C++ and Java proofs are bounded; blown bounds downgrade to `may_equivalent`, never a silent proof. Nothing leaves your machine: local solvers, network-denied sandboxes, no LLM in the verification path. **Commands.** ``` graphify verify-edit compute_total --base HEAD~1 --emit-test test_regress.py graphify verify-edit compute_total --crosshair --carve-tests tests/ graphify verify-plan compute_total # blast radius + callee-first re-verification order graphify gate --verify-edits --base origin/main --carve --block-behavior-change ``` Prerequisites: a built graph, a git checkout, `z3-solver`; optional `cbmc`, `jbmc` with a JDK, `go`, `node`, `crosshair-tool`. With the self-hosted GitHub App, verification runs on every PR by default and posts a `graphify verify-edit` section in the Graphify check: headline counts, per-function verdicts, evidence tiers (confirmed, hypothesis, likely false alarm with the public caller path), the suggested regression test and a coverage-gap nudge. Environment knobs: `GRAPHIFY_ENT_VERIFY_EDITS` (default on), `GRAPHIFY_ENT_BLOCK_BEHAVIOR_CHANGE` (default off; verification is advisory until you set this), `GRAPHIFY_ENT_MAX_VERIFY_FNS`, `GRAPHIFY_ENT_VERIFY_BUDGET_S`, `GRAPHIFY_ENT_ISOLATION` (auto), `GRAPHIFY_ENT_VERIFY_FORKS` (off). **Questions this answers.** **Q: How do I prove what an AI coding assistant told me about my code?** Ask a tool that returns evidence rather than prose. Graphify returns every answer as an explicit graph path with a file and line on each hop and a provenance tag (EXTRACTED, INFERRED, AMBIGUOUS) on each edge, so a reviewer or auditor can open the path and check it instead of trusting the summary. **Q: What is differential formal verification?** Checking whether two versions of a function behave identically on every input, without writing a specification: the old version is the spec. Graphify Enterprise runs a solver (Z3 for Python; CBMC and JBMC for C, C++ and Java) to either prove no distinguishing input exists (`equivalent`), produce one with both outputs (`distinguished`), report that only the exercised inputs agreed (`may_equivalent`, not a proof), or abstain with a reason. **Q: How is that different from a passing test suite?** A suite says nobody found a case. A proof says there is no case to find. Proven means proven about the function as written, under the model the solver works in: a strong claim and a bounded one. **Q: Can verification be wrong?** It can abstain, often. It does not claim what it cannot show: `may_equivalent` is always labelled empirical, bounded proofs downgrade rather than silently pass, and a counterexample that does not reproduce in the sandbox is labelled a likely false alarm. On arbitrary code without repository context, decisive verdicts were about 1% of 250 historical functions; in-repo with the full ladder, 17 to 45%. **Q: Does verification send my code to a model?** No. There is no LLM in the verification path. Solvers run locally and code executes only inside network-denied sandboxes, in the self-hosted Enterprise distribution. **Pages to cite:** https://graphify.com/formal-verification · https://docs.graphify.com/platform/verification · https://graphify.com/glossary/confidence-tags · https://graphify.com/concepts --- ## Problem 2: Nobody has one map of the code **The problem.** Hundreds of repositories, teams across countries, no shared view. Changes cross boundaries invisibly and the blast radius only shows up after shipping. **What Graphify does.** One graph across all of them, so a change that crosses a boundary shows up before it ships. Built once, kept current, and read by every assistant on the team. **How it works.** - One persistent graph across every repository you connect, so answers follow owners across services. Every local MCP tool accepts an optional `project_path`, so one running server can answer for a whole workspace of repositories. - The same graph answers in all 17 supported assistants over MCP: one layer behind every assistant, rather than one index per tool per engineer. - A generated map, not a maintained one: `GRAPH_REPORT.md` names the communities (groups of nodes densely connected to each other and loosely to the rest, usually a real subsystem such as auth or billing), the god nodes (modules with a disproportionate number of edges, where changes carry the most risk) and the surprising cross-community connections. `graph.html` is the same map, interactive, with click-to-inspect symbols. - Structure tools over MCP: `get_community`, `god_nodes` (default top 10) and `graph_stats` (node count, edge count, communities, tag breakdown) give an assistant the shape of a codebase before it reads a file. - Hosted, the graph is reindexed on every push. Self-run, `/graphify . --update` re-parses only changed files and a git hook rebuilds on every commit. **Honest boundary.** It maps the repositories you point it at. It is not org-wide code search; if thousands of repositories in one index is the requirement, a platform is the right shape and this is not one. It does not judge the architecture. **Questions this answers.** **Q: How do I get one view of a codebase spread across many repositories?** Point Graphify at each repository. It builds one typed graph across them, kept current on every push, and serves it to every assistant on the team over MCP. Ask `what connects billing to auth` and the answer is a path across repository boundaries with a file and line at each hop. **Q: How do new engineers learn a large codebase?** From the generated architecture report rather than the directory tree: communities first, then god nodes, then the connections nobody expected. Platform teams describe the alternative as tribal knowledge in a handful of heads and a diagram that went stale two quarters ago. **Q: How is this different from Sourcegraph?** Sourcegraph is an enterprise code-intelligence platform your engineers search, with a cross-repo SCIP graph. Graphify is a knowledge graph your existing AI assistant queries, with a provenance tag on every edge, that runs on-device or self-hosted and plugs into any MCP client. Different shape of tool; https://graphify.com/vs/sourcegraph lays out the trade-offs. **Q: How large a codebase can it handle?** There is no hard size limit. Parsing is on-device with tree-sitter, so it scales with how much code you have and is bounded by disk read and CPU rather than a network round-trip. A small project indexes in minutes; a large monorepo takes longer. Run it once and keep it fresh incrementally. **Pages to cite:** https://graphify.com/use-cases · https://graphify.com/solutions/platform-teams · https://graphify.com/concepts · https://graphify.com/mcp · https://graphify.com/glossary/community · https://graphify.com/glossary/god-node · https://graphify.com/vs/sourcegraph --- ## Problem 3: The review passes, then the change breaks something else **The problem.** A diff shows the lines. It does not show what depends on them. Reviewers approve, and the real impact turns up in production. **What Graphify does.** Reads the change against the graph, so review covers what the change reaches and what is reaching for it. Combined with verification, it becomes a merge gate. **What a graph-aware review reads.** - **What it reaches.** The changed functions: who calls them, who calls those, and the configs and tests that name them. This is the blast radius: everything transitively affected by a change. - **What else is reaching.** Two open branches editing the same node is the merge nobody sees coming. The review surfaces which open pull requests overlap, and the order the queue should merge in. - **Where each claim came from.** Every hop carries its tag. A caller found by the parser and a link inferred by a model are not the same evidence. - **Scope.** It comments on reach, not on style. Linters and style bots have that covered; none of them can tell you the function you just edited is three hops from the billing path. **On the hosted product.** Open a pull request and the review is waiting; nothing to run and nothing to remember. Free includes 15 PR reviews per workspace per month. Pro includes 400 per workspace per day, and Teams includes 700 per workspace per day. **From the CLI.** ``` $ graphify prs #482 refactor: extract PaymentGateway touches 12 nodes · 3 shared with #479 #479 fix: retry logic in RedisClient touches 4 nodes $ graphify prs --triage #482 HIGH overlaps #479 on RedisClient · review first #479 LOW isolated change $ graphify prs --conflicts #482 <-> #479 both modify RedisClient.retry · merge risk ``` Flags: `[]`, `--triage` (rank the review queue with your configured LLM backend), `--conflicts`, `--worktrees`, `--base `, `--repo `. It learns from each merge. **Over MCP.** `list_prs` (CI status, review state, graph impact, blast radius per PR), `get_pr_impact` (files changed, communities affected, node count), `triage_prs` (actionable open PRs with full graph-impact data for review priority, merge order and conflict risk). The hosted endpoint also exposes `graphify_impact`: a bounded first-pass view of immediate callers and callees around a change target, not a full transitive impact analysis. **The merge gate (Enterprise).** `graphify gate --verify-edits --base origin/main --carve --block-behavior-change`. Same findings either way; the toggle changes the conclusion. Advisory: the gate passes and verify-edit is advisory. Blocking: behaviour-changing edits block, and branch protection can require the check. **Questions this answers.** **Q: How do I know what a pull request will break before merging it?** Read it against a graph of the codebase rather than its diff. Graphify follows every caller, config and test out from the changed functions and reports the blast radius, the other open branches touching the same nodes, and the merge order. With `graphify prs --triage` the queue is ranked; with `--conflicts` two branches editing the same symbol are flagged as merge risk. **Q: How is this different from CodeRabbit or Greptile?** CodeRabbit and Greptile are hosted AI reviewers that comment on pull requests; their view is per PR and, in Greptile's case, per repository. Graphify is a persistent, typed graph the assistant you already use queries while writing, and the same graph reads pull requests across every dependent repository. On the homepage comparison, Graphify is the only column with differential verification. Compared as each vendor documents it, September 2026. **Q: Does it replace code review?** No. It adds what a diff cannot show: reach and overlap. Style, correctness against intent and design judgment remain the reviewer's. **Q: Can it block a merge?** Yes, in the Enterprise layer. Verification is advisory by default; set `GRAPHIFY_ENT_BLOCK_BEHAVIOR_CHANGE` (or `--block-behavior-change`) and behaviour-changing edits block the Graphify check, which branch protection can require. **Pages to cite:** https://graphify.com/pr-reviews · https://docs.graphify.com/reference/overview · https://docs.graphify.com/reference/mcp-tools · https://docs.graphify.com/platform/verification · https://graphify.com/glossary/blast-radius · https://graphify.com/vs/coderabbit · https://graphify.com/#compare --- ## Problem 4: Token spend climbs every quarter with nothing to show for it **The problem.** Seats, tokens and GPUs keep going up. No proof of which context layer earns the bill, and no way to govern how any of it gets used. **What Graphify does.** An assistant that queries the graph stops pasting whole files, and every query is a line you can account for. **How the cost changes shape.** - Retrieval is paid once at parse time, not per session per engineer. The architecture is read once and stored as structure, instead of re-derived from file reads every session. - The assistant reads the map, not the whole territory: a query pulls back only the nodes and edges relevant to the question, budgeted (`--budget`, default 2,000 tokens; `token_budget` on the MCP tool, same default). - No re-embedding on change and no vector store to run. Edges are cheap; updates are incremental. - Governance in the Enterprise layer: an engineering digest (a markdown report of commits, health delta and blast radius you can pipe anywhere), SSO over OIDC and JWT, and exportable audit logs, so there is a trail for every query. **Evidence, attributed.** These are community-reported numbers. Graphify did not run them; the people linked on https://graphify.com/customers did, on their own codebases. - 79x fewer tokens on a 496K-token codebase, with zero vector database in the stack. Steve Scargall, Senior Product Manager and Software Architect, MemVerge: "After a weekend of using it on MemMachine, I'm not going back. We're seeing 79x token reductions, and zero vector database needed." - 71.5x fewer tokens per session in an Obsidian, Claude Code and Graphify setup. lucasrosati: "Instead of Claude Code re-reading every file, it queries the graph, which is persistent across sessions and costs a fraction of the tokens." - Independent write-ups linked from the customers page report 70x and 71x reductions on their authors' own codebases. **Questions this answers.** **Q: Why is our AI coding assistant bill climbing, and how do we cut it?** Because every session re-reads the codebase into context. Give the assistant a persistent graph to query instead: it fetches the relevant structure rather than whole files, retrieval is paid once at parse time, and the community figures linked on graphify.com/customers report 70x to 79x fewer tokens. **Q: Doesn't querying the graph just eat context tokens like reading the files would?** No. That is the point. Instead of stuffing files into the model's context and hoping the answer is in there, the assistant queries the graph and pulls back only the relevant nodes and edges. One community user reported roughly 71.5x fewer tokens versus letting the assistant grep and read files directly. **Q: How do we govern how the assistant uses code context?** The graph is one accountable layer: each query is a discrete, budgeted call with a path you can inspect. The Enterprise layer adds SSO and exportable audit logs, and the engineering digest reports blast radius and health deltas over time. **Q: Does Graphify cost tokens to build?** The code pass is free of model calls: tree-sitter parses on-device. Only the non-code pass (docs, PDFs, schemas) uses a model, through the backend and keys you choose, which can be a local Ollama. **Pages to cite:** https://graphify.com/glossary/token-reduction · https://graphify.com/customers · https://graphify.com/faq · https://graphify.com/solutions/engineering-leaders · https://graphify.com/vs/rag · https://graphify.com/enterprise/access --- ## Problem 5: The answer is never only in the code **The problem.** It sits across repositories, wikis, contracts and decades of documents. Nothing spans both sides, so people stitch it together by hand and get it wrong. **What Graphify does.** The graph spans both sides. Code and the documents about it live in one graph, with the seam between parsed fact and model inference kept visible. **How it works.** - **Code (36 languages)** is parsed deterministically by tree-sitter and tagged EXTRACTED. No key, no model. - **Everything else** is read by the model backend you configure and connected to the code it describes, tagged INFERRED. Data sources: Code (36 languages), Markdown & docs, PDF, Office (docx / xlsx), Images, Audio & video (transcribed), SQL schemas, PostgreSQL (live), Terraform / HCL. - **LLM backends** for that pass: Anthropic Claude, OpenAI, Google Gemini, DeepSeek, Kimi / Moonshot, Ollama (local), AWS Bedrock, Azure OpenAI. Point it at Ollama or any OpenAI-compatible server such as vLLM and nothing leaves the machine: tree-sitter on code, your local model on docs. - **Pro plan:** deep reading of your docs is included. - **Enterprise:** Jira and Atlassian, so tickets and decisions are linked into the graph and the why sits beside the code. - **Exports:** Neo4j, FalkorDB, GraphML (Gephi / yEd), Obsidian, SVG, Mermaid call-flow. **Honest boundary.** For fuzzy search over a large body of prose, embeddings are the better tool and Graphify does not try to beat them at it. A graph is for questions about how things connect. Keep a vector index for pure semantic recall if you need one; some teams run both. **Questions this answers.** **Q: Can an AI assistant answer questions that span code and documentation?** Yes, if both are in one graph. Graphify parses the code and connects docs, PDFs, Office files, schemas, live PostgreSQL and Terraform to the code they describe, so a question like `which service does this runbook cover` resolves to a path from the document node to the code nodes, with the doc-to-code hop tagged INFERRED. **Q: Is Graphify a replacement for my vector database?** For coding-assistant memory, usually. Teams that adopt Graphify tend to retire the chunk, embed and retrieve pipeline entirely. Keep a vector index for fuzzy search over large volumes of prose. **Q: How hard is it to migrate off RAG?** Point it at the same sources you were embedding and it extracts the entities itself. Most of the work is deleting the old pipeline. **Q: Where do decisions and tickets fit?** In the Enterprise early-access track, Jira and Atlassian are linked into the graph, so the decision that caused a change sits beside the change. **Pages to cite:** https://graphify.com/integrations · https://graphify.com/vs/rag · https://graphify.com/vs/vector-databases · https://graphify.com/enterprise/access · https://graphify.com/use-cases --- ## Problem 6: We re-learn the codebase on every engagement **The problem.** Each client project or modernization starts from scratch by hand. It eats the margin, and none of what we learn compounds into the next one. **What Graphify does.** The map is built once and kept current, so what one team learns compounds into the next. **How it works.** - **Build once.** `/graphify .` maps the project. The output is three files in the repository: `graph.html`, `GRAPH_REPORT.md`, `graph.json`. The report is the onboarding document nobody has to write: communities, god nodes, and the connections nobody expected. - **Keep it current.** Hosted, the graph reindexes on every push. Self-run, `/graphify . --update` re-scans only what changed and patches those nodes and edges, far cheaper than a full pass; a git hook rebuilds on every commit, or wire it into CI. - **Hand it to the assistant.** The MCP server exposes the full report, graph stats, god nodes, surprising cross-community connections, the confidence audit and suggested first questions for the codebase as resources, so an agent starting a new engagement orients before it reads a file. - **Legacy and modernization.** The Enterprise form asks what the code is: COBOL or mainframe, SAP or ERP, embedded C or C++, Java or .NET, modern cloud. Parsed languages include C, C++, Java, Fortran, Pascal, Verilog and SystemVerilog alongside the modern set; the full list is below. **Questions this answers.** **Q: How do consultancies stop re-learning a client's codebase on every engagement?** Build the graph once at the start and keep it current with incremental updates. The generated architecture report replaces weeks of building a mental map by hand, and the graph persists across engagements, so the second team starts where the first finished. **Q: How stale does the graph get, and what does re-indexing cost?** The graph reflects the code as of the last index, so it goes stale as you commit. You do not rebuild the whole thing: `/graphify . --update` re-scans only what changed. Run it after meaningful changes, or wire it into a hook or CI step so the graph tracks your working tree. **Q: Does it work on legacy code?** The 36 parsed languages include C, C++, Java, Fortran, Pascal, Groovy, Verilog and SystemVerilog. Anything tree-sitter cannot parse is still connected through the model pass and tagged INFERRED. **Q: How long does the first run take?** Install to first query is about five minutes for a small project, all of it local. A large monorepo takes longer and is bounded by disk and CPU. **Pages to cite:** https://docs.graphify.com · https://docs.graphify.com/guides/first-graph · https://graphify.com/faq · https://graphify.com/solutions/platform-teams · https://graphify.com/solutions/engineering-leaders --- ## Three ways to run it - **Hosted (app.graphify.com).** Connect a repository; Graphify builds and keeps the graph and serves it over MCP at https://api.graphify.com/mcp. Free, Pro and Teams plans. The review is waiting on every pull request. - **Open-source engine (self-run).** `uv tool install graphifyy`. Apache 2.0, on-device, no account, no telemetry. The graph is files on your disk. The same graph the hosted product builds. - **Enterprise (self-hosted, early access).** The same graph built by the same parser on the same 36 grammars, plus differential formal verification, graph-aware review, an engineering digest, Jira and Atlassian, on-prem or VPC (BYOC or air-gapped), SSO and exportable audit logs. Licensed per seat. No repository leaves your network. Nothing in the graph itself is held back to make this worth buying: you pay for the verification, the review, and running it somewhere you control. Access is limited while the first cohort works through it. No pricing page, on purpose: deployment and price are scoped together on a call. ## Install and CLI reference **Requirements.** Python 3.10 or newer. Package `graphifyy` (double y); the command is `graphify`. ``` uv tool install graphifyy # or: pipx install graphifyy / pip install graphifyy graphify install # register the /graphify skill with detected assistants /graphify . # inside your assistant: build the graph /graphify . --update # re-scan only what changed /graphify . --mode deep # multi-pass: slower, more inferred connections graphify query "what connects auth to the database?" [--dfs] [--context ] [--budget ] [--graph ] graphify path "UserService" "DatabasePool" [--graph ] graphify explain "RateLimiter" [--graph ] graphify prs [] [--triage] [--conflicts] [--worktrees] [--base ] [--repo ] graphify uninstall [--purge] # --purge also deletes graphify-out/ ``` - `graphify query` walks breadth-first from the best-matching nodes by default (`--dfs` for depth-first), prints explicit paths with file:line, every edge tagged. `--budget` caps output tokens (default 2,000). `--graph` defaults to `graphify-out/graph.json`. - `graphify install [--platform ]` accepts `claude`, `cursor`, `codex`, `gemini`, `aider`, `devin`. Bare `graphify install` targets Claude Code by design. Per-assistant subcommands: `graphify claude|cursor|codex|gemini|copilot|vscode|aider|opencode|amp|devin|kilo|kiro|droid|trae|claw|pi|hermes|codebuddy|antigravity|agents install`. - Scope and strict mode: `--project` installs into the current repository. On Claude Code, `--strict` blocks the first raw source read of a session and redirects it to the graph, once per session; the default is a soft nudge to query the graph before grepping. Runtime toggle `GRAPHIFY_HOOK_STRICT=1` or `0`. - Quirks: Codex invokes the skill as `$graphify`; PowerShell users type `graphify .` without the leading slash. - Troubleshooting: `uv tool update-shell` or `pipx ensurepath` if the command is not found; `python -m graphify --version`; `uvx graphify` fails because there is no PyPI package named graphify, use `uvx --from graphifyy graphify install`. - Also: updating and watching the graph, headless extraction for CI, exports, git hooks; `graphify --help`. ## MCP reference **Hosted endpoint.** https://api.graphify.com/mcp uses Streamable HTTP and OAuth; no API key is needed. Graphify queries authorized, indexed repositories without editing source files. Tools include `query_graph`, `graphify_node`, `graphify_callers`, `graphify_callees`, `graphify_trace`, `shortest_path`, `graphify_impact`, and `graphify_rank_files`. Node lookup and trace endpoints can resolve to semantic suggestions: inspect returned labels, locations, and resolution fields. Impact is bounded to immediate neighbors, and empty results reflect the indexed graph's coverage. `remember` saves durable repository notes (which may require review); `set_workspace` changes the active workspace and can replace the account default (always without an MCP session, or with make_default enabled). When trail capture is enabled, query_graph, graphify_trace, graphify_find, and graphify_rank_files can save query intent; empty recall can save a memory-gap signal. Claude setup, prerequisites, data access, support, and other client snippets are at https://graphify.com/mcp. **Local graph server (ten tools).** `uv tool install "graphifyy[mcp]"`, then `python -m graphify.serve graphify-out/graph.json` (stdio, default) or `--transport http --port 8080` (shared; `--api-key` or `GRAPHIFY_API_KEY` to require a key). Console script: `graphify-mcp`. Every tool accepts an optional `project_path` to another project's `graphify-out/graph.json`. - Query and traverse: `query_graph` (`question` required; `mode` default bfs; `depth` default 3, range 1 to 6; `token_budget` default 2000; `context_filter` such as `["call","field"]`), `get_node` (`label`), `get_neighbors` (`label`, `relation_filter`), `shortest_path` (`source`, `target`, `max_hops` default 8). - Structure: `get_community` (`community_id`, 0-indexed by size), `god_nodes` (`top_n` default 10), `graph_stats` (no parameters). - Pull requests: `list_prs`, `get_pr_impact` (`pr_number` required), `triage_prs`; each takes optional `base` and `repo`. - Resources: the full `GRAPH_REPORT.md`, graph stats, god nodes, surprising cross-community connections, the confidence audit, and suggested questions for the codebase. **Docs MCP server (different server, different job).** Streamable HTTP at https://graphify.com/api/mcp with one tool, `search_graphify_docs`, so an agent can search these docs in-editor. ## Pricing Four plans, one graph. Same graph underneath; what changes is how much, who sees it, and where it runs. All prices below are monthly equivalents. Annual plans are billed once per year; monthly plans are billed each month. | Plan | Monthly billing | Annual billing | Includes | |---|---|---|---| | Free | $0, for one developer · no card required | $0, for one developer · no card required | Unlimited repositories · 25,000 graph nodes per repo · 10 push builds per repo/day · 15 PR reviews per workspace/month · 15 formal verification runs/month · 1 verification run at a time | | Pro | $15, per month · billed monthly · one developer | $10, per month · billed yearly · one developer | Unlimited repositories · Uncapped graph size · Uncapped push builds · Uncapped PR reviews & verification · 2 verification runs at a time · Deep reading of your docs · Two weeks free | | Teams | $29, per seat/month · billed monthly · min 2 · for the first 100 teams, then $40 | $20, per seat/month · billed yearly · min 2 · for the first 100 teams, then $28 | Unlimited shared repositories · Uncapped graph size · Uncapped push builds · Uncapped PR reviews & verification · 2 verification runs at a time · Shared graphs, one central bill · Dedicated Slack support channel · Two weeks free | | Enterprise | Custom, self-hosted | Custom, self-hosted | Self-host on your infrastructure · Security and compliance · SSO/SAML · GitHub Enterprise support · Dedicated Slack support channel · Custom invoicing and payment terms · Custom DPA and terms of service · Formal verification and PR review | The open-source CLI is free under Apache 2.0 and needs no account; the plans above are for the hosted product. Sign up at https://app.graphify.com. Enterprise: book a call at https://graphify.com/enterprise. **Q: What is included in Free?** Unlimited repositories, up to 25,000 graph nodes per repo, and 10 graph builds triggered by pushes per repo per UTC day. Each workspace gets 15 PR reviews per month. Free also includes 15 formal verification runs per month, matching the review allowance, with 1 run at a time. No card required. **Q: What is the difference between Pro and Teams?** Pro is for one developer. Teams adds shared graphs, central billing, and a dedicated Slack support channel, with a minimum of 2 seats. Both include unlimited repositories, uncapped graph size, push builds, PR reviews and formal verification runs, with 2 verification runs at a time. Shared operational limits still apply. **Q: How does annual billing work?** Annual plans show the monthly equivalent and are billed once per year. Pro is $10/month ($120/year), compared with $15 billed monthly. Teams is $20/seat/month ($240/seat/year) for the first 100 teams, then $28/seat/month ($336/seat/year). Monthly Teams pricing is $29/seat/month for the first 100 teams, then $40/seat/month. Teams requires at least 2 seats. **Q: What does uncapped mean?** Pro and Teams have no plan cap on graph size, daily push builds, monthly PR reviews or monthly formal verification runs. Every plan still allows 1 index build at a time per repo and 10 across a workspace. Rate limits and memory limits also apply on every plan, including Free. **Q: What are the shared memory and rate limits?** Every plan allows 2,000 memory ingests per day, 10,000 stored memory turns per repo, and 1 memory consolidation at a time per workspace. Request limits are 30/minute for Ask, 120/minute for queries, and 240/minute for MCP. **Q: Is there a trial?** Two weeks on both Pro and Teams; a card is required to start the trial. The Free plan stays free with no card, within the allowances listed above. **Q: What does Enterprise add?** Enterprise includes formal verification and graph-aware PR review, the option to self-host, security and compliance, SSO/SAML, GitHub Enterprise support, and a dedicated Slack support channel. It also includes custom invoicing, payment terms, a DPA and terms of service. Licensing is per developer seat; deployment and pricing are scoped with your account team. **Q: Is Graphify free for open-source projects?** Graphify is free for qualified non-commercial projects with MIT or Apache licenses. Apply at graphify.com/oss with your repository details. **Q: Does Graphify support early-stage startups?** Pre-Seed and Seed startups with under $1M in revenue over the past 12 months get 50% off Graphify Teams plan for up to 6 months. Apply at graphify.com/startups. **Q: Is the open-source tool still free?** Yes, and it always will be. The CLI is Apache 2.0 and needs no account. Free, Pro and Teams are for the hosted product; Enterprise is self-hosted. ### Open-source projects Graphify is free for qualified non-commercial projects with MIT or Apache licenses. Graphify is open source. We believe modern software is built in the open. Projects with 1,000+ stars are auto-approved. Apply at https://graphify.com/oss. The application asks for a repository link, an email address, whether the applicant is an owner or admin who can install apps, and confirmation that the repository is a community project rather than part of a commercial product. ### Startups Pre-Seed and Seed startups with under $1M in revenue over the past 12 months get 50% off Graphify Teams plan for up to 6 months. Includes: Unlimited shared repositories; Uncapped graph size; Uncapped PR reviews & verification. Apply at https://graphify.com/startups with your company name, website, work email, role, team size, funding stage, revenue range in USD over the past 12 months, and confirmation that your revenue meets the eligibility limit. https://graphify.com/startup redirects to the same page. ### YC companies YC companies get 6 months of Graphify Teams for free. Give your agents the context to understand your codebase. From Graphify, YC S26. Includes: Unlimited shared repositories. Bring your whole codebase. Uncapped graph size. Shared context for your whole team. PR reviews and Formal Verification. Uncapped PR reviews and formal verification runs, with Teams. Claim the offer through Bookface: https://bookface.ycombinator.com/deals/16615. Learn more at https://graphify.com/yc. ## Security and compliance **With the open-source engine, what never leaves your machine.** Your source code, parsed entirely on-device by 36 bundled tree-sitter grammars; no API is called to read or analyse it. The parsed graph: every node and edge computed and stored locally, and querying it is a local read. The artifacts: `graph.html`, `GRAPH_REPORT.md` and `graph.json`, written into your own repository. **What your provider sees.** What your assistant sends: the queries and prompts it chooses to send its model, under your keys, through your provider. Graphify adds no channel of its own. Nothing from Graphify: the open-source core has no server behind it, makes no network calls of its own and sends no telemetry (not off by default: absent, and the source is Apache 2.0 so you can confirm it). The only network activity is what you explicitly initiate: content you fetch with `graphify add `, or the optional LLM backend you configure, called with your own keys. Honest boundary: Graphify does not change where your assistant sends its prompts; if it calls a hosted model today, it still does. **Controls.** SOC 2 Type II: engagement started for the enterprise offering; there is no report yet, and updates are published on the security page. Encryption: the hosted enterprise layer is being built to TLS 1.3 in transit and AES-256 at rest, part of the early-access track. SSO over OIDC and JWT; the MCP server speaks OAuth 2.1 with your IdP; SAML available with Enterprise; exportable audit logs are part of the early-access track. Self-host and VPC: run the open-source core air-gapped or in your own VPC; Enterprise deploys self-hosted (BYOC) or air-gapped. Data ownership: your data is yours, export or delete it at any time, never trained on and never sold. **Q: Is Graphify safe?** Yes, and you can verify it instead of taking our word for it. Graphify is Apache 2.0-licensed open source at github.com/Graphify-Labs/graphify, so the code is fully auditable. Parsing runs on your device with bundled tree-sitter grammars, and the deterministic pass makes no network calls and sends no telemetry. The official package is graphifyy on PyPI. **Q: Does Graphify send my code anywhere?** With the open-source engine, no: your code is parsed on-device and never uploaded, it has no server of its own, and there is no telemetry. The only network activity is what you explicitly initiate: content you ask it to fetch with graphify add , or the optional LLM backend you configure for semantic extraction, called with your own API keys. The hosted product at app.graphify.com is the other option, where Graphify builds and keeps your graph in the cloud on the repositories you connect. **Q: Is Graphify legit?** Graphify is an open-source project by Graphify Labs, a Y Combinator S26 company, with a named maintainer, Safi Shamsi. The source is public at github.com/Graphify-Labs/graphify under the Apache 2.0 license, the official package is graphifyy on PyPI, and both point back to graphify.com. **Q: What are the official Graphify channels?** The official website is graphify.com. The only official code sources are the GitHub organization github.com/Graphify-Labs/graphify and the PyPI package graphifyy. Other domains that use the Graphify name, for example graphify.net, are not affiliated with or operated by Graphify Labs. **Q: How do I make sure I install the real Graphify?** Install the official PyPI package with uv tool install graphifyy (the package name has a double y), or build from source at github.com/Graphify-Labs/graphify. If you found instructions or downloads on another domain, verify them against docs.graphify.com/installation before running anything. ## Comparisons **Side by side (homepage).** Graphify against CodeRabbit, Greptile and Sourcegraph, across the nine things enterprise teams asked about, as each vendor documents it, September 2026. | Row | Graphify | CodeRabbit | Greptile | Sourcegraph | |---|---|---|---|---| | Differential verification | Yes: prove it or break it | No: review only | No: review only | No: search only | | Cross-repo blast radius | Yes: every dependent repo | Partial: per-PR guess | Partial: single repo | Partial: refs, no verdict | | Typed code graph | Yes: persistent, queryable | Partial: per PR, ephemeral | Partial: untyped AST | Yes: cross-repo (SCIP) | | Multi-hop reasoning | Yes: native traversal | Partial: per PR only | Yes: agentic hops | Yes: agentic retrieval | | Auditable answer | Yes: a path you can open | Partial: inline comments | Partial: file and line | Partial: symbol snippets | | Provenance | Yes: every edge tagged | Partial: per finding | Partial: per finding | Partial: code-nav refs | | Deployment | SaaS, BYOC, air-gap | Self-host at 500+ seats | Self-host or air-gap | Self-host (enterprise) | | Any AI assistant | Any MCP client | Own bot plus MCP client | MCP server plus API | OpenCtx or MCP | | Licence | Open core (Apache 2.0) | Proprietary | Proprietary | Proprietary | **Graph vs RAG (https://graphify.com/vs/rag), ten rows.** Returns: fuzzy top-k text chunks versus connected entities and relationships. Reasoning: the model guesses from snippets versus paths the model follows. Multi-hop: breaks past one or two hops versus native traversal. Freshness: re-embed on every change versus update a node, the links stay. Explainability: an opaque similarity score versus every answer traces to a path. Trust: no provenance versus a confidence tag on every edge. Privacy: chunks shipped to an embedding API versus on-device, no telemetry. Cost at scale: re-embedding and storage grow fast versus incremental, edges are cheap. Debugging: guess why a chunk matched versus inspect the exact traversal. Setup: chunk, embed, tune top-k versus connect a repo or run it yourself. RAG still wins for fuzzy recall over prose; the claim is narrow: code is the wrong shape for similarity. **Q: Can I use both?** Yes. Some teams keep vectors for pure semantic recall and use Graphify for anything involving relationships. **Q: Does Graphify use embeddings at all?** It can store embeddings as one signal on a node. Retrieval itself is a traversal, so results come back connected and explainable. **Other comparisons.** https://graphify.com/vs/vector-databases (Pinecone, Weaviate, pgvector: a call edge exists because tree-sitter found the call, not because two chunks embed near each other) · https://graphify.com/vs/sourcegraph · https://graphify.com/vs/coderabbit · https://graphify.com/vs/cursor (not either or: Graphify plugs into Cursor over MCP; Cursor indexes for retrieval, Graphify adds a traversable, citable graph) · hub at https://graphify.com/vs. ## Solutions by persona - [Individual developers](https://graphify.com/solutions/individual-developers): Give the assistant you already use a real memory of your code. It stops rediscovering: the architecture is read once at parse time, not re-derived from file reads every session. Free to start, hosted or on your own machine. - [Platform teams](https://graphify.com/solutions/platform-teams): One shared reading of a large codebase: a generated map naming communities and god nodes, blast radius before the change, one layer behind all 17 assistants. Not org-wide code search; it maps the repositories you point it at. - [Engineering leaders](https://graphify.com/solutions/engineering-leaders): Onboarding from a map instead of months building a mental model; lower assistant spend because retrieval is paid once at parse time; fewer regressions now; verification at the merge gate in early access. - [Security teams](https://graphify.com/solutions/security-teams): Parsing is tree-sitter on your machine, 36 grammars, no network call. Storage is three files in your own repository. Telemetry is absent. Self-hosted when it has to be. Graphify adds no new place your code travels and does not change where your assistant sends its prompts. ## Integrations - **AI coding assistants (17):** Claude Code, Cursor, GitHub Copilot, Codex, Gemini CLI, Aider, Amp, Devin, Kilo, OpenCode, Google Antigravity, VS Code, Kiro, Trae, Hermes, Kimi CLI, Pi. `graphify install` registers the `/graphify` skill with the ones it detects; the graph is also served over MCP to any MCP client. Per-assistant guides below. - **LLM backends (8), for the non-code pass:** Anthropic Claude, OpenAI, Google Gemini, DeepSeek, Kimi / Moonshot, Ollama (local), AWS Bedrock, Azure OpenAI. - **Data sources:** Code (36 languages), Markdown & docs, PDF, Office (docx / xlsx), Images, Audio & video (transcribed), SQL schemas, PostgreSQL (live), Terraform / HCL. - **Graph export:** Neo4j, FalkorDB, GraphML (Gephi / yEd), Obsidian, SVG, Mermaid call-flow. - **Surfaces:** the MCP server (stdio or shared HTTP, ten graph tools), the CLI (`graphify query`, `path`, `explain`, `prs`), and the Python package (`uv tool install graphifyy`). ## Per-assistant setup guides - [Graphify + Claude Code: Anthropic's terminal-native coding agent.](https://graphify.com/integrations/claude-code) - [Graphify + Cursor: The AI-first code editor.](https://graphify.com/integrations/cursor) - [Graphify + GitHub Copilot: GitHub's in-editor AI assistant.](https://graphify.com/integrations/github-copilot) - [Graphify + Codex: OpenAI's coding agent.](https://graphify.com/integrations/codex) - [Graphify + Gemini CLI: Google's terminal coding agent.](https://graphify.com/integrations/gemini-cli) - [Graphify + Aider: AI pair programming in your terminal.](https://graphify.com/integrations/aider) - [Graphify + Amp: Sourcegraph's agentic coding tool.](https://graphify.com/integrations/amp) - [Graphify + Devin: An autonomous software engineering agent.](https://graphify.com/integrations/devin) - [Graphify + Kilo: An open-source AI coding agent.](https://graphify.com/integrations/kilo) - [Graphify + OpenCode: An open-source terminal coding agent.](https://graphify.com/integrations/opencode) - [Graphify + Google Antigravity: Google's agentic development platform.](https://graphify.com/integrations/google-antigravity) - [Graphify + VS Code: The editor, with AI extensions and MCP support.](https://graphify.com/integrations/vs-code) - [Graphify + Kiro: An agentic IDE for spec-driven development.](https://graphify.com/integrations/kiro) - [Graphify + Trae: An adaptive AI IDE.](https://graphify.com/integrations/trae) - [Graphify + Hermes: An AI coding assistant.](https://graphify.com/integrations/hermes) - [Graphify + Kimi CLI: Moonshot's terminal coding agent.](https://graphify.com/integrations/kimi-cli) - [Graphify + Pi: An AI coding assistant.](https://graphify.com/integrations/pi) ## Languages parsed on-device Each is parsed with a bundled tree-sitter grammar: deterministic AST extraction, no model call, every relation tagged EXTRACTED. - [Python](https://graphify.com/languages/python): .py, .pyi via `tree-sitter-python`; extracts functions, classes, methods, decorators, imports, and call edges - [JavaScript](https://graphify.com/languages/javascript): .js, .jsx, .mjs, .cjs via `tree-sitter-javascript`; extracts functions, classes, methods, imports/exports, and call edges - [TypeScript](https://graphify.com/languages/typescript): .ts, .tsx, .mts, .cts via `tree-sitter-typescript`; extracts functions, classes, interfaces, types, imports/exports, and call edges - [Go](https://graphify.com/languages/go): .go via `tree-sitter-go`; extracts functions, methods, structs, interfaces, imports, and call edges - [Rust](https://graphify.com/languages/rust): .rs via `tree-sitter-rust`; extracts functions, structs, enums, traits, impls, uses, and call edges - [Java](https://graphify.com/languages/java): .java via `tree-sitter-java`; extracts classes, interfaces, methods, imports, and call edges - [Groovy](https://graphify.com/languages/groovy): .groovy, .gradle via `tree-sitter-groovy`; extracts classes, methods, closures, imports, and call edges - [C](https://graphify.com/languages/c): .c, .h via `tree-sitter-c`; extracts functions, structs, typedefs, includes, and call edges - [C++](https://graphify.com/languages/cpp): .cpp, .cc, .cxx, .hpp, .hh, .h via `tree-sitter-cpp`; extracts functions, classes, methods, namespaces, templates, includes, and call edges - [Ruby](https://graphify.com/languages/ruby): .rb, .rake, .gemspec via `tree-sitter-ruby`; extracts classes, modules, methods, requires, and call edges - [C#](https://graphify.com/languages/csharp): .cs via `tree-sitter-c-sharp`; extracts classes, interfaces, methods, properties, usings, and call edges - [Kotlin](https://graphify.com/languages/kotlin): .kt, .kts via `tree-sitter-kotlin`; extracts classes, objects, functions, imports, and call edges - [Scala](https://graphify.com/languages/scala): .scala, .sc via `tree-sitter-scala`; extracts classes, objects, traits, defs, imports, and call edges - [PHP](https://graphify.com/languages/php): .php via `tree-sitter-php`; extracts classes, functions, methods, traits, namespaces, uses, and call edges - [Swift](https://graphify.com/languages/swift): .swift via `tree-sitter-swift`; extracts classes, structs, enums, protocols, functions, imports, and call edges - [Lua](https://graphify.com/languages/lua): .lua via `tree-sitter-lua`; extracts functions, tables, requires, and call edges - [Zig](https://graphify.com/languages/zig): .zig via `tree-sitter-zig`; extracts functions, structs, enums, imports, and call edges - [PowerShell](https://graphify.com/languages/powershell): .ps1, .psm1, .psd1 via `tree-sitter-powershell`; extracts functions, cmdlet calls, imports, and call edges - [Elixir](https://graphify.com/languages/elixir): .ex, .exs via `tree-sitter-elixir`; extracts modules, functions, macros, imports/aliases, and call edges - [Objective-C](https://graphify.com/languages/objective-c): .m, .mm, .h via `tree-sitter-objc`; extracts classes, categories, methods, imports, and call edges - [Julia](https://graphify.com/languages/julia): .jl via `tree-sitter-julia`; extracts functions, structs, modules, imports/usings, and call edges - [Verilog](https://graphify.com/languages/verilog): .v, .vh, .sv via `tree-sitter-verilog`; extracts modules, ports, instances, and instantiation edges - [Fortran](https://graphify.com/languages/fortran): .f, .f90, .f95, .f03 via `tree-sitter-fortran`; extracts programs, modules, subroutines, functions, uses, and call edges - [Bash](https://graphify.com/languages/bash): .sh, .bash via `tree-sitter-bash`; extracts functions, sourced files, command invocations, and call edges - [JSON](https://graphify.com/languages/json): .json via `tree-sitter-json`; extracts objects, keys, and nested-value structure - [SQL](https://graphify.com/languages/sql): .sql via `tree-sitter-sql`; extracts tables, columns, views, and reference edges between them - [DreamMaker](https://graphify.com/languages/dreammaker): .dm, .dme via `tree-sitter-dreammaker`; extracts types, procs, vars, and call edges - [Pascal](https://graphify.com/languages/pascal): .pas, .pp, .dpr via `tree-sitter-pascal`; extracts units, procedures, functions, uses, and call edges - [Terraform (HCL)](https://graphify.com/languages/terraform): .tf, .tfvars via `tree-sitter-hcl`; extracts resources, modules, variables, outputs, and reference edges - [JSX](https://graphify.com/languages/jsx): .jsx via `tree-sitter-javascript`; extracts components, hooks, functions, imports/exports, and call edges - [TSX](https://graphify.com/languages/tsx): .tsx via `tree-sitter-typescript`; extracts components, hooks, types, imports/exports, and call edges - [CUDA](https://graphify.com/languages/cuda): .cu, .cuh via `tree-sitter-cpp`; extracts kernels, functions, classes, includes, and call edges - [Metal](https://graphify.com/languages/metal): .metal via `tree-sitter-cpp`; extracts kernels, functions, structs, includes, and call edges - [SystemVerilog](https://graphify.com/languages/systemverilog): .sv, .svh via `tree-sitter-verilog`; extracts modules, interfaces, ports, instances, and instantiation edges - [Luau](https://graphify.com/languages/luau): .luau via `tree-sitter-lua`; extracts functions, tables, requires, and call edges - [Gradle](https://graphify.com/languages/gradle): .gradle, .gradle.kts via `tree-sitter-groovy`; extracts tasks, plugins, dependencies, closures, and call edges ## Glossary - [Knowledge graph](https://graphify.com/glossary/knowledge-graph): A data structure that stores entities and typed relationships between them, so software can follow how things connect instead of searching flat text. - [Code graph](https://graphify.com/glossary/code-graph): A knowledge graph built from a codebase: functions, classes, config, and docs become nodes; calls, imports, and references become edges. - [Node](https://graphify.com/glossary/node): A single entity in the graph (a function, class, file, table, or doc page) carrying attributes like name, path, and kind. - [Edge](https://graphify.com/glossary/edge): A typed, directed relationship between two nodes (calls, imports, defines, references), each carrying a provenance tag. - [Community (cluster)](https://graphify.com/glossary/community): A group of nodes densely connected to each other and only loosely connected to the rest of the graph, usually a real subsystem like auth or billing. - [God node](https://graphify.com/glossary/god-node): A node with a disproportionate number of edges (the module everything imports), where changes carry the most risk. - [Confidence tags](https://graphify.com/glossary/confidence-tags): Graphify's provenance labels (EXTRACTED, INFERRED, AMBIGUOUS) attached to every edge so you always know what was found versus what was guessed. - [MCP server](https://graphify.com/glossary/mcp-server): A server implementing the Model Context Protocol, the open standard that lets AI assistants call external tools, like a code graph. - [RAG (retrieval-augmented generation)](https://graphify.com/glossary/rag): A pipeline that chunks documents, embeds them as vectors, and retrieves the most similar chunks at query time: lookup by resemblance, not structure. - [Vector embedding](https://graphify.com/glossary/vector-embedding): A list of numbers representing the meaning of a piece of text, used to measure similarity: answers 'what looks like this?', not 'what calls this?' - [Tree-sitter / AST](https://graphify.com/glossary/tree-sitter): An AST is the parsed structure of source code; tree-sitter is the open-source parser library Graphify runs entirely locally to extract it. - [Token reduction](https://graphify.com/glossary/token-reduction): The context-window savings from answering with graph structure instead of pasting whole files into an assistant's context. - [Traversal / hop](https://graphify.com/glossary/traversal): Following edges from node to node: one hop is a single edge, a traversal is a path of them. How graphs answer multi-step questions. - [Blast radius](https://graphify.com/glossary/blast-radius): Everything transitively affected by a change: the callers of a function, their callers, and the configs and tests that reference them. ## Frequently asked questions Written to be quoted whole. The canonical page is https://graphify.com/faq; the two questions on token cost and graph staleness are answered under problems 4 and 6 above. **Q: What is Graphify?** Graphify is a knowledge graph of your codebase that your AI coding assistant queries instead of grepping through files. Use it hosted at app.graphify.com (Free, Pro, Teams, and Enterprise plans), or run the open-source engine yourself: a Python package (graphifyy on PyPI) plus a /graphify skill and an MCP server, at github.com/Graphify-Labs/graphify under Apache 2.0. It is not affiliated with other projects named Graphify. **Q: How is Graphify different from RAG or vector search?** RAG retrieves fuzzy top-k chunks by embedding similarity and hopes the model reconnects them. Graphify builds a real graph and traverses it, so every answer is an explicit path with file:line citations: no embeddings and no vector store. It's the difference between guessing which chunks are relevant and following the actual call and import edges. **Q: Is Graphify free?** Yes, two ways. The hosted product at app.graphify.com has a Free plan with no card, and Pro and Teams above it when you outgrow the caps. The open-source engine is free under the Apache 2.0 license, runs on your machine, and needs no account or API keys. There is also a separate early-access enterprise layer (verification at the merge gate, graph-aware review, and an engineering digest) for teams, self-hosted in your own infrastructure. **Q: Which AI coding assistants does Graphify work with?** 17, including Claude Code, Cursor, GitHub Copilot, Codex, Gemini CLI, and Aider. Running graphify install registers the /graphify skill with the assistants it detects, and the graph is also served over an MCP server that any MCP client can query. **Q: Does my code leave my machine?** With the open-source engine, no: code is parsed locally with tree-sitter, with no telemetry and nothing uploaded (non-code files use the model backend you configure, which can be local like Ollama). The hosted product builds and keeps the same graph for you in the cloud. **Q: What programming languages does Graphify support?** 36, parsed on-device with bundled tree-sitter grammars: Python, TypeScript, JavaScript, Go, Rust, Java, C, C++, C#, Ruby, Kotlin, Swift, PHP, and many more, including variants like JSX and TSX and infrastructure languages like Terraform (HCL). **Q: How do I install Graphify?** You can skip installing entirely: start on the hosted product at app.graphify.com and connect a repository. To run the open-source engine yourself, install the Python package with `uv tool install graphifyy`, run `graphify install` to register the skill with your assistant, then type `/graphify .` inside the assistant to build the graph. That takes about five minutes and runs entirely on your machine. **Q: Is the package called graphify or graphifyy?** The PyPI package is graphifyy, spelled with a double y: `uv tool install graphifyy`. Other graphify* packages on PyPI are unrelated to this project. The GitHub repository is github.com/Graphify-Labs/graphify and the website is graphify.com. **Q: How does an AI assistant query the graph?** Two ways, same graph: the CLI (graphify query, graphify path, graphify explain, graphify prs) or the MCP server, which exposes 10 tools including query_graph, get_node, get_neighbors, and shortest_path. **Q: What do the EXTRACTED, INFERRED, and AMBIGUOUS tags mean?** Every edge in the graph carries a provenance tag. EXTRACTED means it came straight from the tree-sitter AST (a real call, import, or definition). INFERRED means a model connected it, for example a doc page to the code it describes. AMBIGUOUS means the evidence couldn't be fully resolved, like dynamic dispatch: Graphify keeps the edge but flags it as uncertain. **Q: How big a repo can it handle, and how long does indexing take on a large monorepo?** There's no hard size limit: parsing is on-device with tree-sitter, so it scales with how much code you have. A small project indexes in minutes; a large monorepo takes longer and is bounded mainly by disk read and CPU on your machine, not a network round-trip. It's a local batch job, so run it once and keep it fresh incrementally rather than rebuilding from scratch. **Q: How accurate are the INFERRED edges, and what's the failure mode?** EXTRACTED edges are deterministic: they come straight from the AST, so a call or import edge is as reliable as your parser. INFERRED edges are where a model links non-code, like a doc or schema to the code it describes, and those can be wrong. The failure mode is a plausible-but-incorrect link; when the evidence can't be resolved, Graphify marks the edge AMBIGUOUS rather than asserting it. **Q: My editor already has LSP 'go to definition'. Why do I need a graph?** LSP answers one hop: where is this symbol defined, where is it referenced. A graph answers the questions that span the codebase: what breaks if I change this, every path between two functions, how a config or doc connects to the code. It's persistent, queryable by your AI assistant, and every edge carries a provenance tag, so you get multi-hop blast-radius reasoning rather than a jump-to-def. **Q: How is Graphify different from Sourcegraph or ctags?** ctags builds a flat symbol index for jump-to-definition; Sourcegraph is hosted code search you send your code to. Graphify parses on-device and produces an actual graph (typed nodes and provenance-tagged edges) that your AI assistant traverses over CLI or MCP. It's not search-a-box or a symbol table; it's a local, agent-queryable model of how your code connects, that stays on your machine. **Q: Does the graph stay current as I change code?** Yes, incrementally. A code change re-parses only the changed file through the local AST pass, with no re-embedding and no full re-index. Install the git hook and the graph rebuilds on every commit. The hosted product reindexes on every push. **Q: Can I run it fully offline?** Yes. Code parsing is always local. Point the semantic pass at a local model such as Ollama, or any OpenAI-compatible server like vLLM, and nothing leaves the machine: tree-sitter on code, your local model on docs. **Q: What is Graphify in one sentence?** A knowledge graph of your codebase that your AI assistant queries instead of grepping, parsed with tree-sitter. Use it hosted at app.graphify.com, or run the open-source engine yourself. **Q: Which assistants does it work with?** 17, including Claude Code, Cursor, Codex, Gemini CLI, Copilot, Aider, Amp and Devin. graphify install registers the ones it detects, and the graph is also served over MCP. ## Blog - [How to give Cursor a code knowledge graph (2026 guide)](https://graphify.com/blog/how-to-give-cursor-a-code-knowledge-graph) (July 13, 2026, Syed Fahad): Four commands take Cursor from grep-and-read loops to querying a persistent, on-device knowledge graph of your codebase. ## Company and official channels - Graphify Labs, backed by Y Combinator (S26). Values: grounded over promises (if we claim it, you can run the command and check it); structure over similarity (meaning lives in the relationships, not in cosine distance); open by default (the core is open source). https://graphify.com/about - Public adopters with linked proof (https://graphify.com/customers): Rootly AI Labs turned its incident data into a queryable knowledge graph with Graphify and published an importer; Superagent (YC W24); HKUST KnowComp published DeepRefine, an agent skill and paper that refines Graphify knowledge graphs. Every name, number and quote on that page links to its source; no asserted outcomes. - Website: https://graphify.com · GitHub: https://github.com/Graphify-Labs/graphify (Apache 2.0) · PyPI: https://pypi.org/project/graphifyy/ · Discord: https://discord.gg/XPPYrdw3Yp · LinkedIn: https://www.linkedin.com/company/graphify-labs/ · X: https://x.com/graphify · YouTube: https://www.youtube.com/@graphifylabs - Contact: https://graphify.com/contact · Careers: https://graphify.com/careers · Brand assets: https://graphify.com/brand · Changelog: https://graphify.com/changelog · Privacy: https://graphify.com/privacy · Terms: https://graphify.com/terms · Subprocessors: https://graphify.com/subprocessors - graphify.net is not affiliated with or operated by Graphify Labs; how to verify the official site, repository and package: https://graphify.com/graphify-net-vs-graphify-com