Project releases

RuVector KGE 0.2.0: local knowledge-graph link prediction with HolE and RotatE

@ruvector/kge, published to npm on September 28 from RuVector PR #1057, trains knowledge-graph embeddings locally and answers which entity most plausibly completes a fact. It ships a CLI, a TypeScript API and a localhost server, with a native addon and a WASM fallback. We installed it and ran its first commands.

GitHub activity: · Published:

What this is about

A knowledge graph stores facts as triples: subject, relation, object, such as Ada, bornIn, London. Knowledge-graph embeddings learn a vector for every entity and relation so you can score facts you have never seen and ask which object most plausibly completes Ada, bornIn, ?.

@ruvector/kge does that locally, with no network calls and no per-query cost, as part of the RuVector vector database family.

What changed

  • Two scorers. HolE is the default; the README explains it is the same model class as ComplEx with half the parameters, which is what allows fast approximate search over candidates. RotatE is opt-in and is the one that can compose relations, such as bornIn then locatedIn.
  • Operations: predict, similar relations, compose (RotatE only), train, evaluate, build index and optimize. After building the HNSW index, predict switches from exhaustive to approximate search and says so in its output.
  • Saved models carry a SHA-256 and optionally an HMAC signature, and loading fails closed if the file was tampered with.
  • A localhost server, kge serve, binds to 127.0.0.1 and logs method, path, status and time only, never request bodies.
  • Platforms: the README says the native addon is bundled for linux-x64 and other platforms use the WASM fallback for now. It was published from merge commit e7b0bc37 of PR #1057, the npm gitHead.

Get started

Prerequisites: Node.js. We used Node 22 on linux-x64 in an empty folder. No MCP server or skills package is documented for KGE.

npm install @ruvector/kge@0.2.0

Create facts.jsonl with one fact per line:

{"s":"Ada","r":"bornIn","o":"London"}
{"s":"London","r":"locatedIn","o":"England"}
{"s":"Alan","r":"bornIn","o":"London"}
npx kge import --triples facts.jsonl --out model.json --scorer hole --dims 64
npx kge predict --model model.json --s Ada --r bornIn -k 3

What we saw: import printed {"added":3,"entities":4,"relations":2,"triples":3,"out":"model.json"}. predict returned three candidates with scores and "exact":true, "ann":false. The ranking was not meaningful, and that is expected: the README says that before training the tables are a deterministic seed, so scores are structurally valid but not yet meaningful. Train with the TypeScript API before trusting rankings.

Use it today

Practical case: you keep a small graph of services, owners and dependencies and want suggestions for missing links. Input is a JSONL export of known triples. Workflow: import, train for a few dozen epochs through the API, build the index, then predict owner for each service with no owner. Output: ranked candidates you review rather than accept blindly.

Acceptance test: hold back ten known triples before training. After training, the true object should appear in the top ten predictions for most of them. If it does not, the graph is too small or training too short.

Push it further

Experimental commentary. Pairing graph embeddings with the same HNSW index RuVector uses for vectors means link prediction can scale like vector search. The built-in optimize step, which promotes a new model only when it beats the incumbent on a held-out test, is a careful way to tune without fooling yourself.

Limitation: we ran untrained commands on a toy graph; we did not train, evaluate or measure accuracy. Falsifiable test: train on a public benchmark split and compare the reported MRR with published HolE results; a large gap would mean a problem in training or evaluation.

Read the original on GitHub pull request

PR #1057, merge commit e7b0bc37 (npm gitHead for @ruvector/kge 0.2.0)

RuVector repository

Back to the newsroom