Case Study · the final reveal

MIMIR² & Ogham

MIMIR² gives coding agents a governed map of the systems they work inside. Ogham preserves the decisions, provenance and operational context that code alone cannot explain. Together they turn comprehension and institutional memory into infrastructure.

Neo4jPostgres + pgvectorPython FastAPIMCPDocker ● In Production · self-hosted
Tap or click any box to focus it and trace its connections · full screen adds pan, zoom, search and exportOpen the interactive diagram ↗
Two peer systems beneath the agents. The enforcement layer (rules, hooks and per-repo contracts) is what makes agents use them rather than fall back to reading files. Both run on hardware I operate; neither has a product UI.
5
Steps in the comprehension ladder
100+
Repositories in the graph
0
Product UIs — the agent is the interface

02The Break

The reasoning behind a system rarely lives in the system. It lives in a chat transcript, a scattered note, or one person’s head. Knowledge that can’t be queried isn’t knowledge the organization actually has. The next person, or the next agent, re-derives it, and sometimes gets it wrong.

This site had that failure. In August the boundary between these two systems was settled in an owner-approved design: the code graph stops at shared semantics, and memory belongs to Ogham. The public copy never heard. Until this week it still described MIMIR² as a “memory engine” that ingested “every repo”. The decision existed; nothing carried it to the place the next change was made.

03The Approach

  1. Comprehension first. MIMIR² answers structural questions about the code (where something lives, what it calls, what calls it) from a graph, so an agent does not load files to find out.
  2. Memory second. Ogham keeps what the code cannot say: why a decision was made, what was rejected, what broke last time. Each memory carries its source, controlled tags and a retention class.
  3. Enforcement third. Rules, tool-call hooks and a per-repo contract make both behaviors routine across agent clients: ask the graph before reading raw files, and write a decision down before the session that made it ends.

04Origin → Evolution

  1. Dec 2025Scattered understanding. MIMIR² begins as a knowledge-graph engine forked from the open-source Mimir and Cognee lineage. Cognee has since been fully absorbed.
  2. Jul 2026Graph-first comprehension. The five-step ladder ships over MCP (14 Jul), and the first controlled evaluations run against it (20–21 Jul).
  3. Jul–Aug 2026Durable decision memory. Ogham joins as the memory half (13 Jul). On 8 Aug the boundary is fixed: code facts in the graph, decisions in memory, and neither reaches into the other.
  4. Jul 2026 →Governed context as infrastructure. An audit on 16 Jul found the deployed graph had zero use. Agents read files out of habit. The enforcement layer exists because of that audit.

05System Map

The map at the top of this page is the deployed shape. AI coding clients (Claude Code and Codex today) are the moment of work. Beneath them sits the enforcement layer, and beneath that two private systems side by side. MIMIR²’s graph MCP is a thin client over a FastAPI core that scopes every query to a tenant and tracks freshness, over a Neo4j code graph built from a manifest of repositories. Ogham is an MCP server over Postgres with pgvector. Both call external embedding APIs for vector search.

Public / private: this page, the diagrams and the walkthrough are public. Both engines, the graph’s contents and the memory store are private, and the walkthrough runs on a public repository with an empty memory profile so nothing private appears on screen.

06Walkthrough

Two separate Claude Code agent sessions, recorded on 27 Sep 2026 against the live graph and memory services, on a public repository (ops-command-center-showcase) with an empty memory profile. Session A has never seen the repo. Asked whether a phone can approve a write-back, it climbs the graph to the answer (no, by design) and confirms the breakpoint. Asked what happens in landscape, it works out from the repo's own device tables that rotating most phones switches the desk-only gate off, and that no test covers it. It writes the evidence and a decision to Ogham. Session B is a different agent that shares nothing with A. Asked an ordinary question before a change, it recalls both. Asked why it should be trusted, it verifies the claim against the current code.

Walkthrough (2 min 51 s, subtitled, silent). Rendered from the two sessions’ transcripts rather than screen-captured. Every prompt is exactly what was sent. Every tool call appears in order with its real session time. Results are excerpted but never rewritten, failed calls are kept, and each reply is quoted verbatim. One result is withheld on screen: a graph call that listed the whole private fleet. Local file paths are shown as the GitHub repository they check out. The subtitles are narration, not the agents’ words.
Tap or click a participant or a message to focus it · full screen adds pan, zoom, search and exportOpen the interactive diagram ↗
The loop, as the recorded sessions ran it. Structure first, with exact lines only to confirm the numbers. The evidence is written before the decision that cites it. Then a different session recalls both and checks them against the code before relying on them.

07Interface

No UI. The LLM is the interface, and these are the real tools it calls.

MIMIR² · the comprehension ladder
mapThe shape of a repository (files, symbols, entry points) before anything is read.
findWhere a symbol or concept lives, without pulling whole files into context.
explainA symbol’s signature, callers and callees.
neighborsThe structural relationships around a symbol: what it calls, what calls it, what it imports.
readExact source, only when the rungs above are not enough.
Ogham · memory
store_memoryDurable operational context with its source, controlled tags and retention.
store_decisionA decision with its rationale and the alternatives rejected, linked to the evidence it rests on.
hybrid_searchRecall by meaning and by keyword together, with a note when the results are stale, low-confidence or contradicted.
find_relatedFollow the edges from a memory to what supports or contradicts it.

08Design Decisions

  • Knowledge is infrastructure. It is queried by agents at the moment of work, not filed for people to find later.
  • Provenance is non-negotiable. Every memory records where it came from. A decision links to its evidence, and recall says when what it found is stale, weak or contradicted.
  • Surface context at the moment of work. Hooks fire on the tool call itself. A reminder in a document nobody opens is not enforcement.
  • Structure and memory have separate owners. A code fact copied into memory goes stale on the next commit, so code facts stay in the graph and decisions stay in memory. When they disagree, an authority order settles it: the standards outrank the graph, and the graph outranks memory.

09Proof in Use

Bounded evaluations with a date and a sample size, not universal savings.

MIMIR² · live frozen run · 2026-09-21 · 12 selected symbols explain 12/12 locations resolved · 12/12 signatures source-supported callers 12/12 locations resolved · every concrete caller claim source-supported retrieved context vs a documented file/search baseline (cl100k_base tokens) all 24 questions 90,819 → 1,724 98.1% less definition / signature 25,505 → 836 96.7% less · median 16.3× (95% CI 10.7–36.7×) callers 65,314 → 888 98.6% less · median 58.1× (95% CI 22.1–112.3×) Ogham · owner-MCP recall · 2026-09-21 · unchanged 26-label set, after a ranking fix top-10 25/26 top-5 19/26 first place 3/26 the fix repaired a regression (top-10 had fallen to 16/26); the remaining miss stays visible

What these are not: token savings on every task, a reduction in anyone’s bill, or evidence of adoption. Activity telemetry for both systems is built and switched off until it passes its own certification gate, so this page makes no usage claims.

10What It Prevents

  • A decision re-derived wrong weeks later, because its reasoning lived in a chat.
  • Agents scanning whole repositories because no structural map was available.
  • The same gotcha rediscovered by every new session.
  • Confident answers built from stale or unattributed context.
  • Memory quietly overriding the code or the standards it was meant to serve.
Lineage: Ogham began as a fork of the open-source ogham-mcp and has since diverged: activity telemetry, recall ranking and store-time controls are mine. MIMIR² descends from the open-source Mimir and Cognee projects, and Cognee has been fully absorbed. Both engines are private and available to walk through in a working session.
← Janus All projects →