Every AI you talk to forgets you the moment you close the tab. You re-explain your job, your projects, your tone, your context, every single time. The model is brilliant for ninety seconds and then a stranger again.
Mneme is the layer that fixes that, and it is not one app. It is four sibling repositories with their own remotes, their own release cadence and no shared tooling: the private engine, a zero-dependency open package, the hosted API, and the public site. Siblings, not a monorepo, and one rule holds the arrangement together: product source never enters the site repo. Not a helper, not a type. The site can be public because there is nothing in it to leak.
The engine indexes email, calendar, chat and files into an entity graph, then retrieves the way the research says works best: Personalized PageRank over that graph, the HippoRAG-2 algorithm reimplemented in pure Python. Feed the same Brain to Claude, GPT, Gemini or a local model. You are never locked to one lab.
The chain matters more than any link in it. Brain.load() picks the best backend available at every step and falls all the way down to offline ones, so the quickstart runs with zero configuration and zero keys. Adding a real backend changes no code, only an environment variable. What it changes is honesty: when the chain falls through to the fake embedder it logs what it tried and why each candidate was skipped, because fake vectors make retrieval meaningless.

The site explains the product the way the engine actually thinks: people (who matters, and why), promises (what you said you'd do), threads (the shape of a conversation), and context (the thing under the thing). The demo questions say it best: who's waiting on me, what did we decide about pricing, who introduced me to Maya.

For developers, build_context returns ranked, deduplicated, cited context sized to a token budget, and why traces any answer back to the entities that carried it and the triples behind them. Everything pluggable is a Protocol: bring your own extractor, embedder, chunker or store.
mneme mcp serves a Brain to Claude Desktop or Cursor over stdio, read-only. mneme hook does the opposite and matters more. MCP is pull memory, where the model has to decide to call a tool; the hook is push memory, wired to Claude Code's UserPromptSubmit so recall runs as code before the model wakes and arrives as additional context. Zero LLM calls on the prompt path, an eight-second budget, and a named failure rather than meaningless recall when the embedder does not match the store's recorded vector identity.

Two auth lanes, on purpose. Every /v1 route accepts either a key or a session, but minting keys and reading usage are session-only: you never provision credentials with a credential. Keys are stored as sha256 hashes and the plaintext is returned exactly once, at creation.
The 404 is the detail I care about most. Every per-Brain route runs the ownership check before it touches a store, and a Brain you do not own answers 404, not 403, so another account's workspace is indistinguishable from one that never existed. A 403 confirms something is there. Underneath, row-level security is on for every table, so authorization is enforced at the database rather than trusted to whatever route is in front of it.
The engine stays local. It is alpha and not on PyPI: pip install mneme there fetches an unrelated package from 2014, so it installs from source while the name is resolved through PEP 541.
The open package is the one meant to be given away: stdlib only, one SQLite file, FTS5 for text fallback, heuristic entities, a co-occurrence graph and a hand-rolled Personalized PageRank, in about a thousand lines across seven files you can read in a sitting. No keys, no model downloads, nothing leaves the machine. It installs from git for the same naming reason.
The cloud API builds without local Docker. Cloud Build takes the workspace root as its context, because the Dockerfile copies the sibling engine source and Docker cannot reach across sibling contexts, with the timeout raised to 1,800 seconds because pip-installing engine and service overruns the ten-minute default. The root is not a git repo, so a tracked ignore file is copied in first and the context is verified to exclude .env before every submit. No agent ever runs gcloud; that step is owner-run and live. The hosted API is not yet deployed.
The site is Next.js on Vercel, and it is the only public repo.
The page leads with one sentence, "Give your AI a brain", set over a slow bloom of petals, and everything else stays out of the way.
The engine is pure Python 3.10 and up, with every answer traceable back to its sources and a cost ledger you can read. The hosted service is FastAPI over Supabase, Postgres and pgvector, containerized through Cloud Build into Artifact Registry. The site is Next.js with Tailwind and Framer Motion.
Model intelligence is converging. What will differentiate one assistant from another is how well it knows you, and whether you trust it with that. Mneme is the bet that both questions have one answer: a memory layer that lives close to home, that you can inspect, and that you own. forget and export are verbs in the CLI, not support tickets. The open package is on GitHub at github.com/saranshahuja/mneme.
