markdown you own
the memory is plain files on your disk. open them, read them, git diff them, leave any time. indexes are derived and disposable. it's your folder.
your agent ran, reported success, and the thing you asked for did not happen. usually there is no record to contradict it. homestead-memory records what your agent actually did into a local file, where every entry is hash-chained to the one before it, so editing, deleting or reordering any record breaks every hash after it. it is a file, not a platform: no server, no deployment, and an evidence pack a third party can verify with nothing installed. the same tool still owns your memory as plain markdown and catches rot, tampering and poisoning with mechanical checks, not vibes.
pip install homestead-memory
v0.4.0 · needs python 3.10+ · macos, linux, windows.
on a stock mac python3 is still 3.9, so if pip says
"no matching distribution found", run
uvx --from homestead-memory hsm verify --demo instead.
Start with the CLI, verify the demo fixture, then connect the MCP server to the agent harness you already use.
pip install homestead-memory
hsm hook --install # prints a hook. you
# paste it. nothing is
# edited for you.
hsm watch # in order, as it ran
# 0 14:02:11 Bash npm test
# 1 14:02:19 Read src/api/billing.py
# 2 14:02:24 Edit src/api/billing.py
hsm export --evidence # a pack anyone can
# check, no install every entry is hash-chained to the one before it. edit, delete or reorder any record and every hash after it breaks, at the exact index.
npm install -g @tobilu/qmd@2.1.0
hsm init ./my-vault
hsm ingest ./my-vault
hsm qmd start
hsm ask "what did i decide about x?" \
--retrieval balanced
hsm verify ./my-vault claude mcp add homestead-memory \
-- hsm mcp ~/my-vault nine mcp tools: ask, search, verify, history, ingest, distill, sign, remember, resolve. your agent gets a memory it can prove.
that's the real cli. hsm verify exits nonzero on rot, so it gates your
ci and your cron like a test suite. run it yourself:
pip install homestead-memory
the memory is plain files on your disk. open them, read them, git diff them, leave any time. indexes are derived and disposable. it's your folder.
memory rots quietly: notes contradict themselves, sources vanish, stale values shadow current ones. hsm verify scores integrity 0 to 100 and exits nonzero on rot. it gates ci like a test suite.
the optional distilled layer extracts facts with verbatim quotes, checked in code. a claim either cites a real source or gets dropped. contradictions append a changelog line, never a silent overwrite.
the week we shipped this, the biggest podcast in tech spent thirty minutes on the same idea: stop handing your data and your edge to a frontier lab that can turn around and compete with you. all-in, ep 279.
"data retention is your treasure. transfer it at your own peril."
"why would you ever share proprietary data with them? you are mortgaging your future."
they're describing the fortune-500 version: on-prem clusters, a server per employee, roll your own model. homestead-memory is the one you run on your laptop tonight. your memory, kept local, plain markdown you own, and it catches rot, tampering, and poisoning. free, and mit.
and the labs that host your model will happily ship the product that competes with you. ask figma. ask cursor. use claude code all you want. just don't hand it your memory.
cloud memory bills you per turn, forever. this is $0 to write, and it's a folder you already own.
the number that's actually ours. mechanical integrity checks over the store, no llm judging its own homework. recall is a crowded lane; nobody else scores whether the memory itself was tampered with or silently rewritten. self-scored on our vaults, so break it: the harness is open.
did the memory surface the right evidence? full 500-question LongMemEval set, 48-session haystacks with distractors. reader-independent.
honest and mid. others self-report higher on harnesses you can't run. ours you can, every failed run included.
the cost axis. verbatim memory, $0 write-time cost.
every number above comes from a harness you can run yourself, judged with the official LongMemEval methodology. recall is a crowded lane and ours is honest but mid. rotbench is the number that's actually ours: locomo and longmemeval measure whether your agent remembers. rotbench measures whether what it remembered can be trusted, that it wasn't corrupted, poisoned, or silently rewritten. nobody else scores that. we publish the failed runs too. the full run history is public.
the integrity score is only credible if it survives adversaries. build a memory store that's obviously rotten but scores intact, or an intact one that false-positives, and we merge your fixture and fix the check. the scoreboard of merged breaks lives in the repo. trust you can inspect, and break, is the one thing a duopoly can't ship.
owasp named this risk ASI06, memory & context poisoning, and its own reference project says the definition still lacks an implementation. rotbench is the missing number: a 0 to 100 score over the store itself.
the rotbench standard → read the spec on github →