JodyOS · knowledge architecture

Markdown is the database. Embeddings are the index.

A personal knowledge base where every fact lives in a plain-text note a human can read and edit, and a local embedding pipeline makes all of it — notes, scanned documents, twenty years of email — searchable by meaning. No cloud services touch the content. Nothing leaves the tailnet.

embed qwen3-embed-8b · 4096-d rerank bge-rerank-v2-m3 store numpy + JSON notes Obsidian vault triage Qwen3.6-35B, local MLX
Architecture

Three layers, one contract

The vault follows the Karpathy "LLM wiki" pattern: raw material is kept immutable, distilled knowledge lives in curated notes, and a schema file teaches any agent how to behave inside the vault.

Layer 1 · Raw

Sources

Immutable imports — exported Evernote notes, scanned documents, mail extracts. Written once at ingest, never edited. Lives in _sources/ and outside the vault (~/JodyOS-mail/).

Layer 2 · Wiki

Curated notes

Hand- and agent-maintained Markdown with YAML frontmatter and [[wikilinks]]. Projects, entities, domains, decisions, reference. This is the layer that gets read, linked, and trusted.

Layer 3 · Schema

The rulebook

A CLAUDE.md at the vault root defines note types, required frontmatter, linking rules, and the security policy (no secrets in notes — only credentials_ref pointers).

Data structure

The vault on disk

Numbered folders sort by workflow stage; underscore folders are infrastructure and excluded from the wiki index. It's a standard Obsidian vault — everything works with any text editor.

FolderHolds
00-inbox/Unsorted captures awaiting triage
01-projects/Active projects, tasks inline in the note
02-domains/Life areas — businesses, health, properties
03-entities/People, organizations, systems, vendors
04-research/Deep dives and articles
05-ideas/Ideas and brainstorms
06-plans/Goals and roadmaps
07-reference/How-tos, cheatsheets, generated indexes
_sources/Raw imports (immutable, not embedded as wiki)
_templates/One template per note type (10 total)
_meta/Dashboards, triage reports, the vector index

Every note declares itself

Frontmatter carries type, lifecycle, and — critically for an agent-maintained vault — provenance: who wrote it and whether a human has verified it.

type:         project | entity | decision | …
status:       active | someday | done | archived
stage:        captured | normalized | reviewed
authorship:   human | imported | llm-assisted
verification: unreviewed | reviewed | trusted
updated_by:   human | agent
source:       evernote | manual | document | …

[[wikilinks]] form the graph: every note links at least one other, and any note that mentions an entity links its entity page.

Embeddings

How Markdown becomes searchable

Three index builders share one embedding client and one 4096-dimension vector space, so their outputs stack into a single search at query time. Vectors are plain .npy arrays with JSON sidecars — derived data, gitignored, rebuildable from the notes at any time.

wiki

One vector per note — no chunking. Frontmatter stripped, title prepended, text capped at 8,000 chars. Curated notes are short enough that whole-note vectors beat fragments.

sources

One vector per triaged document, built from the classifier's output: name + doc type + date + entities + a ≤25-word summary. The summary is embedded, not the raw file.

mail

Real chunking: 6,000-char windows with 200-char overlap over bodies and attachment text. Per-account shards, incremental (already-embedded chunks skipped), atomic writes with crash reconciliation.

All vectors are L2-normalized on save, so cosine similarity is a single dot product across the stacked matrix. The embedding model runs on a GPU host on the private tailnet, with a llama.cpp fallback on-device — a health probe picks the backend per process, so search keeps working when the GPU box is off.

Query lifecycle

What one search does

One CLI command searches everything: vault_index.py search "…". Scopes are selectable with --scope wiki|sources|mail|all.

Load scopes
vectors.npy + sources.npy + mail shards → one matrix

Wiki, sources, and every per-account mail shard are memory-mapped and stacked. They share the same normalized 4096-d space, so they compare directly.

Embed the query
qwen3-embed-8b · 4096-d · L2-normalized

The query becomes a vector via the same model that embedded the corpus.

GPU host unreachable? Falls back to local llama.cpp automatically.
Cosine sweep
sims = vecs @ q → top 20 candidates

One dot product against the whole corpus. No approximate index needed at this scale — exact search stays instant.

Rerank
bge-rerank-v2-m3 scores the 20-candidate pool

A cross-encoder reads query and candidate text together — much sharper relevance than cosine alone. If reranking fails, results degrade gracefully to raw cosine order.

Top-k results
default k=5, labeled by origin

Each hit is tagged with its kind and disposition — wiki, doc/contract·keep ⚠️financial, mail/2009-03-08 — so sensitivity is visible before anything is opened.

# everything, by meaning
vault_index.py search "the server consolidation decision"

# just documents, more results
vault_index.py search "insurance policies" -k 8 --scope sources

# rebuild after editing notes
vault_index.py build && vault_index.py build-sources
mail_index.py build --all        # nightly via launchd
Ingest

How content gets in

Document triage

A local 35B model (MLX, on-device) walks document trees read-only: extracts text (native tools, PDF parsing, Vision OCR), then classifies each file — type, value, entities, and a sensitivity flag (PII / financial / medical / credentials). Output is a resumable JSONL manifest; the summaries become the sources index. Originals are never touched.

Mail pipeline

IMAP accounts mirror to local storage nightly; messages become permanent JSON text extracts (fast native pass, then OCR for scans), which the mail indexer chunks and embeds into per-account shards. Email search spans decades without a mail client.

Agent-mediated writes

Agents draft; humans approve. Frontmatter records authorship and verification on every note, and outbound actions (like email replies) go through a render-review-approve loop before anything sends.

Nightly maintenance

Scheduled jobs re-mirror mail, digest new dev projects, rescan cloud drives, and rebuild the semantic index — so search reflects yesterday by breakfast.

Principles

Why it's built this way

Local-first, tailnet-only

Embedding, reranking, classification, and OCR all run on machines the owner controls. Private documents and email never reach a third-party API.

Plain files outlive software

Markdown + YAML + numpy. Every layer is inspectable with a text editor, versioned with git, and rebuildable — the vector index is derived data, never the source of truth.

Non-destructive by default

Ingest reads; it never moves or edits originals. Raw sources are immutable. Anything derived can be deleted and regenerated.

No secrets in the vault

Notes carry credentials_ref pointers into a password manager, never the credentials themselves — so the whole vault can be embedded, searched, and shared with an agent safely.