Choose a path · Quick start · How it works · Documentation · Русский
AI assistants are excellent at the current conversation and unreliable at remembering the work around it. AI Knowledge Engine is a set of instructions, templates, and reference scripts that gives a Markdown-capable coding agent a durable local memory.
It can start as a compact codebase index or grow into a full knowledge pipeline with document conversion, provenance, NLP enrichment, review queues, health checks, and deliberate AI-assisted reflection. Your files stay in your project; there is no hosted service or database to operate.
The visual above is the actual storage model, not a conceptual cloud architecture:
raw/,processed/,review/, andinteractions/stay out of the generated AI index.knowledge/contains reviewed, reusable Markdown and is safe to index.assets-index/holds searchable descriptions while original binaries remain private..repomix/output.xmlis a rebuildable context artifact, not the source of truth.
| Lite — codebase context | Full — knowledge engine | |
|---|---|---|
| Start here | quick-start/ |
knowledge-base/ |
| Best for | Giving an agent a reliable project map | Building a long-lived personal or team knowledge system |
| You get | Stack-aware indexing, secret checks, token controls, optional Git hook | Raw-first ingest, NLP, provenance, review queues, reflection, linting, watchers |
| Typical setup | About 5 minutes | About 30 minutes |
| Runtime | Repomix + Node.js | Python pipeline + an indexer; Repomix is the included default |
If you only need better code context, start with Lite. Choose Full when decisions, source material, recurring work, or cross-session learning should become durable knowledge.
Copy the deployment kit into the project where the knowledge base should live:
git clone https://github.com/bionicle12/ai-knowledge-engine.git
cp -R ai-knowledge-engine/knowledge-base /path/to/your-project/setup
cd /path/to/your-projectThen send this to your coding agent:
Read setup/README.md and setup/00_OVERVIEW.md, then deploy a
knowledge base for [your role] inside ./knowledge-base/.
When kb_doctor passes, run setup/shell/finalize.sh to flatten
the base into the project root.
The agent asks about your role, configures the pipeline, creates a role-specific knowledge structure, verifies it with kb_doctor.py, and promotes the finished system to the project root.
PowerShell copy command
git clone https://github.com/bionicle12/ai-knowledge-engine.git
Copy-Item -Recurse ai-knowledge-engine\knowledge-base C:\path\to\your-project\setup
Set-Location C:\path\to\your-projectImportant
In every new agent session, begin with: “Read AGENTS.md and use it as the primary instruction for everything that follows.” The knowledge base is local; the agent must be told where its operating instructions live.
npm install -g repomix
cp -R ai-knowledge-engine/quick-start /path/to/your-project/docs/ai-initThen tell your agent:
Read docs/ai-init/INIT_GUIDE.md and set up context indexing for this project.
The guide has the agent inspect the stack, define safe include/exclude patterns, enable Repomix security checks, and generate the first context snapshot.
The project ships two coordinated layers:
- Instruction modules tell the agent what to ask, which choices matter, and how the system should behave.
- Reference implementations handle deterministic work such as ingest, hashing, linting, scheduling, population, and upgrades.
The agent adapts the deployment to your role and project. It does not improvise core pipeline code from scratch.
Original material is preserved before interpretation. Conversion, NLP metadata, provenance, and review happen before anything becomes trusted knowledge. Complex or sensitive material can be routed to review instead of being silently indexed.
The baseline pipeline runs on local files and CPU. It does not require a separate SaaS account, database, or pipeline API key. Optional AI operations use the coding agent you already brought to the project.
Repomix works out of the box, but indexing is an adapter boundary. The durable layer is ordinary Markdown plus explicit metadata, links, and provenance.
source material
↓
raw/ immutable originals, never indexed
↓ kb_ingest.py
processed/ Markdown + extraction metadata, never indexed
↓ routing and review
knowledge/ clean linked notes, indexed
assets-index/ searchable descriptions of binary assets
↓ context indexer
.repomix/output.xml rebuildable AI context
The pipeline adds five controls around that path:
- Provenance — hashes and source references trace knowledge back to its origin.
- Review queues — ambiguity, low-confidence conversion, and sensitive material stop for a decision.
- Health checks — linting finds stale pages, broken links, orphans, and structural drift.
- Reflection — accumulated facts can be synthesized into higher-level insights on a schedule or by command.
- Upgrades —
kb_upgrade.pyrefreshes reference files while preserving user customizations when possible.
Read the contributor-level architecture overview for the full deployment sequence and invariants.
| Capability | What happens locally |
|---|---|
| Documents | PDF, DOCX, PPTX, spreadsheets, and text are converted to Markdown |
| Audio and video | Optional faster-whisper transcription; no system FFmpeg required for the default backend |
| Images and scans | Optional rapidocr-onnxruntime OCR; Tesseract remains an alternative |
| Archives | ZIP and TAR inputs are unpacked with a size safety cap and re-ingested |
| NLP | Entities, keywords, and routing hints are extracted before AI review |
| Knowledge graph | Markdown pages use [[wikilinks]], routing tables, lifecycle metadata, and provenance |
| Automation | Cross-platform reindexing, watchers, health checks, and importance-based reflection triggers |
| Portability | The complete system is files and folders that can be copied, synced, versioned, or inspected directly |
| Multi-machine | Two deployments of one base exchange knowledge through bundles, with content-hash dedup, fast-forward, and conflicts routed to review |
These commands are messages to your AI agent, not shell commands.
| Command | Purpose | Typical AI cost |
|---|---|---|
!view |
Open the local read-only knowledge graph viewer | 0 tokens |
!save |
Capture decisions and insights from a productive session | ~2K tokens |
!reflect |
Synthesize higher-level patterns from accumulated knowledge | ~15K tokens |
!review |
Resolve items waiting in review queues | ~5–30K tokens |
!audit |
Deep review for contradictions, gaps, and merge candidates | ~50–100K tokens |
!populate |
Regenerate role-specific placement examples | ~50 tokens |
!export |
Pack this base into a bundle for another machine | ~100 tokens |
!import |
Merge bundles waiting in sync/inbox/ |
~200 tokens |
!merge |
Finish an import: resolve conflicts, audit contradictions, cross-link | ~5–40K tokens |
!super on/off |
Switch between Python-first and AI-first operation | 0 tokens |
| Mode | Behavior | Expected active usage |
|---|---|---|
default |
Python-first processing with throttled AI decisions | ~3–4K tokens/day |
super |
AI reasoning for surprise detection, annotations, entity resolution, and frequent review | ~50–200K+ tokens/day |
Token figures are planning estimates, not benchmarks. Actual usage depends on document size, activity, model, and agent behavior.
Two deployments of the same base drift apart. One laptop accumulates knowledge
about your tools and gear; the other accumulates analysis notes; both edited the
same page last week. Copying folders around loses one side's work, and rsync
cannot tell an edit from a stale copy. So the engine ships an explicit merge
path: export a bundle, import it on the other machine, let the agent settle
what only a human-level reader can settle. Full contract in
16_MERGE.md.
machine A machine B
┌──────────────┐ ┌──────────────┐
│ knowledge/ │ │ knowledge/ │
└──────┬───────┘ └──────▲───────┘
│ export │ import
▼ │
sync/outbox/bundle.zip ───── copy ─────► sync/inbox/bundle.zip
│
├─ safe cases → applied automatically
└─ ambiguous → review/needs-merge/
│
!merge (agent)
Each machine needs its own name. Open kb.config.yml on both and set a
different sync.label — it is stamped onto every page that travels, which
is what makes a merge traceable later:
sync:
label: "studio-laptop" # on the other machine: "work-laptop"If your bases were deployed before this feature existed, pull it in with the
upgrader (see Updating deployed knowledge bases) —
it adds the scripts, the launchers, the sync/ folder and the config section,
defaulting label to the base's folder name.
| OS | How |
|---|---|
| macOS | Double-click export.command |
| Windows | Double-click export.bat |
| Linux / any terminal | ./shell/export.sh |
| In the AI chat | !export |
A file appears in sync/outbox/:
kb-bundle__studio-laptop__2026-07-31.zip
Useful flags (append them in the terminal, or ask the agent):
./shell/export.sh --since 2026-06-01 # only knowledge touched since a date
./shell/export.sh --with-assets # include the binary originals too
./shell/export.sh --only knowledge # knowledge pages and nothing else
./shell/export.sh --dry-run # show what would be packedCopy that one .zip into the other machine's sync/inbox/ folder — USB stick,
cloud folder, scp, anything. Nothing here touches the network on its own.
| OS | How |
|---|---|
| macOS | Double-click import.command |
| Windows | Double-click import.bat |
| Linux / any terminal | ./shell/import.sh |
| In the AI chat | !import |
Every bundle sitting in sync/inbox/ is processed. You get a summary in the
terminal and a full report in sync/reports/.
Read AGENTS.md and use it as the primary instruction for everything that follows.
!merge
The agent resolves each queued conflict by folding the two versions together —
keeping every fact that is true in either — records genuine contradictions in
knowledge/open-questions/ with both sources instead of silently picking a
winner, cross-links the imported pages into the existing graph, refreshes
routing, lints and reindexes.
Run the same four steps in the other direction and both machines hold the union of the knowledge.
It only does what is provably safe and leaves judgement to the agent:
| Situation | What happens |
|---|---|
| Page exists only in the bundle | Added, stamped with merged_from: provenance |
| Same content, same place | Skipped |
| Same content, richer metadata | Tags, counters and missing fields merged; body untouched |
| Same content under a different name | Skipped, reported as a duplicate |
| Local copy untouched since the last import | Fast-forwarded, previous version backed up |
| Changed on both machines | Local file untouched; incoming version + diff queued in review/needs-merge/ |
| Different name, ~85%+ overlap | Added, plus a merge-candidate note for consolidation |
Two pages count as the same knowledge when their bodies match; frontmatter is bookkeeping and merges instead of conflicting, so simply reading a page on one machine never looks like an edit on the other. Re-importing the same bundle changes nothing.
- Nothing is overwritten silently. Every file the import touches is copied to
sync/backups/<timestamp>/first. - Bundles carry knowledge, not raw material.
raw/,processed/,review/andsync/never travel. Binary assets are opt-in via--with-assets; without them the importedassets-index/entries are annotated as "original file not present in this base" rather than left dangling. sync/is excluded from the AI index and from git.--dry-runclassifies everything and writes nothing.- Sharing the base with someone else? Export with
--only knowledgeand read the result first —interactions/holds session history and the provenance metadata holds original filenames.
| Symptom | What it means |
|---|---|
| Import printed "N conflicts waiting" | Normal. Say !merge in the AI chat. |
Import exits with code 1 |
Same thing — 1 means "merged, conflicts queued", not a failure. |
| "No bundle to import" | The .zip is not in sync/inbox/ on this machine. |
| Both machines produce identically named bundles | Their sync.label is the same. Give each its own. |
| Want to check before committing to it | Run the import with --dry-run. |
Always run the upgrader from your current ai-knowledge-engine checkout. Start
with --dry-run; remove it only after reviewing the plan.
python C:\path\to\ai-knowledge-engine\scripts\kb_upgrade.py `
--kb-root C:\path\to\kb-name --dry-run
python C:\path\to\ai-knowledge-engine\scripts\kb_upgrade.py `
--kb-root C:\path\to\kb-namepython3 ~/path/to/ai-knowledge-engine/scripts/kb_upgrade.py \
--kb-root ~/path/to/kb-name --dry-run
python3 ~/path/to/ai-knowledge-engine/scripts/kb_upgrade.py \
--kb-root ~/path/to/kb-namepython3 /path/to/ai-knowledge-engine/scripts/kb_upgrade.py \
--kb-root /path/to/kb-name --dry-run
python3 /path/to/ai-knowledge-engine/scripts/kb_upgrade.py \
--kb-root /path/to/kb-nameThe first upgrade installs scripts/kb_update.py into the KB. After that,
run python scripts/kb_update.py --dry-run (Windows) or
python3 scripts/kb_update.py --dry-run (macOS/Linux) from the KB root. The
launcher finds the source repo automatically; use --repo-root PATH or
AI_KNOWLEDGE_ENGINE_HOME when it lives elsewhere.
For a directory of bases, use
--all-root /path/to/parent (immediate kb-* children). If a file is truly
safe to replace, prefer repeatable --accept FILE over global --force.
See the upgrading guide for examples and safety rules.
Full mode includes 15 starting configurations in knowledge-base/examples/:
| Template | Role | Focus |
|---|---|---|
b2b-strategic-product-owner.yml |
B2B Strategic Product Owner | SaaS strategy, roadmaps, risks, sales-ready PRDs |
battle-rap-producer.yml |
Battle rap producer & lyricist | Lyric craft, punchlines, vocal stacks, mixing, battle prep |
content-creator.yml |
Content creator | Voice, audience, formats, publishing, monetization |
creative-hybrid.yml |
Creative Hybrid | Software, music production, indie game development |
fiction-writer.yml |
Fiction writer | Craft theory, voice studies, story development, draft critique |
founder.yml |
Startup founder | Product, investors, hiring, decisions, company building |
marketing-director.yml |
Marketing Director | Strategy, brand, campaigns, audience analysis |
music-video-director.yml |
Music video writer-director | Treatments, production, shot planning, edit rhythm |
product-manager.yml |
Product Manager | Prioritization, metrics, research, product requirements |
programmer-senior.yml |
Senior Software Engineer | Architecture, debugging, stack knowledge, engineering principles |
psychologist-gestalt.yml |
Gestalt-oriented psychologist | Ethics, anonymized cases, supervision, interventions |
researcher.yml |
Researcher / Analyst | Literature, hypotheses, evidence, methodology |
russian-software-engineering-student.yml |
Software engineering student in Russia | Coursework, labs, exams, internships |
startup-opportunity-explorer.yml |
Startup Opportunity Explorer | Market gaps, validation, idea scoring, web-app MVPs |
viral-short-form-veo.yml |
Viral short-form video director | Hooks, storyboards, Veo 3 prompt chains, ad learning |
A role file defines useful entities, folder routes, placement examples, recurring queries, and review priorities. If none fits, the deployment agent can derive a custom configuration from your work.
ai-knowledge-engine/
├── quick-start/ Lite mode: codebase indexing guide
├── knowledge-base/
│ ├── 00_…16_*.md 17 ordered instruction modules
│ ├── scripts/ Python reference pipeline and tests
│ ├── shell/ macOS/Linux wrappers + Windows launchers
│ ├── templates/ config, agent, structure, and dependency templates
│ └── examples/ 15 role blueprints
├── scripts/ upgrades, translation checks, repository maintenance
├── i18n/ru/ Russian documentation and instruction set
└── docs/ architecture, troubleshooting, upgrades, roadmap
| Component | Minimum | Needed for |
|---|---|---|
| Python | 3.11+ | Full-mode ingest, NLP, linting, watchers, and health checks |
| Node.js | 20+ | Included Repomix indexer |
| Git | Any recent version | Versioning and optional hooks |
| AI coding agent | Markdown + file and shell access | Guided deployment and AI-assisted commands |
Core Python dependencies are listed in requirements.txt. Speech-to-text and OCR are deliberately separated into requirements-media.txt.
The instructions are designed for Claude, Codex, GPT, Gemini, Cursor, and other Markdown-capable agents that can inspect files and run commands. Exact behavior still depends on the host application's permissions and tool access.
- This repository is a deployment template, not a hosted application or sync service.
- The agent does not remember the knowledge base automatically; each new session must load
AGENTS.md. - Local-first storage protects against an external service by default, but repository permissions, backups, and model-provider data policies remain your responsibility.
- Repomix security checks reduce accidental secret inclusion; they are not a replacement for secret management or a security audit.
- AI review and reflection can be expensive in
supermode and should be enabled deliberately.
| Need | Read |
|---|---|
| Understand the full deployment | knowledge-base/README.md |
| See system architecture | docs/ARCHITECTURE.md |
| Fix a failed setup | docs/TROUBLESHOOTING.md |
| Upgrade a deployed base | docs/UPGRADING.md |
| Merge two machines' bases | knowledge-base/16_MERGE.md |
| Maintain or contribute | docs/MAINTENANCE.md and docs/CONTRIBUTING.md |
| Translate the project | docs/TRANSLATING.md |
| Review planned work | docs/ROADMAP.md |
| Read in Russian | i18n/ru/README.md |
Install the test dependencies and run the suite locally:
python -m pip install -r knowledge-base/templates/requirements-dev.txt
python -m pytestOn Windows, tests for POSIX launchers require a working WSL environment, and the finalize-script fixtures require permission to create symbolic links. The Python pipeline tests run natively.
Contributions are welcome, especially new role blueprints, clearer instruction modules, cross-platform fixes, translations, and pipeline tests. For substantial changes, open an issue first and follow the contribution guide.
MIT — use it for personal or commercial work.
Build knowledge your agent can find, verify, and carry forward.