Language LLMs are math. Markdown files are math. LLMs are expensive. Markdown files are cheap. Let's all be hippies and live in the markdown LLM world.
A reference architecture for compositional generation from markdown files instead of model weights. Zero LLM at inference time. No GPU. No tokenizer. No embeddings. The graph is the model; traversal is the forward pass; markdown is the substrate. The output is whatever the renderer emits — text, audio, music, video, anything a computer can serialize.
This repository contains:
runner/— a ~700-line Python runtime that loads markdown nodes, walks them by pattern matching, and composes responses by template substitution. Imports nothing from any LLM library.orchestration/— markdown files that describe processes the runtime executes (the "score" the conductor plays).corpus/— atomized markdown nodes with YAML frontmatter and[[wikilink]]edges. The example corpus is self-documenting: it explains this framework.tests/— six lie detectors that enforce the architectural contract through static analysis and behavioral checks.frontend/index.html— a single-page UI that lets you call any orchestration, see the output, and click through every cited source file.
pip install fastapi uvicorn pydantic pyyaml httpx pytest
# Verify the runtime is honest — no LLM imports, no embeddings, no model subprocesses
python3 -m pytest tests/ -v
# Serve the example
python3 runner/server.py
# → http://localhost:8042Open the browser, ask "what is provenance?" — get an answer in ~8 ms, with the cited corpus files in the right panel. Click any file to read its raw markdown. The framework documents itself.
- No LLM imports in
runner/*.py(test_no_llm_imports.py) - No embedding / similarity calls in
runner/*.py(test_no_embedding_calls.py) - No shelling out to model binaries or API endpoints in
runner/*.py(test_no_model_subprocess.py) - Every response cites real, existing corpus files (
test_traceability.py) - Deleting a corpus file visibly changes behavior (
test_modularity.py) - The HTTP API serves the inference endpoint and returns provenance (
test_endpoint.py)
If any of these fail, the architecture is compromised. Fresh implementers cannot rationalize past them.
A Markdown Language Model is a compositional generation system. Markdown nodes hold the building blocks; orchestrations describe the process for combining them; a dumb runtime walks the graph and composes the output. The same architectural shape as a transformer — input → frozen graph → output — but on a substrate humans can read, edit, version-control, and reason about.
The medium is whatever you write a composer and renderer for.
- Text (what this example ships with): nodes hold prose chunks. Composer concatenates with templates. Renderer writes the result.
- Audio: nodes hold waveform fragments, envelopes, instrument samples. Composer mixes by time and channel. Renderer encodes WAV/MP3.
- Music: nodes hold motifs, chord progressions, song structures. Composer sequences them under a key/tempo/meter. Renderer writes MIDI.
- Video: nodes hold frame specs, transitions, audio refs. Composer interleaves. Renderer encodes MP4.
- 3D / geometry / animation / image generation / anything else a computer can produce: nodes hold the medium's atomic units; composer/renderer know how to combine and serialize.
The runtime doesn't care what the nodes hold. It's a graph walker. The corpus is the model regardless of medium. The output is whatever the renderer emits.
Common misconceptions worth addressing:
- "This only works for structured knowledge." No — it works for any domain you can articulate in writing as a graph of typed units. Knowledge is one substrate; sound is another; geometry is another.
- "Open-ended creative generation requires neural sampling." No — a combinatorial walk over a corpus of motifs/phrases/structures with constraint rules produces novel output by composition. Explicit instead of opaque. Predates LLMs by decades (algorithmic composition, generative music, procedural content).
- "Fluid conversation across arbitrary topics requires an LLM." No — session-aware multi-classifier routing with graceful "I don't know" nodes and an optional speech organ for surface polish handles it.
- "Reasoning under genuine ambiguity requires probabilistic interpolation." No — multi-classify against alternate interpretations and return both branches, or ask to disambiguate. The branches are explicit instead of silently averaged into one probable answer.
These are orchestration design problems, not substrate limitations. The architecture's reach is bounded by the corpus author and the composer/renderer for the chosen medium — not by anything intrinsic.
LLMs are still permitted at build time — extract knowledge from sources, generate atomic nodes, draft templates. The constraint is at inference time only. See corpus/concepts/build_time_vs_inference_time.md.
- Replace the corpus. Throw away
corpus/concepts/*.mdandcorpus/comparisons/*.md. Write your own nodes — one.mdper atomic concept in your domain. Match the frontmatter schema incorpus/examples/example_corpus_node.md. - Write orchestrations. One
.mdper endpoint. Use the actions documented incorpus/patterns/. Reference shape incorpus/examples/example_orchestration.md. - Don't change the runtime.
runner/*.pyis generic. Adding domain logic to the runtime means the orchestration DSL isn't powerful enough yet — fix the DSL instead. - Run the tests. All six must pass. If they fail, you've drifted off the architecture.
MarkdownLanguageModel/
├── README.md # this file
├── runner/ # generic runtime — do not modify
│ ├── engine.py # orchestration executor
│ ├── parser.py # markdown frontmatter + wikilink + steps parsing
│ ├── traversal.py # keyword-overlap node scoring
│ ├── composer.py # template substitution
│ ├── templater.py # {var.path} placeholder rendering
│ ├── provenance.py # cited-file tracking
│ └── server.py # FastAPI HTTP server
├── orchestration/ # processes the runtime executes
│ └── answer.md # example: answer a question
├── corpus/ # the knowledge graph (markdown nodes)
│ ├── identity/ # voice, tone, persona
│ ├── concepts/ # framework concepts (self-documenting)
│ ├── patterns/ # orchestration DSL action reference
│ ├── comparisons/ # MLM vs RAG, fine-tuning, transformer
│ └── examples/ # corpus node + orchestration templates
├── frontend/
│ └── index.html # single-page UI with provenance viewer
└── tests/ # six lie detectors enforcing the contract
├── test_no_llm_imports.py
├── test_no_embedding_calls.py
├── test_no_model_subprocess.py
├── test_traceability.py
├── test_modularity.py
└── test_endpoint.py
To prove a point most of the AI industry has incentive not to articulate: that for a large class of "AI" deployments, the underlying knowledge is structured enough to live in human-readable files. The inference can then be deterministic and instant, and the result 100% auditable. No 70B model. No GPU. No hallucination.
The architecture isn't new. Symbolic AI tried it in the 1980s and failed because there were no LLMs to populate the corpus. What's new is the synthesis: use LLMs once at build time to extract and atomize, then serve forever from the markdown. Build-time LLM, inference-time zero. That's the framework.
Apache License 2.0 — see LICENSE.