CODE.md format guide

What is CODE.md?

CODE.md is a single, structured markdown file that gives AI coding assistants a parser-generated map of your repository: languages, structure, entry points, callgraphs, UI flows, TODOs, and known gaps, all extracted from the actual code rather than guessed. Drop it into your repo and every AI agent that reads it starts from truth, not exploration.

CODE.md gives AI assistants a reliable, parser-generated understanding of the codebase, helping them answer with more accuracy, reduce hallucinations, and save developers time by avoiding repeated repo exploration.

Where CODE.md fits

CODE.md is one of three core agent files. AGENTS.md tells an agent how to behave, CLAUDE.md gives it the right context, and CODE.md — autogenerated by CodeMD.dev — explains the codebase itself.

Diagram of the three agent core files: AGENTS.md tells AI agents how to behave, CLAUDE.md gives Claude the right context, and CODE.md explains the codebase to the AI, autogenerated by codemd.dev, together producing smarter AI output.

AGENTS.md = the rules · CLAUDE.md = the briefing · CODE.md = the source map

Reference @CODE.md from your AGENTS.md (or CLAUDE.md) file so every agent session loads it automatically instead of re-discovering the repo from scratch.

The CODE.md template

Below is the full CODE.md structure. Each section is a plain markdown heading with a small block underneath it, and each one exists to answer a specific question an AI agent would otherwise have to guess at or spend tokens discovering by reading source files directly.

Overview

Why this helps LLMs: LLMs perform better when they start with a high-level mental model. This reduces misinterpretation and prevents the model from wasting tokens trying to infer the project's purpose.

## Overview
This repository contains the source code for <project_name>.
Its primary purpose is <brief_description>.
Key capabilities include <capabilities>.

Evidence Policy

Why this helps LLMs: LLMs hallucinate when they assume missing context. This section tells the model exactly what was extracted, what was excluded, and what should not be inferred.

## Evidence Policy
- Scope of analysis: <folder or repo scope>
- Artifact root: <path>
- Only direct extraction artifacts used: <true/false>
- LLM-generated content included: <true/false>
- Excluded artifacts: <list>
- Notes: <clarifications about missing semantics or intent>

System Summary

Why this helps LLMs: Knowing the languages, file counts, and scale helps the model reason faster and avoid incorrect assumptions about architecture.

## System Summary
- Primary languages: <Python %, JS %, etc.>
- Total source files: <count>
- Total lines of code: <count>
- Description: <short system description>

Repository Structure

Why this helps LLMs: LLMs waste tokens scanning file trees. A structured summary lets them jump directly to relevant areas.

## Repository Structure
- Folders:
  - /src — core application logic
  - /features — feature detection and analysis
  - /static — UI assets
  - /tests — automated test suite

- Sample files:
  - /src/main.py — entry point
  - /features/detector.py — feature extraction logic
  - /static/dashboard.html — main UI

Modules & Responsibilities

Why this helps LLMs: LLMs often struggle with modular boundaries. This section prevents confusion and improves reasoning about dependencies.

## Modules
- Module: <name>
  - Responsibility: <description>
  - Allowed imports: <list>
  - Forbidden imports: <list>

Callgraph Summary

Why this helps LLMs: Instead of parsing thousands of lines of code, the model gets a pre-computed flow of how functions interact. This dramatically reduces token usage.

## Callgraph
- Node count: <count>
- Edge count: <count>
- Entry points:
  - main.search
  - main.analyze_repo
  - api.search
- Top connected nodes:
  - <function> — <degree>

Filegraph Summary

Why this helps LLMs: Shows architectural hotspots and dependency clusters without requiring the model to scan every file.

## Filegraph
- Core files:
  - main.py — central orchestrator
  - static/autotrack.js — analytics tracking

- Example edges:
  - main.py → features/helpers.py
  - module/static → static/autotrack.js

UI Graph

Why this helps LLMs: LLMs can trace UI → JS → API flows without manually parsing HTML or JavaScript.

## UI Graph
- Node count: <count>
- Edge count: <count>
- Example interactions:
  - dashboard.html.button_1589 → js.exportDashboardPdf
  - dashboard.html.analyzeRepoButton → js.analyzeRepo

Source Inventory

Why this helps LLMs: LLMs can quickly locate TODOs, missing logic, and areas needing improvement.

## Source Inventory
- Function count: <count>
- Comment count: <count>
- TODOs:
  - main.py:148 — Fix callgraph rendering
  - scim.py:2744 — Show neighbors only on click

Behavior & Constraints

Why this helps LLMs: Prevents incorrect assumptions about how the system behaves.

## Behavior
- Known invariants: <list>
- Constraints: <list>
- Error rules: <list>

Drift Analysis

Why this helps LLMs: Helps the model understand whether older documentation may be outdated.

## Drift
- Structure drift: <unknown/low/high>
- Semantic drift: <unknown/low/high>
- Timeline: <list>

Validation

Why this helps LLMs: Improves reasoning about correctness and expected behavior.

## Validation
- Critical flows: <list>
- Invariants: <list>
- Browser tests: <list>

Autoheal / Self-Healing

Why this helps LLMs: Helps the model understand how the system repairs itself or handles errors.

## Autoheal
- Fix patterns: <list>
- Safe rules: <list>
- Risky areas: <list>
- Validation steps: <list>

Why This Template Helps LLMs (and Saves Tokens)

Why this helps LLMs: This section explains the purpose of the entire CODE.md file.

## Why This Template Helps LLMs
- Gives the LLM a structured map of the repository.
- Reduces the need to scan thousands of lines of code.
- Prevents hallucinations by clarifying what is known vs unknown.
- Cuts token usage by providing pre-computed metadata.
- Improves accuracy, speed, and reliability of code reasoning.

Fewer tokens, fewer hallucinations, less of your time

Every section above exists to answer a question an agent would otherwise burn tokens exploring the repo to find out. That has a direct, compounding effect on cost and reliability.

Lower token usage Pre-computed callgraphs, filegraphs, and structure summaries replace thousands of lines of source the model would otherwise have to read and re-read across turns.
Fewer hallucinations The Evidence Policy and Drift sections tell the model exactly what is known, what is excluded, and what should not be inferred, so it stops guessing at missing context.
Less developer time lost code.md gives AI assistants a reliable, parser-generated understanding of the codebase, helping them answer with more accuracy, reduce hallucinations, and save developers time by avoiding repeated repo exploration.

Generate CODE.md for your repo

CodeMD.dev runs this exact template against your repository's real structure, callgraph, and UI graph, so you don't have to fill in the template by hand.

Generate CODE.md now