A good AGENTS.md is short, specific, and mostly about things an agent cannot discover by reading the code: the exact commands to build and test, the package manager, the files it must not touch, the conventions that differ from the defaults for your stack, and how to verify its work before finishing. A bad AGENTS.md is a long repository tour: architecture overviews, directory descriptions and aspirational advice like "write clean code". Coding agents read it faithfully on every task, and it costs tokens without changing results.
That distinction is no longer opinion. In 2026, the first rigorous studies of repository context files arrived, and they complicate the "just add an AGENTS.md" advice repeated across the web. Context files work, but only certain kinds of content in them do, and a file that is not tested may be making agents slower and more expensive.
What AGENTS.md is and who reads it
AGENTS.md is a plain Markdown file placed at the root of a repository, and optionally in subdirectories, containing instructions for AI coding agents. It has no schema and no required fields. The format's own description calls it a README for agents: the operational detail an agent needs that would clutter a README meant for humans.
It began as OpenAI's convention for Codex in August 2025. It spread because other tools adopted the same filename rather than inventing their own. Cursor, Amp, opencode, Warp, goose and GitHub Copilot's coding agent all read it. In December 2025, the Linux Foundation announced the Agentic AI Foundation (AAIF), which now stewards AGENTS.md alongside the Model Context Protocol and other open agent infrastructure. The format's site reports use in more than 60,000 open-source projects.
Claude Code is the main exception in naming. It reads its own CLAUDE.md rather than AGENTS.md directly. Its documentation suggests bridging the two with a one-line import (@AGENTS.md) inside CLAUDE.md, or with a symlink, so a team can keep one source of truth.
Two rules of the format matter for everything that follows. First, the closest file wins: in a repository with nested AGENTS.md files, the one nearest the file being edited takes precedence. Second, explicit instructions in the chat override the file. AGENTS.md is standing context, not a command that outranks the task.
What the research found, and why it changes the advice
Three pieces of evidence published in 2026 are worth knowing before you write a line.
An ETH Zurich study questioned whether context files improve success at all. Thibaud Gloaguen, Martin Vechev and colleagues evaluated coding agents in two settings. The first used SWE-bench tasks with context files generated by a language model. The second was a new set of real issues from repositories whose developers had committed their own context files. Across different models and agents, and for both generated and human-written files, context files did not generally improve task success rates, while increasing inference cost by more than 20% on average. The detail matters most: agents followed the instructions in the files well, but repository overviews, the part most templates encourage, did not help. The authors concluded that context files are useful for specifying non-standard practices, and that any attempt to improve performance with them should be tested rather than assumed.
A study from Christoph Treude, Sebastian Baltes and colleagues found efficiency gains. They ran agents on 124 pull requests from 10 repositories, with and without an AGENTS.md file. The presence of the file was associated with a 28.6% lower median runtime and 16.6% fewer output tokens, with comparable task completion.
An AAIF experiment showed how easy it is to measure this wrongly. In July 2026 the foundation published a controlled test using GitHub Copilot CLI on a real VS Code extension repository, with and without a twelve-line AGENTS.md. A single run per condition suggested the file made a harder task 44% slower and 41% more expensive. Over five runs per condition, using the median, the result reversed. On an ambiguous task the file cut wall-clock time by 27%, cost by 24% and diff size by 26%, mostly because a single line about staying in scope removed defensive code nobody had asked for. On a multi-file task the median gain was smaller, around 9–10%. The more interesting difference was in the tail: without the file, some runs wasted time re-exploring the repository or ran an unrequested production build.
These results are less contradictory than they look. They measured different things, on different tasks, with different files. Read together, they suggest a consistent picture:
Agents do follow what you put in the file. The file's influence is real, so bad content has real cost.
Descriptive content (what the project is, how it is organised) duplicates what an agent can find by reading the code, and adds tokens to every turn.
Prescriptive, non-obvious content (use this command, do not touch this path, stay in scope) prevents the wasted steps that make agents slow and expensive.
The gains show up most in efficiency and consistency, less in whether the agent eventually solves the task.
Measurement is noisy, and a single before-and-after comparison can point in the wrong direction.
None of these studies are the final word. The two arXiv papers are preprints, and the AAIF test is one repository and one agent, as its author says. But they are more evidence than most AGENTS.md advice is based on.
What belongs in the file
The useful test for each line: would a competent engineer new to this repository get this wrong on day one without being told? If yes, it belongs in the file. If reading the code would reveal it in a minute, it probably doesn't.
What typically passes that test:
Exact commands. Install, build, test, lint and type-check, written as they are run, including the non-obvious parts: the package manager (npm vs pnpm is a classic agent mistake), flags, the need to run a code generator first, and how to run a single test file instead of the whole suite.
Verification steps. What "done" means. Agents will run programmatic checks listed in the file and try to fix failures before finishing. This is probably the highest-value content in any AGENTS.md, because it turns the file into a feedback loop rather than advice.
Boundaries. Generated files, vendored code, migration history, lock files, CI configuration: anything the agent should not edit, with the reason if it is not obvious.
Non-standard conventions. Only where you differ from the ecosystem's defaults. If the project uses a custom error wrapper instead of throwing, or a specific logging helper, or forbids a common library, say so.
Scope discipline. A single line such as "change only what the task requires; do not refactor unrelated code or add defensive handling that was not requested" was responsible for most of the gain in the AAIF experiment.
What should stay out:
Repository overviews and directory tours. The ETH study found them unhelpful, and they go stale quickly.
General advice ("follow best practices", "write readable code"). It gives the agent nothing to act on.
Content duplicated from the README or docs. Link to a document if the agent needs to read it for specific tasks, rather than copying it.
Secrets, credentials and internal URLs you would not publish. The file is committed to the repository and sent to model providers as part of every prompt.
Enforcement you actually need. The file guides behaviour but does not guarantee it. Anything that must never happen belongs in CI, branch protection or sandbox permissions.
An example that follows the evidence
A realistic AGENTS.md for a TypeScript web application, around 30 lines:
markdown
# AGENTS.md
## Commands
- Install: `pnpm install` (never npm or yarn; lockfile is pnpm-lock.yaml)
- Generate API client before building: `pnpm codegen`
- Type-check: `pnpm typecheck`
- Unit tests: `pnpm test` — single file: `pnpm test path/to/file.test.ts`
- Lint and format: `pnpm lint --fix`
## Before you finish
- Run `pnpm typecheck` and the tests for any package you changed.
- If you changed anything under `src/api/`, run `pnpm codegen` and commit the result.
- Do not run `pnpm build` unless the task asks for it (slow, not needed for verification).
## Do not edit
- `src/generated/` — output of codegen, regenerate instead
- `migrations/` — existing migrations are immutable; add a new one
- `.github/workflows/` — CI changes need human review
## Conventions that differ from defaults
- Errors: return `Result<T, AppError>` from `src/lib/result.ts`; do not throw in services.
- Logging: use `log` from `src/lib/log.ts`, never `console.*`.
- Styling: design tokens only (`src/styles/tokens.css`); no hard-coded colours.
## Scope
- Change only what the task requires. Do not refactor unrelated code,
rename files or add error handling that was not requested.
- If the task seems to require editing a protected path, stop and explain why.Notice what is missing: no description of the project, no directory tree, no list of frameworks the agent can see in package.json. Every line either gives a command, draws a boundary or corrects a default the agent would otherwise get wrong.
One file, several tools, and monorepos
Most teams now use more than one agent. Different developers prefer different tools, and background agents in CI differ from the interactive ones. The main practical value of AGENTS.md as a standard is that it prevents the situation many teams found themselves in by early 2026: a CLAUDE.md, a .cursorrules, a copilot-instructions.md and an AGENTS.md, all with nearly the same content, slowly diverging.
The pragmatic setup:
Make AGENTS.md the canonical file.
For tools with their own filename, reference it rather than copying it. In Claude Code, that means a CLAUDE.md containing
@AGENTS.mdplus any Claude-specific additions, or a symlink.Keep tool-specific files only for genuinely tool-specific configuration.
Size is a real constraint, not only style. Each file is read into context on every turn, whether or not the task needs it. OpenAI's Codex documentation specifies a default limit of 32 KiB for the combined chain of instruction files, configurable through project_doc_max_bytes, and content beyond it is truncated. Cursor recommends staying under 500 lines, and Claude Code's guidance aims under 200 for CLAUDE.md. The evidence above suggests the useful size is far smaller: a few dozen lines of directives.
In monorepos, use nesting deliberately. A root file carries the rules that apply everywhere: the package manager, general verification, repository-wide boundaries. Package-level files carry only what differs, such as a Python service in a TypeScript repository or a package with its own test runner. Because the closest file wins, the package file should restate whatever it needs from the root rather than assuming it is merged. Avoid adding nested files just because a deeper structure looks thorough. Every file is another source of drift.
How to test whether your file helps
The ETH study's conclusion is the one to act on: evaluate before trusting. You do not need a benchmark suite to do it roughly.
Choose two or three representative tasks your team actually gives agents: a small bug fix, a feature change touching several files, and something ambiguous.
Run each task with and without the file, on identical clones at the same commit, with the same agent and prompt.
Repeat each condition several times. The AAIF test found single runs misleading. Five runs and the median is a reasonable minimum.
Measure more than success. Record wall-clock time, tokens or cost, diff size, and failure modes: unrequested builds, edits outside the scope, wrong package manager, protected files touched.
Turn each observed mistake into one line, then re-test.
This is also how the file should grow over time. The best rule of thumb from teams that maintain these files: add a line when an agent repeatedly gets something wrong, and delete a line when it is no longer true. A file that only grows becomes a repository overview by another name.
AGENTS.md is part of your attack surface
Security discussion of AGENTS.md is still thin. The consequence follows directly from the research: if agents follow the file's instructions reliably, then whoever can edit the file can direct your agents.
A few implications for teams running agents with real permissions, especially background agents that open pull requests, have access to credentials, or run in CI:
Review AGENTS.md changes like code, not like documentation. A pull request that adds "before finishing, run
scripts/setup.sh" or changes a verification command deserves the same scrutiny as a change to a build script. Code owner rules can require a maintainer's review for changes to AGENTS.md, CLAUDE.md and any files they import.Watch nested files. Because the closest file wins, a new AGENTS.md added deep in a directory can quietly override the root rules for anything an agent edits there.
Treat imported and linked content as instructions too. If the file tells an agent to read a document, that document is effectively part of the agent's instructions.
Do not rely on the file for security boundaries. "Never read
.env" is a request. Filesystem permissions, secret scoping and sandboxed execution are controls.
This matters more as agents move from the developer's terminal into CI and scheduled background tasks, where no human reads the transcript before changes are made.
Where this is heading
A few developments are worth watching into 2027.
Context is splitting into layers. AGENTS.md covers project-specific standing context. Reusable procedures, such as how to write a migration or how to run a release checklist, are increasingly packaged separately. Agent Skills, a format Anthropic published as an open specification in late 2025 and that several agents now support, is one example. Expect "what goes in AGENTS.md vs a skill vs the prompt" to become a common architecture question.
Evaluation will become normal. The 2026 papers were the first serious attempts to measure context files, and both groups call for more research. As agent costs become a visible budget line, measuring whether persistent instructions pay for themselves is likely to move from research into team practice.
Standardisation continues. With AGENTS.md and MCP under the same foundation, and the AAIF's first AGNTCon+MCPCon conference in San Jose on October 22–23, 2026, the format may gain more formal conventions. For now it remains deliberately minimal: plain Markdown, no required structure.
The practical conclusion doesn't depend on how that plays out. Write the file for the specific ways agents fail in your repository, keep it short enough that every line pays its way, test it against runs without it, and review changes to it as seriously as the code it governs.