llms.txt and Schema for AI Crawlers: What Actually Moves Citations
An llms.txt file and clean schema help AI crawlers read your site, but they do not earn citations on their own. Here is what each does and where they fit.

Key Highlights
- llms.txt is a plain-text file at your root that points AI crawlers to your most important, answer-ready content.
- Schema (structured data) labels what your content means so an LLM can extract claims confidently.
- Neither one earns citations by itself. They remove friction so quotable content can actually be found and lifted.
Teams keep asking whether adding an llms.txt file or more schema will get them cited by AI engines. The honest answer: these are enablers, not causes. They make your best content readable and extractable for AI crawlers, which matters, but a model still has to find something worth quoting. Here is what each does, what it does not do, and the order to do them in.
What llms.txt actually is
llms.txt is a markdown file you place at the root of your domain (yoursite.com/llms.txt). It is a curated map: a short description of who you are, plus links to the pages you most want AI systems to read and cite. Think of it as a hand-built index for language models, the way robots.txt and sitemap.xml serve traditional crawlers.
It does two useful things. It signals which content is canonical and answer-ready, so a crawler does not have to guess. And it gives smaller or newer sites a clean entry point instead of hoping a model stumbles onto the right page.
What it does not do is force citations. Publishing the file does not put you in answers. It just makes sure that when a model does look, it finds your strongest, clearest pages first. If you want one without hand-writing the markdown, the free llms.txt generator builds a valid file from your site in a couple of minutes.
What schema does for AI extraction
Schema, or structured data, is machine-readable labeling (usually JSON-LD) that tells a crawler what a piece of content means: this is an FAQ, this is the question, this is the answer, this is the author, this is the organization.
For AI engines, schema lowers extraction risk. When a model can see that a block is explicitly marked as an answer to a specific question, it can lift that answer with more confidence than guessing from raw prose. FAQPage and Article schema are the two that pull the most weight for AEO, because they map directly onto how people query AI assistants: in questions.
The table below sorts what each layer is for.
| Layer | What it is | What it does | What it does not do |
|---|---|---|---|
| llms.txt | Root file mapping key pages | Points crawlers to answer-ready content | Guarantee a citation |
| Schema (JSON-LD) | Machine labels on content | Helps LLMs extract claims confidently | Replace good content |
| Answer capsules | A 40-60 word direct answer on the page | Gives the model a clean sentence to quote | Work if buried below the fold |
| Authority signals | Mentions on trusted sources | Tell the model you are worth trusting | Happen without earned media |
The mistake: treating these as the whole job
The pattern we see most: a team adds llms.txt and schema, waits a month, and sees no movement in AI answers. The fix was never going to come from markup alone. Schema labels content; it does not create a quotable claim. If the underlying page does not open with a tight, standalone answer, there is nothing clean for the model to lift.
That is why the structural work comes first. The page has to answer the real question in the first 40-60 words, then back it with clear H2 sections, a table where it fits, and real entity clarity. We break the full pattern down in what content structure actually gets cited by AI assistants. Schema and llms.txt then make that good content easy to find and safe to extract.
The right order of operations
Run it in this sequence and each step pays off the next.
- Identify the real questions buyers ask AI engines in your category, and where competitors already get cited. The method is in our guide to getting cited by AI engines.
- Write answer-first content for those exact questions: capsule, structure, table, FAQ.
- Add FAQPage and Article schema so the answers are labeled and extractable.
- Publish an llms.txt that points to those pages.
- Keep the feed fresh. Models favor sources that publish answer-shaped content consistently, which is the job the AI Feed Engine runs on autopilot.
- Measure citation share monthly so you know what is working.
You can see how the whole engine fits together on the how OnlyAEO works page, and what it costs to run on the OnlyAEO pricing page. For a worked example of this exact sequence moving citations, the FastTrackr AI case study walks through the before and after.
Get your free AI visibility audit
Generate a valid llms.txt for your site in minutes, then see which pages AI engines should be reading first.
Build my llms.txtFrequently asked questions
Frequently Asked Questions
Will adding llms.txt get my brand cited by AI?+
Which schema types matter most for AEO?+
Where do I put the llms.txt file?+
Do I still need schema if I have an llms.txt file?+

OnlyAEO
Expert insights on Answer Engine Optimization and AI visibility strategy.
Related Articles

Your Client's AI Mentions Dropped. Here Is How to Diagnose It Before the Call
A drop in AI citations is usually noise, a platform-wide event, or a competitor, and rarely your work. Here is the five-branch diagnostic agencies can run in 45 minutes, plus what to say on the client call.
Read article
How Long Does AEO Take to Work? A Timeline by Engine
AEO does not run on one clock. Here is the realistic timeline from publish to crawl to first citation to stable share, broken out by ChatGPT, Perplexity, Gemini, and Claude, plus what to measure while you wait.
Read article
What an AI-Sourced Lead Is Actually Worth (and Why the Studies Disagree)
Published studies put AI referral conversion anywhere from 0.3x to 23x organic. Here is why they disagree, the four-number formula for your own value per AI-sourced session, and how to use it without overclaiming.
Read article