AI citation reporting is the practice of asking the AI engines the questions your prospects ask, on a schedule, and recording whether your firm is named, how it is described, and which firms are named instead. It exists because none of your existing reporting can tell you this. Analytics sees a visit after the decision has been made, and Search Console does not break out AI surfaces at all. Without it, a firm's AI visibility is whatever someone happened to notice in ChatGPT last week.
Why Search Console cannot tell you this
Because Google does not report AI features separately. Its documentation on AI features in Search states that performance data for AI Overviews and AI Mode is included in the Search Console Performance report under the "Web" search type, mixed in with everything else. There is no filter that isolates it. A firm that grew its AI Overview presence and a firm that lost it can produce the same chart.
The same page is worth reading for a second reason. Google states there are no additional requirements or special optimizations for appearing in these features, and that you do not need to create new machine readable files, AI text files, or markup to be eligible. That is a useful check against a whole category of pitch, and the argument is made at more length in GEO vs SEO for law firms.
Outside Google there is no reporting at all. ChatGPT, Claude, and Perplexity do not send you a dashboard. If a prospect asked one of them for a personal injury firm in your city this morning and it named three firms that were not you, nothing in your stack recorded it. That silence is the gap citation reporting fills.
How to track law firm mentions in ChatGPT and Google AI Overviews
Treat it as a repeatable measurement rather than a search. The value is in running the same thing the same way every month, because a single answer on a single day tells you nothing about a system that varies between sessions.
Build a fixed prompt set
Start from how people ask an assistant, which is not how they type into Google. Nobody asks ChatGPT for "personal injury lawyer Los Angeles". They ask who to call after a rear-end collision on the 405, whether they need a lawyer for a slip and fall in a supermarket, or what a good injury firm in Pasadena would charge. Write thirty to fifty of those, spread across your case types, your geography, and the stages of the question, from "do I even have a case" through to "who should I hire". Then freeze the list, because a prompt set you keep editing produces numbers that cannot be compared.
Run it under controlled conditions
Same prompts, same engines, same cadence, no logged-in session carrying your own history into the answer. Personalization will happily show you your own firm and tell you nothing. Cover the engines your prospects use: ChatGPT, Google's AI Overviews and AI Mode, Perplexity, Claude, Gemini, and Bing. Run each prompt more than once, because these systems are not deterministic and one run is an anecdote.
Record four things, not one
Whether you were named is the obvious one and the least informative. Also record how you were described, because a firm named with the wrong practice areas or the wrong city has an entity problem rather than a visibility problem. Record which competitors were named, the most actionable column in the report. And record the sources the engine cited, where it shows them, because that is the list of pages you need to be on.
Watch what sits behind the answer
Citations lean heavily on a small set of sources: the legal directories, bar association pages, regional press, and a handful of publishers the engine trusts for legal questions. When a competitor is named consistently and you are not, the explanation is usually in that source list rather than on either firm's website. That is where earned coverage stops being a brand exercise and becomes a measurable input.
Want the baseline before you decide anything?
The free audit runs a prompt set for your firm across the major engines and reports where you are named, how you are described, and who is named instead.
Get My Free AI Visibility AuditWhat the monthly report should contain
Citation share by engine, so you can see that you are strong in Perplexity and absent in ChatGPT rather than averaging the two into a number that describes neither. Citation share by prompt group, broken out by case type and geography, because being named for estate planning while missing every motor vehicle prompt is a specific and fixable problem. The wording of how you were described, quoted rather than summarized. The competitor set, with movement since last month. The sources behind the answers. And the change list: what happens next month and which prompts it targets.
What it should not contain is a single composite visibility score. Those are comforting and unfalsifiable. A number that goes up without telling you which prompt moved is not a measurement, it is a mood.
It also has to sit next to the Google numbers rather than in its own deck. A prospect given three names by an assistant then opens Google and searches two of them, so a citation you win and a ranking you lose is a case you still lose. One report, both channels, is how we run it inside AI search visibility engagements.
What to do when the answer comes back wrong
There are three failure modes and they have different fixes. If you are not named at all, the problem is usually authority and third party coverage, and the fix is slow: earn mentions on the sources the engine already draws from for your case types and geography.
If you are named but described wrongly, the problem is entity data: inconsistent firm names, a stale address, contradictory practice areas across your site, your Google Business Profile, and the legal directories, or structured data that disagrees with the visible page. That is a faster fix, covered in schema markup for personal injury law firms.
If you are named for some case types and invisible for others, the problem is content depth: the case types where you are absent are the ones without a page thorough enough to be quoted. That is the most common pattern in injury work, where a firm has one strong motor vehicle page and thin pages for everything else, and it is why case-type depth sits at the center of our personal injury lawyer SEO program.
How often to run it
Monthly, for reporting and for decisions. Weekly reads noise as signal, since these systems change their answers between runs for reasons unrelated to anything you did, and quarterly is too slow to catch a competitor moving on your case types. The exception is a rebuild, a rebrand, or a merger, where running the entity prompts more often catches the engines describing you as two firms.
Expect movement on the same timeline as the rest of the work. Measurable movement typically begins within 3 to 6 weeks of campaign start, and competitive terms in major metro markets typically take 2 to 4 months. Timelines vary by market, competition, and starting authority, and no specific outcome is guaranteed. What changes immediately is that you stop guessing.
Frequently asked questions
It is the practice of running a fixed set of prompts against the AI engines on a schedule and recording whether your firm is named, how it is described, which competitors are named instead, and which sources the answer drew on. It turns AI visibility from an anecdote into a number you can compare between months.
Not separately. Google includes AI Overviews and AI Mode performance in the Search Console Performance report under the "Web" search type, mixed with the rest of your search traffic, with no filter that isolates it. That is why prompt-based citation tracking exists as a separate exercise.
No. Google states there are no additional requirements or special optimizations for AI Overviews and AI Mode, and that you do not need to create new machine readable files, AI text files, or markup to be eligible. The page has to be indexed, eligible for snippets, and compliant with search policies. The work that earns a ranking is the work that earns the citation.
Thirty to fifty works for most firms: enough to cover each case type, each geography you serve, and the stages of the question, without becoming a list nobody maintains. What matters more than the count is that the set stays frozen, so the numbers remain comparable between months.
Because these systems are not deterministic, and personalization, session history, and location all affect what comes back. That is why citation reporting runs each prompt more than once, from a clean session, under the same conditions every month. One answer on one day is not a measurement.