You have two or three agency calls booked, and every agency on the list is using the same three words. GEO. AEO. Somewhere in the deck, LLM optimization. They all describe the same goal, which is getting your brand named when a buyer asks ChatGPT, Perplexity, or Google’s AI answers a question.
But herein lies the problem: most agencies get bookings by relabeling what they already sold. Heck, they probably used the same checklist built for the SEO era, just packaged differently to suit your needs.
And the worst part? Since showing up on LLMs is a relatively new practice, you’re not sure which questions to ask these agencies to ensure that they know their stuff. This lack of clarity regarding GEO/AEO/LLM optimization makes deciding which agency to get help from all the more difficult.
To make things much easier for you, here are 29 questions to ask an AI search agency before you sign, grouped in the order that actually decides an engagement. AI search visibility first, because that is where the relabeled pitches come apart. Money and people last.
And just as important, we answer all 29 ourselves, in writing, right here: real prices, real contract terms, real case numbers, and an honest answer about when you should not hire us. If we are going to hand you a vetting standard, it is only fair that we go through it first.
Why the Hiring Questions Changed in 2026
For centuries, alchemy meant turning lead into gold. The instinct behind the old agency checklist still holds: find someone who can turn attention into revenue. But the science and mechanics underneath it have changed.
Before, a buyer used to run a search, scan ten blue links, and click two of them. That same buyer now types a full paragraph into ChatGPT, Perplexity, or Google’s AI answers and gets back a single synthesized response with maybe three brands named in it. If yours is not one of the three, there is no second page for you to be on. The click simply never happens.
This changes what you are buying from an agency, which is presence inside an answer that a model assembles on the fly, differently every time, from sources it decides are worth trusting. The 29 questions test whether a vendor can measure that presence and actually move it.

If you want the longer version of what we believe about this work, it is on our what we mean by alchemy page, and the practice itself lives under AI SEO. This piece is the short version.
The 8 Questions That Decide It
The other 21 are useful, but these eight do most of the separating.
| # | Question | Why it decides it |
| 1 | Against what prompt set do you measure our AI visibility? | No fixed prompt set means no measurement, just anecdotes |
| 2 | How do you separate a mention from a citation from a recommendation? | A vendor who treats these as one thing is guessing |
| 3 | What are you doing that is not classic SEO with a new label? | The single fastest way to find a repackager |
| 4 | Do you run a scored diagnostic, or do you audit by feel? | A rubric can be checked. A vibe cannot |
| 5 | Do you audit at scale, or spot-check twenty pages? | Twenty pages of a 4,000-page site is a sample, not an audit |
| 6 | How do you attribute revenue to organic and AI search separately? | If they cannot split it, they cannot prove either one |
| 7 | If we leave, what do we keep? | Reveals who actually owns the work |
| 8 | What would make you tell us not to hire you? | The only question with no rehearsed answer |
Write those eight down and ask them in that order. Then use the rest of this piece to dig into whichever answers wobbled.
Group A: Questions About AI Search Visibility
There is a moment on most vendor calls where you can feel the ground get soft under the pitch. It usually arrives right after you ask how a model decides which brands to name in an answer.
Start your questioning here rather than at pricing to show agencies you mean business. This is the ground we work on under generative engine optimization, and it is where the gap between saying GEO and doing GEO shows up first.
Questions 1-4 are about measurement. A vendor who cannot measure your AI visibility today cannot promise to change it, whatever the deck says. The next three test whether they can change what happens on it, which is a different skill. Finally, questions 8-10 in this group are the repackager detectors. A vendor selling renamed SEO will answer the first seven acceptably and then get vague right about here.
1. How Do You Measure Our Visibility in AI Answers Today, and Against What Prompt Set?
What to listen for: A specific number of prompts, built from your category and your buyers, with a stated method for how the list was assembled. Vague talk about “monitoring AI mentions” usually means somebody is typing searches in by hand.
Our answer: We build a prompt set from your commercial keywords, the objections your sales team hears most, and the questions your buyers actually ask before they purchase. Most sets land between 100 and 300 prompts. The average is around 150, and enterprise accounts with wide catalogs run into the thousands. We run the set across engines on a fixed schedule, log every result, and hand you the documented set itself.
2. Is the Prompt Set Fixed and Repeatable, or Does It Change Between Reports?
What to listen for: A locked baseline, with additions tracked separately. If the prompt list quietly changes from month to month, every trend line built on it is fiction.
Our answer: The baseline set is locked at the start of the engagement. When your category shifts we add prompts, and the additions are versioned and tracked on their own so the baseline trend stays clean. You can see exactly what changed and when it changed.
3. How Do You Separate a Mention From a Citation From a Recommendation?
This is the question where inflated reporting likes to hide, so listen closely.
What to listen for: Three distinct definitions, given without hesitation. These are three different business outcomes, and a vendor who blurs them together is padding your results, whether they mean to or not.
Our answer: A mention is your brand appearing in the answer text. A citation is a link back to your domain in the sources. A recommendation is the model naming you as the answer rather than one option among several. We report all three separately, every time, because only recommendations reliably move revenue. Mentions and citations usually move first, so we watch them as early signals. We just do not let them headline a report.

4. How Do You Handle Answer Volatility?
What to listen for: An acknowledgment that the same prompt returns different answers on different days, followed by an actual method: repeated runs, share-of-answer scoring, some form of averaging.
Our answer: We run each prompt multiple times per cycle and report share of answers rather than one snapshot. If you appear in six out of ten runs, that is 60% presence, a number you can watch month over month. One screenshot offered as proof is a coin flip.
5. Which Engines Do You Actually Optimize For, and Which Do You Only Report On?
What to listen for: An honest split. Nobody has real levers on every engine, and better to hear that now than in month five.
Our answer: We work directly on the surfaces where source selection responds to changes we can make. Today that means Google’s AI answers, ChatGPT’s browsing and index-backed responses, and Perplexity. The rest we report on. We will tell you which bucket each engine sits in before you sign.
6. Have You Ever Moved a Brand Into an Answer Set Where It Was Absent? Prove Causation.
What to listen for: A before state, a specific change, an after state, and a stated reason the two are connected. “Traffic went up” is not causation.
Our answer: Yes. Our clearest public example is an aerospace gearbox engineering firm, a technical B2B manufacturer in a category most people assume AI search never touches. Over twelve months the account earned 158 new AI citations across generative search platforms, qualified quote requests rose 228% year over year, and organic traffic grew 21% with no paid spend anywhere in the mix.

Those numbers only make sense read together. A 21% traffic lift is nothing special on its own. A 228% lead lift sitting on top of it means the traffic changed character, not just volume. The citations are the mechanism. Every proof we show a client has that same shape: logged baseline, dated change, logged result, prompt set held constant across all of it. We also run the same relentless strategy and testing on our own properties first, which is how we know a tactic works before your budget is the one paying to find out.
7. What Happens When an Engine States Something False About Our Brand?
What to listen for: A named process and a realistic timeline. Anyone promising a same-week fix does not understand how model training and retrieval work.
Our answer: We trace the source the model is pulling from, correct it or displace it, strengthen the accurate source with clearer structured signals, and use whatever feedback channels the engine provides. Retrieval-driven errors often correct within weeks. Errors baked into training data can take a full model cycle to clear. Part of the job is telling you which kind you have.
8. What Are You Doing That Is Not Classic SEO With a New Label?
What to listen for: Concrete artifacts that did not exist in an SEO scope of work. If the whole answer is “well, good SEO is good AEO,” you have found a repackager.
Our answer: Prompt-set construction and tracking. Passage-level restructuring so an extractable answer sits inside the first 100 words of a section. Entity consistency work across the properties models trust. Citation-source displacement, which means going after the pages currently being quoted instead of yours. None of that appears on a 2019 SEO scope.
9. The Engines Change Constantly. What Is Your Process When a Model Update Wipes Out a Tactic?
What to listen for: A detection method, not a promise to stay current. Press on this. How would they even know a tactic had died?
Our answer: The fixed prompt set is the detection system. When presence drops across many prompts at once and nothing changed on your site, that points to an engine-side shift. We isolate it, test replacements on our own properties, and then roll the fix out to clients. We would rather tell you a tactic stopped working than quietly keep billing for it.
10. How Does Our Existing Content Library Get Evaluated for AI Extractability?
What to listen for: Page-level scoring against stated criteria. Structure, answer placement, entity clarity, source credibility.
Our answer: Every page gets scored on how cleanly a model can lift a self-contained answer out of it. Most content libraries are full of pages that rank fine and extract terribly, usually because the actual answer is buried in paragraph six. That is ordinary matter waiting for the right conditions, to borrow our own phrase, and fixing it is the cheapest win available to most brands. The answers already exist. They are just in the wrong place on the page.
Group B: Questions About How They Diagnose Your Site
Case studies are written after the fact, by the people who want your business, about the engagements that went well. Nothing wrong with that, but they shouldn’t be the only evidence you get to inspect.
A diagnostic is a different kind of evidence. It is a standing method you can inspect before you spend anything, and it tells you what a vendor believes matters, in what order, and whether they have done this enough times to have a rubric rather than an opinion. Six questions get you to the bottom of it.
11. Do You Run a Scored Diagnostic, or Do You Audit by Feel? What Is the Rubric?
What to listen for: A named system with a scoring scale. “We do a deep audit” is not a system. A useful follow-up: ask what a bad score looks like.
Our answer: We run the AlchemyLeads AI SEO diagnostic on our AURIC AI SEO Bot. Seven dimensions, each rated on a defined scale, with the criteria written down. What comes out is a maturity stage and a ranked fix list, not a 90-page PDF of everything that is technically wrong. The full scope is on our digital marketing audit page.
12. Walk Me Through All Seven Dimensions You Score.
What to listen for: They can name every dimension without checking their notes, and they can say what data each one requires. A vendor who scores four things and calls it seven will stumble on this one.
Our answer: Here they are, along with the access each one needs.
| Dimension | What gets scored | Data we need from you |
| Indexation | How much of your site engines can find, crawl, and keep | Google Search Console, sitemaps, live crawl |
| Technical Health | Speed, rendering, structured data, crawl waste. Detail on our technical SEO page | Crawl, Core Web Vitals, server logs |
| Content Depth | Coverage against buyer questions, and extractability per page | CMS export or full crawl |
| Authority | Link quality, brand mentions, and how much models trust your sources | Backlink data, brand mention scan |
| AI Visibility | Presence across the fixed prompt set, split by mention, citation, recommendation | Our prompt-set runs |
| Local Presence | Profile accuracy and citation consistency, where location drives revenue | Google Business Profile access |
| Measurement Readiness | Can we tie any of this to money today, and if not, what is broken | GA4, CRM or revenue reporting |
Five of those seven have nothing to do with AI search directly, and that is deliberate. AI visibility built on a site that engines cannot crawl properly does not hold. We learned to be stubborn about this the expensive way.
An industrial cleanroom manufacturer makes the point better than any argument could. Average order value north of $50,000, real demand in the market, and a site quietly strangling itself: 933 pages tangled in mixed hreflang localization, plus 307 broken 404 errors. None of that sounds like an AI search problem. All of it was blocking one. Clearing it inside 90 days, alongside 120 mapped topics and 32 new posts, took monthly quote requests from 20 to 54 and moved new monthly sales opportunities from roughly $1M to $2.7M. Nobody rewrote a single answer for ChatGPT. The foundation was the work.
13. What Data Access Do You Need, and Specifically Why for Each One?
What to listen for: A justification per source. A vendor who asks for everything with no reason attached has not thought hard about what they are going to do with it.
Our answer: Search Console for query and indexation truth. GA4 for behavior and conversion paths. CMS access for content structure at scale. Server logs for what crawlers actually do, as opposed to what everyone assumes they do. CRM or revenue data so the measurement connects to money. If you cannot grant one of those, we will tell you which dimension goes unscored. A guess dressed up as a score helps nobody.
14. Do You Audit at Scale, or Spot-Check Twenty Pages and Extrapolate?
What to listen for: A number. How many pages, and by what mechanism.
Our answer: At scale. Every crawlable URL gets scored programmatically, and then we hand-review the pages carrying commercial weight.
Scale is not an abstraction here. One 3PL warehousing client went from 40,000 to 180,000 monthly organic visits and now holds more than 8,000 keywords on page one. An online pet pharmacy sits at 13,459. Positions at that volume do not get found, fixed, or defended by sampling. Twenty pages out of a 4,000-page site tells you about those twenty pages, and in our experience the expensive problems live in the corners of a site nobody has opened in three years.
15. What Does the Diagnostic Cost, and Is It Credited Against the Engagement?
What to listen for: A real number, and a clear answer on crediting. A free audit is usually a sales document with an audit’s haircut. A priced audit is usually a deliverable.
Our answer: The AlchemyLeads AI SEO diagnostic runs on our AURIC AI SEO Bot and costs $2,500 per site. If you move forward with an engagement, the $2,500 is credited against your first month. You can also buy the diagnostic, take the output, and never speak to us again. Some clients do exactly that, and we are fine with it.
16. What Is the Deliverable at the End, and Can I Take It to Another Agency?
What to listen for: Yes, without qualification. A diagnostic you cannot act on elsewhere is not a diagnostic but a lock-in device.
Our answer: You get the scores, the rubric, the ranked fix list, the prompt set, and the raw data. All of it is yours. Hand it to another agency and they can execute against it tomorrow. We would rather win the work on the strength of the plan than on your inability to take it anywhere else.
Group C: Questions About Results and Measurement
If you have been burned by an agency before, it probably was twelve months of reporting that looked busy and never connected to a number your CFO recognized. Sessions rose and rankings improved but nobody could say what any of it was worth.
These six questions make that particular failure very hard to hide. They are about the shape of the reporting and not the promises attached to it. Questions 20-22 are the ones vendors like least, which is exactly why they earn their place. Each one invites the vendor to say something unflattering about themselves.
17. What Does Success Look Like at 90 Days, 180 Days, and 12 Months?
What to listen for: Different metrics at each stage. If 90-day success and 12-month success are the same metric, one of them is wrong.
Our answer: At 90 days, leading indicators: indexation fixed, extractability scores up, baseline presence established. At 180 days, presence and qualified organic sessions moving together. At 12 months, revenue attributable to organic and AI search, tracked separately.

Sometimes revenue shows up faster than that, and when it does we will happily say so. The cleanroom manufacturer’s quote volume moved inside 90 days because the bottleneck was technical and the fix was immediate. The aerospace account took the full twelve months. What we will not do is promise you the fast version before we have seen your site. Anyone promising revenue at 90 days on a discovery call is selling you a timeline, not a plan.
18. How Do You Attribute Revenue to Organic and AI Search Separately?
What to listen for: An admission that AI search attribution is partly dark, plus a real method for the part that is not.
Our answer: Referral data captures what the engines pass through, which is incomplete, and we say so upfront. We fill the gap with tracked landing paths, self-reported source capture at the form, and correlation between presence changes and direct or branded demand. Our analytics work sets all of this up before campaigns start, because retrofitting attribution after the fact never comes out clean.
19. What Is Your Reporting Cadence, and Who Presents It?
What to listen for: A named person with a title, not “your account team.” It is also worth asking whether the person presenting actually did the work.
Our answer: Monthly, presented live by the strategist running your account. You get the deck and the underlying data, and when you ask why a number moved, the person in the room can tell you, because they are the one who moved it.
20. What Is Outside Your Control, and How Does That Show Up in Reporting?
What to listen for: A real list. A vendor claiming full control of outcomes either has not been doing this long or is not being straight with you.
Our answer: Core algorithm updates. Model retraining cycles. Your competitors’ spend. Your own product, pricing, and sales follow-up. We flag external events on the reporting timeline so a drop caused by an update does not get blamed on the work, and, just as important, so a lift we did not earn does not get claimed as ours.
21. Show Me an Engagement Where Results Were Slower Than Expected.
What to listen for: A specific story with a diagnosis and a change made afterward. A vendor with no slow engagements has either not worked long enough or is editing the story.
Our answer: We have had them, and the pattern repeats: a site with deep technical debt, where the first three or four months go entirely to foundation work that produces nothing a client can see. Nine hundred broken localization pages do not photograph well in a monthly report.
What we changed afterward was the reporting, not the plan. We started showing leading indicators weekly, so a client can watch progress accumulate before revenue catches up. Setting that expectation on day one turned out to be the real fix, and we now do it on every account where the diagnostic comes back heavy on technical debt.
22. What Is the Metric You Would Fire Yourself Over?
What to listen for: One metric, named quickly. Hesitation here means they have never defined failure for themselves, and a vendor who cannot define failure cannot recognize it either.
Our answer: Qualified pipeline from organic and AI search. If that number is flat at twelve months against a clean baseline, and no external event explains it, then we did not do our job. Rankings can rise and presence can rise, and if the pipeline never moved, none of it counted.
Group D: Questions About Pricing, Terms, and Who Does the Work
The commercial questions come last on the call and first in your actual decision. Everyone knows this, everyone pretends otherwise, and so the money conversation gets crammed into the final six minutes while the rep watches the clock.
Ask these plainly, and ask them early enough in the call to get real answers. A vendor who gets cagey about a price range will be just as cagey about a missed month.
23. What Does This Actually Cost? Give Me a Range Before the Call.
What to listen for: A number. An agency that will not price anything before a discovery call is protecting a number they already know you will push back on.
Our answer: Our engagements start at $6,000 per month, and most sit above that depending on site scale, competitive pressure, and how much of the work is content production. The number is published on our pricing page, on purpose, so you can qualify us out before either of us spends an hour on a call.
24. How Long Is the Contract, and What Is the Exit?
What to listen for: The exit terms, specifically. A long contract is not automatically a problem. A long contract with a painful exit is.
Our answer: We recommend a 90 to 120 day setup period, then month to month after that. Leaving takes 30 days written notice. That is the whole policy.
We built it this way because a twelve month lock-in mostly protects the agency. The setup window is there because foundation work genuinely takes that long to land. After that, the thing keeping you with us should be momentum, and if it ever stops being momentum, thirty days and you are out.
25. If We Leave, What Do We Keep?
What to listen for: Everything, listed out loud. Hedging about “proprietary methods” means part of what you paid for walks out the door with the agency.
Our answer: All of it. The content, the technical fixes, the dashboards, the prompt set, the tracking configuration, the documentation, the diagnostic output. It was built for your business with your budget. There is no version of this where we hold your own prompt set hostage.
Pricing and terms tell you what an engagement costs. The last four questions tell you what you are actually buying, which is people.
26. Who Specifically Works on Our Account, and How Many Others Do They Carry?
What to listen for: Names, titles, and an account load number. The pitch team and the delivery team being different people is normal in this industry. Nobody warning you about it is not.
Our answer: You meet your strategist before you sign, and that person stays on the account. We cap account loads deliberately, because a strategist carrying too many accounts has no room left to think about yours. If the people in your pitch meeting will never touch your account again, you deserve to hear that from the vendor rather than discover it in month two.
27. Is Any of the Work Subcontracted or Offshored? Say So Plainly.
What to listen for: A direct yes or no. Subcontracting is not disqualifying, but hiding it is.
Our answer: Strategy, technical work, and analysis are done in-house. Where we use outside specialists for production capacity, you are told which parts, and the output is reviewed by your strategist before it ever reaches you. No mystery layer between you and the work.
28. an I Speak to Two Clients, Including One Who Left?
What to listen for: Genuine willingness on the second half. Every agency has happy references lined up. The client who left is where you learn something.
Our answer: Yes to both. Published reviews and case studies are on our reviews page with the numbers attached, and we will connect you with a current client and a former one. Ask the former client why they left. Sometimes the answer is us. Sometimes they built the capability in-house, which is a decent outcome for everyone.
29. What Would Make You Tell Us Not to Hire You?
This is the last question and the most useful one. Ask it, then stop talking and let the silence sit there.
What to listen for: A real list, offered without defensiveness. How a vendor handles this one question tells you more than the previous 28 combined.
Our answer: We publish the whole list on our who we don’t work with page. Three items from it matter most in this context.
If you are under roughly $2M in revenue with an average order value below $300, the math does not work yet, and we will tell you so on the call. If you have been through four agencies already, we will want to talk about why, because one bad agency is bad luck and four is a pattern. And if you are shopping for the cheapest option, that is not us. We never save pennies to lose dollars.
The rest of the list covers reporting theater, treating your website as a catalog nobody visits, permanent emergencies, marketing treated as a cost line, and plain old fit. We are blue collar about the work and we like working with people we like.
Red Flags in the Answers You Get Back
Some warning signs cut across every question on the list, so watch for them no matter what you asked.
Guarantees are the loudest one. Nobody can guarantee placement inside a generated answer, so a vendor offering one has turned the guarantee itself into the product. Close behind it is black-box language, where “proprietary process” gets used as a reason you cannot see the method. Proprietary should describe how something is packaged, never why it cannot be explained.
Then there is the total absence of downside. No failed engagements, no stated limits, no clients who ever left. Nobody’s record is that clean. Just as common, and quieter: metrics that never touch money. Impressions, rankings, and mentions, month after month, with no line ever drawn to pipeline, until nobody in the room remembers what the budget was supposed to produce in the first place.
Watch the room too. If the senior people pitching you today will be replaced by unnamed people in six weeks, that is worth knowing before you sign anything. And pay attention to how a vendor reacts to the eight decisive questions above. A wrong answer is survivable. Visible discomfort is the more useful signal, because it means the question landed somewhere unprepared.
None of these disqualifies a vendor on its own. Two or more in the same call usually should.
Ask Us the Same 29 Questions
You now have a standard. Use it on every vendor on your list, and use it on us.
That is why all 29 answers are sitting above in writing. A checklist an agency will not answer about itself is marketing. A checklist an agency publishes its own answers to is something you can hold us against in month seven, when the honeymoon is over and the numbers are doing all the talking.
We hold one belief about this work that does not fit in a metric: growth is the byproduct, and the transformation of the people inside it is the point. That only happens when the work is real and the reporting is honest, which is what the 29 questions are for.
Bring the hardest question on the list. Book a strategy call with AlchemyLeads.




