How to choose an AEO agency: 12 questions
Twelve due-diligence questions, a weighted scorecard, and clear red flags for choosing an AEO agency without buying guarantees or dashboard theater.
Choosing an AEO agency is difficult because the category has more confident claims than settled methods. Answer products change, outputs vary, and no agency controls whether an external system retrieves, cites, mentions, or recommends a page. That does not make AEO work useless. It makes evidence, scope, and experimental discipline unusually important.
A credible provider should be able to explain what it observes, what it infers, what it can change on your properties, and what remains outside its control. It should connect AEO to technical SEO, content quality, entity consistency, source evidence, and measurement rather than sell a separate layer of “AI tricks.”
Use the following questions in an actual buying conversation. Ask for samples with client-sensitive details removed. The quality of the answers will tell you more than a polished visibility dashboard.
First, decide what kind of help you need
“AEO agency” can describe very different engagements. One provider may sell a diagnostic audit. Another may run ongoing technical, editorial, digital PR, and monitoring work. A software vendor may call its dashboard an agency service even when it provides little analysis or implementation.
Write down the immediate decision before requesting proposals:
- Do you need a baseline of current answer visibility?
- Are systems repeating incorrect facts about the company?
- Do you need technical and content defects diagnosed?
- Does your team need implementation capacity?
- Are you planning a site migration or repositioning?
- Do executives need defensible measurement rather than a one-time report?
A bounded audit fits when the team can implement findings and needs a prioritized starting point. An ongoing engagement fits when the site changes frequently, multiple owners must be coordinated, or repeated experiments and content work are required. A monitoring tool fits a team that already has the method and mainly needs collection.
Leaf explains the boundary between technical, content, and answer testing in its SEO and AEO audit guide.
Questions 1–3: how does the agency establish a baseline?
1. Which answer products do you test, and why?
The agency should name the products, interfaces, market, language, and account state where relevant. “We track AI everywhere” is not a method. Coverage should reflect where your buyers research, not which logos look best in a proposal.
2. How do you build the prompt set?
Look for prompts derived from real buyer stages: discovery, comparison, objections, implementation, and branded factual accuracy. Neutral prompts should not name your company. Branded prompts should be reported separately. Ask whether you can review and approve the panel before baseline collection.
3. What exactly do you preserve from each response?
At minimum, expect the exact prompt, date, product or engine, output, cited URLs, brand mentions, competitors, and factual errors. Where model labels, location, personalization, or account state are visible, those should be recorded too. A percentage without raw or reviewable evidence is difficult to audit.
These questions matter because generated responses are variable. A provider may repeat prompts to estimate stability, but it should disclose the repetition count and denominator. It should not present a small chosen panel as universal market share.
Questions 4–6: can it diagnose causes rather than report symptoms?
4. How do you distinguish access, retrieval, citation, mention, and recommendation?
These are different events. A crawler being allowed to fetch a page does not prove the page is indexed. Indexing does not prove retrieval for a prompt. A mention is not necessarily supported by a citation, and a citation is not automatically an endorsement or recommendation.
Google’s official guidance says that the same SEO fundamentals apply to its AI features. OpenAI separately documents crawler controls for its products in its official bots documentation. Those are verified platform facts. Neither source promises that allowing a crawler or following SEO basics will secure a citation.
5. How do you investigate a missing or incorrect answer?
A good answer should include technical access, page ownership, claim accuracy, internal linking, visible evidence, third-party corroboration, and competitor-source review. A poor answer jumps directly from “not cited” to “publish more articles.”
Ask the provider to walk through an anonymized example from observation to recommendation. It should show why a proposed change is plausible and what alternative explanations remain.
6. How does core SEO fit the engagement?
AEO cannot responsibly ignore status codes, canonicals, robots directives, rendering, internal links, structured data accuracy, page intent, and conversion measurement. The agency does not have to deliver every SEO service itself, but it must identify dependencies and coordinate with whoever owns them.
If you need a more detailed technical buying list, use Leaf’s technical SEO audit checklist.
Questions 7–9: are the recommendations safe and implementable?
7. How do you verify company and product claims?
Ask who approves claims, how dates and sources are recorded, and what happens when the website conflicts with product documentation or third-party listings. The agency should not invent proof, manufacture author biographies, or turn internal estimates into public facts.
Google’s people-first content guidance asks whether content demonstrates first-hand expertise and serves an intended audience. It does not provide a guaranteed ranking recipe. A responsible agency will use such guidance as a quality standard, not as a promise.
8. What does an implementation-ready recommendation contain?
Expect the affected URL or template, captured evidence, proposed change, business rationale, likely owner, dependency, and retest. “Improve authority” or “make this answer-ready” is not an implementable instruction.
Ask whether implementation is included, optional, or explicitly excluded. Also ask who owns code, copy, design files, and reporting data at the end of the engagement.
9. How do you separate fixes from experiments?
Correcting a broken canonical can have a deterministic pass/fail test. Rewriting a comparison section to improve retrieval is an experiment whose outcome depends on external systems. The proposal and backlog should label these differently. Otherwise, the agency can call every shipped change a success regardless of observed results.
Questions 10–12: can the agency report honestly over time?
10. Which metrics stay separate?
At minimum, the agency should separate mentions, cited answers, unique cited URLs, factual accuracy, recommendation appearances, referral visits where observable, qualified conversions, and implementation completion. A composite score may be a convenience, but the underlying checks and weights must remain visible.
11. How will you retest, and how do you handle product changes?
Ask whether the original prompt panel remains fixed, how often key prompts are repeated, and how changes to prompts, markets, engines, or account settings are versioned. Trend lines become misleading when the denominator changes silently.
12. What will you not promise?
The right answer explicitly refuses to guarantee rankings, citations, traffic, leads, or revenue. Search and answer platforms control their outputs. Agencies can improve owned assets, remove defects, strengthen evidence, and run measured experiments. They cannot sell placement they do not control.
This question also exposes commercial maturity. A provider should be willing to say when your problem is primarily product positioning, analytics, reputation, engineering, or legal review rather than AEO.
Use a weighted agency scorecard
Score each proposal from 0 to 3 and multiply by the weight. This is a buyer heuristic, not an industry standard. Adjust the weights to your risks before sending the request for proposal so a charismatic presentation cannot rewrite your criteria afterward.
| Criterion | Weight | 0 | 1 | 2 | 3 |
|---|---|---|---|---|---|
| Prompt method and preserved evidence | 20 | Undisclosed | Vague sample | Defined but incomplete | Reviewable, versioned method and evidence |
| Technical SEO depth | 15 | Ignored | Tool export only | Key checks covered | Template-aware diagnosis and dependencies |
| Claim and source governance | 15 | No process | Informal review | Sources recorded | Owners, approval, dates, and qualifications |
| Recommendation quality | 15 | Generic advice | Page-level ideas | URLs and priorities | Owners, dependencies, implementation, retests |
| Measurement discipline | 15 | One mystery score | Basic mention count | Separate useful metrics | Denominators, caveats, baselines, business events |
| Scope and ownership clarity | 10 | Ambiguous | Major gaps | Inclusions mostly clear | Deliverables, exclusions, IP, access, and handoff clear |
| Commercial honesty | 10 | Guarantees | Heavy implication | Qualified language | Explicit limits and fit boundaries |
To calculate a comparable result, divide each rating by 3, multiply by the weight, and add the rows. Do not let the total override a critical red flag. Fabricated evidence, guaranteed citations, unsafe access requests, or instructions to publish false claims should disqualify a provider regardless of score.
Red flags that should end the sales call
Walk away when an agency:
- guarantees inclusion, citations, or a stable position in an answer product;
- refuses to disclose the prompt set, denominator, or evidence behind a score;
- treats crawler permission as proof of retrieval or recommendation;
- proposes mass publishing before examining technical access and existing page ownership;
- recommends schema types or hidden content that do not match visible facts;
- uses invented customer outcomes, credentials, quotes, or “research”;
- cannot explain how findings become tickets and how fixes are retested;
- bundles SEO, PR, content, monitoring, and development without clear responsibilities;
- claims one tactic works across every model and answer interface.
Be equally cautious with artificial urgency. Platform change is real, but it does not eliminate the need for access controls, stakeholder approval, careful measurement, and a sensible contract.
Compare proposals on the same scope
Provide each shortlisted agency with the same site context, priority markets, buyer segments, known constraints, analytics availability, and implementation capacity. Request a response to the same 12 questions. Then normalize the proposals in a simple table: pages or templates reviewed, technical checks, prompt products and count, repetitions, evidence format, stakeholder interviews, deliverables, implementation, retest, timeline, exclusions, and total fee.
Price cannot be interpreted without this scope. A cheap dashboard export and a professional diagnosis are not equivalent products. Neither is automatically right; the decision depends on whether your team needs collection, analysis, implementation, or all three.
Leaf offers a bounded $875 SEO and AEO audit for established B2B sites, with implementation excluded. That may be suitable when you need diagnosis and can own the fixes. It is not a substitute for ongoing execution when your site, product, and evidence change every week. You can also run Leaf’s free homepage assessment to see how a narrowly scoped automated check differs from a full audit.
The best AEO agency is not the one claiming the most control. It is the one that makes uncertainty visible, protects factual integrity, produces work your team can execute, and agrees in advance on what evidence would count as progress.