Answer engine optimization agencies: evaluation checklist
A short, practical checklist for assessing an AEO agency — capability tests, reporting requirements, and the warning signs that matter.
This article is a checklist for evaluating an answer engine optimization agency. It assumes you already know what AEO involves and need to judge whether a specific provider can do it.
Capability tests
Run these before the commercial conversation.
- Ask them to check three of your buying questions across two engines and send you what they find, before any contract. Competent providers do this in an hour and it reveals more than a deck.
- Ask how they handle answer variance between runs. Anyone who has measured seriously will describe repetition and averaging. Anyone who has not will treat single answers as facts.
- Ask what they would check before writing content. Retrievability and existing page structure should come first; a content calendar first means the sequence is wrong.
- Ask for a report where results declined, and how they responded.
- Ask which surfaces they distinguish between — featured snippets, AI Overviews, and assistant answers behave differently and blending them hides the story.
- Ask what they think schema markup does. An answer that positions it as the mechanism for winning AI answers is a competence signal in the wrong direction.
Reporting requirements to agree upfront
- A frozen, documented question set and competitor set, with all changes logged.
- Raw engine answers stored and accessible to you.
- Per-surface and per-engine breakdowns rather than a single score.
- Both presence metrics and the competitor gap.
- A monthly note of what changed and what it was expected to affect, so cause and effect can be assessed later.
Warning signs
- Guarantees of any kind about placement in answers.
- Proprietary scores whose calculation is not disclosed.
- Content volume as the primary commitment.
- No baseline before work starts.
- Case studies with percentages and no absolute numbers.
- Unwillingness to show raw evidence behind metrics.
Scoping questions for yourself
Before engaging anyone, be clear on three things, because they change what you should buy.
Where your buyers actually are. If your category is researched in traditional search with AI Overviews present, you need a provider strong in SERP features. If it is researched in assistants, you need one strong in prompt-based measurement. These are different competencies.
Whether you need content production. Restructuring existing pages is a different, smaller engagement than producing new content, and mixing them makes proposals incomparable.
Who owns it afterwards. The question set is a strategic asset. Decide in advance whether it lives with you or the agency, and put it in the contract.
A reasonable structure
A bounded diagnostic engagement of four to six weeks — question research, baseline, retrievability audit, gap list — is the lowest-risk way to start. It produces a deliverable you keep regardless of whether the relationship continues, and it makes the subsequent scope arguable from evidence rather than assertion.
Common questions about choosing an AEO agency
- How do I check an agency's claims quickly? Ask them to run three of your buying questions across two engines and send the results before any contract. Competent providers do it in an hour, and how they read the answers tells you most of what you need.
- What is a fair first engagement? A bounded diagnostic of four to six weeks producing a question set, a baseline, a retrievability audit, and a prioritised gap list. You keep all of it regardless of what happens next.
- Do they need to be an AEO specialist? Not necessarily. Established SEO agencies that have genuinely built the measurement capability are often better placed, because the technical foundation overlaps heavily. The measurement test settles it either way.
- Who should own the question set? You. It is a strategic asset, and portability of questions, response history, and briefs belongs in the contract from the start.
- What is the most common way these engagements go wrong? No agreed baseline. Almost every dispute about results traces back to a measurement nobody signed off at the beginning.
- Should I run a paid pilot? It is usually the cheapest way to resolve uncertainty, provided the deliverable is something you keep.
- How do I compare two agencies fairly? Give both the same five buying questions and the same brief, and compare what they come back with. Proposals written against different scopes are not comparable, and agencies will happily scope to their strengths if you let them.
- What if my content team is already stretched? Say so during scoping. An engagement that produces a fifty-page backlog nobody can action is worse than a smaller one matched to your actual publishing capacity.
- Does agency location matter? For the measurement and content work, no. For off-site outreach it can, since relationships with publishers and review sites in your market are what make that half effective.
- How do I hand this over internally later? Insist that the question set, baseline, measurement method, and briefs are documented as you go rather than reconstructed at the end. Handover quality is decided in month one, not month twelve.
Tip: whoever you pick, insist the baseline is delivered and signed off before implementation begins. Almost every dispute about AEO results traces back to a measurement that was never agreed at the start.
Still stuck? We typically reply within 1 business day.
Contact support