Choosing a marketing agency is a high-stakes decision made under incomplete information: every agency's pitch deck looks confident, every case study looks impressive, and most buyers have no consistent way to compare one proposal against another. This article gives a structured way to evaluate candidates side by side, so the decision rests on comparable evidence rather than whichever pitch felt most persuasive in the room.
Start with business fit, before anything else
An agency can be excellent in general and still be the wrong choice for your business. Before evaluating skill, check fit: does the agency have real experience in your business model (B2B versus ecommerce versus local service), your typical deal size, and your sales cycle length? An agency skilled at high-volume ecommerce ad campaigns may be a poor fit for a long B2B sales cycle, and vice versa, regardless of how good their general reputation is.
The ten evaluation criteria
- Business fit: relevant experience with your business model, industry, and typical customer journey.
- Domain competence: depth of actual expertise in the specific channels you need (SEO, paid media, AI search visibility, email, etc.), not generalist claims.
- Case evidence: verifiable examples of past work, ideally with a reference you can speak to directly.
- Creative quality: the actual quality of content, design, or campaigns they've produced, evaluated on samples, not descriptions.
- Measurement practices: how they define success metrics, and whether they distinguish attributed revenue from profit (see article 099 on agency guarantees).
- Account ownership: who specifically will manage your account day to day, and their experience level.
- Team structure: whether work is done by the team you met, a different internal team, or outsourced further.
- Reporting: format, frequency, and whether you get access to raw data or only summarized dashboards.
- Pricing: fee structure (retainer, project, performance-based) and what's included versus billed separately.
- Operating process: how change requests, strategy reviews, and escalations are actually handled month to month.
| Criteria | Agency A | Agency B |
|---|---|---|
| Business fit | Strong: 3 referenceable clients in your sector | Weak: mostly ecommerce, you are B2B |
| Account ownership | Named senior strategist, verified on call | Pitched by a director, unclear who manages it |
| Measurement practice | Distinguishes attributed revenue from profit | Leads with a headline ROAS, no definitions |
| Reporting | Monthly call plus raw dashboard access | Quarterly PDF summary only |
| Pricing | Retainer plus clearly scoped project fees | Retainer with vague 'additional services' clause |
Build a weighted scorecard
Not every criterion matters equally for every business. A weighted scorecard assigns each criterion a weight (reflecting its importance to you) and a score per agency (how well they performed on it), then multiplies and sums for a comparable total.
Spotting unsupported AI and results claims
Many agencies now market 'AI-powered' services. This can mean genuinely useful applications, such as AI-assisted ad creative testing, research, or reporting automation, or it can mean little more than marketing language attached to standard work. Ask specifically what AI tools are used, for what task, and what a human still reviews or decides. Be skeptical of claims that AI guarantees rankings, guarantees inclusion in AI-generated answers, or guarantees sales outcomes; no agency, AI-assisted or not, controls a third-party platform's algorithm or a buyer's independent decision (see article 099 for more on guarantee claims generally).
Twenty practical interview questions
Use these during vendor calls to surface how an agency actually operates, not just what their proposal states.
- Which of your current or recent clients is most similar to our business, and can we speak with them directly?
- Who specifically will manage our account day to day, and what is their experience level?
- What happens if our assigned strategist leaves or changes roles?
- Walk us through a campaign that underperformed. What did you change, and what did you learn?
- How do you define a 'qualified lead' or 'success' for a client like us?
- Do you distinguish attributed revenue from profit in your reporting? How?
- What attribution model do you use, and why?
- What access will we have to raw data versus summarized reports?
- How often do we meet, and what's covered in each meeting?
- What is included in the retainer, and what is billed separately?
- How do you handle a request to change strategy mid-contract?
- What specific AI tools do you use, for which tasks, and where does a human still make the final call?
- Can you guarantee any outcome? If so, exactly how is that guarantee defined and verified?
- What is your typical contract length and exit process if it's not working out?
- How many accounts does each strategist typically manage at once?
- What creative or content samples can you show us that are directly comparable to our industry?
- How do you stay current on platform changes (for example, in search, social, or AI advertising)?
- What's an example of a client request you declined because it wasn't the right fit for their business?
- How do you measure and report on both short-term activity and longer-term outcomes?
- What would the first 90 days with us actually look like, week by week?
- Weighted scorecard completed for all shortlisted agencies using the same criteria
- At least one reference call completed with a comparable current or former client
- Measurement definitions (qualified lead, revenue, profit) agreed and written into the contract
- Account ownership and team structure confirmed, not assumed from the pitch meeting
- Pricing model and what's included versus billed separately fully understood
- No unverifiable guarantee claims accepted at face value
Common mistakes
- Choosing based on the most polished pitch deck rather than verified evidence and fit.
- Skipping reference calls because the case studies 'looked convincing' on their own.
- Not clarifying who actually does the work versus who presented the proposal.
- Accepting vague AI or guarantee claims without asking how they're implemented or defined.
- Comparing agencies on price alone without normalizing for what's actually included in each proposal.
- Failing to agree on measurement definitions before signing, leading to disputes about whether the engagement is 'working.'
When this is not the right tactic
If you have very limited budget and the realistic choice is between a freelancer and no outside help at all, rather than between multiple agencies, a full weighted scorecard may be more process than the decision warrants; focus instead on a shorter trial project and direct reference checks. If you already have a strong in-house team and are only evaluating a narrow, well-defined task (for example, a single landing page build), a lighter-weight comparison focused on relevant portfolio samples and price may be more proportionate than this full framework.



