Eighty-six percent of B2B companies outsource some content. The agency evaluation process, for most teams, comes down to three data points: portfolio samples, pricing, and references. That process reliably identifies good writing services. It almost never identifies strategic content partners.

A writing service produces articles. A strategic partner builds a content operation that compounds. The difference in business impact is roughly 5x over three years — and it's detectable in the first 30-minute call if you ask the right questions.

14%
of B2B teams do everything in-house. The other 86% will evaluate an agency at some point — and most will choose based on criteria that don't predict performance.
Siege Media, 2026 (n=353)

Here are the 12 questions, grouped by what they evaluate: strategic capability, operational reliability, and measurement philosophy. Each includes what a good answer sounds like and the red flag you're listening for.

Strategic capability questions

QUESTION 01

"Walk me through how you'd structure a topic cluster for our space."

Do not give them advance notice of this question. A strategic partner can do this live — they understand cluster architecture as a discipline, not as a deliverable they prepare for.

Good answer

They ask 2-3 clarifying questions about your ICP and buying committee, then sketch a cluster structure in 3-5 minutes. The structure has a pillar page, 8-12 cluster articles, and a rationale for the ordering. They explain why certain topics belong in this cluster versus a different one.

Red flag

"We'd need to do a full content audit and strategy engagement to answer that. It's a 4-6 week process." Translation: we don't think in clusters, or we gate strategic thinking behind a separate engagement.

QUESTION 02

"Show me a piece of content you recommended killing — and why."

Every agency can show you their best work. The signal is in what they chose not to publish. An agency that never recommends killing a topic is optimizing for revenue-per-article, not client outcomes.

Good answer

They describe a specific client situation where a requested topic was too broad for the search intent, too competitive for the domain authority, or too disconnected from the revenue conversation. They explain the alternative they proposed and why it performed better.

Red flag

"We let the client decide what to publish. We execute on their direction." Translation: we're a production shop. We'll write whatever you ask for, whether it will perform or not.

QUESTION 03

"How do you think about content distribution — and what do you actually do after publish?"

Content without distribution is a library with the lights off. Most agencies stop at publish. The ones that don't are rare and worth the premium.

Good answer

They describe a post-publish process: social copy, newsletter inclusion, community sharing, repurposing into other formats, internal linking from existing content, outreach for backlinks or citations. They have a calendar, not a checklist.

Red flag

"We share every article on our clients' social channels and include it in their newsletter." Translation: we do the minimum. One tweet and one email is not distribution — it's notification.

QUESTION 04

"What's your perspective on AI in content creation — and where's the line?"

Eighty-seven percent of B2B marketers use AI for content. The question is not whether the agency uses AI — it's how. Specifically, where do they draw the line between AI assistance and human authorship.

Good answer

They use AI for research synthesis, outline generation, and SEO metadata. Writing, framework development, and original analysis are human-only. They can articulate why — AI produces grammatically correct text that is substantively shallow, and substantive depth is the only thing that performs over time.

Red flag

"AI lets us produce more content at lower cost." Translation: we generate AI drafts and lightly edit them. Your content will sound like everyone else's content — because it literally is.

Operational reliability questions

QUESTION 05

"Who writes my content — and will I know their name?"

Many agencies use a pooled writer model. You get whoever is available. Quality varies by writer. Some agencies actively conceal writer identity to prevent you from hiring them directly.

Good answer

You get a named writer or a named team of 2-3 writers who learn your space. You can speak with them directly. They stay on your account for 12+ months. If someone leaves, there's a handoff process — not a silent substitution.

Red flag

"We have a team of experienced writers who collaborate on every piece." Translation: we assign articles to whoever has capacity. You'll get different voices, different quality levels, and no institutional knowledge.

QUESTION 06

"What does your briefing process look like — and show me a real brief."

The brief is the single most important document in the content production process. A bad brief guarantees a bad article regardless of writer quality. Most agencies don't brief — they take a topic and a keyword and start writing.

Good answer

They share an actual brief from a current client. The brief includes: target reader and their context, the specific question the article answers, the argument structure, 3-5 source references, internal links to include, and the conversion goal. The brief is the strategy document for each piece.

Red flag

"We work from the topic and target keyword. Our writers handle the research." Translation: there is no brief. The writer googles the topic and writes what they find. Your content will be a summary of the first page of search results.

QUESTION 07

"How do you handle subject-matter expertise when you're new to a space?"

Every agency claims they can write about any B2B topic. The question is how they bridge the gap between generalist writing ability and domain-specific knowledge.

Good answer

They describe a structured onboarding: interviews with your SMEs, review of your sales calls and customer conversations, immersion in your competitive landscape. The ramp period is 4-6 weeks, during which they expect to require more review and revision. After ramp, the revision cycle shortens.

Red flag

"Our writers are quick studies. Give us your style guide and we'll be up to speed in a week." Translation: we will learn your space by reading your competitors' blogs. Your content will be a remix of what's already published.

QUESTION 08

"What happens when I'm unhappy with an article?"

Every agency has a revision policy. The question reveals whether revision is treated as standard operating procedure or as an exception that triggers internal friction.

Good answer

Two rounds of revision are standard and included. After two rounds, there's a retro conversation about the brief — because if two rounds didn't solve it, the problem is upstream from the writing. They don't charge for revisions. They investigate the root cause.

Red flag

"We offer one free revision with a 3-day turnaround." Translation: revisions are a cost center. We'll discourage you from using them through slow turnaround and limited scope.

Measurement philosophy questions

QUESTION 09

"What content metrics do you report — and what do they actually tell me?"

Most agencies report pageviews, time on page, and bounce rate. These are output metrics. They tell you what happened, not whether it mattered.

Good answer

They report traffic growth rate by cluster (not total traffic), conversion rate by article, and citation frequency. They can explain why traffic growth rate matters more than absolute traffic (because it reveals whether content is compounding or plateauing) and why conversion rate by cluster matters more than site-wide conversion (because different clusters serve different buying stages).

Red flag

A dashboard with 15 metrics that all boil down to "traffic went up this month." More metrics is not better measurement — it's obfuscation that makes everything look like progress.

QUESTION 10

"How do you measure content's contribution to pipeline and revenue?"

This is the hardest question in content measurement. Most agencies can't answer it. The ones that can are the ones you want.

Good answer

They have a methodology: multi-touch attribution with content as a touch (not the only touch), first-touch attribution for top-of-funnel content, assisted conversion tracking. They're honest about the limitations — content is rarely the only touch in a B2B sale, so single-touch attribution overstates or understates its contribution.

Red flag

"Content builds brand awareness. It's not a direct-response channel." Translation: we can't measure it, so we claim it's unmeasurable. Content can and should be tied to revenue — it just takes more instrumentation than most teams have.

QUESTION 11

"What's the lead time before we see attributable results?"

Agencies that promise results in 90 days are selling you SEO shortcuts that work temporarily and then backfire. Content is a compounding channel — the timeline is real and should be communicated honestly.

Good answer

6-9 months for first attributable leads, 12-18 months for consistent pipeline contribution, 24-36 months for content to become a primary revenue channel. They explain the compounding curve — months 1-6 are building the architecture, months 6-12 are the acceleration phase, months 12+ are when the compounding becomes visible.

Red flag

"We typically see results within 90 days." Translation: we're going to target low-competition keywords with thin content that ranks quickly and converts nothing. The traffic chart will go up. The pipeline will not.

QUESTION 12

"If we stopped working together after 12 months, what would we keep?"

This is the question most buyers skip. It reveals whether the agency builds assets or dependency.

Good answer

You keep the content library, the topic cluster architecture, the editorial calendar framework, the measurement dashboard structure, and the distribution playbook. The agency transfers operational knowledge during the offboarding. The goal is for you to be more capable after the engagement than before it.

Red flag

"You keep the content we produced." Translation: you keep the articles but lose everything that made them work — the strategy, the distribution system, the measurement framework. You'll be back in 3 months because the content will plateau without the operating system behind it.

The one question that matters most

If you only have time for one question, make it Question 02: "Show me a piece of content you recommended killing." The answer reveals more about how the agency thinks than any other single response. An agency that has never recommended killing a topic is an agency that has never made a strategic judgment. They're taking orders, not building outcomes.

87/39
87% report AI productivity gains. Only 39% see performance improvement. The gap is not about tool access — it's about whether the agency uses AI to accelerate thinking or to replace it.
CMI, 2026

The agency you choose determines whether your content compounds or flatlines. The 12 questions above separate the two outcomes in a 30-minute call. Use them.