A CMO told us in April that her team had bought three products with the word “agent” in the name and none of them had moved a single number on her board deck. She wasn’t wrong about the products. She was wrong about what she expected them to do.
That gap is most of what’s happening with AI agents for sales right now. The category is real, the buying is running ahead of the operating knowledge, and a lot of teams are paying agent prices for what is functionally a scheduled job with a language model bolted to the front of it. So before anything else, the definition needs to be tight enough to be useful in a budget conversation.
An agent decides. A workflow just runs.
Your team already runs automation. LeanData routes leads on rules you wrote. HubSpot fires a workflow when a form hits a threshold. Outreach steps a prospect through a nine-touch sequence whether or not that prospect has done anything interesting. Clay enriches a row the moment it lands. All of this is deterministic. Someone sat down, mapped the if/then, and the system executes it identically every time.
An agent is different in one specific way: it is given a goal and it chooses the steps.
Concretely, a workflow says “when the lead score crosses 80, assign to the territory owner and enroll in sequence 14.” An agent says “determine whether this account resembles our closed-won profile, and if it does, assemble the account brief and put it in front of the AE who owns that patch.” To get there it might pull firmographics from ZoomInfo, check 6sense intent, read the last two Gong calls with a similar account, and query SFDC for existing relationships. Nobody wrote that sequence in advance. The agent picked it, ran it, evaluated what came back, and ran more.
That’s the upside and the risk in the same sentence. Workflows fail loudly. An agent fails quietly and with total confidence, and it writes the failure to your CRM. Hold that thought, because it comes back later.
Where agents are actually earning their keep
The honest answer is that agents are working today in exactly the places where the output is checkable and the cost of being wrong is one wasted minute.
Account research is the clearest win we’ve seen. A BDR carrying 90 named accounts in a $60K ACV motion spends somewhere between six and eight hours a week on manual research: funding, headcount changes, tech stack, who moved where, what the 10-K says about the initiative you sell into. An agent that reads across ZoomInfo, Demandbase, LinkedIn, and the company’s own site and returns a two-paragraph brief with sourced claims gives most of that time back. Not all of it. The rep still reads the brief and throws out the third of it that’s noise.
Meeting prep is the same shape of problem. Pull the last four touches from Salesloft, the last call summary from Gong, the open opportunity fields from SFDC, and drop a prep note into Slack fifteen minutes before the call. That’s genuinely multi-step, it requires deciding what matters and what doesn’t, and if the agent misses something the AE catches it in the room.
The rest of the working set:
- Enrichment and data hygiene, where an agent reconciles duplicates, fixes account hierarchies, and normalizes the industry field that seven years of reps have destroyed
- First-pass list building in Clay or Apollo, producing a candidate list a human then cuts by 40 percent
- CRM updates from call transcripts, where the agent proposes field changes and next steps rather than committing them
- Summarizing calls and threads into something a CMO can read in the Monday pipeline review
- Anomaly flagging against Clari, where the agent notices a stalled stage-3 deal before the forecast call does
Notice the pattern. Every one of these is retrieval, synthesis, and structured output. None of them is a judgment call that lands in front of a customer.
The work agents are still bad at, and it is not getting fixed this quarter
Outreach is the big one. Raw agent-written email reads like agent-written email, and buyers pattern-match it in about a second and a half. We’ve watched teams push fully generated sequences and land at 0.4 percent reply rates against a 2 to 3 percent human baseline, then respond by increasing volume, which makes the domain reputation problem worse than the reply rate problem ever was.
The failure is not grammar. Agent-written copy is grammatically cleaner than what your SDRs send. It fails because it has no point of view. It observes that the prospect raised a Series C and that scaling is presumably a priority, which is true of every Series C company that has ever existed and is therefore worth nothing.
Account fit judgment is the second gap. An agent can tell you an account matches the ICP on paper. It cannot tell you that this particular logo burned your CS team two years ago, or that the new CRO came from a competitor and will not take the meeting, or that the deal is technically qualified and strategically pointless. That knowledge lives in the heads of four people on your team and in none of your systems.
And then there’s taste. Which of these 60 accounts deserves the executive dinner. Whether this quarter’s message should lead with cost or with risk. What to cut from the board deck. Agents produce competent, average output at volume, and average is precisely the thing that loses competitive deals.
Where the human actually belongs
The two popular answers are both wrong. “Human reviews everything” means you’ve built a slower version of the manual process and added a review queue nobody staffs. “Human reviews nothing” means you find out about the problem from a customer.
The human belongs at two fixed points, and mostly nowhere else.
First, at the judgment gate. The agent builds the list, the human cuts it. The agent scores fit, the human overrides on the accounts where institutional memory beats the data. This step takes a RevOps lead or a senior AE about twenty minutes for a 200-account list and it’s the highest-value twenty minutes in the whole workflow.
Second, at the point of contact. Anything a prospect will read gets a human hand on it. In practice this doesn’t mean writing from scratch. It means the agent assembles the research and the structure, and the rep writes the first two sentences and the ask. Those are the parts that carry a point of view. Everything between them is scaffolding and the agent can have it.
We go deeper on the sequencing of this in When an AI Marketing Agent Makes Sense, because the calculus shifts once you’re talking about campaign-level output instead of one-to-one motion.
Who owns the agent when it misfires
Ask a room of GTM leaders who owns the agent and you get silence, then everyone looking at RevOps.
RevOps is usually the right answer, but only if it comes with actual authority. An agent that writes to SFDC is a system-of-record change, which means it belongs to whoever owns the system of record, not to whoever bought the tool. We’ve seen Demand Gen stand up an agent inside HubSpot that quietly rewrote lifecycle stages for six weeks, which broke the MQL-to-pipeline reporting the CMO was presenting to the board. Nobody had done anything wrong. Nobody owned it either.
Three things need a name attached before an agent touches production. Who reviews its output on a set cadence. Who can shut it off inside five minutes without filing a ticket. And whose number is affected if it’s wrong, because that person will find out first regardless of what the org chart says.
This is a staffing question as much as a governance one, and it’s the reason the GTM engineer role has moved from novelty to line item, which we cover in What Every B2B Leader Should Know About GTM Engineering. Somebody has to own the connective tissue between the tools, and that someone is not a marketer with a Zapier login.
Writing to your system of record is a different risk than reading from it
Read-only agents are close to free from a governance standpoint. Let them read everything. Give them Gong, SFDC, Clari, the data warehouse, the website analytics.
Write access is where teams get hurt. Our default is that agents propose and humans commit, at least for the first quarter. Practically that looks like writes landing in a staging object or a custom field, with a source stamp and a confidence value, and a human or a rule promoting them into the fields your forecast actually depends on. Stage, amount, close date, and account owner should be the last four fields you ever hand over.
Keep an audit trail that distinguishes agent writes from human writes. When the pipeline number looks strange in month three, and it will, that field is the difference between a twenty-minute diagnosis and a two-week one. This connects directly to the reporting integrity problem we cover in AI GTM: What Changes When AI Runs Inside Your Go to Market Motion, and it’s the same Unified Funnel argument we’ve been making for years: if you can’t trace where a number came from, you can’t defend it in front of the board.
The SDR replacement claim
We’d push back on it hard.
The claim that AI agents replace SDRs assumes the SDR job is research plus sending. It isn’t. Research and sending are the visible parts, and agents are eating a real share of both. The job that’s left is the harder part: reading a lukewarm reply and deciding whether to press or wait, getting a gatekeeper to route you, holding a five-touch conversation across three weeks without sounding like a follow-up robot, and knowing when an account is dead.
What we do see is the ratio changing. Teams that ran eight SDRs per pod are running five, with those five carrying larger territories and spending their week on conversations instead of tab-switching. Pipeline per rep goes up. Headcount goes flat or down slightly. That’s a meaningful shift and it’s worth planning for, but it isn’t replacement, and the vendors selling it as replacement are selling to the CFO, not to you.
The teams getting the most out of this right now are the ones who picked two workflows, instrumented them properly, and left everything else alone. What have you actually put into production, and where did it break?