Resources
Article GTM Strategy 10 min read

AI Sales Tools: Why Revenue Teams Get Stuck

Clay, Claude, Codex, agents: why AI sales tools stall in revenue teams, and a framework that decides which tool does which job, built bottom up.

Explore with AI
Share on LinkedIn

Key takeaways

  • Most revenue teams already use AI; very few turn it into measurable results, because they buy tools before they know which job each tool is for.
  • Clay, Claude, Codex and AI agents are not substitutes: they do different jobs in the revenue system.
  • The Revenue AI Stack sorts every tool into four layers (Store, Find, Think, Act), plus a build rail on the side and humans on top.
  • Build bottom up: nothing above the CRM works if the HubSpot record underneath it is incomplete.
  • Our recommendation: fix one bottleneck at a time and count the meetings a workflow creates, not the emails it sends.

Everyone is using AI. Few are getting paid for it.

Adoption is no longer the question; impact is. Almost every sales organisation now uses some form of AI, yet only a minority can point to a financial result, and the tools that promised to replace entire sales roles are the ones teams cancel fastest.

What the research foundSource
87% of sales organisations use AI, and 54% of sellers have already used AI agentsSalesforce, State of Sales 2026
88% of organisations use AI, but only 39% see any EBIT impact and nearly two-thirds haven't started scalingMcKinsey, The state of AI in 2025
AI SDR tools see 50–70% annual churn, as teams that bought full rep replacement cancel within monthsRework, Are AI SDRs worth it? (2026)

The gap isn't access to AI. It's turning AI into a working revenue system.

Why is implementing AI in sales so hard?

Because most rollouts start in the wrong place. Teams pick a tool before they have named the problem, run it on data nobody trusts, and then judge it by activity instead of pipeline. Each of the seven pain points below is common on its own; together they explain why so many AI projects stall after the pilot.

  1. Everything is called "AI". Clay, Claude, Codex and AI SDRs do different jobs, yet they get pitched and compared as if they were substitutes.
  2. Tools before problems. Teams start with "we should use Clay" instead of "our reps lose too much time researching every account."
  3. Broken data. In SalesPlaybook client projects, 40–60% of CRM records are often incomplete or outdated, according to our B2B Go-To-Market Efficiency Guide 2026. AI just runs bad data faster.
  4. Stack sprawl. Every new tool adds another login, another copy of the data and another sync that breaks.
  5. Nobody owns the system. The GTM engineer who connects it all is a new role most mid-market teams don't have. A Claude Code script on one laptop isn't a system.
  6. Automating a weak motion. A message that fails when sent by hand fails faster at scale, and burns your sending domains on the way.
  7. Measuring activity. Emails sent and rows enriched look good in a dashboard and say nothing about pipeline.

How do you decide which AI tool does which job?

Sort every tool by the job it does in the revenue system, then build in order. We call it the Revenue AI Stack: four layers, a build rail on the side and humans on top. Once each tool has one job, the question stops being "Clay or Claude?" and becomes "which layer is missing?"

Human layer

Your reps and CSMs: discovery, demos, negotiation

Approve what goes out, own strategic accounts, build trust

4 · Act

Lemlist · HubSpot workflows · AI agents

Route leads fast, run signal-triggered sequences

3 · Think

Claude (chat, API, inside Clay)

Briefings, lead scoring, context-based copy

2 · Find

Clay · Claygent · Fibbler · BuiltWith

Enrichment, TAM mapping, buying signals

1 · Store

HubSpot: single source of truth for Marketing, Sales, CS

Every layer reads from and writes back to it, no duplicates

▲ Build order: bottom up

Build rail

AI coding agents, workflow tools, APIs

Connect the layers: prototype fast, then run what the team relies on in a shared platform.

1 · Store: HubSpot as the single source of truth. One platform for Marketing, Sales and Customer Success, so every contact, deal, signal and conversation lives in one record. Every other layer plugs into it: Clay enriches into HubSpot, Claude and Jev read from it, and sequences and workflows run off it. If that record is 40–60% incomplete, nothing above it can be trusted. This is the foundation, and it's where most AI projects quietly fail.

2 · Find. Clay and its research agent Claygent pull firmographics, technographics, contacts and signals from many sources. The output is not a list; it's a trigger: a champion changed jobs, a target account is hiring, a prospect is migrating to S/4HANA.

3 · Think. Claude turns raw data into judgment and language: a short account briefing, a lead classification, a first line that references the actual signal. The structure of a message is a template; the relevance lives in the variables, and that's exactly what an LLM is good at filling.

4 · Act. Sequencers, HubSpot workflows and agents execute: route the lead in minutes, run the multi-touch sequence, stop on reply, hand over to a rep with the briefing attached.

Build rail. Claude Code and Codex are not sales tools. They're how a GTM engineer or technical RevOps person builds and maintains the glue between layers. Use them to prototype fast, then move anything the team relies on daily into a platform.

Human layer. Automation doesn't close mid-market or enterprise deals; people do. The system collects, scores and prioritises. Humans hold the conversation, build trust and steer the system.

Our take: we build the stack in this order because every layer inherits the quality of the one below it. A clean record is cheaper than any agent you could put on top of a broken one.

Three rules that make the stack work

  • Build bottom up. Don't buy an agent (Act) before HubSpot (Store) is clean. Every layer inherits the quality of the one below it.
  • Everything writes back to HubSpot. Research, signals, decisions and replies land on the HubSpot record, not in a Clay table, a spreadsheet or someone's inbox.
  • Match autonomy to deal size. Under roughly €10k ACV, automate heavily. From €25k to €250k, run a hybrid model where AI does research and follow-up and humans own demos, closing and complex objections. Above €250k, AI handles admin and research while relationships stay fully human. The ACV bands come from the SalesPlaybook Go-To-Market Efficiency Guide.

How have B2B teams actually used AI in sales?

By fixing one bottleneck each, not by launching an AI programme. No pilots, no "AI strategy" decks. The three teams below each picked the single constraint that held back pipeline, put AI on the research or the first touch, and kept people in the conversation. Each result arrived in under 6 months.

node.energy: from 5–6 to 30–35 demos a month in under 3 months

Before. The Frankfurt SaaS platform for the energy sector ran outbound through classic channels and got almost no response. SDRs called wrong contacts, lacked phone numbers and spent hours researching each account.

How AI was used. Outbound was rebuilt around signals with Clay integrated into HubSpot. Automatic enrichment now finds the qualified account, the reason to reach out right now and the direct decision-maker, phone number included. SDRs receive hundreds of qualified leads ready to call instead of researching them.

Result. Demos grew 5–6x in under 3 months, and outbound demos now achieve a higher win rate than inbound demos.

The lesson: AI took over the research, not the conversation. SDRs got their time back for calling.

Read the node.energy case →

knk Software: a standing outbound pipeline in under 6 months

Before. knk, a software vendor for publishers, wanted to move from resource-heavy enterprise deals to a scalable model for smaller publishers. It had no outbound motion: new business ran through existing email lists, and workflows were manual.

How AI was used. Signal-based outbound with an AI-driven first touch that runs around the clock. The system handles research and the first message. Reps step in once a prospect engages.

Result. A working outbound pipeline in under 6 months, without reps doing first contact by hand.

The lesson: Smaller deals need a lower cost per touch. Automating the first touch is what made the new segment profitable, while people still own the conversations.

Read the knk case →

Workist: about €200,000 saved a year, and a foundation AI can run on

Before. The Berlin AI-SaaS company (about 50 employees, 150+ paying customers) had outgrown its Salesforce-centred stack. It paid over €100,000 a year in licences, needed a full-time role just for maintenance, and had no end-to-end view of the customer journey.

How AI was used. HubSpot became the single platform for marketing, sales and customer success in under 3 months, replacing Salesforce, Zendesk and other tools. Clay plugs straight in, and backend systems connect through webhooks, so enrichment and AI research write into one clean system.

Result. About €200,000 saved per year, the maintenance role gone, and pipeline questions answered in minutes instead of weeks.

"We can now also see in real time whether we are on course." – Markus Bimüller, Workist

The lesson: Every AI workflow is only as good as the system it reads from. Consolidating first made everything after it cheaper and faster.

Read the Workist case →

Which layer of your stack is holding back pipeline?

Free · 60 minutes · no pitch · a clear fit or no-fit answer.

Book Strategy Call

When does an AI step not need an LLM at all?

Whenever the step is a quick judgment with a fixed set of answers. Most AI steps in GTM aren't writing tasks: is this account in our ICP, which segment is this lead, is this reply interested or not. Running an LLM for each of those works, but it's slow and expensive for a yes-or-no job.

It's like hiring a bartender to check IDs at the door. Jev, from TypeSafe AI, is built for exactly that job. It became available in September 2026, starting in early access. It doesn't generate text. You send it input plus the questions you define, and it returns a typed answer (a choice, a score or a yes/no) with a confidence score. All questions are answered in parallel, in one call.

LLM (e.g. Claude)JevSource
OutputOpen-ended textA typed decision + confidence scoreTypeSafe AI
SpeedSeconds70–500 ms end to end, per TypeSafeMarkTechPost
CostFull generation on every call$0.042 per million input tokens, output free (launch pricing)TypeSafe AI
Best forWriting, summarising, reasoning, complex judgment callsClassification, routing, filtering, scoring at high volumeSalesPlaybook

Where it fits in GTM. Any step with a fixed set of answers is a candidate:

  • ICP fit yes/no on every account before you spend enrichment credits
  • Inbound lead routing by segment, intent and deal size
  • Reply triage: interested, not now, wrong person, unsubscribe
  • Signal filtering: is this news actually relevant to what we sell?
  • Job titles mapped to persona buckets

A Clay example: a use case generator. One build currently in progress in Clay works like this:

  1. Ingest a company's website copy.
  2. Jev classifies which GTM motions the company runs, with a probability for each.
  3. An LLM matches those motions to a library of use cases.
  4. Relevant use cases appear within seconds.

Running the whole chain through an agent would be slower and more expensive. Jev makes the decision; the LLM only does the part that needs language. Any marketer could build the same for their own product to show prospects how to use it.

One caveat. Speed and pricing are TypeSafe's launch claims. "Can't hallucinate" means Jev can't answer outside your schema, but it can still be wrong (MarkTechPost). Use the confidence score: act automatically when it's high, and pass low-confidence cases to an LLM or a human.

The rule of thumb: use Jev to filter and route at scale. Use an LLM when the work needs thinking and writing.

What are the mistakes that sink most AI rollouts?

They're rarely about the tool. Most failed rollouts automate a message that never worked, build on data nobody cleaned, add complexity before anything is proven, or celebrate activity that never turns into meetings. Each of the four mistakes below can be caught before the first sequence goes live.

Diagram: the four mistakes that sink AI sales tool rollouts – a weak message, bad data, over-engineering and measuring the wrong thing
  • Automating a weak message. If it doesn't land when sent by hand, it won't land at scale. Fix positioning first.
  • Building on bad data. A signal system is only as good as the CRM it reads from.
  • Over-engineering before validating. Prove the simple version surfaces real opportunities, then add complexity.
  • Measuring the wrong thing. Activity metrics hide failure. Count the meetings the workflow creates.

If you want this built for your own outbound, our AI outbound and pipeline generation pages show how we set it up.

Get our GTM Efficiency Guide

The Revenue AI Stack shows which tool does which job. How to build it step by step in your own go-to-market is the subject of the playbook behind this post.

Cover of the SalesPlaybook B2B Go-To-Market Efficiency Guide 2026

Free guide · PDF · German edition

B2B Go-To-Market Efficiency Guide 2026

  • Six use cases, from CRM consolidation to automated outreach
  • Step-by-step solution designs and the exact tech stack we use
  • A readiness check with scorecard to find your biggest lever in under 10 minutes
Download the free guide

Sequence first, tools second

Start by checking how complete your HubSpot record is, then pick the one bottleneck that holds back pipeline and give it a single layer of the Revenue AI Stack. That order beats any new AI sales tool you could add on top. In a Strategy Call we go through your stack layer by layer and tell you honestly where the first lever is.

Free · 60 minutes · no pitch · a clear fit or no-fit answer.

Authors Florian Lussi

Frequently asked questions

What is the Revenue AI Stack?
The Revenue AI Stack sorts every AI sales tool by the job it does: Store (HubSpot as the single source of truth), Find (Clay and Claygent), Think (Claude), Act (sequencers, workflows, agents), plus a build rail for Claude Code and Codex and a human layer on top. You build it from the bottom up.
Which AI sales tool should a revenue team set up first?
None of the AI tools comes first. The CRM record does. If contacts, deals and signals are incomplete or scattered across tools, every enrichment, briefing and sequence built on top inherits those gaps, so clean up and consolidate HubSpot before you add Clay, Claude or an agent.
What does a GTM engineer do?
A GTM engineer connects data, tools and workflows across the revenue system. They use tools like Claude Code, Codex and APIs to prototype the glue between layers, then move anything the team relies on every day into a shared platform so it doesn't live on one laptop.
When is an AI SDR worth it?
It depends on deal size. Under roughly €10k ACV you can automate heavily. Between €25k and €250k a hybrid model works better: AI handles research and follow-up, humans own demos, closing and complex objections. Above €250k, AI supports admin and research while relationships stay human.

Resource download

Unlock the download.

Enter your email address to start the download of “AI Sales Tools: Why Revenue Teams Get Stuck”.

PDF, immediately after submission · No newsletter, no call

We process your email address in HubSpot to provide this resource. See the privacy policy for details.

Could your company be the next operating system story?

Use a free 60-minute Strategy Call to clarify the revenue constraint, fit or no fit and the right next step. No pitch.

Book Strategy Call