Picking Your First Agentforce Use Case: What Actually Predicts Whether It Survives
Broad, ambitious first Agentforce use cases fail for a predictable reason: they were never one job. Here's how to test a candidate use case before you spend a single Flex Credit building it.
What you’ll learn
- Broad topics like "customer service" routinely produce far higher escalation rates than narrow ones like "refund processing" - the model isn't the variable, the scope is.
- Data Cloud is not a prerequisite for a first Agentforce pilot. CRM-grounded agents cover most first use cases; Data Cloud earns its place for unstructured content or cross-system identity, not by default.
- Flex Credits pricing (20 credits, $0.10, per standard action) makes a narrow use case cheap to test and a broad one expensive to quietly abandon.
- Seven questions to score a candidate use case before anyone builds anything, plus who should actually own the agent once the pilot window closes.
The most common way a first Agentforce pilot dies isn't a weak model - it's a topic called "customer service" or "general enquiries" that was actually six different jobs wearing one label. Salesforce's own published guidance on writing agent instructions says as much directly: broad, vaguely-scoped topics and actions produce agents that misclassify intent and escalate constantly, while narrowly defined ones - "refund processing," not "customer service"; "order status," not "order help" - stay reliable. The wider AI-pilot data says the same thing at a larger scale. MIT's 2025 State of AI in Business study, built on interviews and analysis across roughly 300 public enterprise AI deployments, found that 95% of generative AI pilots delivered no measurable P&L return - and traced the failure to undefined outcomes and data readiness before anyone built anything, not to model quality.
The short answer
In brief: pick the narrowest job you can find that already has clean, accessible data, a human doing it today who can define what "correct" looks like, and a low-stakes way to fail. That's a stricter filter than "pick something valuable" - a valuable but broad job (full customer service, full sales enablement) is exactly the shape that produces high escalation and gets quietly shelved before it proves anything. Score two or three candidates against the seven questions below before committing to one. If nobody in your org currently has a confident, technical read on which data is actually clean enough to ground an agent, that's the same gap a Salesforce Health & Roadmap review is built to surface - it now explicitly covers Agentforce and AI readiness alongside architecture and data quality, rather than treating AI readiness as a separate exercise.
Why the use case matters more than the model
This is a scoping problem dressed up as an AI problem, and the pricing model now makes the difference expensive to ignore. Since May 2025, Agentforce has priced standard actions under Flex Credits at 20 credits each - $0.10 per action, sold in blocks of $500 per 100,000 credits - alongside the older per-conversation and per-seat models. A narrow, well-scoped pilot with a handful of clean actions is genuinely cheap to run and cheap to kill if it doesn't work. A broad agent wired into a dozen loosely-defined actions racks up consumption on every misclassified intent and every unnecessary escalation loop, and by the time someone notices the pilot isn't working, there's a credit bill and a half-built integration to unwind, not just a decision to make. Scope discipline isn't a nice-to-have here - it's what keeps a failed first pilot cheap instead of expensive.
95%
of GenAI pilots showed no measurable P&L return - MIT, 2025
$0.10
cost per standard Agentforce action under Flex Credits (20 credits)
$500
per 100,000 Flex Credits - the unit Salesforce sells them in
Seven questions that predict whether a use case survives
None of these questions are about whether the use case would be valuable if it worked - almost every candidate on a first list clears that bar. They're about whether it's scoped narrowly enough, and grounded cheaply enough, to actually get a fair test. Run every candidate through all seven before building anything.
1. Does the job description fit in one sentence without an "and"?
"Look up an order and explain the delivery status" is one job. "Handle customer service" is not - it's shorthand for order status, returns, billing disputes, account changes, and product questions, each with a different data source and a different definition of correct. If the description needs an "and" to cover what people actually mean by it, it's not one topic yet. Split it, and pick the narrowest piece as the pilot.
2. Do you already have the data, or are you reaching for Data Cloud out of habit?
Agentforce does not require a Data Cloud project for a first pilot. An agent that reads and writes CRM records is covered by standard actions, Flow, Apex, and prompt templates with record merge fields - the same grounding a page layout already has. Data Cloud earns its place when an agent has to answer from unstructured content (documents, transcripts, knowledge outside structured fields) or needs one identity-resolved profile across systems that don't share a key. Reaching for Data Cloud on a first pilot because it's the newest name in the room usually means signing up for an identity-resolution project before you've proven the underlying use case is worth automating at all.
3. Is there a human doing this today who can define what "correct" looks like?
If nobody can say, concretely, what a right answer looks like for this job today, an agent can't be scored against it either - you'll be debating vibes instead of a pass rate. This is the same gap that shows up industry-wide as "unclear success criteria," and it's specific to each candidate use case, not a one-time setup step: a rep who currently does deal-health scoring from memory can tell you when the agent's summary is wrong; nobody can do that for a job that's never existed as a defined task before.
4. What happens when it's wrong, and who notices first?
A wrong order-status lookup gets corrected by the customer asking again. A wrong account-tier change made without review gets discovered by finance, weeks later, across every account the same misconfigured action touched. Score the blast radius of a wrong answer, not just its likelihood - a low-probability, high-blast-radius action is a worse first pilot than a higher-probability, low-blast-radius one, even though the failure rate alone would suggest the opposite.
5. Is the first action reversible?
Start with read actions, or write actions gated behind explicit confirmation, before handing an agent an unsupervised write path. This isn't caution for its own sake - it's what lets a pilot fail cheaply and informatively instead of expensively and ambiguously. An agent that gave a wrong answer nobody acted on yet, has produced a data point. An agent that already updated forty records incorrectly has produced a cleanup project.
6. Does it touch a system nobody fully owns?
An agent that has to reconcile answers across Salesforce and a poorly-documented integration inherits every ambiguity already sitting in that connection - the same ownership gap that quietly breaks integrations without any AI involved at all will just as quietly produce a confidently wrong agent answer instead. A first pilot grounded entirely inside data Salesforce already owns cleanly is a fairer test of the agent than one that also has to survive an upstream system nobody's fully mapped.
7. Who owns this after the pilot window closes?
A successful pilot with no named owner doesn't get scaled - it gets left running until someone notices its answers have drifted out of date with a data model that's since changed underneath it. Assign ownership before the pilot starts, to a role rather than a name, the same discipline that matters for any other piece of platform infrastructure nobody wants to become a single point of failure for.
| Candidate first use case | What the seven questions reveal | Verdict |
|---|---|---|
| Full "customer service" inbox agent | Several unrelated jobs sharing one topic label, no agreed baseline for "correct," no single owner | Don't start here - split into narrow topics and pilot one |
| Order-status lookup for one channel | One-sentence job, CRM-only data, an existing baseline (current deflection rate), a reversible read action | Strong first candidate |
| Sales meeting-prep summaries grounded in CRM history | Narrow job, data already inside Salesforce, reps already do this manually and can judge quality directly | Strong first candidate |
| Cross-system deal-health scoring | Needs an identity-resolved profile across systems that don't share a key - a Data Cloud modelling project, not a pilot | Defer until the Data Cloud work is scoped as its own engagement |
Candidate list
every "we could automate this" idea from support, sales, or ops
One-sentence test
job description with no "and"
Data check
CRM-only, or does it genuinely need Data Cloud
Reversibility check
read action first, write action behind confirmation
Named owner
assigned before the pilot starts, not after it succeeds
From the field
CloudAvant's own multi-agent Agentforce and Data Cloud work followed this discipline directly: agents for account intelligence, order lookup, and sales activity summaries were grounded in CRM data via Prompt Templates and Apex-backed actions and scoped to one well-defined business question each, rather than shipped as a single general-purpose assistant. Data Cloud was brought in only where a question genuinely needed a unified profile or external retrieval - the public website agent and the logged-in e-commerce agent, each grounded in unified customer profiles - not applied everywhere by default. See the Agentforce and Data Cloud engagement. A related pattern shows up in the Brenntag Service Cloud work, where Agentforce actions were built to pull account history and survey sentiment directly onto a case rather than functioning as an open-ended assistant - a scoped action inside an existing workflow, not a new surface to govern on its own.
What we're advising clients right now
Run the candidate list past the seven questions before a single action gets built, not after the pilot has already stalled. Most orgs that come to us mid-pilot aren't struggling with the model - they're trying to rescue a use case that was too broad, too unowned, or grounded in data nobody had actually checked. Where a first pass through those questions surfaces real uncertainty about data quality or governance, that's exactly the gap a Salesforce architecture review is built to close before licensing spend and integration work are already sunk into the wrong first pilot. If you'd rather get a second, independent read on a specific candidate use case before committing engineering time to it, that's worth doing early, not once the pilot's already stalled.
Questions worth answering before you build anything
Does Agentforce require Data Cloud?
No, not for a first pilot. An agent that reads and writes CRM records is fully covered by standard actions, Flow, Apex, and prompt templates - no Data Cloud modelling required. Data Cloud becomes necessary when an agent has to ground answers in unstructured content or needs one identity-resolved profile spanning systems that don't share a common key. Confirm which case you're actually in before scoping a Data Cloud workstream into a first pilot.
How much does a narrow first pilot actually cost?
Under Flex Credits, a standard action costs 20 credits - $0.10 - and credits are sold in blocks of $500 per 100,000. A genuinely narrow pilot (one topic, a handful of actions, a few hundred conversations a month) can be tested for a fraction of a single credit block. That math is exactly why scope discipline matters: the cost of testing a narrow use case is close to trivial; the cost of running a broad, loosely-scoped one long enough to notice it isn't working is not.
What if the first pilot fails?
That's a reasonable outcome, not proof the technology doesn't work - provided the pilot was scoped narrowly enough that "failed" means something specific: this job, grounded in this data, didn't clear the bar a human doing it today already sets. A narrow pilot that fails is informative and cheap to retire. A broad pilot that fails is usually ambiguous about which part of it failed, and expensive enough that shelving it quietly, rather than learning from it, becomes the path of least resistance - which is a large part of why the industry-wide pilot failure rate is as high as it is.
The bottom line
None of this is an argument against Agentforce, and it isn't an argument that every org needs an AI pilot running this quarter either - some don't, and saying so is a more useful answer than manufacturing a use case to justify the licensing spend. It's an argument that the difference between a first pilot that survives and one that gets quietly abandoned is decided before a single action gets built: in how narrowly the job is defined, whether the data it needs already exists and is trustworthy, and whether a reversible first action gives it a fair, cheap test. Get that right and the model has an actual job to prove itself on. Get it wrong, and no amount of prompt engineering fixes a topic that was never one job to begin with.
Sources and further reading
- MIT Finds 95% Of GenAI Pilots Fail Because Companies Avoid Friction - Forbes
- Salesforce Introduces New Flexible Agentforce Pricing to Accelerate the Digital Labor Revolution - Salesforce
- Agentforce Pricing - Salesforce Help
- How to Write Effective Natural Language Instructions for Agentforce - Salesforce Developers Blog
- Connecting Agentforce to Data Cloud for Grounding With RAG - Salesforce Ben
- Grounding Agents with Data: Key Mechanisms Explained - Salesforce Trailhead
Not sure your first Agentforce use case is scoped to survive its pilot?
Written by CloudAvant Team - Senior Salesforce consultants and architects with hands-on enterprise delivery experience across Sales Cloud, Service Cloud, Experience Cloud, and complex multi-region implementations. More about CloudAvant.
