Measured AI Agents Are the Ones That Scale

The AI agents that get measured are the agents that get to scale, and the encouraging part is that the measurement itself is far more within reach for CRM teams than the current adoption numbers suggest.

SAP published its Value of AI Report in July 2026, drawing on 2,600 business leaders across 13 countries. Average AI return has reached 21 per cent, roughly 6.3 million dollars, up from 16 per cent a year earlier. Sixty-nine per cent of businesses said they were satisfied with the returns they are already seeing. The number that matters most for anyone planning next year’s roadmap is the forward one: expected returns from agentic AI specifically are projected to reach 17.6 million dollars within two years, up from 4.3 million in last year’s estimate. That is not a speculative curve. It reflects organisations that moved past pilots and started counting what came back.

A second study, published by Liferay on 12 August 2026 and based on 500 professionals involved in AI decisions, sharpens the picture. Fifty-four per cent of companies are running agents in production or actively piloting them, while 25 per cent measure the impact against clear KPIs. Read one way, that is a gap. Read practically, it is the most affordable competitive advantage available in enterprise AI this year, because the quarter of teams who can prove what their agents return are the ones who get to fund the next ten.

Why measured agents are the ones that get funded

There is a pattern worth noticing in how agent programmes grow inside large organisations. The programmes that expand are rarely the ones with the most sophisticated technology. They are the ones whose sponsor can walk into a budget conversation with a number. When a revenue leader can say that the qualification agent handled 4,100 inbound enquiries last quarter, that response time on those enquiries fell from nine hours to eleven minutes, and that conversion on agent-handled leads held steady at 8 per cent, the follow-on investment becomes an easy decision rather than an act of faith.

This is the quiet reason the measurement quarter is pulling ahead. Agent capability is now broadly available across Salesforce, HubSpot and Dynamics 365, so capability is no longer what separates fast movers from slow ones. Evidence is. A team that can show its work compounds every quarter, because each proven agent creates the mandate and the budget for the next one. Gartner reported that 80 per cent of enterprise applications shipped or updated in the first quarter of 2026 embed at least one agent, up from 33 per cent in 2024, which means the raw material is already sitting in most technology stacks. The organisations that instrument it early get to compound sooner.

How do you measure the ROI of a CRM AI agent?

Measuring the ROI of a CRM AI agent means comparing a defined before-and-after on a specific process, not estimating value across the whole platform. Pick one workflow the agent owns, such as lead qualification, case triage or quote generation. Capture four things for the 90 days before the agent went live: volume handled, average cycle time, downstream conversion or resolution rate, and fully loaded cost per item. Measure the same four for 90 days after. The difference in cycle time and cost gives you efficiency return, the difference in conversion or resolution gives you effectiveness return, and the two together produce a figure that survives scrutiny from a finance team. Subtract platform consumption costs to reach net value.

The reason this approach works better than a platform-wide ROI model is that it is falsifiable. A single workflow with a clean baseline can be checked by anyone. A blended estimate across eleven agents and three business units cannot, which is why those estimates rarely survive their first serious budget review. Start narrow, publish the result, then widen the frame once the method has credibility.

What does a good agent baseline look like, and where does it come from?

A good baseline is a set of numbers from the period immediately before the agent went live, drawn from systems that were already recording them. For most CRM workflows this means activity timestamps, case or opportunity stage history, response time fields and outcome flags. CRM data quality determines how usable that baseline is: complete, consistently entered records produce a baseline you can defend, while sparse or manually overridden fields produce one that invites argument. The practical test is whether two colleagues querying the same period independently arrive at the same figure. If they do, you have a baseline worth building on.

This is where a lot of teams discover good news they did not expect. Salesforce, HubSpot and Dynamics 365 all retain field history, activity records and stage transitions by default, which means most organisations already hold the before-picture without knowing it. The work is usually a matter of days spent writing the right reports, not months spent building new data capture. SAP’s research found that 73 per cent of companies report incomplete data affecting their AI readiness, and the constructive response is to treat baseline work as the first cleanup with an obvious payoff attached, since every field you tidy in service of measurement also improves what the agent itself can reason over.

Measurement is what turns agent adoption into agent habit

Improving CRM adoption rates has always depended on showing people that the system gives back more than it takes, and agent measurement makes that case unusually easy to make. When a sales team can see that the agent returned an average of 6.4 hours per rep per week, or that agent-drafted follow-ups are opened at a higher rate than manually written ones, adoption stops being a mandate and becomes a preference. The most effective adoption programmes now publish these numbers back to the users who generated them, monthly, by team, without commentary. Visible evidence does more for uptake than any amount of enablement content.

There is a second effect that tends to surprise leadership teams. Once users can see which agents are earning their place, they start telling you which workflow should be automated next, and their suggestions are usually better than the ones on the original roadmap. Measurement turns an AI programme from something done to a commercial organisation into something the organisation actively pulls forward. That shift is worth more over eighteen months than any individual efficiency gain in the spreadsheet.

Start with three metrics, not thirty

The most common reason measurement stalls is ambition. A team sets out to build a full agent performance dashboard covering accuracy, containment, sentiment, deflection, cost per interaction and twelve variations of throughput, and eight weeks later nothing is live because the definitions are still being debated. The teams that succeed start with three metrics per agent and refuse to add a fourth until those three have been stable for a quarter.

A reliable starting trio is volume handled, cycle time and outcome quality. Volume tells you whether the agent is actually being used, which is the question most programmes forget to ask. Cycle time is where efficiency return shows up first and is almost always available from existing timestamps. Outcome quality is the one that needs a local definition, whether that is conversion rate, first-contact resolution, quote accuracy or manual rework rate, and choosing it well is the single highest-leverage decision in the whole exercise. Liferay’s respondents named better accuracy as their most pressing improvement, at 42 per cent, and a rework rate tracked weekly is how accuracy stops being a feeling and starts being a trend line you can act on.

What is an independent CRM partner worth when you are measuring agents?

An independent CRM partner is a consultancy that implements and optimises across multiple CRM platforms without reselling any single vendor’s licences, which means its recommendations are shaped by what the client’s process needs rather than by what a particular vendor is incentivised to sell. In agent measurement specifically, that independence matters because every platform ships its own AI value dashboard, and every one of those dashboards is designed to present its own agents favourably. An independent partner builds the measurement layer on your data, using definitions your finance team accepts, so the resulting numbers hold up whether you expand on Agentforce, HubSpot’s agent tooling, Copilot in Dynamics 365, or a combination of all three.

The practical benefit shows up at renewal and expansion time. An organisation that measures agent value independently negotiates from evidence, chooses where to concentrate consumption spend, and can tell the difference between an agent that is genuinely productive and one that is merely busy. That is a strong position to be in during a period when every vendor in the market is expanding its agent catalogue.

The Sirocco perspective

We think this is one of the more encouraging moments in enterprise AI, because the thing separating the teams that scale from the teams that stall is genuinely within reach. It is not a model, a licence tier or a data science hire. It is a baseline, three metrics and the discipline to publish them. In our work across Salesforce, HubSpot and Dynamics 365, the clients who instrument one workflow properly almost always find that the second and third are faster, because the definitions, the reporting patterns and the internal credibility all carry over.

If your organisation is already running agents, you are in the 54 per cent, and the step into the measuring quarter is smaller than it looks. Choose the workflow where the agent is most active, pull the 90 days before it went live, and put a number on what changed. Teams that do this consistently find that their next agent business case writes itself, and that the conversation with the board shifts from whether AI is working to how quickly the next one can be live. That is a good place to be heading into 2027.

If you would like a second pair of eyes on how your agents are performing, or help building a measurement baseline your finance team will accept, you are welcome to schedule a consultation with our team.

Get in Touch

If you are running AI agents in Salesforce, HubSpot or Dynamics 365 and want to know what they are actually returning, tell us which workflow you would measure first and we will help you build the baseline.

So where do you start?

As your long-term partner for sustainable success, Sirocco is here to help you achieve your business goals. Contact us today to discuss your specific needs and book a free consultation or workshop to get started!