Driving ROI in your AI initiatives

Date:



Over the last four years, artificial intelligence has captured the imagination of technology and business leaders worldwide, holding out the promise of revamped business models, new growth avenues and efficiency gains. In the past 12 months, however, enterprise conversations have undergone a sharp strategic pivot. As the initial AI hype cycle cools, corporate boards and CFOs are asking harder questions regarding implementation costs and bottom-line return on investment (ROI).

Given the scale of capital involved — with Gartner projecting global AI spending by enterprise and vendors to hit $2.59 trillion — the C-suite must take a rigorous look at how AI initiatives are scoped, budgeted and governed.

A recent survey of 500 senior US and UK finance leaders showed that 79% (4 out of 5) of large enterprises missed their AI budgets in the past 12 months. In another survey, 85% of enterprises said that they missed AI forecast by greater than 10% and a quarter of them said that they missed it by 50% or more. These numbers are staggeringly high. The silver lining, however, is that 6% of large enterprises have reported an EBIT impact of greater than 5% driven by an enterprise-wide implementation of AI. This indicates that there is a clear path to ROI, provided one runs a disciplined, prudent and efficient governance model for all AI initiatives.

Simply put, an enterprise can increase the ROI on AI investments by addressing the two components of the equation appropriately:

  1. Defining and amplifying the business value (numerator)
  2. Defining and optimising the cost of AI implementation (denominator). 

Defining and amplifying business value

An enterprise must focus on chasing value, not chasing AI. The board has an important role to play here — it needs to reconsider whether it is asking the right questions. By aggressively pushing for an “AI everywhere” strategy, the board unwittingly shifts the focus to proliferation of AI tools and use cases rather than tangible business value. Pressure from the board incentivizes “tokenmaxxing” by the teams and legitimizes rampant use of AI tokens by those who suffer from AI FOMO. This severely compromises the returns on AI investments, which is later used to question the value of the technology itself.

A few steps an organization could take:

  • Allocate costs appropriately: It would be prudent for organizations to keep AI training and access costs under a separate training budget while tracking ROI only for business-approved use cases. Track ROI only on costs linked to the use case.
  • Avoid force-fitting: Enterprises should resist the temptation to force-fit AI into every solution, and instead apply it only when the business problem requires it. The board and CEO can take a significant step towards ROI simply by resisting FOMO, distinguishing signals from the noise and pushing back on hype while cautiously embracing AI as required.
  • Define ‘kill’ criteria: If AI is deemed fit to solve or reimagine a business problem, it is critical to have clear termination criteria for AI pilots, with an alignment of the metrics to be tracked, their frequency, success or failure criteria and the termination process.
  • Capture true value: It is imperative to ensure that business value generated by AI is correctly captured. For example, If AI leads to a 10% time saving to the sales team, that is equivalent to having an additional salesperson for every team of 10 salespeople. This should result in additional sales and this resultant value must be included as part of the value captured in the numerator of the ROI equation.

Defining and optimizing the cost of AI implementation

The reason most organizations exceed AI budgets is that budgets are still being calculated on a ‘per user’ basis. AI licensing cost per user does not give any indication of the tokens that the user will consume, which is where costs scale exponentially. Since token consumption is linked to the task the agent performs, linking AI budgets to the number and complexity of the tasks in a workflow would give a much clearer view of expected costs.

Cost estimates must also account for realities of production environments which are typically missed out in sanitized test environments, including poor-quality long prompts, hallucinations, retry loops, network latency, employee ‘bad’ behaviour trying to ‘break’ the AI tool, etc. Cost estimates must make reasonable assumptions about a combination of these elements happening all the time, which will make token costs more realistic and higher than in more sanitized, non-production Test environments.

Costs can be further optimized in three more layers:

Optimizing at the prompt layer

It makes little sense to upload five voluminous PDF documents for three queries that a user may have and then get the LLM to process all of it when the relevant information might be in 20% of the data uploaded. Instead, chunking the data, vectorizing it with model embeddings and storing it in a vector database will make model usage a lot more effective, with reduced latency, increased retrieval accuracy and reduced cost. This kind of prompt pruning can be very useful. Prompt caching can also be considered if, via extensive user training, users can be trained to give a standardized prompt.

Alternatively, the enterprise could keep a middleware layer between the user and the LLM, which will append every prompt from every user into a standardized format for the LLM to process. By caching the response, the user can access a stored response rather than pinging the LLM for a re-computation on a prompt that has been asked earlier. Large language model providers provide up to an 80% discount on cached responses. Semantic caching, again achieved through vectorization, will enable the LLM to understand the intent of a prompt rather than look for an exact match, which can again reduce token costs significantly.

Optimizing model usage

Likely, a very large percentage of prompts in an organization do not need a frontier LLM to process them. A mid-tier model which comes at a significantly lower cost can deliver the desired output. This is where model cascading can play a critical role in reducing costs. Proposed by three Stanford researchers, it demonstrates an AI framework with model cascading at its core. By classifying prompts based on complexity, understanding model appropriateness for each prompt, caching responses and by assigning each prompt to the most appropriate model optimized for the trade-off between cost, accuracy and latency, they demonstrated significant savings in AI implementation costs.

Optimizing at the architecture layer

To avoid runaway AI bills, it is important for the CIO to embed cost controls within the architecture. As an example, the technology team can embed hard token caps for workloads, which are activated if token consumption is found to be much higher than anticipated for a given workflow or velocity circuit breakers if there are sudden spikes in token consumption, respectively. In addition, it is possible to embed a lot of the above-mentioned strategies (at the prompt and model usage layer) into the AI architecture, thereby paving the way for architectural FinOps, as distinct from the legacy FinOps that most enterprises have today. As an example, a middleware layer can help in prompt pruning, response caching and model cascading.

Many enterprises also change the architecture by moving the inferencing to an ‘on-premises’ environment to avoid incurring API and cloud costs for every prompt execution.

By incorporating all of it as part of the architecture, the CIO has greater visibility and ability to control costs in near real time.

In summary, technology leaders can drive ROI through a mix of sound business judgment, user awareness and training, and tighter control on processes and architecture.


Share post:

Subscribe

spot_imgspot_img

Popular

More like this
Related

DASH Diet Showed The Strongest Link To Lower Cognitive Decline

For decades, scientists have searched for the “best” diet...

Investors spill what they aren’t looking for anymore in AI SaaS companies

Investors have been pouring billions into AI companies over...

Sensible Default

A Sensible Default is a practice that, absent some...