Cost optimisation in the enterprise: a practical guide
AI cost optimisation reduces spend by analysing entire workflows, not just unit costs, and attributing spend to specific business processes, allowing organisations to make informed decisions about resource allocation and identify areas for cost reduction, such as caching and model routing.
The decision to allocate budget to AI cost optimisation efforts is one that many enterprise leaders will face this week, and it is often based on incomplete or misleading information about where the money is actually going. Most organisations currently rely on rough estimates or simplistic metrics, such as unit costs, to guide their decisions, but this approach can lead to ineffective or even counterproductive optimisation efforts.
The unit cost trap
The distinction between unit cost and total cost of an AI workflow is crucial, yet often overlooked. Unit cost refers to the cost of a single instance of a process, such as the cost of a single prediction made by a machine learning model. Total cost, on the other hand, refers to the overall cost of the entire workflow, including all the instances of the process, as well as any additional costs such as data storage, model training, and maintenance. Focusing solely on unit cost can lead to optimisation efforts that reduce the cost of individual predictions, but ultimately increase the total cost of the workflow.
A good example of this is a company that optimises its machine learning model to reduce the cost of individual predictions, but in doing so, increases the number of predictions required to achieve the same level of accuracy, resulting in a higher total cost. To avoid this trap, organisations must consider the total cost of the workflow and optimise for that, rather than just focusing on unit cost.
Real levers for cost optimisation
There are three main levers that organisations can use to optimise the cost of their AI workflows: caching, batching, and model routing. Caching involves storing the results of expensive computations so that they can be reused instead of recomputed, which can significantly reduce the cost of repeated predictions. Batching involves grouping multiple predictions together and processing them as a single unit, which can reduce the overhead costs associated with individual predictions. Model routing involves directing predictions to the most cost-effective model available, which can help to reduce the overall cost of the workflow.
For instance, a company that uses machine learning to predict customer churn could use caching to store the results of expensive computations, such as customer segmentation, and reuse them instead of recomputing them for each new prediction. This can help to reduce the cost of individual predictions and improve the overall efficiency of the workflow.
The token cost myth
Token cost is often cited as a major contributor to the cost of AI workflows, but in reality, it is rarely the dominant line item. Token cost refers to the cost of processing individual units of data, such as text or images, and is often used as a rough estimate of the overall cost of the workflow. However, this can be misleading, as it does not take into account other costs such as data storage, model training, and maintenance.
In fact, most organisations find that the majority of their costs are associated with data storage and model training, rather than token cost. To get a accurate picture of where the money is going, organisations must look beyond token cost and consider the total cost of the workflow.
Attributing spend to business processes
One of the key challenges in optimising AI costs is attributing spend to specific business processes. Most organisations have multiple AI workflows running in parallel, each with its own set of costs and benefits. To optimise costs effectively, organisations must be able to attribute spend to specific business processes, such as customer service or marketing, and understand how those costs are impacting the overall bottom line.
For example, a company that uses AI to power its customer service chatbots may want to attribute the costs of those chatbots to the customer service department, rather than to a central AI budget. This can help to ensure that the costs are aligned with the benefits and that the organisation is getting the best possible return on investment. To learn more about how AsscherAi can help with this, visit our website.
Common mistakes
Most people get wrong the idea that AI cost optimisation is all about reducing the cost of individual predictions. While this can be an important part of the process, it is not the only consideration. In fact, focusing too much on unit cost can lead to optimisation efforts that ultimately increase the total cost of the workflow.
A better approach is to consider the total cost of the workflow and optimise for that, rather than just focusing on unit cost. This requires a deep understanding of the workflow and its associated costs, as well as the ability to attribute spend to specific business processes.
When to use AI cost optimisation
AI cost optimisation is not always the right choice. In some cases, the costs associated with optimisation may outweigh the benefits, or the organisation may not have the necessary expertise or resources to implement optimisation efforts effectively. In these cases, it may be better to focus on other priorities, such as improving the accuracy or efficiency of the AI workflow.
To determine whether AI cost optimisation is right for your organisation, it is a good idea to speak with an expert. Our team is happy to help you understand your options and determine the best course of action. Visit our contact page to get in touch.
Conclusion of sorts
The decision to allocate budget to AI cost optimisation efforts is a complex one, and requires a deep understanding of the workflow and its associated costs. By considering the total cost of the workflow, rather than just focusing on unit cost, and by attributing spend to specific business processes, organisations can make more informed decisions about where to allocate their resources.
Frequently asked questions
What is the most common mistake made in AI cost optimisation efforts?
Focusing solely on unit cost, which can lead to ineffective or counterproductive optimisation efforts.
How can organisations attribute spend to specific business processes?
By using tools like AsscherAi to track and analyse costs associated with each workflow and business process.
What are some effective levers for reducing AI costs?
Caching, batching, and model routing can help reduce costs by minimising repeated computations and optimising resource allocation.
When is AI cost optimisation not the right choice?
When the costs associated with optimisation outweigh the benefits, or the organisation lacks the necessary expertise or resources to implement optimisation efforts effectively.