Kathy Gibson reports from Nutanix .Next in Sandton – For much of the 20th century, oil was the currency that drove the economy. As we entered the 21st century, businesses started to refer to data as the “new oil”.

Today, organisations are realising that data is only useful if it can be accessed and used – and artificial intelligence (AI) is proving to be the best and fastest way of doing this.

Raif Abou Diab

In fact, Raif Abou Diab, regional director: South Gulf, Oman and sub-Saharan Africa at Nutanix, turns the AI value proposition on its head, explaining that the important thing is not the model, but how it relates to the data.

But, as companies around the world realise the value they can unlock with AI, they are running into a new – and very expensive – issue.

The cost of tokens is threatening to overwhelm many businesses, but there is a seemingly limitless appetite for them.

Jason Langone, global AI director at Nutanix, points out that token consumption has grown by about 17 000-times in the last four years, with a massive 38-trillion of them processed in the first quarter of 2026 alone.

Agentic AI is largely responsible for this growth: while basic chatbots consumed thousands of tokens to perform basic tasks, agents typically use tens of thousand per task, while autonomous agents could consume in the millions for long-run, multi-step workflows.

“And bear in mind that we are still in chapter one of agentic AI,” Langone says. “This consumption will only continue to grow.”

But there are ways to manage token use, and ways to cut it right back to manageable levels – without putting the brakes on innovation, he adds.

He cites Nutanix’s own experience.

The company was consuming as many as 1-trillion tokens per month, primarily for agentic coding, and recognised that it would soon be on track to spend $20-million or $30-million per year.

So it took measures to curtail the costs, putting place prompy hygiene, aggressive coaching and cheaper models for lower-value tasks.

This delivered some savings but the costs were still too big to swallow.

Nutanix did some calculations and decided to invest in its own GPU infrastructure. Coming in at $20-million this was not a trivial investment.

It then started using less expensive open weight models instead of the pricier frontier models for most of its AI operations.

Langone explains that frontier models are still vital for some high-value workloads, but they now account for just 20% of LLM usage, with 80% running on open weight models.

“So the bulk of use is in our own clusters, at a cost that is fixed in advance,” he says. “And the frontier models are metered for the work that genuinely needs them.”

Although the upfront investment in GPU infrastructure was significant, Nutanix had a one-year return on investment – and the value will become more significant over time, Langone says.

“The cost per token will only rise, but once the initial investment in your own GPU infrastructure is realised, thereafter you only have to think about licencing costs.

“If you are in the first phase of your AI journey, this is a great time to think through the next few years’ strategy.

“Already customers are using AI with Nutanix to deliver business value.”

Langone’s advice for organisations struggling with rising token costs is to use the two levers provided by Nutanix to rein them in.

The first lever is ensuring the right tokens are used for the right workload. Designing a birthday card for your cat doesn’t need to use the most expensive tokens; but designing a new business-critical system does.

So a solution that is able to route tokens depending on various factors can be a real boon.

Langone describes the Nutanix Agent Gateway as the steering wheel  for centralised control and model routing.

The solution also offers global rate limiting, unifed observability, unified end points and universal API translation, and lets users navigate risk when the environment changs. Admins can also get granular access to agents, seeing what they are correlating them to user, spend, model and more.

The second lever, decoupling the success curve from the cost curve, involved embracing fixed price inference.

The Nutanix Agent Gateway allows for both token consumption and fixed cost as part of your strategy.

Instead of metering engineers – which could limit the experimentation flywheel – organisations have the choice of frontier models or open weights on-premise.

“Nutanix Private Inferencing is a simplified model serving operations, and is designed for core IT,” Langone says. “We realised AI would become a core workload for IT so we give customers freedom of choice.”

Nutanix’s AI tools sit within the company’s stack of enterprise tools.