Artificial intelligence has become an essential tool, since it is present in search engines, recommendation systems, virtual assistants, etc. In addition, automate tasks, accelerates processes and allows you to create content that, a few years ago, could only make one human.
But that capacity is not free, because every time you interact with an AI system, dozens of servers are activated, thousands of processing cores and cooling mechanisms that consume huge amounts of electricity.
The problem is not only the scale, but the speed at which all this grows. According to the World Economic Forum, the use of generative 50 % annually will grow until 2030. If nothing changes, that implies a Energy demand that could very soon overcome what the current infrastructure can offer.
It is important to mention that it is no longer about optimizing models or training faster neural networks, but that it is about answering an urgent question, how to energetically feed a system that never stops growing?
Custom designed chips for a more efficient AI
Given this situation, some companies are looking for alternatives to traditional chips, dominated by Nvidia. Its GPUs are the most powerful in the market, but also the most expensive and the ones that consume the most energy. In inference tasks —Generate answers in real time – they are not always the most optimal option.
This is where new proposals such as Positron and Groq come into play, two startups that bet on chips designed from scratch to solve this bottleneck. Positron has developed a specific chip for inference tasks. Unlike generalist chips, it is optimized only for a very specific set of operations.
The result? A promise of efficiency between three and six times higher per watt in front of the Nvidia chips. The key is to simplify the hardware, avoid unnecessary tasks, as well as to concentrate all resources to respond quickly with the lowest possible expense.
Gleon the other hand, bet on another architecture, such as chips that directly integrate memory in the processor itself, which reduces data access time and improves performance. According to those responsible, These new processors can perform the same tasks as the current ones consuming up to a sixth part of energy.
These advances are already being tested by companies such as Cloudflare, which has begun to evaluate the real performance of positron chips in real environments.
If the results are confirmed, they could not only implement them on a large scale, but also reduce their dependence on NVIDIA, whose hardware imposes what is already known in the sector as the «NVIDIA Tax»: a profit margin of up to 60 % per chip that drastically raises the operating cost of any AI.
Why does AI need so much energy (and why is that a problem)
Most users think that using a chatbot or generating an image is somewhat light, but each request starts a complex physical network, as data centers distributed throughout the world, high -performance processors working at maximum load, specialized RAM, SSD discs operating without rest and cooling to keep everything at operational temperature.
All this, just to offer a text response or an image. Nevertheless, The most demanding part is the training of modelsbut it is not the only one. Inference, that is, the moment in which you ask the AI and it answers, is the operation that is most repeated.
Every time a user generates an image, writes an automatic email or asks for help to solve a problem, A trained neuronal network is activated that must work instantly, without margin of error or pause. The energy cost of that operation, multiplied by millions of users every minute, is so high that large technological ones are concerned.
Not only for what it implies for the planet, but also because maintaining these operations involves paying millionaire invoices. It is for this reason that if that spiral does not stop, neither the current infrastructure nor the budget of Many companies can support the expansion of AI.
The future of AI also depends on how it feeds
But although these new chips represent a significant advance, they do not completely solve the problem. Artificial intelligence models do not stop growing in complexity. Chatgpt 5, Google, Anthropic or Mistral models are increasingly heavy.
Not only do they require more energy to train, but also to run. In addition, the number of services that make up the IA does not stop increasing: from search engines to gps mail or navigation platforms.
That expansion means that energy demand is not reduced, it is simply distributed more efficiently. Therefore, some companies are going beyond hardware. Google, for example, already studies how to apply nuclear energy or even experimental fusion to feed your future data centers.
The same explores companies such as Microsoft or Amazon, aware that the solution is not just in making more efficient chips, but in producing sufficient electricity for what is coming.
Artificial intelligence continues to evolve, but every step forward also implies a greater consumption of resources. If the hardware improves, the demand grows. If the chips are optimized, the models multiply. The only way to prevent the entire collapse system from being rethinking not only how AI is trained and executed, but also how it feeds.
Know How we work in NoticiasVE.
Tags: Artificial intelligence
