- Circuit design reduces energy consumption, but IP needs special attention since the integrator didn’t design it.
- Individual chips can run at voltages optimized for the chip at a given time, allowing some to operate at lower power than others.
- Data center power utilization can increase by using available reserve power while scaling back if a power feed fails.
Data centers are notorious energy consumers. On the one hand, ongoing efforts attempt to reduce electronics power to keep overall power from growing. On the other hand, data centers are provided with more power than they nominally need, providing redundancy.
These two issues are converging in the ongoing focus on data-center power — minimizing component power while maximizing utilization.
“Idle power is by far the biggest sink of power in systems,” said Arif Khan, vice president of product marketing for design IP at Cadence. This arises because some servers may run at a higher-than-necessary voltage and because power is left in reserve.
Others agree, pointing to the increasing significance of idle power, also known as stranded power. “Customers are talking about this more openly now, especially in large data centers where small inefficiencies add up fast,” said Steven Woo, fellow and distinguished inventor at Rambus. “Workloads vary in how they use resources. Sometimes resources are over-provisioned, and sometimes they’re under-provisioned. Over-provisioned resources can lead to stranded power if data centers plan for all resources to be at or near full utilization. Under-provisioned resources can lead to some resources being fully utilized while others are not, again leading to stranded power.”
Power reduction can be achieved through smart power architectures, as well as by monitoring performance and tweaking voltages to ensure the lowest voltage that still provides the necessary performance. Reducing power improves efficiency, but it adds to the extra provisioned power that lies unused, hurting power utilization. New techniques aim to make the excess power usable.
It starts with power-saving architectures
Any idle circuitry that isn’t providing a function but is being clocked and consuming energy hurts power efficiency. Most custom designs plan for power savings to account for this.
One of the bigger sources of idle power can be IP and its interfaces, whose details are not under the integrator’s control. Since IP is purchased outright, the integrator must know what power-savings opportunities a given piece of IP includes and take advantage of them.
“From an IP perspective, a lot of this has to do with which interfaces are put to sleep, which are woken up, and which channels you are using,” said Khan. “As speeds have gone up, these interfaces have become more power hungry, so there’s a lot of work done to use power-optimized states and power-aware transactions. Now we’ve got multiple levels of standby. You have a larger menu of tradeoffs to make, but that requires the operating system and the orchestration logic to be familiar with the options. If you are running multiple jobs on a system, for example, even though one may not be using an interface, something else might be, so you can power down only to the least common denominator.”
Some IP may have features that can be compiled out, such as an interface that is wider than necessary. In that case, the feature never hits the silicon. But other features do, and the IP typically provides the ability to move into and out of power states, adjust the clock frequency, or adjust the supply voltage.
The AMBA bus, for example, provides a low-power interface (LPI) port to manage interface power. It has a so-called Q channel for simple clock gating and power-domain control, and a P channel for more advanced power management when multiple power states are used. Their granularity determines how many circuits they control.
Clock gating can effectively shut down circuits while leaving power intact. Power gating takes more time, as does powering back up.
“There’s a need to power things down as low as possible,” said Khan. “The flip side is bringing things back up and balancing the resumption time against how much power you want to trade off.”
Dynamic voltage/frequency scaling (DVFS) is another widely accepted way to tweak the operating state, with the system picking different frequency and voltage settings based on temperature or other conditions. Chips may also have an external signal that lets external components, including software, set the operating state based on conditions outside the chip.
However, it’s unlikely that any IP has more impact than the network-on-chip (NoC). “As a NoC is the backbone of an SoC, it is imperative that we implement robust power management schemes to save power for the whole chip,” said Ewald Liess, senior vice president of sales and business development at SignatureIP, which makes extensive use of clock gating and power states on the AMBA interface. “The clock is actively driven inside each cluster. Ten independently gated clocks are in a typical 5+5 nodes configuration. Furthermore, there are power states on each node of a NoC, and quiesce/run events are negotiated over the P-channel.”
Each chip is a special snowflake
Ideally, the power architecture and the ways it is employed should affect all units of a given chip equally. In reality, different chips have different operating characteristics, and a single chip may operate differently after a couple of years than when it was newly manufactured. The architectural and circuit design techniques noted above must be guard-banded to ensure that all chips, regardless of where they are in the life cycle, will operate as promised.
“Power consumption scales with the square of voltage,” said Noam Brousard, vice president, systems at proteanTecs. “But the voltage applied to a device is based on characterizing the worst-case conditions in which it may run. The idea is that more challenging workloads, system and environment conditions, and inherent degradation over time will require higher voltage to maintain performance. However, it is rare that all these circumstances happen at the same time, all the time.”
One example is chip aging. Silicon wearout is a phenomenon that the industry ignored for many years, but it’s an important performance consideration today. A chip must be guard-banded to operate both out of the box and on the last day of expected service. As a result, the supply voltage must be set to a level that will work after the chip has aged. That voltage is typically higher than what a new chip needs.
“For instance, while aging may require voltage increase compensation, this is not needed on day one, when aging is still not in effect, ” said Brousard. “And while the device may run hot, the workloads actually running may be less stressful at that time.”
To address this, in-circuit monitoring provides greater flexibility in circuit operation. During test, an individual chip can be characterized to establish the voltage it requires. This allows a real-time monitoring system to reduce the voltage if it looks like the chip can still meet performance that way. “We have the ability to measure in real time the actual margins to failure of logic throughout the device,” said Brousard. “These margins are directly affected by the aggregation of all the effects that may eventually require the higher voltage for the device to meet its spec. This visibility lets the system instantly optimize voltage to meet the device’s needs at any moment. By reclaiming these guard-bands when they are not needed, the chip runs at a significantly lower voltage while still meeting its operational requirements. When the circumstances that require higher voltage occur, and only when they occur, this capability allows for the scale-up of the voltage to meet the new requirements.”
Just as reducing power can help, at times it’s necessary to raise power. “By running on low margins, there is a risk that a sudden change in operational or environmental conditions may cause the immediate need for a higher operational voltage,” Brousard said. “For instance, a sudden change in workload may cause a dynamic voltage drop (Ldi/dt) that may threaten digital logic timing integrity.”
ProteanTecs has a ‘safety net’ for these scenarios in the form of a built-in hardware-based protection mechanism that enables the host system to identify and mitigate the sudden drop in margins within a couple of cycles. This could involve clock dividing or throttling while readjusting the voltage to safe levels before returning the clock to its normal value.
Even more important is the ability to monitor aging on an ongoing basis so that the voltage can be reduced initially to save power, then boosted occasionally over the chip’s life as it ages.
The data center gets more power than it uses — intentionally
All of these power-saving techniques can lower a data center’s overall energy draw, which is good, but then what? Data centers are provisioned with a certain amount of power, and the flip side of efficiency is utilization. If you’re not using all of your power, you could be if you added more equipment. While additional equipment represents an additional investment, it’s far cheaper than building a new data center or even acquiring more power at an existing data center.
“Getting access to power is one of the biggest time constraints on data center operators’ business,” said Marissa Hummon, CTO of Utilidata. “If you just wanted to run a new transmission line from the substation to your facility, that is expensive, time-consuming, and generally involves lots of permitting and lots of regulations. Even to build on-site, new generation takes a while to permit.”
And while power savings can reduce utilization, they’re not the primary reason power remains unused. Data centers must be reliable and avoid any downtime. Power feeds can fail, so data centers typically overprovision in case a feed goes down, and they can reallocate from working feeds.
One common approach is called 3-4 (often pronounced “three makes four”), where three 100% feeds are necessary, but a fourth is added as backup. All four feeds run at 75% so that when one goes, the remaining three can supply 100% of what’s needed. Other arrangements include 4+2 (an example of a “2n” system), and “n+1” (meaning one feed more than necessary).
“Nvidia recommends a three-makes-four system or an n+1 system, but most of the common old-school data centers go for a full 2n system,” said Hummon. “That means there are two lines coming in. They load both of them at 50% so that if they lost one of those lines, they could fail over to the other 50%.”
Still, redundant power remains unused today. “The data center paid for the capacity. The utility had to put in the capital cost to deliver that capacity, but they’re not selling the kilowatt-hours [for unused capacity],” said Hummon. “A site may be provisioned for 2GW, but the site manager thinks of that as only 1GW [given 2n provisioning]. Then, within that gigawatt, they generally aim for 70 to 80% utilization. When you look at it as a whole, they’re getting about a third of the 2GW to the compute infrastructure.”
If the extra available power is used in the event of a failure, the sum of jobs running would require more power than was available with one feed down. In a 3+1 scenario, running all feeds at, say, 90% would mean that, if one feed goes down, the other three would have to provide 120% of what’s provisioned — meaning someone would have to go hungry. And yet it would be tempting to use that extra power when it’s available while still accommodating failover.
Monitoring system power
To manage this, a logical approach is to monitor power and, if a feed fails, scale back operations to fit the remaining available power. But that’s a difficult task. Jobs underway would have to be unscheduled or checkpointed to ensure nothing failed when the available power dropped. In the shorter term, dropping voltages and clock rates can reduce total utilization down to levels sustainable with a missing feed. But this takes real-time monitoring and responsiveness, a task that Utilidata addresses.

Fig. 1: Processor for managing power sensing and failover management. Source Utilidata
“Before an outage, you can definitely see transients in the system that indicate the outage is coming,” explained Hummon. “If we know something’s coming, we give signals to the jobs that have some flexibility to back down. If you were operating at 70% of your total capacity, you just need to cut 20% to get back to 50% if you had a 2n system.”
The company operates on two control loops. One is the workload scheduler. “AI workloads are very dynamic, and scheduling them takes a global view of power, memory, and compute availability,” said Hummon. “We’re providing a signal into the scheduler so that it can change the way it allocates jobs across GPUs”.
This is a longer timescale loop since scheduling and unscheduling can be complex, depending on the state of computation at the time. “We do not schedule or unschedule work,” noted Hummon. “We can influence the scheduler by providing power capacity forecasts for each rack/row/data hall. Schedulers operate in seconds because moving inference work means draining in-flight requests and then loading model weights onto another replica.”
The second control loop is faster and controls local power. The system can also drive power changes within processor chips. “The other control loop is directly to the server,” said Hummon. “We interface with the DVFS signal (also known as power capping), and so we are changing the power on the millisecond timeframe.”
Here again, the DVFS control isn’t direct. “Our technology reads and acts through standard server interfaces,” explained Hummon. “Nvidia’s management library gives us GPU power behavior, and the BMC (baseboard management controller) gives us server-level power and telemetry. Those control surfaces reach the firmware that does DVFS. We include the interface and its behavior in a rack-level power optimization. Server-only controls must estimate what all of those servers add up to as AC draw at the rack, whereas we measure it directly at a million samples per second. That removes the guesswork.”
A technique like this lets the excess provisioned power be used under normal conditions, with time to back off in an orderly fashion when necessary. “We’re saying you could run on all four lines all the way up to probably 95% or 98% if you had that dynamic system.”
Given that opportunity, a data center may be able to accommodate more servers. Hopefully, the amount of time without full power would be short and infrequent, giving a sizeable amount of new power to work with. In a 2n configuration, you literally double the available power.
Security becomes part of power control
This capability could also create a new way to hack into the system if security is inadequate. “Security has been a design requirement from the start,” said Hummon. “We maintain SOC 2 (Service Organization Control 2), and we are listing Karman [monitoring system] to IEC 62443-4-2, the industrial control system security standard. Every server in an AI data center already has a BMC that can cap its power, and every GPU host exposes power limits to any sufficiently privileged process. The ability to modulate power at scale is already there and already on the network. Our technology coordinates across those interfaces, so it concentrates capability that is currently scattered, and we treat the control unit as a high-value target accordingly. Secure boot and a hardware root of trust limit it to firmware we have signed, and mutual authentication on every interface leaving the device closes any unauthenticated path to power actuation. And the control loop runs entirely on-site and does not depend on a connection to the internet.”
Where efficiency and utilization meet
It’s nice to think that, by saving power in each unit, we’d use less energy overall. But that’s not how it works. Saving power means lower utilization. Lower utilization leaves power on the table that could do more work and generate more revenue.
Designers keep finding ways to do more with less energy. Data-center operators may now have a way to do more with the energy they have available — subject to footprint limitations, of course. This dynamic is likely to persist until we run out of additional work to do in the data center, and that’s not happening anytime soon.
Editor’s note: Since publication, Utilidata has changed its name to Karman
Related Articles
Crisis Ahead: Power Consumption In AI Data Centers
Four key areas where chips can help manage AI’s insatiable power appetite.
Harnessing Silicon Lifecycle Management For Chip Security
Real-time monitoring and proactive risk mitigation can identify vulnerabilities and attacks throughout a device’s lifetime, and much more.
Liquid Cooling Gains Traction In Data Centers
There are numerous ways to remove heat from chips, and more are on the way.