Two critical facilities engineers inspecting cooling infrastructure in a data center.

The Data Center Staffing Shortage is also an Engineering Capacity Problem

Read about our privacy policy.
Thank you! Your file is ready for download:
Download now!
Oops! Something went wrong while submitting the form.

Data center expansion is increasing demand for technical talent while hiring remains difficult. JLL expects nearly 100 GW of new data center capacity to be added globally between 2026 and 2030¹. In the US, Deloitte found that postings for core data center roles increased 64% between 2023 and 2025, with electrical technician postings rising more than 180%².‍ Uptime Institute’s 2026 survey found that more than half of respondents now have difficulty finding qualified candidates for open roles³︎.

Headcount explains only part of the pressure on critical facilities teams. Experienced engineers carry knowledge that takes years to build: which sensors they trust, why a control sequence was changed, how equipment behaves outside design conditions, and what the team learned the last time a familiar fault appeared. Some of that knowledge becomes procedure or survives through shift handover and maintenance records. Some remains tribal knowledge.

At the same time, that expertise is being stretched across more infrastructure. Uptime Institute reports that more operators are seeing peak rack densities of 30 kW or above. At some facilities, critical facilities teams are also integrating liquid-cooled AI infrastructure alongside legacy systems, adding new operating relationships and constraints to an already complex environment. [Sources: Uptime Institute, Global Data Center Survey 2026; etalytics, Next-level cooling for AI data centers, January 2026]  

A new hire does not inherit years of site-specific knowledge on day one. For a growing data center operator, the staffing question therefore extends beyond recruitment: how much infrastructure can the engineering expertise already inside the organization effectively support?

Why do solved problems keep consuming senior engineering time?

Because the conclusion from an earlier investigation is often easier to preserve than the reasoning that produced it. A report may show what was fixed, while the assumptions, comparisons and operating context that led to the decision remain partly with the people who did the work.

A recurring low delta-T problem on a chilled-water loop illustrates the issue. Diagnosing it may require cooling load, supply and return temperatures, differential pressure, valve positions, pump staging, bypass flow and recent control changes. An experienced engineer may recognize the pattern quickly once the right context is assembled. Getting to that point is often where the work accumulates.

When a similar condition appears six months later on another loop or at another site, the previous report and trend charts may still be available. What is harder to preserve is which measurements were considered reliable, which operating points were comparable, which explanations had already been ruled out and which limits shaped the decision.

When that analytical path lives partly in documentation and partly in the experience of the person who performed the investigation, familiar problems still demand senior attention. The specialist is not necessarily being asked to solve a new engineering problem. Much of the effort goes into reconstructing the context needed to apply engineering judgment again.

The same issue affects escalation and mentoring. If a technician brings a senior engineer an alarm and several raw trends, the specialist first has to establish what changed, which evidence matters and where to investigate next. A better starting point leaves more of that interaction for engineering judgment, discussion and learning rather than basic data assembly.

Documentation remains important, but a report and reusable engineering work are not the same thing. Reuse also depends on retaining the structure of the investigation: the relevant assets and measurements, their physical relationships, the conditions that make comparisons valid, expected system behavior and the limits that constrain a decision.

For organizations already competing for scarce technical talent, the opportunity is to spend less expert time reconstructing operational context and more of it on the engineering judgment that actually requires experience.

How can more of that engineering work carry forward?

Data centers already generate large amounts of operational data through BMS, controls, meters and monitoring platforms. Making better use of engineering expertise depends less on generating more data than on connecting existing signals to the physical system they describe.

Operational Intelligence adds that context. In etaONE®, assets and operational data can be structured within physical system topologies, while digital twins model the relevant physical and operational relationships. Current performance can then be assessed in relation to equipment state, load, environmental conditions and expected system behavior. Existing BMS and control infrastructure remain part of the operating environment rather than being replaced. [Source: etalytics, etaONE® platform]  

For a recurring problem such as low delta-T, that system context gives the next investigation a stronger starting point. Engineers can work from current measurements together with the affected equipment, its system relationships and the conditions under which it is operating rather than rebuilding the picture entirely from old trend charts.

Expert knowledge, previous investigations and current system context combine to support engineering judgment.

Directing scarce attention where it matters

Engineering capacity is also consumed by deciding which deviations deserve attention at all.

Pump efficiency, for example, naturally changes as the operating point moves. A change becomes more meaningful when actual performance departs from what would be expected under comparable conditions. etaONE® applies this principle in deviation detection by comparing expected and actual behavior, ranking deviations by operational relevance and linking them to the affected signals, timing and operating context. [Source: etalytics, Automated Risk & Deviation Detection]  

That is important for a stretched critical facilities team because another alert queue does not create engineering capacity. If software generates additional noise, unsupported recommendations or more work to review, it simply moves the bottleneck. Useful Operational Intelligence should help teams establish context faster and direct specialist attention toward conditions that warrant investigation.

For less experienced staff, that means reaching an escalation with better evidence. Across shifts, more of the context behind an investigation remains available. Across facilities, teams have a stronger basis for applying established analytical methods while still accounting for the equipment and operating conditions at each site.

There is also a limit to how far the staffing problem should be treated as an automation problem. Uptime Institute has warned that poorly integrated AI and automation can contribute to skill decay when staff become too removed from operational decision-making. Its guidance emphasizes ongoing training, preserving holistic system understanding and maintaining interaction between junior staff and subject-matter experts. [Source: Uptime Institute, Preventing skill decay as AI use expands, September 3, 2026]  

Recruitment, training and mentoring therefore remain essential. Technology has a narrower and more useful role: reduce avoidable context assembly, preserve more of the structure behind recurring engineering work and help teams focus scarce expertise where it adds the most value.

As data center portfolios grow, engineering capacity will depend partly on how much of a critical facilities team’s existing knowledge and analytical work can be reused without rebuilding each investigation from scratch.

See the underlying technology in live data center operations

At Telehouse Germany’s Frankfurt data center, etaONE® combines digital-twin modeling, predictive analytics and automated setpoint optimization within predefined operating limits. The optimization also accounts for site-specific constraints, including nighttime noise requirements that restrict cooling-tower flexibility.

On one of five cooling units, the project measured a 10.4% reduction in electricity consumption by comparing optimized operation between April 14 and August 5, 2026 with historical periods under similar thermal demand and ambient conditions. [Source: Telehouse Germany & etalytics, Telehouse Germany Cuts Cooling Electricity Use by 10.4% with AI-Supported Cooling Optimization, September 9, 2026]  

Read the Telehouse story

‍

Thank you! Your file is ready for download:
Download now!
Oops! Something went wrong while submitting the form.
Read about our privacy policy.

¹ JLL, 2026 Global Data Center Outlook

² Deloitte Insights, In the AI age, data centers and power companies compete for the same core workforce, March 2026

³︎ Uptime Institute, Global Data Center Survey 2026