Document jBXdXzMv3VkKEYmJwR1nZKa3k
Energy and Buildings 50 (2012) 81-92
Contents lists available at SciVerse ScienceDirect
Energy and Buildings
journal homepage: www.elsevier.com/locate/enbuild
34I, I rdhP- r
Minimizing the thermal impact of computing equipment upgrades in data centers
Jayantha Siriwardanaa'*, Saman K. Halgamuge a, Thomas Schererb, Wolfgang Schottb
a Department of Mechanical Engineering, The University of Melbourne, Parkville, Victoria, Australia b IBM Zurich Research Laboratory (ZRL), Rilschlikon, Switzerland
ARTICLE
INFO
Article history: Received 11 November 2011 Received in revised form 6 February 2012 Accepted 12 March 2012
Keywords: Data center energy efficiency Data center cooling Hot air recirculation Thermal aware equipment upgrading Load spreading Computational fluid dynamics (CFD) Particle swarm optimization (PSO)
ABSTRACT
Upgrading of today's air-cooled data centers (DCs) with high-performance computing, networking, and storage equipment is a challenging task due to typically higher power needs and more adverse cooling requirements of new equipment. To cope with the increase of the power and heat load in DCs, load spreading is commonly applied. This technique distributes the thermal load of new equipment over multiple racks if the power requirement and heat generation of new equipment exceeds the rack's capacity. We present a novel load spreading technique that allows upgrading of the computing equipment with minimal thermal impact on the existing optimized DC cooling environment. Our approach is based on an abstract heat-flow model of the DC, whose parameters are determined by performing a measurement campaign in the DC and with support of computational fluid dynamics simulations. The optimum placement of the new equipment in the racks of the DC is found by applying a particle swarm optimization technique to this model. The effectiveness of our method was assessed based on experiments performed in a production DC. The results show that our holistic approach for optimizing the placement of the upgraded computing equipment in the DC outperforms the conventional load spreading technique.
2012 Elsevier B.V. All rights reserved.
1. Introduction
The successful operation of an air-cooled data center (DC) requires an efficient cooling environment to ensure that the DC operator can provide its services to customers with maximum availability and reliability at minimal operational cost. An efficient cooling system guarantees that the temperatures at the inlets of all devices in the computer racks of the DC never exceed a given threshold value to prevent device overheating, and achieves this goal with a minimum amount of cooling energy.
In today's DCs, a hot-/cold-aisle cooling concept is usually employed to efficiently cool the computing equipment [1]. For this purpose, the racks are arranged in rows to form hot and cold aisles with alternating airflows. In the cold aisles, chilled air from the computing room air conditioners (CRAC5) blown through the raised-floor plenum and perforated floor tiles is directed to the device inlets, while in the hot aisles heated exhaust air from the racks circulates back to the CRAC5. To ensure that the inlet temperatures at the devices never exceed a given threshold value and no cooling energy is wasted, a nominal airflow and temperature distribution is provided in the DC by adjusting various parameters of the cooling system (e.g. supply airflow, supply temperature),
* Corresponding author. Tel.: +61 425177770. E-mail addresses: Wpgrad.unimelb.edu.au,
Q. Siriwardana).
0378-7788/$ - see front matter 2012 Elsevier B.V. All rights reserved. doi:10.1016/j.enbuild.2012.03.026
determining the proper placement of the perforated tiles in the cold aisles, and carefully selecting the best-suited placement of the computing devices in the racks. These steps of setting the DC cooling environment to an optimal operating point have to be carefully performed and successfully verified by using, for example, IBM's Mobile Measurement Technology (MMT) [2] before putting the DC into operation.
Due to the increasing demand for supporting new on-line services such as video on-demand, Internet banking, cloud computing, social networking, etc., DC operators are continuously compelled to upgrade their DC with high-performance computing equipment such as blade servers. These devices are usually smaller and process information at significant higher rates than their predecessors, but typically consume more power and thus dissipate more heat. This additionally generated heat should be efficiently removed from the devices to avoid creating hot spots and other cooling inefficiencies. Hot air exhausted from the rack outlets can recirculate into the cold air stream supplied to the rack inlets, causing the equipment to under-perform, malfunction, or fail all together [3].
To combat the undesired thermal effects of the upgraded equipment on the existing optimized cooling environment, various strategies can be applied: The DC operator can lower the supply air temperature of the cooling system. This approach avoids device overheating, but will significantly increase the annual cost to be spent for cooling energy. Another approach, which is commonly applied in practice, is load spreading [4]. To determine the best placement of the new equipment in the computer racks, the
82
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
maximum allowed power and heat capacity are specified for each rack. Whenever the load of the devices put to a rack exceeds any of theses design parameters, the equipment is split and distributed over multiple racks. This simple approach more evenly distributes the power and cooling demands of the new equipment over the DC, but does not take into consideration the airflow and temperature distribution in the DC when spreading the heat load.
In this article, we present a novel load spreading technique that allows upgrading of the computing, networking, and storage equipment with minimal thermal impact on the existing optimized cooling environment of a DC. Our approach is based on an abstract heat-flow model that describes the nominal operating point of the DC before upgrading the equipment. The parameters of the model are determined by means of performing an extensive measurement campaign and with support of CFD simulations. The model takes into account the temperature distribution within the DC, various airflows, and the power consumption of the computing devices. The best-suited placement of the upgraded equipment in the racks is found by formulating a constrained non-linear optimization problem based on this model and solving it with a particle swarm optimization (PSO) technique [5]. The effectiveness of our method is assessed based on experiments performed in a production DC. The results show that our holistic approach for minimizing the thermal impact on the DC cooling environment in case of an equipment upgrade outperforms the conventional load spreading technique.
The article is organized as follows: Section 2 describes related work. Section 3 presents the proposed strategy for minimizing the thermal impact of a computing equipment upgrade in a DC. Section 4 gives results of a case study, in which we applied our optimization strategy to a production DC. The article concludes with a brief summary.
2. Related work
Recently considerable attention has been paid to the optimization of the cooling performance of air-cooled DCs with a hot-/cold aisle layout. In this section, we give a brief overview on three previously published approaches that focus on minimizing the hot air recirculation in DCs and are relevant to the proposed equipment upgrading methodology.
2.1. Cluster cooling performance optimization
Shrivastava et al. [6] proposed to combine a genetic algorithm with a neural network (NN) model to obtain the optimum layout of a cluster of DC equipment in terms of thermal performance. They assumed that a cluster typically comprises two equal length rows of racks and coolers bounding a common hot aisle. The NN model predicts the cooling performance by observing the power demand for each rack, and the rack inlet and exhaust airflow rates. Based on the observations, the genetic algorithm assesses all possible cluster configurations to find the best distribution of the heat load among the racks or to identify the best physical cluster layout.
The applicability of this method to a production DC is limited as the approach does not take into account the thermal interference between various clusters covering the DC room. Moreover, when upgrading DC equipment, a complete layout change of the DC is undesirable due to the high installation cost and the required long DC down time.
2.2. Heat recirculation minimization
Moore et al. [7] proposed a system-level solution to control the heat generation in a DC through temperature aware workload placement to servers. A server workload placement policy called
minimum heat recirculation (MinHR) has been proposed to assign tasks to servers that have minimum contribution to the hot air recirculation in the DC, and a metric called heat recirculation factor (HRF) to assess the contribution to the hot air recirculation of each server. MinHR maintains a priority list of servers according to their HRF values that is used in allocating tasks.
Even though the scalar dimensionless HRF values can serve as a simple metric for placing tasks on servers, they cannot be used for accurately capturing and modeling thermal effects in the DC room. Moreover, the benefit of using the MinHR server workload placement policy diminishes if the utilization rate of the DC equipment approaches 100%.
2.3. Thermal aware task scheduling
Thermal aware task scheduling as proposed by Tang et al. [8] is a thermodynamic based, but yet simple approach for assigning tasks to servers to minimize the hot air recirculation from the server outlets to the inlets. Tang's method assumes a homogeneous DC with identical equipment. It predicts the temperature at the server inlets with a low-complexity abstract heat-flow model [9] and a power model that linearly maps the CPU utilization to the server's power consumption. The parameters of the heat-flow model can be obtained with CFD simulations. Based on the predicted result, the incoming tasks are distributed among the servers so that the temperature of the chilled air supplied by the CRACs is maximized while keeping the temperature at the server inlets below a given maximum allowed threshold value.
The method has mainly been proposed for temperature aware scheduling of the processing workload among servers to minimize DC cooling requirements. It has not been applied for optimizing the layout of the DC to minimize undesired thermal effects such as hot air recirculation and, in particular, not for spreading the heat load over the DC while upgrading equipment.
3. Strategy for minimizing the thermal impact of server upgrades
The goal of the proposed load spreading strategy is to find the best placement of the upgraded equipment in the DC with minimal thermal impact to the existing cooling environment. This is achieved by formulating a constrained non-linear optimization problem that uses Tang's abstract heat-flow model to quickly predict the outlet and inlet temperatures of the server equipment for various upgrading scenarios, and applies a PSO algorithm to search for the best-suited position for the new servers in the racks of the DC. After having identified the best placement, the servers currently located at these positions are moved to the location of the servers to be replaced before the new servers are installed. In this section, we first review Tang's abstract heat-flow model due to its importance to our work, and then present the proposed PSO-based load spreading technique.
3.1. Abstract heat-flow model
We assume that the DC consists of n nodes. In case of blade servers, a node can be considered as a chassis that houses a number of blade servers as shown in Fig. 1. Each node i thus consist of m servers. Each node i draws air with inlet temperature Tiin and dissipates hotter air with average outlet temperature Toiut . In a typical DC, the difference between inlet and outlet temperatures of a node is approximately 11 C [4]. The outlet temperature of a node is determined by the inlet temperature and the heat load of the equipment inside the node, while the inlet temperature is determined by
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
83
jiQ
j out
cold
equipment
aisle
racks
Nj iiQ iout Ni Q iin
hot aisle
Qsup
cold air supply
air supply plenum
Fig. 1. Schematic of typical thermal cross-interferences among equipment nodes housed in racks in a data center.
the mixture of cold air supplied from the CRAC and hot air recirculated from the node outlets.
3.1.1. Derivation of abstract heat-flow model Considering that the total power drawn by a node is dissipated
as heat, the law of energy conservation can be used to express the steady-state relationship between the power consumption Pi and the in-/outlet temperature of node i as
Pi = ficp(Toi ut - Tiin) or equivalently as
Toi ut = Tiin + Ki-1Pi,
(1)
where is the air density, fi is the airflow rate through the node i, cp is the constant-pressure specific heat of air and Ki = ficp. The heat load Pi thus causes the temperature of the air passing through node i with a constant airflow rate fi to increase from Tiin at the inlet side to Toiut at the outlet side of the rack.
The air recirculation inside the DC is characterized as crossinterference among the nodes. The inlet temperature Tiin of node i can be considered as a mixture of the supplied cold air with tem-
perature Tsup from the CRACs and the recirculated hot exhaust air from other nodes. The total amount of heat exhausted per time unit in the outlet airflow Qoiut is given by
Qoiut = Qiin + Pi = KiToi ut .
(2)
Since the inlet heat Qiin carried by the inlet airflow is a mixture of the heat Qsup of the supplied cold air and the recirculated hot air as illustrated in Fig. 1, the following expression holds:
Qi
n
=
Qj
+ Q
,
(3)
in
ji out
sup
j=1
where the amount of recirculated heat from node j to node i is defined as jiQojut . Hence, the matrix A = [ij]nn defines the crossinterference among all the server nodes and is thus called crossinterference matrix.
From Eqs. (2) and (3), we have
Qoi ut
n j
= jiQout + Qsup + Pi
j=1
n
(4)
= jiKjToj ut + Qsup + Pi.
j=1
If tnhe total recirculatedn airflow rate from all nodes to node i is
j=1jifj, then fi - j=1jifj is the flow rate of supplied cold
airflowdnrawn by node i and, consequently, we have
fi - j=1jifj cpTsup. Thus, we can rewrite Eq. (4) as
K T i = n K T j + K - n K T + P .
i out
ji j out
i
ji j sup i
j=1
j=1
Qsup = (5)
For all the nodes from 1 to n, we can derive a group of linear equations from Eq. (5) above as follows:
KTout = AKTout + KTsup - AKTsup + P ,
(6)
where A is the transpose of A, the diagonal matrix K is defined as
K1 . . . 0 K = ... . . . ... ,
0 . . . Kn
and column vectors Tout , Tsup and P are defined as
Tout = [To1ut , . . . , Tonut ], Tsup = [Tsup, . . . , Tsup], and P = [P1, . . . , Pn],
respectively. Eq. (6) can also be rewritten as
K(Tout - Tsup) = AK(Tout - Tsup) + P ,
(7)
which is referred to as the system function of the DC.
3.1.2. Determining the cross-interference matrix In Eq. (7), we can consider P as the system input and Tout as
the system output. We can eliminate the term Tsup using a known
reference power distribution P ref as follows:
K(Tout - Troeuft ) = AK(Tout - Troeuft ) + P - P ref .
(8)
Eq. (8) denotes a system of n equations with cross-interference matrix A comprising n2 unknown parameters. We can determine the elements of A by generating n systems of Eq. (8) using n different known power distribution vectors P as follows:
K(Tout - Troeuft ) = AK(Tout - Troeuft ) + P - Pref ,
(9)
where the matrices Tout, P, Troeuft , and Pref are defined as
Tout
=
[T 1out
,
.
.
.
,
T
n out
],
P = [P 1, . . . , P n],
Troeuft = [Troeuft , . . . , Troeuft ], and
Pref = [P r1ef , . . . , P rnef ],
respectively. Eq. (9) can now be rewritten as
A = I - (P - Pref )(Tout - Troeuft )-1K-1.
(10)
The above system of n n equations can be used to determine the elements of the cross-interference matrix A under the assumption that the chosen n different power distributions vectors are mutually independent.
3.2. Optimization problem
In general, the cross-interference matrix A is unique for a given physical rack layout of the DC and depends on the airflow rates through the rack enclosures and perforated tiles on the floor. In this paper, we assume that the physical rack layout is kept unchanged and the airflow rates through the equipment are constant during the equipment upgrade. In reality, new DC equipment can increasingly be operated at higher temperatures allowing minimal change to pressure head of cooling fans [10]. In the case of blade servers,
84
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
servers can be swapped with new ones without changing the fans fixed to their chassis. In this case also the pressure head change of fans is minimal since each fan provides cooling for multiple servers within the chassis. Therefore, our assumption can be considered as reasonable for our purpose. As a consequence, the elements of matrix A remain unchanged while performing load spreading.
Moreover, we assume that the equipment consumes a constant amount of power at idle state and its power consumption increases linearly with the utilization rate. The power consumption of node i can thus be expressed as
Pi = bi + aiFi(t),
(11)
where bi is the power consumption at idle state, ai is the additional power required to achieve 100% server utilization, and Fi(t) is the utilization rate of equipment i at time t. Note that 0 Fi(t) 1. In this work we assume that Fi(t) is the average day utilization rate of equipment i at time t. In most cases, new equipment requires more power than their predecessors and dissipates more heat than the equipment to be replaced. We can account for this fact by increasing bi and ai as necessary, and introducing a new expected utilization rate function Fi(t) in Eq. (11).
When the cross-interference matrix A is known, the following equation can be derived from Eqs. (1) and (7) to predict the inlet temperatures of all equipment nodes:
Tin = Tsup + [(K - AK-1) - K-1]P .
(12)
We then define the optimization function, which is later referred to as fitness function as
min{avg(Tiin)}
(13)
after having changed the power profile vector P according to the properties of the upgraded equipment. Our goal is thus to minimize the average inlet temperature of the equipment nodes in the racks by altering the position of the new equipment to be installed.
Since the rack inlet airflow is a mixture of cold supply air and hot air recirculated from other racks, minimizing the average rack inlet temperature means minimizing hot air recirculation. Therefore, we are able to obtain the best possible cooling performance for the given upgrade. In other words, we place new equipment at locations where it is less prone to recirculating hot air back to rack inlets.
3.3. Particle swarm optimization
We use a particle swarm optimization (PSO) technique for minimizing the fitness function Eq. (13) to optimally spread the increased heat load of the upgraded equipment over the racks of the DC. PSO is a population-based self-adaptive optimization technique that was introduced by Kennedy and Eberhart [5] in 1995 and since then went through several modifications [11-13]. Even though this technique cannot always find the optimum solution, we have chosen this biology-inspired approach because of its simple applicability to the load spreading optimization problem, and the easiness to incorporate non-linear constraints into the iterative search algorithm. PSO technique has also been applied for optimization tasks in related topics such as district heating systems [14] and HVAC systems [15].
Similar to other population-based optimization methods, the proposed PSO algorithm starts with randomly initializing a population of candidate solutions, which are commonly called particles, and computes for these particles the corresponding value given by the fitness function. To further improve the initial solution, the PSO algorithm iteratively moves around the particles in a pre-defined search space by imitating the social behavior of animals in a swarm, whereas the movement of each particle is influenced by its own best
known position and also the best known position found by other particles. The swarm characterized by the locations and velocities of the particles at a given time instant is often referred to as a generation. The velocity of the particles can be used as a termination condition of the algorithm, as the velocities of the particles approach to zero if they reach the global optimum of the fitness function.
In our PSO algorithm, the trajectory of each particle in the search space is adjusted by dynamically altering the velocity of each particle according to its best known position and the best known positions found by other particles in the search space. Let us denote the position and velocity vector of the ith particle in a
d-dimensional space by X i = [xi1, . . . , xid] and V i = [vi1, . . . , vid],
respectively. According to the fitness function, the best fitness value obtained by the particle up to time t is Si = [si1, . . . , sid] and the fittest particle found in the entire swarm up to time t is Sg = [sg1, . . . , sgd]. Then, the new velocities and the position of the particles of the next generation at time t + 1 can be calculated as
vij(t + 1) = vij(t) + c1R1[sij - xij(t)] + c2R2[sgj - xij(t)],
(14)
xij(t + 1) = xij(t) + vij(t + 1),
(15)
where c1 and c2 are known as acceleration coefficients, R1 and R2 are two independent random variables uniformly distributed in the range [0, 1] and j is an integer in the range [1, d].
To optimize the load spreading problem with the PSO technique, we model the change of the location of the new equipment during the heat load optimization as a change of the positions of particles in a PSO search space. The position vector X i = [xi1, . . . , xik] of particle i contains the physical locations xij, j = 1, . . ., k, of the new equip-
ment, while the elements vij of the velocity vector V i = [vi1, . . . , vik]
reflect the rate of the position change of equipment being upgraded represented by the elements of X i. The fitness value for each new equipment configuration X i can be computed by using Eqs. (12) and (13). Note that the dimension of the optimization is given by the number k of equipment nodes to be upgraded and is always less than the total number n of nodes available in the DC.
4. Case study
We applied the novel load spreading technique for minimizing the thermal impact of computing equipment upgrades that was introduced in the previous section to a data center at the IBM Zurich Research Laboratory (ZRL). Fig. 2 shows the floor plan of the DC, which has a raised floor area of 286 m2 and houses 27 racks with servers. A summary of the equipment in the DC is given in Table 1.
We investigated the effect of replacing a number of servers in the existing DC with new ones that had similar airflow characteristics but considerably higher heat load. To apply our optimization approach, we first determined the parameters of the heat-flow model. We considered 10 racks in the DC for this case study. The area of interest is highlighted in Fig. 2. Each rack houses a number of servers. This number varies from rack to rack since there are various types of servers of different sizes. To reduce the complexity of the model, we divided each rack into 3 blocks. Each of these blocks corresponds to one node in the heat-flow model with a total of 30 nodes. The calculation of the cross-interference matrix for the heat-flow model required measurement data for 30 different heat load scenarios in the DC. We collected one set of measurement data for one distinct heat load scenario and obtained the rest of the data with support of CFD simulations. We then built a detailed 3dimensional model of the room and used the measurement data to define the boundary conditions for the CFD simulations. We used the commercial CFD software ANSYS CFX [16] to determine the temperature distribution for the different heat load scenarios. A similar approach is described in [17]. Based on the measurement
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
85
Fig. 2. Floor plan of the data center.
data and CFD simulation results, we calculated the parameters for the heat-flow model. Finally, we applied the novel load spreading technique to determine the optimal location for the new servers and evaluated the results.
Section 4.1 describes the measurement campaign that we conducted in the IBM ZRL DC to determine the temperature distribution within the room, airflow rates through perforated floor tiles, ceiling vents, server racks and CRACs, and power consumption of the servers in the DC. In Section 4.2, we give an overview of the CFD modeling and compare simulation results with actual measurement data for verification. Parameter identification for the heat-flow model is described in detail in Section 4.3. Finally, we
Table 1 Overview of data center layout.
Item
DC room raised floor area DC room height Server racksa Racks with networking equipment and storage equipment Computer room air conditioners (CRACs) Power distribution units (PDUs) Perforated floor tiles Ceiling vents
a Includes two high performance computers (HPCs).
Value
18 m 15.9 m 2.7 m 27 9 6 4 92 25
illustrate the optimization approach in Section 4.4 and discuss the results in Section 4.5.
4.1. Measurement campaign
We used IBM's Mobile Measurement Technology (MMT) to obtain the temperature distribution in the data center. This technology has been successfully used to obtain a temperature distribution in DCs [18]. Fig. 3 shows a MMT measurement cart with a total of 72 temperature sensors. Nine sensors are mounted on each of the eight quadratic metal frames to measure the temperature information at the respective height. The cart is on wheels and can thus be readily moved from tile to tile to record the temperature information above every unoccupied floor tile. This allowed us to create a 3-dimensional temperature map of the DC with a spatial resolution of 20 cm in horizontal x- and y-direction and 30 cm in vertical z-direction.
To measure the volumetric airflow rates through perforated floor tiles, ceiling vents, server racks and CRACs, we used a Shortridge Instruments FlowHood CFM-88L [19]. This air capture hood is equipped with a 16-point flow-sensing grid to measure the dynamic pressure in front of an air inlet or outlet. A floor tile of size 60 cm 60 cm is entirely covered by the hood, while two measurements are required to cover a ceiling vent, and three measurements are required to cover the rear side of a rack. The three
86
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
Fig. 4. Model of the data center used for CFD simulations.
Table 2 Material properties of fluid domain in CFD simulations.
Item
Material Reference state Density Molar mass Specific heat at constant pressure Dynamic viscosity Thermal conductivity
Value and/or description
Ideal gas 25 C and 1 atm 1.2 kg m-3 28.96 kg kmol-1 1.0044 kJ kg-1 K-1 1.831 10-5 kg m-1 s-1 2.61 10-2 W m-1 K-1
Fig. 3. MMT measurement cart [2] used to obtain the 3-dimensional temperature distribution in the data center.
measurements at the rear of the racks correspond to the airflow through the three blocks as defined at the beginning of this section. When an air capture hood is in place, the airflow is always reduced to some degree. The hood is equipped with flaps that allow us to take two separate measurements with defined air resistance for compensating this effect. Based on the two measurements, the device automatically calculates the volumetric airflow rate. The power consumption per rack was read from the power meters at the power distribution units. For our model, we assumed that the power consumption of the individual servers was proportional to the corresponding measured airflow rates.
4.2. Model for computational fluid dynamics simulations
Computational fluid dynamics based models can be used to simulate the complex airflow patterns and temperature distribution in data centers as evidenced by number of previous studies [20,21]. We used the CFD software ANSYS CFX to determine the temperature distribution for 30 different heat load scenarios in the DC. The model was validated by comparing simulation results with actual measurements.
Fig. 4 shows the geometry of the DC. The fluid domain had to be discretized into a number of control volumes to solve the governing equations numerically; we created a mesh of 627,076 equilateral hexahedral elements. Increasing the number of elements did not change the simulation results significantly, thus grid independence was achieved. The properties of fluid domain are given in Table 2. Table 3 gives an overview of the measurement data used to define
the boundary conditions of the CFD simulations. The measured volumetric airflow out of the DC was about 20% higher than the total measured airflow into the DC. This difference can be explained by air leaking through openings in the floor that were not covered by our measurements. In particular, a significant amount of air streams from the plenum directly into the racks because of openings underneath the racks used for cabling. We therefore extended Eq. (1) to account for this leakage airflow, i.e.
T i = fiin T i + flieak T + Pi , out foiut in foiut leak cpfoiut
(16)
where fiin is the airflow rate through the front of the rack into node i, flieak is the airflow rate through the floor into the node, and foiut is the airflow rate out of the node. The temperature of the air leaking into the racks Tleak was assumed to be the same as the supply air temperature Tsup. With this assumption, the rest of the heat-flow
Table 3 Data used to define boundary conditions for the CFD simulations.
Parameter
Value
Minimum airflow rate at the server inletsa Maximum airflow rate at the server inletsa Total volumetric airflow into the DC through perf. floor tilesa Total volumetric airflow into the DC through HPCsa Total volumetric airflow out of the DC through ceiling ventsa Total volumetric airflow into the DC through leakageb Flow rate of air leaking into the racksc
Supply air temperature Total power consumption by IT equipmentd
0 m s-1 0.73 m s-1 4.725 m3 s-1 7.670 m3 s-1 15.720 m3 s-1 3.325 m3 s-1 0.139 m s-1
16.5 C
219.73 kW
a Measured value. b Calculated value: total outflow minus total inflow. c Calculated value: total volumetric flow divided by total rack outlet area. d Including 140 kW consumed by two HPCs.
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
87
Fig. 5. Comparison of CFD simulation results (left) with measurement data obtained with the MMT cart (right).
model remains unchanged and Eq. (10) for calculating the crossinterference matrix is still valid.
ANSYS CFX was used to numerically solve the Reynolds Averaged Navier-Stokes equations with the Shear Stress Transport (SST) turbulence model to determine the airflow field and temperature distribution. The convergence criterion for the maximum normalized values of the equation residuals was set to 0.0001. The model converged in roughly 20 h on a workstation with a 3 GHz dual core processor.
Fig. 5 shows a comparison of CFD simulation results and measurement data. The temperature distributions match well except for two locations where the simulation indicates significantly higher temperatures around some rack outlets. This is partly because of the simplified model with 3 blocks of servers per rack.
As described in the previous section, the power consumption was measured per rack and then distributed among the 3 nodes based on the assumption that the heat load of the individual servers is proportional to the measured airflow.
4.3. Identification of cross-interference matrix
To calculate the cross-interference matrix that describes the heat-flow among the 10 racks in the lower right corner of the data center, we used the CFD model introduced in the previous section. We set the power consumption of each node to Pmax = 2000 W at full load and to Pidle = 500 W at idling state. The other parameters remained unchanged to match the operating conditions that
88
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
Table 4 Overview of heat-flow model.
Item
Number of racks Number of equipment nodes per rack Total number of nodes Power consumption of a node at full load Power consumption of a node at idle state
Value
10 3 30 2000 W 500 W
were determined during the measurement campaign. Table 4 gives a summary of the parameters for the heat-flow model.
We conducted a CFD simulation to obtain the temperature distribution for the reference state, in which all servers are idling, i.e.
P ref = [Pidle, . . . , Pidle].
(17)
To calculate the cross-interference matrix with dimension 30 30, the node inlet and outlet temperatures have to be determined for 30 distinct operating points, which are defined by linearly independent power distribution vectors. Furthermore, to minimize the influence of small errors in the CFD simulation on the resulting cross-interference matrix, it is essential that the operating points sufficiently differ from each other. In each case J = 1, . . ., 30, we therefore set the power consumption of exactly one node to maximum load, i.e. PJ = Pmax, while the other 29 nodes were idling, i.e. Pj = Pidle, where 1 j 30 and j =/ J. The power distribution vector P J for case J therefore was
P J = [Pidle, . . . , Pidle, PJ = Pmax, Pidle, . . . , Pidle].
(18)
By using the CFD simulation results from the reference state to define the initial conditions, we significantly reduced the simulation time, since we only changed the power consumption of one node at a time: The simulation for the reference state converged in roughly 20 h, while the subsequent simulations only required about 6.5 h each.
The cross-interference matrix of the heat-flow model could then be calculated by using Eq. (10). The resulting coefficients are shown in Fig. 6. It is clear that along the diagonal the peaks are higher than in other areas of the bar chart. This implies that the hot air recirculation of a node causes the most impact on inlet temperatures of adjacent nodes. Fig. 7 gives an analysis on the recirculation that occurs at each layer of equipment nodes. This indicates that more hot air recirculation occurs at higher layers of server nodes. These results are consistent with the temperature maps obtained during the measurement campaign.
4.4. Optimization
We applied the novel load spreading technique to determine the optimum placement for upgraded equipment in the data center. As shown in Table 5, we investigated 5 different upgrading scenarios S1, S3, S5, S7, and S10, replacing 1, 3, 5, 7, and 10 out of 30 nodes.
For each upgrading scenario we considered 4 node positioning vectors Un,k, k = 1, 2, 3, 4, where n = 1, 3, 5, 7, 10 is the number of nodes to be upgraded. Each element of Un,k corresponds to one of the 30 available node locations shown in Fig. 8 for the new equipment. Table 5 gives an overview of the upgrading scenarios and includes the number of possible combinations for the node positioning vector for each scenario. For each upgrading scenario, the first three node positioning vectors, Un,1, Un,2, Un,3, were selected randomly for comparison, while Un,4 is a result of the particle swarm optimization of the proposed load spreading technique.
The parameters for the PSO are given in Table 6. The power consumption of each new node was set to 500 W at idling state and to 3000 W at full load. We assumed a constant server utilization rate of 60%. In each iteration step of the PSO, the fitness of all particles was assessed by using the fitness function stated in Eq. (13). We defined some penalty functions in order to avoid undesired solutions. For example, the particle with the lowest average inlet temperature could represent an undesired positioning vector, i.e. lead to a situation where some node inlet temperatures heavily
0. 2 0.18 0.16 0.14 0.12 0.1 0.08 0.06 0.04 0.02
0
5
Serv10 to heear nod 15
t receirsc con 20 ulatitornibuting 25 30
30
25
15
2c0eiving
10 ver noldaetesdreheat
5
Serrecircu
Fig. 6. Cross-interference matrix A calculated using CFD simulations.
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
89
0.2
0.2
0.2
0.18
0.18
0.18
0.16
0.16
0.16
0.14
0.14
0.14
0.12
0.12
0.12
0.1
0.1
0.1
0.08
0.08
0.08
0.06
0.06
0.06
0.04
0.04
0.04
0.02
0.02
0.02
0
0
0
5
Stoerve1r0 heat nreode1s5 20 circuclaontrib 25 tion uting
7 101316s19r2e22c5e28iavitng S30erv1e4rcinrcoudleated he
re
5
tSoerve10 heart rneod1e5s circulcaotn2t0rib 25 ion uting
8 11141s720r2e32c6e2a9ivting S30erv2ec5rirncoudlaeted he
re
5
Stoerver 10 heatnroedes15 circuclaotniotri2b0ut n ing25
9 1215182r1e24c2e73i0ving
30 3 6
erver
cnuoldaetesd
heat
S recir
(a) Bottom layer
(b) Middle layer
(c) Top layer
Fig. 7. Analysis of hot air recirculation at each layer of equipment nodes by decomposing the cross-interference matrix.
Table 5 Equipment node upgrading scenarios.
Upgrading scenario S1
S3
S5
S7
S10
Number of combinations 310 = 30 30 3 = 4060 30 5 = 142, 506 370 = 2, 035, 800 3100 = 30, 045, 015
Positioning vectors for new equipment
U1,1 = [8]
U1,2 = [16] U1,3 = [23] U1,4 = [3]
U3,1 = [7, 11, 16]
U3,2 = [25, 28, 29] U3,3 = [16, 17, 20] U3,4 = [3, 6, 9]
U5,1 = [7, 10, 11, 13, 16]
U5,2 = [2, 15, 23, 25, 26] U5,3 = [3, 13, 17, 20, 30] U5,4 = [3, 6, 9, 18, 21]
U7,1 = [2, 12, 15, 23, 25, 26, 27]
U7,2 = [4, 7, 20, 22, 24, 28, 30] U7,3 = [7, 10, 11, 13, 16, 23, 26] U7,4 = [1, 3, 6, 9, 15, 18, 21]
U10,1 = [17, 20, 23, 24, 25, 26, 27, 28, 29, 30]
U10,2 = [8, 11, 13, 14, 16, 17, 18, 20, 22, 23] U10,3 = [1, 7, 10, 11, 13, 16, 17, 23, 26, 28] U10,4 = [1, 2, 3, 4, 6, 9, 12, 15, 18, 21]
exceed the maximum allowed value. We therefore added a penalty to the average inlet temperature in Eq. (13) in order to make it very unlikely that such particles were selected as the global best particle. Repeating elements in the positioning vector also lead to
Table 6 Parameters for particle swarm optimization (PSO).
Parameter
Number of particles Problem dimension Maximum number of iterations Range for each element of position vector Range for each element of velocity vector Acceleration coefficient c1 Acceleration coefficient c2
Value
30 1, 3, 5, 7, 10 500 1 xi 30 -10 vi 10 0.5 c1 2.5 0.5 c2 2.5
invalid solutions. This was also countered by adding a penalty to the average inlet temperature.
4.5. Evaluation
We ran CFD simulations for all upgrading scenarios to determine the inlet temperatures and compared the results for all positioning vectors defined in Table 5. Except for the power consumption of the nodes, all model parameters were kept as described in Section 4.3. A summary of the results is given in Table 7.
At the bottom layer, where abundant cold air supply is available, the average inlet temperature increase is marginal in all trials. It never exceeds 0.5 C. However, at the middle and even more at the top layer, differences between random upgrading and the proposed load spreading technique become evident. The number of
Table 7 Averaged equipment node inlet temperatures, average temperature increases, maximum temperature increases and the number of occurrences where inlet temperature increase by more than 0.5 C.
Positioning vector
Bottom layer
Middle layer
Top layer
U1,1 U1,2 U1,3 U1,4
U3,1 U3,2 U3,3 U3,4
U5,1 U5,2 U5,3 U5,4
U7,1 U7,2 U7,3 U7,4
U10,1 U10,2 U10,3 U10,4
Average inlet temperature [ C]
17.00 17.01 17.04 17.00
17.02 17.13 17.01 17.03
17.02 17.09 17.03 17.04
17.13 17.04 17.04 17.07
17.15 17.03 17.08 17.13
Average temperature increase [C]
0.01 0.02 0.05 0.01
0.03 0.14 0.02 0.04
0.03 0.1 0.04 0.05
0.14 0.05 0.05 0.08
0.16 0.04 0.09 0.14
Maximum temperature increase [C]
0.1 0.1 0.1 0.1
0.1 0.3 0.1 0.1
0.1 0.2 0.1 0.1
0.3 0.1 0.1 0.2
0.4 0.1 0.2 0.3
No. of occurrences above 0.5 C increase
0 0 0 0
0 0 0 0
0 0 0 0
0 0 0 0
0 0 0 0
Average inlet temperature [ C]
21.74 21.80 21.80 21.78
21.71 21.94 21.79 21.81
21.90 21.94 21.80 21.90
22.06 21.79 21.79 21.93
22.54 21.98 21.97 21.93
Average temperature increase [C]
0.08 0.14 0.14 0.12
0.05 0.28 0.13 0.15
0.24 0.28 0.14 0.24
0.4 0.13 0.13 0.27
0.88 0.32 0.31 0.27
Maximum temperature increase [C]
0.7 0.9 0.3 0.4
0.3 0.6 0.7 0.3
0.7 0.6 0.5 0.4
0.8 0.4 0.4 0.8
1.6 1.5 0.8 0.7
No. of occurrences above 0.5 C increase
1 1 0 0
0 3 1 0
1 2 1 0
3 0 0 1
9 2 1 2
Average inlet temperature [ C]
26.94 26.97 26.98 26.98
26.99 27.45 27.01 26.97
27.22 27.18 27.10 27.00
27.34 27.20 27.15 27.12
27.61 27.37 27.62 27.27
Average temperature increase [C]
0.11 0.14 0.15 0.15
0.16 0.62 0.18 0.14
0.39 0.35 0.27 0.17
0.51 0.37 0.32 0.29
0.78 0.54 0.79 0.44
Maximum temperature increase [C]
0.9 1 0.5 0.4
0.8 1.4 0.6 0.3
1.3 0.6 1 0.4
0.9 0.9 0.7 0.6
1.9 1.1 1.3 0.7
No. of occurrences above 0.5 C increase
1 1 1 0
1 6 1 0
3 4 2 0
6 3 4 2
7 4 8 7
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
90
18 21
17
24
20
16
23
19 22
3 6
2 5
1 4
27 30
26 29
25 28
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
91
9
12
8
15
11
7
14
10
13
air-cooled data center. Whenever computer equipment has to be upgraded, care has to be paid to avoid significant changes of the airflow and the thermal distribution in the DC room. For this purpose, we developed a novel upgrading strategy that allows replacing the equipment in the computer racks with minimal thermal impact on the existing optimized DC cooling environment. Our approach is based on an abstract heat-flow model of the DC, whose parameters are determined by means of performing a measurement campaign in the DC and with support of CFD simulations. The optimum placement of the new equipment in the racks of the DC is found by applying a particle swarm optimization technique to this model. The effectiveness of our method has been assessed by comparing it to the conventional necessity based equipment upgrading. Based on experiments performed in a production DC and simulations, we demonstrated that it is possible to systematically distribute the inlet temperature increases among equipment nodes in such a way that no hot spots are formed. The proposed method outperforms the commonly applied load spreading technique because our holistic approach takes into account not only the increased heat load of the upgraded equipment, but also the airflow and temperature distribution in the DC room.
Fig. 8. Equipment nodes in the area of interest in the DC (cf. Fig. 2).
occurrences where the inlet temperature increases by more than 0.5 C gets higher as the number of replaced nodes increases. For upgrading scenarios S1, S3 and S5, the results for the positioning vectors obtained with proposed load spreading technique show no inlet temperature increase above 0.5 C. It is also apparent that the proposed method does not lead to inlet temperatures that exceed the maximum allowed inlet temperature. On the other hand, we notice that random equipment upgrades can cause hot spots even if just one server is replaced: For instance, positioning vector U1,2 leads to an inlet temperature increase of more than 0.5 C at the middle and top layers. For the upgrading scenario S7, the result obtained with the proposed load spreading technique is among the best results in terms of average inlet temperature increase. Given that there are more than 2 million possible ways to place the new equipment in that case, PSO achieves a good result after only 500 iterations. For the upgrading scenario S10, large temperature increases at the inlets are inevitable. However, the proposed method spreads the temperature increase evenly across all the nodes. Even though we have a high number of nodes with temperature increase above 0.5 C in the top layer, the proposed method allows us to significantly reduce the maximum temperature increase compared to random upgrading.
It is important to maintain the inlet temperatures within the manufacturer's specified temperature ranges to ensure proper operation of the equipment. Hot spots caused by unplanned equipment upgrades are often countered by DC operators by reducing the supply air temperature. This leads to an overall lower cooling efficiency of the DC. With the proposed novel load spreading technique, DC operators can maintain a stable cooling environment. Cooling system performance deterioration and risk of equipment malfunction is minimized by keeping a uniform equipment inlet temperature distribution.
5. Conclusions
The provision of an efficient cooling environment is important for the reliable and economically successful operation of an
Acknowledgments
The authors would like to acknowledge the support of G08/DAAD grant (2010/2011) on optimization in wireless sensor networks for travel support.
References
[1] R. Schmidt, E. Cruz, M. Iyengar, Challenges of data center thermal management, IBM Journal of Research and Development 49 (2005) 709-723.
[2] H.F. Hamann, J.A. Lacey, M. O'Boyle, R.R. Schmidt, M. Iyengar, Rapid threedimensional thermal characterization of large-scale computing facilities, IEEE Transactions on Components and Packaging Technologies 31 (2008) 444-448.
[3] J. Siriwardana, W. Schott, S. Halgamuge, The power grabbers, IEEE Power and Energy Magazine 8 (2010) 46-53.
[4] N. Rasmussen, Cooling strategies for ultra-high density racks and blade servers, APC White Paper 46 (2009).
[5] J. Kennedy, R. Eberhart, Particle swarm optimization, in: Proceedings of IEEE International Conference on Neural Networks, vol. 4, 1995, pp. 1942-1948.
[6] S.K. Shrivastava, J.W. VanGilder, B.G. Sammakia, Optimization of cluster cooling performance for data centers, in: Proceedings of 11th IEEE Intersociety Conference on Thermal and Thermochemical Phenomena in Electronic Systems, 2008, pp. 1161-1166.
[7] J. Moore, J. Chase, P. Ranganathan, R. Sharma, Making scheduling "cool": temperature-aware workload placement in data centers, in: Proceedings of the Annual Conference on USENIX Annual Technical Conference, USENIX Association, Berkeley, CA, USA, 2005, pp. 61-75.
[8] Q. Tang, S.K.S. Gupta, G. Varsamopoulos, Energy-efficient thermal-aware task scheduling for homogeneous high-performance computing data centers: a cyber-physical approach, IEEE Transactions on Parallel and Distributed Systems 19 (2008) 1458-1472.
[9] Q. Tang, T. Mukherjee, S. Gupta, P. Cayton, Sensor-based fast thermal evaluation model for energy efficient high-performance data centers, in: Proceedings of 4th International Conference on Intelligent Sensing and Information Processing, 2006, pp. 203-208.
[10] M.K. Patterson, The effect of data center temperature on energy efficiency, in: Proceedings of 11th Intersociety Conference on Thermal and Thermomechanical Phenomena in Electronic Systems ITHERM 2008, 2008, pp. 1167-1174.
[11] A. Ratnaweera, S. Halgamuge, H. Watson, Self-organizing hierarchical particle swarm optimizer with time-varying acceleration coefficients, IEEE Transactions on Evolutionary Computation 8 (2004) 240-255.
[12] F. van den Bergh, A. Engelbrecht, A cooperative approach to particle swarm optimization, IEEE Transactions on Evolutionary Computation 8 (2004) 225-239.
[13] S. Patchararungruang, S. Halgamuge, N. Shenoy, Optimized rule-based delay proportion adjustment for proportional differentiated services, IEEE Journal on Selected Areas in Communications 23 (2005) 261-276.
[14] K. Steer, A. Wirth, S. Halgamuge, Control period selection for improved operating performance in district heating networks, Energy and Buildings 43 (2011) 605-613.
92
J. Siriwardana et al. / Energy and Buildings 50 (2012) 81-92
[15] Y. Bichiou, M. Krarti, Optimization of envelope and HVAC systems selection for residential buildings, Energy and Buildings 43 (2011) 3373-3382.
[16] ANSYS CFX, 2011. http://www.ansys.com/Products/Simulation+Technology/ Fluid+Dynamics/ANSYS+CFX.
[17] P. Biller, P. Chevillat, F. de Lorenzi, T. Scherer, W. Schott, R. Ullmann, C. Vmel, Efficient cooling of data centers, in: Proceedings of the 2011 World Engineers' Convention, 2011.
[18] H.F. Hamann, T.G. van Kessel, M. Iyengar, J.-Y. Chung, W. Hirt, M.A. Schappert, A. Claassen, J.M. Cook, W. Min, Y. Amemiya, V. Lpez, J.A. Lacey, M. O'Boyle,
Uncovering energy-efficiency opportunities in data centers, IBM Journal of Research and Development 53 (2009) 10:1-10:12. [19] FlowHood CFM-88L, Shortridge Instruments, 2011. http://www. shortridge.com/flowhood specs.htm. [20] J. Cho, T. Lim, B.S. Kim, Measurements and predictions of the air distribution systems in high compute density (internet) data centers, Energy and Buildings 41 (2009) 1107-1115. [21] J.F. Karlsson, B. Moshfegh, Investigation of indoor climate and power usage in a data center, Energy and Buildings 37 (2005) 1075-1083.