New: failure-mode coverage for transformers and switchgear. See the split →
Data center cooling & power

Power fails instantly.
Cooling warns you first.

That is the single most important fact about your fluid plant. Electrical faults trip. Cooling degrades, and degradation is a measurable signal weeks out, in telemetry you already collect. PRISM reads it, names the failure, and tells you how many days you have.

150+ days
Median warning, measured in water
4-8 wks
Export to calibrated model
$0
To evaluate, with no live connection

Performance figures are measured on Firstlook production water deployments and are labeled as such throughout this page.

The incident everyone remembers
A drain valve halted global futures trading.
AHEAD OF THE FREEZE Cooling tower drain procedure not performed 3:40 AM First cooling units fault. Cascade begins. 6:19 PM Every chiller offline. Hall passes 100°F. NEXT DAY Global markets halted for more than 11 hours. FIFTEEN HOURS OF TELEMETRY. NOBODY WAS READING IT FOR FAILURE.

A 450,000 sq ft, 109 MW campus. Not a small or unsophisticated facility. Listed among the Uptime Institute's ten example Category 5 outages of 2025, and the only cooling failure on that list.

The stakes

One outage in five now costs over $1M.

Half of all major outages clear six figures and one in five clears seven. That is the second year running at that level, drawn from operators reporting on their own most recent incident.

Uptime Institute, n=94 operators ITIC 2024 hourly cost survey
📈

20% clear $1M

Of most recent major outages: 43% under $100K, 37% between $100K and $1M, and 20% above $1M. At that level a single prevented failure pays for years of monitoring.

91% say an hour costs over $300K

Among mid-size and large enterprises, with 41% putting a single hour between $1M and $5M or more. That is one organization's loss, not a whole multi-tenant facility.

A number we refuse to use

The famous “$9,000 a minute” traces to a vendor-sponsored study of 63 US sites and has not been refreshed in a decade. If a vendor quotes it to you this year, they have not checked.

Where cooling ranks

Cooling causes fewer outages than power, and more claims than anything.

Counting outages makes cooling look like the second problem. Counting what insurers actually pay out makes it the first.

Roughly one outage in five

Cooling has held a 13% to 19% band across recent years, second only to power. And power's share fell to 45% from 54%, so cooling's relative weight is rising, not falling.

By incident count

💧

Water damage leads by volume

Across 221 data center claims worth roughly $782M, water damage led at 21%, ahead of willful acts at 19% and fire at 14%. Fire drives severity; water drives frequency.

By claim frequency

🔥

Liquid cooling is already 24% of loss costs

Direct liquid cooling accounts for nearly a quarter of total data center loss costs, before the technology has even finished arriving.

By loss cost

The trend

Density up four times. Ride-through down to seconds.

Air gave you thermal mass and time. Direct-to-chip gives you the coolant in the loop, and nothing else.

  • Under ten seconds of ride-through on a cold plate at full AI load. An industry estimate rather than a benchmark, but not a window a human can act inside.
  • 46% of operators cannot find qualified candidates, and 37% cannot retain the ones they have.
  • 92% of outages have human error as a contributor. A vigilance problem, and vigilance does not scale with density.

If the reaction window has closed, the only remaining move is to open the prediction window instead.

Average rack density, kW
Not to scale.
7 2021 16 2025 27 2026 600 2027 ROADMAP THE FLUID PLANT IS NOW THE SYSTEM MOST LIKELY TO TAKE THE FLOOR DOWN.

Published vendor rack roadmaps and industry density surveys. Bars are indicative rather than proportional.

The opening

Chiller faults trend three to eight weeks out.

Roughly 78% of chiller failure modes give advance warning through parameter trending. That warning window is the whole opportunity, and it is being spent waiting for a threshold alarm.

78%

Of chiller failure modes trend before they fail

3 to 5 times

Cost of the same mechanical repair done as an emergency

28-52 wks

Lead time on a centrifugal chiller. A surprise costs two quarters of capacity

$12K-$45K

An emergency chiller repair event, before downtime or secondary damage

What that warning window is worth. Spent, the repair happens as an emergency at three to five times the cost, with the outage risk attached and the part ordered against a lead time you no longer control. Used, it is a scheduled job in a planned window, at a fraction of the cost, with the chiller on order months before you need it. The repair is the same either way. The only variable is when you found out.

Cooling plant

The control loop masks the decline.

As heat transfer fades, the control system quietly raises flow to hold temperature. Outputs look fine for months while the loop degrades underneath. Then the pumps hit their ceiling and the temperature runs away with no usable warning.

That masked-degradation signature, a loop spending more effort to deliver the same result, is exactly what PRISM already detects in water. Watching effort instead of output is what turns a months-long hidden decline into a flagged, plannable event.

  • Chillers, cooling towers, CRAH and CRAC units, chilled and condenser water pumps, over BACnet or Modbus via the BMS.
  • CDUs and coolant loops over Redfish REST: flow, supply and return temperature, delta pressure, heat removed, coolant level, pump redundancy, leak detection.
  • Coolant degradation and contamination, read as a decline in transfer efficiency rather than a level alarm.
  • Pump and fan bearing wear across the plant, the asset class the models have the deepest production history on.

See the full integration surface

Masked thermal decline
Temperature flat. Effort climbing.
LOOP SUPPLY TEMPERATURE · WITHIN SPEC THERMAL RUNAWAY PUMP FLOW SPENT TO HOLD IT · CLIMBING PUMPS AT CEILING PRISM FLAGS THE MASKED DECLINE

Most tools watch the result. PRISM watches the effort behind it: the pump working harder to hold the same temperature.

Power chain

Redundancy is not the same as warning.

N+1 protects you from the first fault. It does nothing to tell you that fault is coming, or that the backup you are counting on has been quietly degrading since the last transfer test.

🔋

Backups degrade out of sight

UPS strings and standby paths sit idle until they are needed. By the time a transfer exposes a weak string, the load is already at risk. The decline was readable long before the test.

Switchgear gives little notice

Thermal drift, loose connections and insulation breakdown build slowly, then trip suddenly. Threshold alarms catch the trip, not the months of drift that led to it.

A power event is a cluster event

An unplanned interruption does not just drop a feed. It stalls or loses the workload riding on it, and the cost of the lost compute dwarfs the cost of the part.

Power is read from your existing EPMS and DCIM over SNMP and Modbus, alongside cooling, on one platform. You can plan a window around the whole picture rather than one system at a time. The next section takes a transformer and a switchboard apart mode by mode, and says exactly which ones your existing telemetry already sees. See the coverage split

Electrical asset coverage

What your existing telemetry already sees.

Two asset classes, taken from real reliability libraries. For each one, the documented failure modes split three ways. Some have a continuous signature in data your EPMS already streams. Some need one added sensor. Some genuinely need a person or an outage. We are specific about which is which, because that split is the scoping conversation.

Oil-filled transformer

Continuous, no new hardware One added sensor Lab Outage
Failure modeSignature we readCoverage
Cooling system failureFan and pump status with the residual between measured oil-temperature rise and the rise predicted from load and ambient. A seized pump or fouled radiator is a step change in thermal resistance.Existing data
Overload and overheatingLoad current, winding and top-oil temperature, ambient. Accumulated loss of life computes directly from the IEEE thermal model.Existing data
Core saturation and DC biasMagnetising-current harmonics, second-order content and collapsing power factor. All standard outputs of a power-quality meter.Existing data
Insulation ageingCumulative consumption from load and hot-spot temperature. This gives the ageing rate; degree of polymerisation gives the state.Existing data
Thermal fault, bulkTop-oil rise drifting against the modeled rise. A localized hot spot of a few hundred watts will not move bulk oil temperature. That one needs gas analysis.Existing data
Partial discharge, arcingDissolved gas (hydrogen for discharge, acetylene for arcing) or a partial-discharge sensor. Incipient discharge dissipates milliwatts and moves nothing in facility telemetry.Added sensor
Moisture, bushing condition, clampingOnline moisture, bushing leakage current off the capacitance tap, tank accelerometers. Vibration only interprets correctly against load current, which you already have.Added sensor
Oil degradationInterfacial tension, acid number, dielectric strength. No electrical signature until it becomes a cooling problem. But the driver is cumulative thermal exposure, which is predictable from loading history.Lab
Shorted turns, winding deformationTurns ratio, winding resistance, frequency response analysis. Gross cases raise a thermal flag first; confirmation needs de-energisation.Outage

Insurance data on 94 large-transformer failures puts insulation failure at roughly a quarter of events but over half of paid losses. It is the most common failure mode, the most expensive one, and the one with the longest detectable lead time. Annual oil sampling observes the transformer on one day in 365. It catches slow-developing faults well, and fast ones only by chance.

Main switchboard

Continuous, no new hardware One added sensor Truck roll Outage
Failure modeSignature we readCoverage
Protection setting driftTrip-unit and relay settings read over Modbus and diffed continuously against the approved coordination study. Catches a mis-set breaker the day it is changed, not at the next three-year test. The most under-exploited signal on the board.Existing data
Harmonic distortion, power qualityTHD, individual harmonics and demand distortion measured against IEEE 519 limits continuously rather than at a survey.Existing data
Breaker duty and contact wearOperation counts, trip history with cause, accumulated interrupted current. Where a protective relay is present, trip and close coil signatures expose mechanism and lubrication state.Existing data
Protection and control availabilityTrip-circuit supervision, relay watchdog and self-diagnostics, comms heartbeat, control power, settings-change log. Availability of the trip path, continuously. Calibration still needs injection testing.Existing data
Earthing and insulation leakageA rising standing residual on ground-fault protection is a real insulation-leakage trend. Electrode resistance itself is invisible.Existing data
Moisture and condensationSpace-heater circuit integrity from existing current data. A failed heater is detectable today. Enclosure humidity itself needs a sensor.Existing data
Loose and high-resistance connectionsOne temperature input per critical joint, trended as temperature rise normalized by current squared. See below. This is the important one.Added sensor
Partial discharge, contact breakdownTransient earth voltage, high-frequency CT or ultrasonic. No electrical signature at the metering point until it flashes over.Added sensor
Protection calibration, mechanism conditionSecondary injection and functional operation. A breaker that never operates has a hidden failure; only operating it reveals hardened lubrication.Truck roll
Insulation and contact resistanceInsulation resistance testing, contact resistance, servicing and lubrication.Outage
🌡
Load current you already have. Temperature you do not. Loose and high-resistance connections are the most frequent electrical failure mode, and they fall just outside what existing data covers. An annual infrared scan is a snapshot taken at whatever load happened to be running. A joint that would be critical at full load can read clean at thirty percent, and joints behind barriers or boots are not visible at all. One temperature input per critical joint converts that into a continuous trend that is load-normalized by construction. It is the highest-leverage sensor addition in the asset class, and NFPA 70B’s 2026 edition now permits permanently installed thermal monitoring in place of the field thermography task.

Sources. Failure modes and task intervals from published reliability-centered maintenance libraries for these asset classes. Diagnostic mapping against IEEE C57.91 (thermal and loss of life), IEEE C57.104 and IEC 60599 (dissolved gas), IEEE 519 (harmonics) and ANSI/NETA MTS (thermographic severity). The coverage classification, meaning which modes we can read from existing telemetry, is our own engineering judgment. We will walk through it asset by asset on a call.

One scoping detail worth knowing early. Harmonics and power quality are only exposed by higher-tier trip units. On basic trip units you get current, voltage, power, trip history, breaker status and settings, but no power quality. Harmonic-driven modes need a separate meter on those boards. We establish this on the scoping call rather than discovering it in week three.

Size it yourself

Priced on your hourly number.

We will not hand you a dollars-per-minute rate, because the honest ones do not exist and the famous one is a decade stale. You know what an hour costs you. Put it in.

The second half of the answer is the part nobody counts: asset-level failures caught before they escalate. In the validated water deployment that was worth $1,616 per monitored asset per year, and almost none of it came from avoided outages. It came from exactly the category redundancy hides in a data center.

That number was earned in an environment where a failure costs far less than it does on your floor. We use it as a floor, not a target.

Per-asset rate: operator verified Downtime exposure: your inputs

Prevented-downtime calculator

Two components: the downtime you avoid, priced at your own hourly cost, and the maintenance cost you avoid, priced at the rate we actually measured.

Exposure per event$0
Downtime exposure avoided, year one$0
Maintenance & emergency-repair cost avoided$0
Total avoided cost, year one$0
Five-year cumulative$0
Downtime hours prevented0

Basis. Downtime exposure is entirely your input: hourly cost, duration, events avoided. The presets are anchored to published survey bands: 91% of mid-size and large enterprises put an hour above $300K, and 41% put it between $1M and $5M or more. Maintenance avoidance is anchored to $1,616 per monitored asset per year, reconciled across 391 assets over nine months in a water deployment, and applied here unchanged as a conservative floor. Excludes secondary damage, expedited freight, capacity loss from chiller lead times, and insurance effects.

Asset-class mapping

It is the same equipment, in a better-instrumented building.

Your cooling plant is motor-driven pumps and fans moving a fluid around a closed loop, under a control system that compensates as performance fades. That is the exact asset class PRISM has the deepest production history on. Your instrumentation is denser and faster than a rural lift station on a radio link, so the models have more to work with here, not less.

Validated in waterData center equivalentTransferWhat changes
Centrifugal pump, motor-drivenChilled and condenser water pumps, CDU pumpsDirectNothing material. Same device, same signature, different loop.
Motor and VFD driveCRAH fans, tower fans, pump drivesDirectA fan in a server and a fan in a treatment plant are the same device at different scale.
Closed loop with a circulating mediumCoolant loop, chilled water loopDirectCoolant chemistry replaces water chemistry. Masked-degradation signature is identical.
Control loop masking declineBMS raising flow to hold setpointEasierDenser instrumentation and faster sampling than a remote lift station. More signal, not less.
Current through a distribution assetUPS strings, switchgear, distributionEasierContinuously monitored on-site, rather than over an intermittent rural radio link.

What we will not do is quote you a number. The per-asset rate on this page was measured on someone else's fleet, in an environment where failures cost less than they cost you. It anchors the conversation. Your figure comes out of your own history, in the retrospective, and that is the only one worth deciding on.

How it works

Ingest once. Decide five times.

PRISM reads the BMS, CDU, EPMS and DCIM telemetry you already collect, under a read-only credential, and delivers the decision into the system your team already runs.

The pipeline
Ingest once. Decide five times.
SOURCES SCADA / historian BMS / DCIM EPMS / RTU Work-order text READ-ONLY PRISM ENGINE DetectSignature drift, multivariate ClassifyProbable failure mode ForecastDays to failure, with band Triage & rank DISPATCH CMMS work order SMS to on-call Voice escalation Email / chat digest API / webhook NO WRITE PATH TO YOUR CONTROLS

One-directional ingest in, decisions out. Nothing in the product can change a setpoint, close a valve or start a machine.

See the engine in detail

The evaluation

Four to eight weeks, offline, at no cost.

Designed to clear your security review by not needing it. Nothing connects, nothing installs, nothing writes.

Call · 45 min

Scope the export and agree the asset list

One conversation with operations. What you run, what you already trend, and which failures you remember well enough to score us against.

Week 1

Export

Six to twelve months of BMS trend logs and DCIM history as flat files, under NDA. No connection to your environment, no credential issued, no agent installed.

Weeks 2 to 4

Calibrate

Models tuned to your plant, your naming conventions and your asset classes, with sensitivity set against your real cost of intervention.

Weeks 4 to 6

Look back

A retrospective against the incidents in your own history. What would PRISM have said, and when would it have said it?

Weeks 7 to 8

Score, then decide

Hits, lead times and misses on the same page, against incidents you can check us on. No exclusivity, no obligation, and no invoice for any of it. Design-partner terms are materially better than list.

🔒
Certifications, in writing. We will state our current SOC 2 and ISO 27001 position directly and in writing at the scoping call, rather than putting a badge on a web page, because the honest answer belongs in a document you can hold us to. In production, access is a read-only credential on the historian or API. No write path exists in the product.
FAQ

Common questions from operators

No. PRISM is software only. It reads the BMS, CDU, pump and loop telemetry you already collect. There is nothing to install in the white space and no capital cost to deploy. A chiller on a current controller typically exposes fifty to sixty usable points, which is more than enough.

It cannot touch your controls, and that is deliberate: no write path to plant exists in the product. What it does is get the decision to the people and systems that act. That means a work order in your CMMS or EAM with priority and scope, an SMS to the on-call engineer, a voice escalation when severity demands it, or a webhook into your dispatch queue. Your compute and facilities teams get the lead time; the intervention stays under human control.

BMS alarms fire on thresholds, which is to say after the fact, and OEM tools stay tied to their own hardware. PRISM is vendor-neutral across your whole mixed fleet, models the joint signature rather than single channels, and catches the masked decline before any threshold trips. It also reduces alert volume rather than adding to it. Sites routinely carry hundreds of active nuisance alarms, and an untuned alert stream only compounds that.

Yes, and commissioning is arguably the best time to start. PRISM establishes the healthy signature bands at go-live. That validates the new plant against how it was designed to behave, and it gives the model a baseline from day zero instead of waiting a year to learn what normal looks like.

You should not take it on our word. That is what the retrospective is for, and it runs on your data at no cost. The argument is that the asset class is the same: motor-driven pumps and fans moving a fluid around a closed loop, with a control system compensating for decline. And your instrumentation is denser and faster than a remote lift station on a radio link. The data is better, not worse. The stakes are what changes.

Reference rights, two or three quotable metrics, and a case study you review and approve before anything is published. In exchange, design-partner terms are materially better than list. If the retrospective does not convince you, you owe us nothing and we do not publish anything.

One call. Everything after that is measurement.

Forty-five minutes with operations, six to twelve months of history under NDA, and a scored retrospective against incidents you already remember.

$0, offline, no exclusivity, no obligation.

Sources

Outage and cost data. Outage cost distribution and cause ranking from Uptime Institute Annual Outage Analysis 2026 (n=94 operators) and Global Survey 2025. Hourly cost bands from the ITIC 2024 Hourly Cost of Downtime Survey. Claim frequency and loss cost from Allianz Commercial's data center claims review, and from Swiss Re and FM Global on direct liquid cooling. Rack density and staffing from AFCOM State of the Data Center 2026 and published vendor rack roadmaps. Incident detail from contemporaneous public reporting and the operator's own statement.

Diagnostic and maintenance standards. Failure modes and task intervals are drawn from published reliability-centered maintenance libraries for these asset classes. Diagnostic mapping references IEEE C57.91 (thermal and loss of life), IEEE C57.104 and IEC 60599 (dissolved gas), IEEE 519 (harmonics) and ANSI/NETA MTS (thermographic severity). Telemetry standards: DMTF Redfish DSP2064, ASHRAE Guideline 36, BRICK schema. Transformer failure economics from published insurance loss analysis. Equipment lead times from Wood Mackenzie survey data reported in trade press.

Estimates. Chiller warning rates, emergency-versus-planned repair multiples and thermal ride-through are industry and trade-source estimates rather than measured benchmarks. The classification of which failure modes can be read from existing telemetry is Firstlook's own engineering assessment, and we review it asset by asset on request. Performance figures are measured on Firstlook production water deployments.