That is the single most important fact about your fluid plant. Electrical faults trip. Cooling degrades, and degradation is a measurable signal weeks out, in telemetry you already collect. PRISM reads it, names the failure, and tells you how many days you have.
Performance figures are measured on Firstlook production water deployments and are labeled as such throughout this page.
A 450,000 sq ft, 109 MW campus. Not a small or unsophisticated facility. Listed among the Uptime Institute's ten example Category 5 outages of 2025, and the only cooling failure on that list.
Half of all major outages clear six figures and one in five clears seven. That is the second year running at that level, drawn from operators reporting on their own most recent incident.
Of most recent major outages: 43% under $100K, 37% between $100K and $1M, and 20% above $1M. At that level a single prevented failure pays for years of monitoring.
Among mid-size and large enterprises, with 41% putting a single hour between $1M and $5M or more. That is one organization's loss, not a whole multi-tenant facility.
The famous “$9,000 a minute” traces to a vendor-sponsored study of 63 US sites and has not been refreshed in a decade. If a vendor quotes it to you this year, they have not checked.
Counting outages makes cooling look like the second problem. Counting what insurers actually pay out makes it the first.
Cooling has held a 13% to 19% band across recent years, second only to power. And power's share fell to 45% from 54%, so cooling's relative weight is rising, not falling.
By incident count
Across 221 data center claims worth roughly $782M, water damage led at 21%, ahead of willful acts at 19% and fire at 14%. Fire drives severity; water drives frequency.
By claim frequency
Direct liquid cooling accounts for nearly a quarter of total data center loss costs, before the technology has even finished arriving.
By loss cost
Air gave you thermal mass and time. Direct-to-chip gives you the coolant in the loop, and nothing else.
If the reaction window has closed, the only remaining move is to open the prediction window instead.
Published vendor rack roadmaps and industry density surveys. Bars are indicative rather than proportional.
Roughly 78% of chiller failure modes give advance warning through parameter trending. That warning window is the whole opportunity, and it is being spent waiting for a threshold alarm.
Of chiller failure modes trend before they fail
Cost of the same mechanical repair done as an emergency
Lead time on a centrifugal chiller. A surprise costs two quarters of capacity
An emergency chiller repair event, before downtime or secondary damage
What that warning window is worth. Spent, the repair happens as an emergency at three to five times the cost, with the outage risk attached and the part ordered against a lead time you no longer control. Used, it is a scheduled job in a planned window, at a fraction of the cost, with the chiller on order months before you need it. The repair is the same either way. The only variable is when you found out.
As heat transfer fades, the control system quietly raises flow to hold temperature. Outputs look fine for months while the loop degrades underneath. Then the pumps hit their ceiling and the temperature runs away with no usable warning.
That masked-degradation signature, a loop spending more effort to deliver the same result, is exactly what PRISM already detects in water. Watching effort instead of output is what turns a months-long hidden decline into a flagged, plannable event.
Most tools watch the result. PRISM watches the effort behind it: the pump working harder to hold the same temperature.
N+1 protects you from the first fault. It does nothing to tell you that fault is coming, or that the backup you are counting on has been quietly degrading since the last transfer test.
UPS strings and standby paths sit idle until they are needed. By the time a transfer exposes a weak string, the load is already at risk. The decline was readable long before the test.
Thermal drift, loose connections and insulation breakdown build slowly, then trip suddenly. Threshold alarms catch the trip, not the months of drift that led to it.
An unplanned interruption does not just drop a feed. It stalls or loses the workload riding on it, and the cost of the lost compute dwarfs the cost of the part.
Power is read from your existing EPMS and DCIM over SNMP and Modbus, alongside cooling, on one platform. You can plan a window around the whole picture rather than one system at a time. The next section takes a transformer and a switchboard apart mode by mode, and says exactly which ones your existing telemetry already sees. See the coverage split
Two asset classes, taken from real reliability libraries. For each one, the documented failure modes split three ways. Some have a continuous signature in data your EPMS already streams. Some need one added sensor. Some genuinely need a person or an outage. We are specific about which is which, because that split is the scoping conversation.
| Failure mode | Signature we read | Coverage |
|---|---|---|
| Cooling system failure | Fan and pump status with the residual between measured oil-temperature rise and the rise predicted from load and ambient. A seized pump or fouled radiator is a step change in thermal resistance. | Existing data |
| Overload and overheating | Load current, winding and top-oil temperature, ambient. Accumulated loss of life computes directly from the IEEE thermal model. | Existing data |
| Core saturation and DC bias | Magnetising-current harmonics, second-order content and collapsing power factor. All standard outputs of a power-quality meter. | Existing data |
| Insulation ageing | Cumulative consumption from load and hot-spot temperature. This gives the ageing rate; degree of polymerisation gives the state. | Existing data |
| Thermal fault, bulk | Top-oil rise drifting against the modeled rise. A localized hot spot of a few hundred watts will not move bulk oil temperature. That one needs gas analysis. | Existing data |
| Partial discharge, arcing | Dissolved gas (hydrogen for discharge, acetylene for arcing) or a partial-discharge sensor. Incipient discharge dissipates milliwatts and moves nothing in facility telemetry. | Added sensor |
| Moisture, bushing condition, clamping | Online moisture, bushing leakage current off the capacitance tap, tank accelerometers. Vibration only interprets correctly against load current, which you already have. | Added sensor |
| Oil degradation | Interfacial tension, acid number, dielectric strength. No electrical signature until it becomes a cooling problem. But the driver is cumulative thermal exposure, which is predictable from loading history. | Lab |
| Shorted turns, winding deformation | Turns ratio, winding resistance, frequency response analysis. Gross cases raise a thermal flag first; confirmation needs de-energisation. | Outage |
Insurance data on 94 large-transformer failures puts insulation failure at roughly a quarter of events but over half of paid losses. It is the most common failure mode, the most expensive one, and the one with the longest detectable lead time. Annual oil sampling observes the transformer on one day in 365. It catches slow-developing faults well, and fast ones only by chance.
| Failure mode | Signature we read | Coverage |
|---|---|---|
| Protection setting drift | Trip-unit and relay settings read over Modbus and diffed continuously against the approved coordination study. Catches a mis-set breaker the day it is changed, not at the next three-year test. The most under-exploited signal on the board. | Existing data |
| Harmonic distortion, power quality | THD, individual harmonics and demand distortion measured against IEEE 519 limits continuously rather than at a survey. | Existing data |
| Breaker duty and contact wear | Operation counts, trip history with cause, accumulated interrupted current. Where a protective relay is present, trip and close coil signatures expose mechanism and lubrication state. | Existing data |
| Protection and control availability | Trip-circuit supervision, relay watchdog and self-diagnostics, comms heartbeat, control power, settings-change log. Availability of the trip path, continuously. Calibration still needs injection testing. | Existing data |
| Earthing and insulation leakage | A rising standing residual on ground-fault protection is a real insulation-leakage trend. Electrode resistance itself is invisible. | Existing data |
| Moisture and condensation | Space-heater circuit integrity from existing current data. A failed heater is detectable today. Enclosure humidity itself needs a sensor. | Existing data |
| Loose and high-resistance connections | One temperature input per critical joint, trended as temperature rise normalized by current squared. See below. This is the important one. | Added sensor |
| Partial discharge, contact breakdown | Transient earth voltage, high-frequency CT or ultrasonic. No electrical signature at the metering point until it flashes over. | Added sensor |
| Protection calibration, mechanism condition | Secondary injection and functional operation. A breaker that never operates has a hidden failure; only operating it reveals hardened lubrication. | Truck roll |
| Insulation and contact resistance | Insulation resistance testing, contact resistance, servicing and lubrication. | Outage |
Sources. Failure modes and task intervals from published reliability-centered maintenance libraries for these asset classes. Diagnostic mapping against IEEE C57.91 (thermal and loss of life), IEEE C57.104 and IEC 60599 (dissolved gas), IEEE 519 (harmonics) and ANSI/NETA MTS (thermographic severity). The coverage classification, meaning which modes we can read from existing telemetry, is our own engineering judgment. We will walk through it asset by asset on a call.
One scoping detail worth knowing early. Harmonics and power quality are only exposed by higher-tier trip units. On basic trip units you get current, voltage, power, trip history, breaker status and settings, but no power quality. Harmonic-driven modes need a separate meter on those boards. We establish this on the scoping call rather than discovering it in week three.
We will not hand you a dollars-per-minute rate, because the honest ones do not exist and the famous one is a decade stale. You know what an hour costs you. Put it in.
The second half of the answer is the part nobody counts: asset-level failures caught before they escalate. In the validated water deployment that was worth $1,616 per monitored asset per year, and almost none of it came from avoided outages. It came from exactly the category redundancy hides in a data center.
That number was earned in an environment where a failure costs far less than it does on your floor. We use it as a floor, not a target.
Two components: the downtime you avoid, priced at your own hourly cost, and the maintenance cost you avoid, priced at the rate we actually measured.
Basis. Downtime exposure is entirely your input: hourly cost, duration, events avoided. The presets are anchored to published survey bands: 91% of mid-size and large enterprises put an hour above $300K, and 41% put it between $1M and $5M or more. Maintenance avoidance is anchored to $1,616 per monitored asset per year, reconciled across 391 assets over nine months in a water deployment, and applied here unchanged as a conservative floor. Excludes secondary damage, expedited freight, capacity loss from chiller lead times, and insurance effects.
Your cooling plant is motor-driven pumps and fans moving a fluid around a closed loop, under a control system that compensates as performance fades. That is the exact asset class PRISM has the deepest production history on. Your instrumentation is denser and faster than a rural lift station on a radio link, so the models have more to work with here, not less.
| Validated in water | Data center equivalent | Transfer | What changes |
|---|---|---|---|
| Centrifugal pump, motor-driven | Chilled and condenser water pumps, CDU pumps | Direct | Nothing material. Same device, same signature, different loop. |
| Motor and VFD drive | CRAH fans, tower fans, pump drives | Direct | A fan in a server and a fan in a treatment plant are the same device at different scale. |
| Closed loop with a circulating medium | Coolant loop, chilled water loop | Direct | Coolant chemistry replaces water chemistry. Masked-degradation signature is identical. |
| Control loop masking decline | BMS raising flow to hold setpoint | Easier | Denser instrumentation and faster sampling than a remote lift station. More signal, not less. |
| Current through a distribution asset | UPS strings, switchgear, distribution | Easier | Continuously monitored on-site, rather than over an intermittent rural radio link. |
What we will not do is quote you a number. The per-asset rate on this page was measured on someone else's fleet, in an environment where failures cost less than they cost you. It anchors the conversation. Your figure comes out of your own history, in the retrospective, and that is the only one worth deciding on.
PRISM reads the BMS, CDU, EPMS and DCIM telemetry you already collect, under a read-only credential, and delivers the decision into the system your team already runs.
One-directional ingest in, decisions out. Nothing in the product can change a setpoint, close a valve or start a machine.
Designed to clear your security review by not needing it. Nothing connects, nothing installs, nothing writes.
One conversation with operations. What you run, what you already trend, and which failures you remember well enough to score us against.
Six to twelve months of BMS trend logs and DCIM history as flat files, under NDA. No connection to your environment, no credential issued, no agent installed.
Models tuned to your plant, your naming conventions and your asset classes, with sensitivity set against your real cost of intervention.
A retrospective against the incidents in your own history. What would PRISM have said, and when would it have said it?
Hits, lead times and misses on the same page, against incidents you can check us on. No exclusivity, no obligation, and no invoice for any of it. Design-partner terms are materially better than list.
No. PRISM is software only. It reads the BMS, CDU, pump and loop telemetry you already collect. There is nothing to install in the white space and no capital cost to deploy. A chiller on a current controller typically exposes fifty to sixty usable points, which is more than enough.
It cannot touch your controls, and that is deliberate: no write path to plant exists in the product. What it does is get the decision to the people and systems that act. That means a work order in your CMMS or EAM with priority and scope, an SMS to the on-call engineer, a voice escalation when severity demands it, or a webhook into your dispatch queue. Your compute and facilities teams get the lead time; the intervention stays under human control.
BMS alarms fire on thresholds, which is to say after the fact, and OEM tools stay tied to their own hardware. PRISM is vendor-neutral across your whole mixed fleet, models the joint signature rather than single channels, and catches the masked decline before any threshold trips. It also reduces alert volume rather than adding to it. Sites routinely carry hundreds of active nuisance alarms, and an untuned alert stream only compounds that.
Yes, and commissioning is arguably the best time to start. PRISM establishes the healthy signature bands at go-live. That validates the new plant against how it was designed to behave, and it gives the model a baseline from day zero instead of waiting a year to learn what normal looks like.
You should not take it on our word. That is what the retrospective is for, and it runs on your data at no cost. The argument is that the asset class is the same: motor-driven pumps and fans moving a fluid around a closed loop, with a control system compensating for decline. And your instrumentation is denser and faster than a remote lift station on a radio link. The data is better, not worse. The stakes are what changes.
Reference rights, two or three quotable metrics, and a case study you review and approve before anything is published. In exchange, design-partner terms are materially better than list. If the retrospective does not convince you, you owe us nothing and we do not publish anything.
Forty-five minutes with operations, six to twelve months of history under NDA, and a scored retrospective against incidents you already remember.
$0, offline, no exclusivity, no obligation.
Outage and cost data. Outage cost distribution and cause ranking from Uptime Institute Annual Outage Analysis 2026 (n=94 operators) and Global Survey 2025. Hourly cost bands from the ITIC 2024 Hourly Cost of Downtime Survey. Claim frequency and loss cost from Allianz Commercial's data center claims review, and from Swiss Re and FM Global on direct liquid cooling. Rack density and staffing from AFCOM State of the Data Center 2026 and published vendor rack roadmaps. Incident detail from contemporaneous public reporting and the operator's own statement.
Diagnostic and maintenance standards. Failure modes and task intervals are drawn from published reliability-centered maintenance libraries for these asset classes. Diagnostic mapping references IEEE C57.91 (thermal and loss of life), IEEE C57.104 and IEC 60599 (dissolved gas), IEEE 519 (harmonics) and ANSI/NETA MTS (thermographic severity). Telemetry standards: DMTF Redfish DSP2064, ASHRAE Guideline 36, BRICK schema. Transformer failure economics from published insurance loss analysis. Equipment lead times from Wood Mackenzie survey data reported in trade press.
Estimates. Chiller warning rates, emergency-versus-planned repair multiples and thermal ride-through are industry and trade-source estimates rather than measured benchmarks. The classification of which failure modes can be read from existing telemetry is Firstlook's own engineering assessment, and we review it asset by asset on request. Performance figures are measured on Firstlook production water deployments.