Water utilities can be running and still not be recovered

Photos for You via Getty Images
COMMENTARY | Questions around whether water is safe to drink tell only half the story. Public reporting does not show if trusted operational authority has been re-established.
On July 26 and 27, more than 30 Minnesota water sites reported cyber incidents; most confirmed attacks involved technology used to remotely monitor and control equipment. Some utilities shifted to manual operation, and an intrusion in Braham briefly disabled the controls for the city’s well and treatment plant.
On July 30, the Cybersecurity and Infrastructure Security Agency warned utilities nationwide to pull exposed industrial control equipment off the internet, citing activity that had locked operators out, disrupted facilities and triggered boil-water notices. On August 1, Michigan reported cyberattacks involving nine of its water systems and said all continued operating safely with no known threat to public health.
Federal officials said related incidents had been reported in at least seven states. The FBI is investigating and has not publicly identified the responsible actor. The incidents followed federal warnings about Iran-linked actors targeting water and wastewater operational technology, but those warnings provide context, not attribution.
CISA’s alert told utilities what to do to reduce their exposure: disconnect internet-facing programmable logic controllers, replace default passwords, restrict remote access to trusted devices. That is the right instruction, and it is overdue. But it is prevention, and prevention is only half of what state and local officials will need from their utilities over the coming weeks. The other half has barely been discussed.
Public reporting has focused, appropriately, on availability: whether the water is safe, whether plants are operating, and how quickly affected systems can return to normal control. Those are the right questions to ask first. They are incomplete, however, because public reporting offers no standard measure for a harder one: whether trusted operational authority has been re-established.
Service can resume, or never stop, without the utility having demonstrated that the digital authority controlling the plant was rebuilt from a base the intruder never reached. Water can be flowing while it remains unclear which remote-access credentials are still valid, which controller logic is authoritative, which engineering workstation is clean, and who is permitted to reconnect automated control. Availability is visible on a dashboard. The integrity of control authority is not visible anywhere.
None of that implies the Michigan or Minnesota operators failed to work through those questions. Public reporting does not tell us either way, and it would be irresponsible to assume the worst about utilities that appear to have handled a difficult week competently. The point is narrower and more durable: operating safely and independently re-establishing trusted control authority are two different recovery claims, and only the first is routinely visible outside the utility.
The federal cybersecurity agencies have already articulated this distinction, though in the language of corporate networks rather than water treatment.
Joint guidance from CISA, the National Security Agency and partner agencies on Active Directory compromises records that many persistence techniques survive ordinary remediation, that capable intruders can remain resident for long periods, and that evicting the most determined of them can require resetting credentials at scale or rebuilding the directory itself. CISA’s eviction guidance for networks compromised through SolarWinds asks organizations to map their trust boundaries and sever federation trusts before restoring affected systems. Both documents describe recovery as an exercise in re-establishing trust rather than an exercise in restoring files.
A water utility is not an enterprise network, but the underlying recovery principle still transfers: a compromised environment cannot, by itself, establish that its own control path is clean. If the recovery path authenticates against the same identity infrastructure an intruder controlled, or restores a controller from a configuration file nobody has independently validated, it does not deliver recovery. It delivers a second copy of the incident, on a schedule.
Incident exercises across critical infrastructure commonly emphasize detection, containment and service restoration. They less consistently demonstrate how privileged operational authority will be reconstructed from components known to lie outside the compromise. Four requirements would make that capability testable, though implementing them may take real architectural and operational work.
The first is an independent recovery authority: a small, audited set of break-glass credentials, known-good configurations and recovery tools held apart from the production identity infrastructure, so that a clean starting point exists before anyone needs it.
The second is a known-good operational baseline. For a water system that means validated reference copies of controller logic, human-machine interface configuration, setpoints and firmware, so that an operator can answer what the correct pre-incident state actually was rather than inferring it from the system under suspicion.
The third is a reconstitution exercise with a defined objective. Rotate remote-access credentials, revalidate controller configuration, re-establish a vendor connection, reissue privileged roles, and record how long each step took and what broke. A compromised signing key or certificate-authority key deserves the same treatment, because fraudulent credentials will look valid throughout whatever portion of the environment trusts that key.
The fourth is a set of reduced-trust operating rules written in advance. Which functions may continue in manual mode, which require independent physical verification of what a sensor or screen reports, and which must remain offline until the trust chain is rebuilt. That is a management judgment about public safety, and it should not be improvised at three in the morning by whoever is on shift.
There is also a measurement gap worth naming. Public recovery reporting foregrounds service restoration, availability and water-quality compliance. Those measures matter, but none of them directly establishes whether trusted operational authority has been rebuilt. The missing companion metric is time to re-establish trusted operational authority: how long it takes an operator to show which identities, controller configurations, credentials and remote-access paths are trustworthy again, and that the basis for that judgment stayed outside the compromise.
As Congress and state regulators consider their response, they should resist asking only for faster restoration. Requiring utilities to demonstrate that they can rebuild control authority from an uncontaminated base, and that they have practiced it before they need it, would make the sector harder to keep compromised and safer to restore.
Rehearsing that capability creates a planned disruption. Discovering the missing dependencies during an incident turns it into an uncontrolled one.
Burak Oktenli is completing an MPS in Applied Intelligence at Georgetown University and research authority, autonomy and assurance in critical systems. His commentary has appeared in RUSI, the Modern War Institute at West Point, RealClearDefense and The Space Review.
NEXT STORY: How cyber ranges can help build trust in AI




