9.2 million US homes still have a lead pipe. Every utility must find and fix theirs by November 2027. I built a self-serve tool with Claude that turns 120Water's internal lead science into something any utility can run on its own.
Lead pipes cluster in older homes, and the EPA is clear on who that hits hardest: lower-income communities. That's why federal funding sets money aside for these systems specifically. Not a bonus. A requirement.
Budget doesn't change the deadline.
120Water already had a model that could score each unknown line. It just lived outside the actual product. Someone on 120Water's team had to run it by hand, one utility at a time.
Real machine learning, wrapped in a white-glove service. White-glove doesn't scale to every utility racing the same deadline.
Two utility customers, two internal reviewers, one live prototype. These three confusions showed up in every single session.
The same person showed up every time: a non-technical utility employee, wearing several hats, no data science background, and no budget to hire one.
Utility Operator, User Testing Session
"It sucks to review 700 lines, but I'd rather take the time to review them one at a time than mess up 700 locations."
Internal Product Discussion, on why bulk actions stay off by default for high-risk data
Each area replaces a manual step. Data Quality shipped first, on purpose. Those same checks help every customer, not just PM buyers.
Never say "predictive" to a free customer. Hover to reduce complexity, never to remove information. One color means one thing, everywhere. Everything else followed from those three.
A new utility starts here. Clean the data first, unlock the planner next.

Materials, data quality, and lead prediction, together on one page.

Three buckets instead of a raw score. Straight from finding one.

Approve, edit, or delete, right in the drawer. No spreadsheet round-trip required.

Every other suggestion here is eligible for Accept All. A Lead classification never is.

Outliers and duplicate coordinates, shown exactly where they sit.

One question starts the plan: how many lines can your team verify a month?

Review the list, split it across teams, send it out.

Review progress, then plan the next batch or change the strategy.

Built conversationally, 41 versions in. Best on desktop.
Every session started from a real transcript. Never memory, never a guess.
v41, functionally complete across all three areas. Tested with four people, two customers, two internal. Developer handoff is next.
First time running user testing here. The rule: don't build past what's confirmed.
The same three problems kept showing up, in different sessions, with different people. It's why four short tests beat one long one.