Know which roads go under — before the rain stops.
A rainfall total is not a decision. FloodGrid carries the storm all the way through: a predictive rainfall grid from the next fifteen minutes to the next three days, machine-learned stage and inundation for your catchments, and an alerting engine that names the asset, the crossing and the lead time — then publishes how often it was right.
1 km rainfall grid · 15-min forecast step · 0–72 h horizon · machine-learned stage & inundation · threshold, escalation and audit trail · published hit rate and false-alarm ratio
Grid resolution, cadence and horizon are platform specifications. Delivery latency is a design target measured from threshold evaluation to hand-off at the notification provider — not an end-to-end guarantee, since carrier and paging delivery are outside any vendor's control.
Most flood tools stop one step short of the decision.
Radar tells you where it is raining. A gauge tells you what one point measured. A model run tells you what a hydrograph did last night. None of them tell the duty officer at 02:40 whether to close Mill Creek Road, and none of them tell them how much time they have. That last step — from rainfall to consequence to a specific instruction to a specific person — is the whole product.
A schematic of the decision chain, not a product comparison. Every organisation covers some of this already — the question worth asking a vendor is which stage their output actually ends at, and who is expected to do the remaining steps at 3 a.m.
One system from the radar sweep to the phone that rings.
FloodGrid is a single pipeline, not a bundle of point tools. The same quality-controlled rainfall grid that feeds the forecast feeds the machine-learned stage model, the inundation surrogate, the threshold engine and the after-action record — so the number in the alert and the number in the report are the same number.
Predictive rainfall
Radar nowcasting blended into high-resolution numerical weather prediction, on a 1 km grid at 15-minute steps out to 72 hours — with the blend weights published rather than hidden.
Nowcast · blend · NWP ensembleLearned catchment response
A sequence model trained on your own gauge and SCADA history maps rainfall to stage and flow at the points you care about, including the ones with no rating curve and no hydraulic model.
Per-site models · retrained monthlyConsequence, not just level
Forecast stage is resolved onto terrain and your asset layers: which crossings overtop, which pump stations are exposed, which parcels and critical facilities sit inside the wet edge.
Depth grid · asset exposureNotification that escalates
Thresholds you set, an escalation ladder you control, deduplication so one storm is not forty pages, quiet hours, on-call rotations, acknowledgement tracking and a complete audit trail.
SMS · voice · email · webhook · CAPVerification in the open
Hit rate, false-alarm ratio, critical success index and median lead time — computed per site, per threshold, per season, and visible to you whether the number flatters us or not.
Scored against your sensorsBuilt to be integrated
REST and webhooks, CAP 1.2 for public-alerting stacks, OGC services for GIS, Esri-ready layers, and time series shaped for InfoWorks ICM, SWMM, HEC-RAS and your historian.
API-first · no scraping requiredBuilt for the people who have to make the call.
Water & wastewater utilities
Wet-weather overflow risk, pump-station and lift-station exposure, force-main surcharge, treatment-plant inflow surges, and defensible records of what was known when.
Counties & municipalities
Low-water crossings, culvert and storm-drain capacity, repetitive-loss neighbourhoods, and public messaging that is timed to the hazard rather than to the news cycle.
Flood-control & drainage districts
Channel and detention performance under a live storm, gate and pump pre-positioning, and inter-basin awareness when the next basin over is about to hand you its water.
Emergency management
Lead time on the assets that matter — shelters, hospitals, evacuation routes, staging areas — with CAP output that drops into the alerting stack already in use.
State & local DOTs
Route-level closure risk, maintenance-crew staging, and a defensible timeline showing when a segment first crossed a closure threshold.
Dam & levee safety
Inflow forecasting into the reservoir, emergency action plan trigger tracking, and an unbroken, timestamped record of every threshold crossing and every notification sent.
Five stages, every five minutes, whether or not anyone is watching.
Ingest and quality-control
Radar mosaics, dual-pol fields, gauge networks, your own rain gauges, level sensors and SCADA tags arrive on their own cadences. Each is range-checked, flatline-checked and cross-checked against its neighbours. A sensor that has failed is marked failed — never quietly interpolated into a zero.
Build the rainfall grid
Radar is bias-corrected against gauges that survived quality control, producing one 1 km analysis at 5-minute steps. The correction method, the gauges used, and the gauges rejected are all recorded per cell so the number can be defended later.
Forecast forward
Radar-extrapolation nowcasting dominates the first ninety minutes, hands over to blended high-resolution NWP through the first day, and to ensemble guidance beyond it. The handover is gradual and the weights are published on the page, not buried in a methods appendix.
Turn rain into consequence
Per-site sequence models produce stage and flow with uncertainty bands. An inundation surrogate — trained against full hydraulic runs — expands the forecast stage onto terrain in seconds rather than hours, and intersects it with your assets, roads and critical facilities.
Decide, notify, record
Every threshold you defined is evaluated on every cycle. Crossings escalate along your ladder, notifications are deduplicated and routed to whoever is actually on call, acknowledgements are tracked, and the whole sequence is written to an immutable event record you can export.
Everything an agency can actually do requires a head start.
Flood losses are only partly a function of depth. They are also a function of how many of the cheap actions were taken before the water arrived. Each of these has a minimum notice below which it stops being possible — which is why FloodGrid reports lead time as a first-class metric next to accuracy, rather than as a footnote.
Indicative planning ranges, not measurements. The exact notice each action needs is specific to your crews, your geography and your procedures — capturing those real numbers is part of configuring the escalation ladder during onboarding, and they are what the alerting thresholds are then tuned against.
Put your catchments on the grid.
Onboarding starts with your sensors and your asset layers, not with a generic demo. Within a pilot you get a scored record for your own basins — including the events where the model was wrong.
From the next fifteen minutes to the next three days.
No single technique is good at every lead time. Radar extrapolation is unbeatable for the first hour and useless by the sixth. Numerical weather prediction is the only thing that works at day two and is far too coarse and too laggy to catch the cell currently sitting over your interceptor. FloodGrid runs all of them and blends them on a published schedule — so the forecast never quietly changes character without telling you.
Radar nowcast
Optical-flow motion fields extrapolate the observed radar mosaic forward, with growth and decay estimated from the last several sweeps. Updated every five minutes. This is the regime where a warning still changes what a crew can do.
Blended short range
Nowcast skill decays and convection-allowing NWP takes over on a weighted ramp. The blend is bias-corrected against the last several days of local performance, so a model that has been running wet over your basin is discounted before it reaches you.
Ensemble outlook
Multi-member, multi-model guidance carried as a distribution rather than a single trace. At this range the honest output is a probability of crossing a depth, not a number to two decimal places.
An illustrative schedule, not a fixed constant. The real weights are re-estimated continuously from recent verification over your own basins: a nowcast that has been performing well in a slow-moving stratiform event keeps weight longer than one chasing fast discrete cells. The important property is that the ramp is smooth — a forecast that jumps when a source switches produces threshold crossings that are artefacts of the blend rather than of the weather.
The useful question is not “how much rain”. It is “what are the odds it crosses my number”.
A deterministic forecast of 2.4 inches is a statement with no risk content. An operations manager needs the chance that the basin exceeds the depth that has historically put water over Mill Creek Road — and needs it to move honestly as the storm approaches. Set a threshold below and watch the exceedance probability evolve with lead time.
Synthetic ensemble for demonstration. The shape is the point: a probability that climbs early and holds is a storm you can plan around; one that oscillates between 20% and 70% every cycle is telling you the atmosphere has not made up its mind, and that pre-positioning is a better response than evacuation. FloodGrid records both the probability it issued and what subsequently happened, so the reliability of these curves is measurable rather than asserted — see Verification.
Radar sees everywhere. Gauges measure truth. Neither is sufficient alone.
A forecast is only as good as the analysis it starts from. Radar reflectivity is an inference about drop-size distribution, not a measurement of depth; it drifts with calibration, beam blockage, bright band and range. A rain gauge is a direct measurement of one square foot of the planet. FloodGrid corrects one with the other, and writes down what it did.
Radar first guess
Dual-polarization fields give a better rain-rate estimate than reflectivity alone, particularly in heavy rain and mixed precipitation. Mosaicked so the seams between radars are not mistaken for storm structure.
Gauge correction
Multiplicative bias fields and kriging-with-external-drift, computed on rolling windows. More than one method is run; where they disagree materially, the disagreement is surfaced rather than averaged away.
Quality control that admits gaps
Stuck tips, blocked funnels, frozen sensors and comms outages are detected and flagged. A gap is reported as a gap. Nothing is silently backfilled with zero, because a zero is a measurement and a gap is not.
Synthetic storm. The failure mode is not that the gauge is wrong — it is right about the ground under it. The failure is treating one point as representative of a basin whose response is driven by where the core actually sat.
Why this matters for alerting
A threshold defined on a single gauge inherits that gauge's blind spots. If the storm core sits three kilometres away, the threshold never trips and the crossing floods anyway. If the gauge is under the core and the rest of the basin is dry, the threshold trips and a crew is dispatched to a dry road — which is how organisations learn to ignore their own alerts.
What FloodGrid does instead
Thresholds are defined on catchment-integrated grid rainfall, on modelled stage, or on both with an AND condition. The gauge is still used — as a truth source for correcting the grid, and as an independent check that raises a data-quality flag when it disagrees with the grid over its own cell.
And when the radar is down
Coverage degradation is a first-class state. The platform reports reduced confidence, widens the uncertainty bands, and continues to alert on gauge and sensor evidence — while telling you plainly that it is doing so.
See the forecast on your own basins.
A pilot runs on your catchment boundaries, your gauges and your thresholds — with a hindcast over your last few significant storms so you can see how the forecast would have behaved before you rely on it.
Machine learning where it earns its place — and physics where it does not.
“AI-powered” is not a specification. What matters is which component is learned, what it was trained on, what it does when it meets a storm larger than anything in its training set, and who is accountable when it is wrong. FloodGrid publishes that for every model it runs. Some parts of this pipeline are learned because learning genuinely beats the alternative. Other parts are deliberately not, because a conservation law you can audit is worth more at 3 a.m. than a correlation you cannot.
Sensor quality control & anomaly detection
Learned Rule-checkedEvery incoming gauge, level sensor and SCADA tag is scored against its own history and its neighbours. A gradient-boosted classifier flags stuck tips, drift, freeze, siphoning, submerged transducers and comms dropout; deterministic range and rate-of-change rules run alongside it and can veto. Bad data is quarantined before it reaches the rainfall analysis, because a single stuck gauge reading zero through a storm will drag a bias correction across an entire basin.
The classifier proposes; the rules dispose. Any sensor the model wants to exclude is logged with the reason, and an operator can reinstate it — which is itself recorded as a labelled example.
Rainfall analysis & forecast blending
Physical / statistical Learned weightingThe radar-to-rainfall step is physical and statistical: dual-polarization rain-rate relations, mosaicking, and gauge-based bias correction by kriging with external drift. What is learned is the blending — how much weight nowcast, convection-allowing NWP and ensemble guidance should each carry, conditioned on lead time, storm regime, season and recent local verification.
This is a deliberate boundary. Learning the rain-rate relation from scratch would discard decades of radar meteorology; learning the blend weights recovers exactly the local, seasonal, regime-dependent behaviour that no published constant can capture.
Catchment response — rainfall to stage and flow
LearnedThis is where machine learning is unambiguously the right tool. A recurrent sequence model — LSTM-family, trained per site with regional pooling — maps a window of catchment-integrated rainfall, antecedent wetness, recent observed stage and static basin attributes onto stage and flow at 15-minute steps. It is trained on your own historical record, which means it learns your basin's actual behaviour, including the parts no one ever wrote into a model: the culvert that has been half-blocked since 2019, the way the interceptor surcharges once the tide is above a certain level, the diurnal irrigation signal.
Crucially it works at sites with no rating curve and no calibrated hydraulic model — which is most sites. Regional pooling lets a site with two years of record borrow structure from neighbours with fifteen.
Inundation surrogate — stage to depth on the ground
Learned surrogate Physics-trainedFull two-dimensional hydraulic simulation is the gold standard and is far too slow to run on every forecast cycle. So it is run offline, hundreds of times, across a designed sweep of boundary conditions — and a convolutional surrogate is trained to reproduce the resulting depth grids from terrain, the drainage network and the forecast boundary stage. At runtime the surrogate produces a depth grid in under a second.
The surrogate never invents hydraulics; it interpolates a physics-based response surface. Where a forecast stage falls outside the sweep it was trained on, it says so and the cell is returned as out-of-envelope rather than as a confident number.
Threshold evaluation & alert decision
Deterministic rulesDeliberately not a learned component. Whether an alert fires is a policy decision an agency owns, expressed as explicit rules on explicit quantities: this stage, at this site, with this probability, sustained for this long. A learned classifier here would be unauditable, would drift without anyone noticing, and would make it impossible to answer the only question that matters in an after-action review — why did this alert fire?
Machine learning informs this layer by supplying the forecast and its uncertainty. It does not get to decide the policy.
The operator
Human in the loopAn operator can suppress, escalate, extend or override any alert, and every one of those actions is recorded with who, when and why. Overrides are not treated as noise to be engineered away — they are the highest-quality supervision signal the system receives, and they feed directly into the next retraining cycle and into the threshold review.
If a site is repeatedly overridden, that is a finding about the threshold or the model, and it appears in the monthly review rather than waiting for someone to notice.
A forecast is a distribution that narrows as the storm arrives.
Below is a single synthetic site through one synthetic event. Move the issue time to see what the model was saying at each point — how wide the band was, where the central estimate sat, and when the threshold crossing first became likely. This behaviour, not a single headline accuracy figure, is what determines whether a forecast is usable in an operations room.
Synthetic site and synthetic event. Note what the early issue times do: the median is close but the band is wide, which is a correct and useful statement — it says prepare, do not commit. A model that reported a narrow band eight hours out would look more impressive and be considerably more dangerous.
No hidden inputs. No mystery features.
Illustrative attribution for one synthetic catchment. The pattern is the transferable part: at short lead times the model is mostly reading the river, and at long lead times it is mostly reading the sky. A model whose attribution did not shift this way with lead time would be a warning sign.
Dynamic inputs
Catchment-integrated rainfall (observed and forecast), antecedent rainfall over 3, 7 and 30 days, observed stage and rate of change at the site and at upstream sites, upstream reservoir releases where telemetered, tide or downstream boundary stage in coastal basins, and seasonal terms.
Static attributes
Drainage area, time of concentration, imperviousness, slope and hypsometry, soil hydrologic group, land cover, channel and storm-network density, and detention storage. These are what let a site with a short record borrow from a hydrologically similar neighbour.
Deliberately excluded
Anything that would not be available at forecast time, anything derived from the target it is predicting, and anything that encodes the answer through a back door. Data leakage is checked by re-running validation on a strictly time-forward split — never a random one.
Every model ships with its limits written down.
A model card travels with each deployed model and is visible in the application, not filed in a report. It states what the model was trained on, how it was validated, where it is known to be weak, and when it will be retrained. If a model card cannot be produced for a component, that component does not go into production.
Catchment response — stage
Learned- Purpose
- 15-minute stage and flow forecast at instrumented sites, 0–72 h.
- Training data
- Site record plus regionally pooled neighbours; minimum 24 months, target 60+ months of 15-minute stage and matched grid rainfall.
- Validation
- Strictly time-forward split. Held-out final season never seen in training or tuning. Scored on peak timing and peak magnitude error, not only on mean error, because a model can win on RMSE while missing every crest.
- Known weaknesses
- Events materially larger than anything in the training record; step changes after construction, channel works or a new detention facility; sites where the stage sensor itself was relocated mid-record; frozen-precipitation and snowmelt-driven response where the training record is thin.
- Retraining
- Monthly on a schedule, and out of cycle after any event exceeding the prior record or after physical works in the basin.
- Degradation behaviour
- Widens quantile bands and raises a confidence flag rather than narrowing on thin evidence. Falls back to persistence plus routed rainfall if the sequence model is unavailable.
Inundation surrogate — depth
LearnedPhysics-trained- Purpose
- Sub-second depth grid from forecast boundary stage, for asset and roadway exposure.
- Training data
- Offline two-dimensional hydraulic simulations across a designed sweep of boundary stages, durations and initial conditions on the client's own terrain and drainage network.
- Validation
- Held-out simulations plus, where available, observed high-water marks and documented closure records from past events.
- Known weaknesses
- Boundary conditions outside the trained sweep — returned as out-of-envelope, never extrapolated. Blockage-driven flooding the underlying hydraulic model did not represent. Terrain changes since the survey. Levee and floodwall breach behaviour, which is explicitly out of scope.
- Retraining
- On terrain or network updates, on hydraulic model revision, and annually.
- Degradation behaviour
- Falls back to a static stage–depth lookup at instrumented sites and reports the reduced fidelity in the alert payload.
Sensor anomaly detection
LearnedRule-vetoed- Purpose
- Flag failed or degraded sensors before their values reach the analysis.
- Training data
- Historic maintenance records, operator overrides and synthetic fault injection.
- Validation
- Precision and recall against held-out maintenance tickets; tuned to favour recall, since a missed bad sensor corrupts a basin while a false flag costs one review.
- Known weaknesses
- Slow drift that resembles genuine seasonal change. Novel fault modes not present in the record. Newly installed sensors with no personal history.
- Retraining
- Quarterly, and whenever a fault class is added to the maintenance taxonomy.
- Degradation behaviour
- Deterministic range, flatline and rate-of-change rules continue to run independently and can quarantine a sensor on their own.
Forecast blending weights
Learned- Purpose
- Set the weight given to nowcast, convection-allowing NWP and ensemble guidance by lead time and regime.
- Training data
- Rolling local verification of each source against the gauge-corrected analysis over the client's own footprint.
- Validation
- Continuous — the blended product is scored against the same analysis it is trying to predict, on a strictly forward-looking basis.
- Known weaknesses
- Regimes rare in the local record, such as a first tropical remnant in an inland basin. Periods immediately after an upstream model upgrade, where recent skill is no longer representative.
- Retraining
- Continuous, with a floor on how fast weights are allowed to move so a single bad day cannot swing the blend.
- Degradation behaviour
- Reverts to a conservative published default schedule when recent verification is unavailable or unrepresentative.
A model that is not being watched is not in production. It is just running.
Input drift and performance drift are monitored separately. Input drift catches a changing world — a new subdivision, an upgraded upstream model, a relocated sensor — often before performance visibly degrades. Performance drift catches everything else. Either one crossing its control limit opens a review; neither silently retrains the model on its own.
New model versions run in shadow before they run in anger. A candidate scores against live data alongside the incumbent for a full retraining cycle. It is promoted on a documented margin across peak timing, peak magnitude and threshold-crossing skill — not on aggregate error alone, which is the metric most easily gamed.
Every promotion is reversible. Model versions are immutable and addressable; the version that produced any historical alert can be re-instantiated and re-run against the exact inputs of that moment. An after-action review never has to reconstruct what the model "probably" said.
Out-of-envelope is a state, not an error. When conditions exceed the range a model was trained on, the platform says so in the alert payload and in the interface. It does not extrapolate confidently into a regime it has never seen — which is precisely the regime that produces record floods.
Overrides and misses are reviewed monthly, with people in the room. Every missed event, every false alarm and every operator override is listed. Some are model problems, some are threshold problems, and some are correct behaviour that looked wrong. Telling them apart is not something to automate.
Ask us the hard questions about the models.
We would rather spend the first call on training windows, validation splits and failure modes than on a slide deck. Bring your hydrologist.
An alert that does not say what to do is just anxiety with a timestamp.
The hardest problem in flood notification is not detection. It is credibility. An organisation that has been paged forty times for a storm that produced nothing will not answer on the night it matters. Everything in this layer — deduplication, escalation, quiet hours, acknowledgement, the published false-alarm ratio — exists to protect the one thing that actually saves property: whether people still believe the alert.
Specific
Names the site, the asset, the road segment, the forecast depth and the time. Not “flooding possible in your area”.
Timed
States the lead time remaining and the confidence, so the recipient can tell a prepare from a commit.
Deduplicated
One storm, one thread. Updates amend the existing incident instead of starting a new page every cycle.
Accountable
Delivery, acknowledgement and every override are recorded, exportable and admissible in an after-action review.
Write the rule in the language of the decision.
Thresholds are authored as explicit, versioned conditions on quantities an engineer can defend — not as a sensitivity dial. Change the rule below and watch the message that would be delivered change with it.
Trigger definition
Rules are versioned. Editing one records the change, the author and the reason, and the previous version stays retrievable — so a review can establish what rule was in force on the night in question.
Illustrative rendering of an SMS or push payload. The same incident is delivered simultaneously as an email digest, a webhook payload and, where configured, a CAP 1.2 message — all carrying the same incident identifier so downstream systems can thread them.
The right person, at the right hour, on the channel they will actually answer.
Escalation is configured per site and per level, because the person who needs to know that a detention basin is filling is not the person who needs to know that a crossing is about to overtop. Quiet hours apply to advisory levels and never to emergency ones.
Nothing is sent. The record is still kept.
Conditions are logged continuously whether or not anything crosses. When an event is reviewed six months later, the quiet hours before it are part of the record too.
Duty officer, one message, no siren.
Sent to the on-call duty officer and the operations channel. Subject to quiet hours: between configured hours this becomes a queued digest rather than a page, unless it escalates.
Paged, and chased until acknowledged.
Primary on-call is paged immediately. If no acknowledgement arrives inside the configured window, the alert escalates to the secondary, then to the supervisor. Quiet hours do not apply.
Everyone on the list, simultaneously, no suppression.
Fan-out to the full notification group with no deduplication delay and no quiet-hour handling. CAP output is emitted for downstream public-alerting systems. Suppression is disabled at this level by design — an operator can annotate, but cannot silence.
Every false alarm spends credibility you cannot buy back.
A threshold is a trade. Lower it and you catch more events and cry wolf more often; raise it and every alert is real but some floods arrive unannounced. There is no setting that avoids the trade — there is only choosing it deliberately, with the numbers in front of you, rather than by accident. Move the threshold below and watch both sides move.
Synthetic five-year record for one site. Notice that lead time falls as the threshold rises — the certainty you gain by waiting is paid for in the time your crews lose. The right setting is not a technical answer; it depends on what the action costs and what the miss costs, which is a conversation FloodGrid structures rather than one it makes for you.
One incident per storm per site. Subsequent cycles amend the open incident — raising or lowering its level, updating the forecast — instead of paging again. Recipients see a thread, not a flood of messages.
Alerts clear on a lower threshold than they set on. Without hysteresis a forecast hovering on the boundary produces a set-clear-set-clear cycle that is worse than no alert at all.
A crossing must persist for a configured number of cycles. One noisy cycle does not page anyone. This is the single most effective false-alarm control available, and it costs only the sustain window in lead time.
Configurable per level, never applied to emergency. Advisory traffic queues into a digest overnight; life-safety traffic never does, and the interface makes that distinction impossible to misconfigure by accident.
On-call schedules, handovers and escalation chains are first-class. The alert goes to whoever is actually on duty at that minute, with automatic escalation when acknowledgement does not arrive.
During a declared event the posture changes. Sustain windows shorten, digests are suspended, and everything routes live — because the cost of a delayed alert during an active storm is not the same as it is on a quiet Tuesday.
Six months later, someone will ask what you knew and when.
Every incident carries a complete, timestamped, immutable chain: the inputs that were available, the forecast that was produced, the model version that produced it, the rule that fired, the messages that were sent, who acknowledged and when, every override with its stated reason, and what was observed afterwards. It exports whole, in machine-readable form, without a support ticket.
MC9-WATCH-v4 · model catch-lstm 2026.07.2 · sent to duty officer via SMS + TeamsIllustrative record for a synthetic event. The details are invented; the structure is not. Every field shown here is captured on real incidents, including the ones where the forecast was wrong — especially those, because an error record that only contains successes is not a record.
Bring your current thresholds. We will show you what they would have done.
A hindcast over your last three years scores your existing triggers against what actually happened — hits, misses, false alarms and lead time — before a single new alert is configured.
A stage forecast becomes useful the moment it lands on something you own.
“Mill Creek will reach 9.6 feet” means something to a hydrologist and nothing to a maintenance supervisor. “Route 9 low crossing goes under at 05:10, the Ridgeview lift station access road follows twenty minutes later, and the elementary school stays dry” is the same forecast expressed as a decision. The translation between them is terrain, a drainage network, and your own asset layers.
Entirely synthetic geography, synthetic assets and synthetic depths. Nothing on this map corresponds to a real place. A production deployment runs on your terrain, your drainage network and your own asset register — and reports out-of-envelope cells rather than guessing at them.
The list a supervisor can act on without opening a GIS.
Crossings are ranked by time to threshold, not by depth, because time is what determines the order in which a limited number of crews can get anywhere. Each row carries the forecast, the confidence, and the threshold that was used — so the decision can be defended and, later, scored.
| Crossing | Status | Time to threshold | Forecast depth over deck | Confidence | Threshold basis |
|---|
Synthetic scenario. “Confidence” is the share of ensemble members crossing the threshold within the stated window, not a subjective rating. Rows update with the forecast valid time selected on the map above.
Your assets, not a generic layer
Pump and lift stations, treatment facilities, cabinets and control panels, manholes and regulators, critical customers, shelters and evacuation routes — loaded from your own register with your own identifiers, so an alert names the asset the way your work-order system names it.
Thresholds tied to real elevations
Deck elevations, low-chord clearances, panel and vent heights, door sills and access-road grades. A crossing threshold is a surveyed number, not an assumption — and where the elevation is unknown, the platform says the threshold is provisional rather than presenting it as certain.
Honest about what it cannot see
Blockage-driven flooding, surcharge from a collapsed line, breach behaviour at levees and floodwalls, and depths outside the trained hydraulic envelope are all outside scope and are reported as such. A model that never says “I don't know” has simply hidden the cases where it doesn't.
Hydraulic fidelity, at forecast speed.
Build the response surface offline
Your calibrated two-dimensional hydraulic model is run across a designed sweep of boundary stages, storm durations and antecedent conditions. This is the expensive step and it happens once, on a schedule — not on every forecast cycle.
Train the surrogate on those runs
A convolutional encoder–decoder learns to reproduce the simulated depth grids from terrain, the drainage network, and the boundary conditions. It is validated on held-out simulations and, where they exist, on surveyed high-water marks and documented closures from past events.
Evaluate it every cycle
At runtime the surrogate turns the current forecast stage into a depth grid in under a second, for every valid time in the horizon. That is what makes a live, forward-looking inundation view possible at all.
Intersect with the asset register
The depth grid is overlaid on crossings, facilities and parcels; exposure is computed against surveyed thresholds; and the resulting list is what feeds the alert payload and the operations view.
Refuse to extrapolate
Any cell whose boundary condition falls outside the trained sweep is returned as out-of-envelope. It appears on the map as an explicit hatched state, is counted in the exposure tiles, and is never rendered as a confident depth.
Why not just run the hydraulic model live?
Because a two-dimensional run over a meaningful area takes minutes to hours, and a forecast cycle is five minutes. Running it live would mean either a much coarser mesh, a much smaller area, or a forecast that is stale before it is delivered. The surrogate keeps the fidelity of the full model and pays for it in offline compute instead of in latency.
What if we don't have a 2-D model?
Many agencies do not. Deployment then starts from stage–depth relationships at instrumented sites plus terrain-based flood extent, which is a coarser product and is labelled as such. Building the hydraulic basis is a normal part of a phased rollout rather than a precondition for getting value.
Does the surrogate drift?
It drifts when the world changes — new construction, a rebuilt culvert, a revised terrain survey, a recalibrated hydraulic model. Each of those triggers a retrain. It does not drift on its own, because unlike the catchment-response model it is not learning from a moving observational record.
Put your crossings and assets on the map.
Send an asset register and a terrain surface and we will show you the exposure view for a storm you already remember — including which thresholds we could not establish from the data provided.
Do not take our word for it. Read the record.
Every forecast FloodGrid issues is stored with the time it was issued and the model version that produced it. When the observation arrives, the two are matched automatically. Nothing is re-scored after the fact, nothing is quietly dropped for being an outlier, and the resulting numbers are visible to you continuously — not assembled for a renewal meeting. This page is the part of the product we would want to see first if we were buying it.
Both panels are synthetic demonstrations of the format, not achieved results. A reliability curve below the diagonal means the forecast is over-confident — when it says 70% it should be right about 70% of the time, and if it is right 45% of the time that is a defect to be corrected, not a presentation to be reframed. Baselines are shown deliberately: a flood model that cannot beat persistence at three hours and climatology at three days has not demonstrated anything, and any vendor unwilling to show those two lines should be asked why.
Named, defined, and computed the same way every time.
Flood verification is full of numbers that sound similar and mean different things. These are the definitions FloodGrid uses, stated once so that a score can be compared across sites, seasons and vendors without a translation step.
| Metric | Definition | What it is good for | How it misleads |
|---|---|---|---|
| POD Probability of detection | Hits ÷ (hits + misses) | The share of real events you were told about. The number a public-safety audience cares about most. | Trivially driven to 1.0 by alerting constantly. Meaningless without FAR beside it. |
| FAR False alarm ratio | False alarms ÷ (hits + false alarms) | The share of alerts that were wrong. The number that predicts whether people keep answering. | Trivially driven to 0 by never alerting. Meaningless without POD beside it. |
| CSI Critical success index | Hits ÷ (hits + misses + false alarms) | Combines both sides into one number. Ignores correct quiet periods, which is appropriate for rare events. | Hides the trade-off it summarises. Two very different operating points can share a CSI. |
| Lead time Median, on hits | Time from first qualifying alert to observed threshold crossing | The only metric that maps directly onto what a crew can actually do. | Computed on hits alone, so a system that misses hard events can post a flattering lead time. |
| Peak timing error | Forecast crest time minus observed crest time | Reveals systematic early or late bias that mean error conceals entirely. | Undefined for events with a flat or multi-peaked crest; reported with the crest shape noted. |
| Peak magnitude error | Forecast crest stage minus observed crest stage | What matters for depth, exposure and closure decisions. | Small absolute errors can be large relative errors on shallow, flashy sites. |
| Brier skill score | Brier score against a climatology reference | Scores the probability itself, not just the yes/no decision derived from it. | Dominated by the many quiet periods unless events are stratified out. |
| Reliability | Observed frequency within each forecast-probability bin | Tests whether a stated probability means what it says. Directly checks over-confidence. | Needs a large sample per bin; thin bins produce noisy curves that invite over-reading. |
Including the ones we got wrong.
Every event above a configured significance is scored individually and kept. Misses and false alarms are shown at the same size and in the same list as hits, because a record that surfaces only successes tells you nothing about what will happen on the night that matters.
Synthetic composite. Seasonal structure like this is normal and worth stating plainly: summer convection is harder to forecast than a winter frontal passage, so error grows and lead time shrinks. Thresholds and staffing postures should differ by season for exactly this reason, and FloodGrid reports the seasonal breakdown rather than a single annual average that conceals it.
Your record stays empty until your catchments are on the grid.
Every figure on this page is a synthetic demonstration of a format. The only verification numbers worth anything to your agency are the ones computed on your basins, your sensors and your thresholds — and those do not exist until FloodGrid has run through a season with you. We will not show you someone else's scores and imply they are yours.
Nobody needs another dashboard nobody opens.
FloodGrid has an operations interface, and on a quiet Tuesday nobody will be looking at it. So the platform is built API-first: everything visible in the interface is available as a documented endpoint, a webhook, a standards-compliant message or a file in a format your existing tools already read. The goal is for FloodGrid to show up inside the systems your people are already watching.
Webhooks
Signed JSON on every incident open, amend, escalate, acknowledge and close. At-least-once delivery with retry and replay, so a downstream outage does not silently lose an event.
REST API
Forecasts, observations, thresholds, incidents, exposure and verification. Cursor-paginated, versioned, OpenAPI-described, with sensible rate limits and a sandbox that returns synthetic data.
CAP 1.2
Common Alerting Protocol output for public-alerting stacks, IPAWS-compatible pipelines and EOC systems — emitted under your authority and your approval workflow, never ours.
OGC services & Esri
WMS, WMTS and OGC API – Features for the rainfall grid, depth grid and exposure layers; plus feature services and file geodatabase exports for ArcGIS Pro and Enterprise.
Hydraulic & hydrologic
Rainfall time series and grids shaped for InfoWorks ICM (CSV and RED), EPA SWMM, HEC-HMS and HEC-RAS boundary conditions, plus NetCDF and GeoTIFF for anything else.
SCADA & historian
OPC UA and MQTT for inbound telemetry; forecast tags written back to the historian so operators see the forecast next to the measurement on screens they already trust.
Comms & paging
SMS, voice, email, Microsoft Teams, Slack, and pager or incident-management integrations with on-call rotation, acknowledgement and escalation handled natively.
SSO & access control
SAML 2.0 and OIDC single sign-on, SCIM provisioning, and role-based access down to the site and threshold level, with every configuration change attributed and reversible.
Every payload carries its own provenance block. A downstream system — or an auditor two years later — can establish exactly which rule fired, which model versions produced the numbers, whether the inputs were complete, and which sources were degraded at the time. Synthetic example; field names are illustrative.
Your data stays yours, and leaves with you.
Your sensor data, asset register and thresholds remain your property. FloodGrid holds a licence to process them for the purpose of delivering the service, and for nothing else. They are not pooled into a shared training corpus and not used to train models for other customers unless you specifically agree in writing.
A full export is available at any time, without a support request. Historical forecasts, observations, incidents, configuration, verification records and model cards, in open formats. An exit that requires a professional-services engagement is not an exit.
Managed cloud by default; single-tenant and government-cloud options where procurement requires it. Air-gapped installation is possible for the alerting and threshold layers, with the caveat that forecast quality depends on continuously ingested public data feeds.
Alerting degrades gracefully rather than failing closed. If the forecast pipeline is unavailable, threshold evaluation continues on live observations alone, at reduced lead time, and says so explicitly in every message it sends.
Role-based, attributed and reversible. Who may author a threshold, who may suppress an alert and who may only read are separate permissions. Every configuration change is versioned with an actor and a reason.
Set by you, not by us. Retention periods for observations, forecasts and incident records are configurable to match your records-management schedule, including indefinite retention where a regulator requires it.
A pilot that produces a defensible answer, not a demo.
Scope — one basin, real assets
Pick a catchment that already causes trouble. Provide the sensor list, the asset register, the terrain surface and whatever thresholds are currently in use, even if they are informal.
Hindcast — score what you already do
Before configuring anything new, your existing triggers are replayed against three years of history. That produces the hit rate, false-alarm ratio and lead time of the status quo, which is the only fair baseline for anything that follows.
Build — models, thresholds, ladder
Catchment-response models are trained on your record, the inundation basis is established at whatever fidelity your data supports, thresholds are authored with your engineers, and the escalation ladder is configured against your actual on-call reality.
Shadow — alerts nobody has to act on
The system runs live for a full season, generating alerts into a shadow channel. You see what it would have sent, and when, without anyone being paged. Thresholds are tuned against real storms rather than against assumptions.
Cut over — and keep scoring
Live notification begins when the shadow record justifies it. Verification does not stop at go-live: it is the permanent operating record, reviewed monthly, and it is what tells you whether the system is still earning its place.
Send us your asset register and a bad storm.
The fastest way to evaluate this is a hindcast on an event you remember well, scored against what actually happened.
Frequently asked — answered plainly.
Does FloodGrid replace National Weather Service warnings?
How much history do you need before the models are useful?
What happens in a storm bigger than anything in your training data?
Do we need a calibrated 2-D hydraulic model first?
How do you keep us from drowning in alerts?
What is actually machine learning here, and what is not?
Can we see your accuracy numbers before we buy?
Who owns the data, and what happens if we leave?
What happens when a radar goes down mid-storm?
How does this relate to GridRain?
What does it cost?
Is any of the data on this website real?
Shared vocabulary, so meetings go faster.
Nowcast
A very short-range forecast produced by extrapolating current observations — here, motion fields derived from successive radar sweeps — rather than by solving atmospheric equations. Dominant skill from 0 to roughly 90 minutes.
QPE / QPF
Quantitative precipitation estimate (what fell) and quantitative precipitation forecast (what will fall). QPE is an analysis of the past; QPF is a prediction. Conflating them is a common source of confusion in specifications.
Antecedent moisture
How wet the catchment already was when the storm began. The single largest reason the same rainfall produces very different responses on different days, and a primary input to the catchment-response model.
Action / warning stage
Stage thresholds at a site: action stage is where preparation begins, warning stage is where impacts are expected. FloodGrid treats these as configurable per site rather than inheriting a default.
Surrogate model
A fast approximation trained to reproduce the output of a slow, physically-based simulation. Used here so a full 2-D hydraulic response can be evaluated every forecast cycle instead of every few hours.
Out of envelope
A condition outside the range a model was trained on. FloodGrid reports this state explicitly instead of extrapolating — which matters most in exactly the record-breaking events where extrapolation is most tempting.
Hysteresis
Clearing an alert at a lower threshold than the one that set it, so a forecast hovering on the boundary does not produce a set-clear-set-clear cycle.
CAP 1.2
Common Alerting Protocol — the OASIS standard message format used by public alerting systems. FloodGrid emits CAP so your existing stack can consume it under your authority.
See your watershed on the grid.
Tell us the basin that keeps you up at night. The most useful first conversation is usually thirty minutes with whoever actually gets called at 2 a.m., plus whoever owns the asset data.
What a good first call looks like
Thirty minutes, three questions. Which events do you wish you had known about earlier? What do you currently do when a storm is coming, and who does it? What data already exists — sensors, asset register, terrain, hydraulic models — and what state is it in?
What we will send afterwards
A written scope for a pilot on one basin, with deliverables, the data we need from you, what we will measure, and exit terms. No pricing table dressed up as a proposal.
Bring the sceptic
The most productive sessions have included the engineer who does not believe machine learning belongs anywhere near a flood warning. Those are the questions this product was designed to survive, and the model cards exist to answer them.