Onlink notes
Why intermittent internet outages are so hard to diagnose
The hardest internet problems are often not the ones that stay broken.
A connection that fails for a few seconds or minutes and then recovers can be much harder to explain than a complete, persistent outage. By the time you open a diagnostic tool, everything may look normal.
Recovery destroys the most useful moment
Suppose a page stops loading at 8:42 PM. You open a speed test at 8:44 PM and it reports a healthy connection.
That does not mean the earlier problem was imaginary. It means the test happened after recovery.
The missing information is the state of the connection while the incident was happening.
A useful incident has a timeline
For an intermittent problem, three windows matter:
Before — Was the connection healthy just before the failure?
During — Which parts of the connection path degraded or stopped responding?
After — Did the same path recover normally?
That timeline helps distinguish a momentary local Wi-Fi problem from DNS trouble or a wider internet failure.
Onlink v1.2 can preserve before/during/recovery evidence around an incident when enough data is available. The incident remains readable even when evidence is incomplete.
The likely cause should reflect confidence
Short incidents do not always produce complete evidence. Some checks may succeed while others are unavailable. A flap may recover before every probe finishes.
Good incident diagnosis needs to expose that uncertainty.
Onlink uses evidence strength and can report the cause as unclear instead of assigning a confident label without enough support. Older incidents recorded before v1.2 automatic forensics show Cause unavailable rather than receiving invented historical evidence.
Repeated incidents are more valuable than isolated anecdotes
One outage can happen for many reasons. A pattern is more actionable.
If supported incidents repeatedly point toward the same category, or problems cluster in the same part of the day, longer-term analysis becomes useful. That is the idea behind Onlink’s What Onlink noticed insights.
The app deliberately avoids calling one event a recurring problem.
Do not run heavy tests automatically
It can be tempting to launch a speed test every time the connection looks bad. That creates two problems: it consumes substantial bandwidth, and it can change the network conditions you are trying to observe.
Onlink’s automatic incident diagnosis therefore stays lightweight. It does not launch a Cloudflare or M-Lab throughput test, and it does not run the manual loaded-responsiveness transfer.
Heavy tests remain explicit actions.
For more on that design choice, see How much data should an internet health monitor use?.
Preserve evidence before contacting support
If a problem keeps returning, a timeline with timestamps, duration, likely cause, evidence strength, and repeated patterns is much more useful than “it dropped again.”
How to document outages before contacting your ISP explains what to keep and what to leave out.
Intermittent outages are difficult because the network heals faster than humans can investigate it. The answer is not a more aggressive test after the fact; it is better evidence captured at the right moment.