When things go wrong
Every way Farsight can fail that we know of: how it notices, what it does, what you see, and the test that pins it.
| Situation | Farsight does | You see | Test |
|---|---|---|---|
| A target fails once | Counts it, re-checks after the retry interval, sends nothing | "Failing 1 of 2" on the monitor | engine::state::first_failure_only_counts |
| A target is down | After confirm_after failures: Down, an incident, one alert per channel | Red monitor, open incident, an alert | confirm_after_opens_incident |
| A whole server goes down | All its monitors fail; alerts within 20 s are merged | One message listing them all | alerts::group_window_merges_burst |
| A target flaps | After 4 changes in 30 min: one "flapping" alert, then quiet until stable for 15 min | A flapping badge | flapping_detected_and_settled |
| The path between Farsight and a target breaks (seen 2026-10-04: Google us-west1 could not reach Oracle Chicago for a while) | A failure where the target never answered asks a Cloudflare Worker for a second opinion; if Cloudflare gets the answer the monitor expects (a port for tcp/ping, the expected status for http), the failure is not counted, for 30 min at most | The monitor stays as it was with a note; one "network path problem" line on the timeline per host per hour; after 30 min the failures count with a note | confirm::tests::*, live on oregon |
| The Worker answers but the target is really broken (a Cloudflare 403, a redirect, an error page) | Fails closed: anything other than the expected status means the failure counts | A normal Down | confirm::tests::http_needs_the_expected_answer |
| Farsight's own internet fails | The canary (3 targets, 2 must fail) marks it blind; failures become Unknown | A banner; afterwards one "Farsight was offline" notice | canary::canary_two_of_three, blind_failures_are_unknown |
| A server answers headers, then trickles its body | The whole check ends at the timeout, naming the phase | "no answer in 10 s (reading the body)" | probe::http::http_body_trickle_times_out |
| A huge response | Reads 1 MB at most | Rules judge the first 1 MB | http_body_capped |
| Redirect loop | Stops after 5 hops | "more than 5 redirects" | http_redirect_loop_fails |
| Redirect to another site | Follows without credentials | Normal result | redirect_to_other_origin_drops_credentials |
| Expired or wrong certificate | Fails the check (unless TLS checks are off) | "the certificate has expired" | https_trusted_root_and_cert_expiry |
| Certificate near expiry | One warning per threshold (14, 7, 3, 1 days) | A cert alert | cert_thresholds_once_each |
| A monitor points at a private or metadata address | Refuses (private addresses can be allowed by the owner) | "is a private address" | probe::tests::blocklist, private_targets_blocked_by_default |
| Farsight restarts mid-incident | Reloads state; no second Down alert; the incident closes once | Nothing unusual | engine::tests::restart_mid_incident |
| A heartbeat is late | Down after interval + grace | "no ping in 1 h 1 m" | scheduler::tests::heartbeat_late_goes_down_and_fail_endpoint |
| Brevo is down | The same email goes through Resend | Delivery log shows "resend" | alerts::email_brevo_then_resend_fallback |
| Every provider fails | The outbox retries for 24 h with backoff | The channel shows "failing" | retry_backoff_schedule_and_four_xx_stops_after_three |
| An alert storm | Email capped at 80 a day, one notice | "Email alerts paused until tomorrow" | daily_email_cap_sends_notice_once |
| The disk fails | Results stay in memory (10,000 at most), alerts keep working | Diagnostics: writer failing | writer (manual) |
| Password guessing | Per-address blocks, a global breaker, Turnstile | Security events and emails | auth::tests::* |
| Slow or idle clients try to fill the server (slowloris, idle keep-alives) | Headers within 10 s, a TLS handshake within 10 s, 512 connections in all, 24 per address, 30 min per connection, 75 s per request, 1 MB bodies (4 KB for heartbeats), 64 live streams | The dashboard and status page keep answering | serve::tests::*, live on oregon |
| A session cookie is stolen | Changing the password or the authenticator signs out every other session at once (a reset signs out all); Settings also lists and revokes sessions | "Password changed. 1 other session signed out." | auth::tests::password_or_authenticator_change_signs_out_other_sessions |
| A leaked API key | Scoped; cannot mint keys, open private targets or touch sign-in | Revoke it in Settings | e2e keys cannot mint keys |
| Someone edits a monitor to send its stored password elsewhere | Masked secrets are kept only for the same origin | "enter the password again" | masked_secret_not_sent_to_a_new_host |
| The process dies | systemd restarts it in 3 s | A short gap in the charts | |
| The VM dies | Nothing inside Farsight can tell you | Use an outside watcher (for example a free heartbeat elsewhere) |