# When things go wrong Every way Farsight can fail that we know of: how it notices, what it does, what you see, and the test that pins it. | Situation | Farsight does | You see | Test | |---|---|---|---| | A target fails once | Counts it, re-checks after the retry interval, sends nothing | "Failing 1 of 2" on the monitor | `engine::state::first_failure_only_counts` | | A target is down | After `confirm_after` failures: Down, an incident, one alert per channel | Red monitor, open incident, an alert | `confirm_after_opens_incident` | | A whole server goes down | All its monitors fail; alerts within 20 s are merged | One message listing them all | `alerts::group_window_merges_burst` | | A target flaps | After 4 changes in 30 min: one "flapping" alert, then quiet until stable for 15 min | A flapping badge | `flapping_detected_and_settled` | | The path between Farsight and a target breaks (seen 2026-10-04: Google us-west1 could not reach Oracle Chicago for a while) | A failure where the target never answered asks a Cloudflare Worker for a second opinion; if Cloudflare gets the answer the monitor expects (a port for tcp/ping, the expected status for http), the failure is not counted, for 30 min at most | The monitor stays as it was with a note; one "network path problem" line on the timeline per host per hour; after 30 min the failures count with a note | `confirm::tests::*`, live on oregon | | The Worker answers but the target is really broken (a Cloudflare 403, a redirect, an error page) | Fails closed: anything other than the expected status means the failure counts | A normal Down | `confirm::tests::http_needs_the_expected_answer` | | Farsight's own internet fails | The canary (3 targets, 2 must fail) marks it blind; failures become Unknown | A banner; afterwards one "Farsight was offline" notice | `canary::canary_two_of_three`, `blind_failures_are_unknown` | | A server answers headers, then trickles its body | The whole check ends at the timeout, naming the phase | "no answer in 10 s (reading the body)" | `probe::http::http_body_trickle_times_out` | | A huge response | Reads 1 MB at most | Rules judge the first 1 MB | `http_body_capped` | | Redirect loop | Stops after 5 hops | "more than 5 redirects" | `http_redirect_loop_fails` | | Redirect to another site | Follows without credentials | Normal result | `redirect_to_other_origin_drops_credentials` | | Expired or wrong certificate | Fails the check (unless TLS checks are off) | "the certificate has expired" | `https_trusted_root_and_cert_expiry` | | Certificate near expiry | One warning per threshold (14, 7, 3, 1 days) | A cert alert | `cert_thresholds_once_each` | | A monitor points at a private or metadata address | Refuses (private addresses can be allowed by the owner) | "is a private address" | `probe::tests::blocklist`, `private_targets_blocked_by_default` | | Farsight restarts mid-incident | Reloads state; no second Down alert; the incident closes once | Nothing unusual | `engine::tests::restart_mid_incident` | | A heartbeat is late | Down after interval + grace | "no ping in 1 h 1 m" | `scheduler::tests::heartbeat_late_goes_down_and_fail_endpoint` | | Brevo is down | The same email goes through Resend | Delivery log shows "resend" | `alerts::email_brevo_then_resend_fallback` | | Every provider fails | The outbox retries for 24 h with backoff | The channel shows "failing" | `retry_backoff_schedule_and_four_xx_stops_after_three` | | An alert storm | Email capped at 80 a day, one notice | "Email alerts paused until tomorrow" | `daily_email_cap_sends_notice_once` | | The disk fails | Results stay in memory (10,000 at most), alerts keep working | Diagnostics: writer failing | `writer` (manual) | | Password guessing | Per-address blocks, a global breaker, Turnstile | Security events and emails | `auth::tests::*` | | Slow or idle clients try to fill the server (slowloris, idle keep-alives) | Headers within 10 s, a TLS handshake within 10 s, 512 connections in all, 24 per address, 30 min per connection, 75 s per request, 1 MB bodies (4 KB for heartbeats), 64 live streams | The dashboard and status page keep answering | `serve::tests::*`, live on oregon | | A session cookie is stolen | Changing the password or the authenticator signs out every other session at once (a reset signs out all); Settings also lists and revokes sessions | "Password changed. 1 other session signed out." | `auth::tests::password_or_authenticator_change_signs_out_other_sessions` | | A leaked API key | Scoped; cannot mint keys, open private targets or touch sign-in | Revoke it in Settings | e2e `keys cannot mint keys` | | Someone edits a monitor to send its stored password elsewhere | Masked secrets are kept only for the same origin | "enter the password again" | `masked_secret_not_sent_to_a_new_host` | | The process dies | systemd restarts it in 3 s | A short gap in the charts | | | The VM dies | Nothing inside Farsight can tell you | Use an outside watcher (for example a free heartbeat elsewhere) | |