When things go wrong

Every way Farsight can fail that we know of: how it notices, what it does, what you see, and the test that pins it.

SituationFarsight doesYou seeTest
A target fails onceCounts it, re-checks after the retry interval, sends nothing"Failing 1 of 2" on the monitorengine::state::first_failure_only_counts
A target is downAfter confirm_after failures: Down, an incident, one alert per channelRed monitor, open incident, an alertconfirm_after_opens_incident
A whole server goes downAll its monitors fail; alerts within 20 s are mergedOne message listing them allalerts::group_window_merges_burst
A target flapsAfter 4 changes in 30 min: one "flapping" alert, then quiet until stable for 15 minA flapping badgeflapping_detected_and_settled
The path between Farsight and a target breaks (seen 2026-10-04: Google us-west1 could not reach Oracle Chicago for a while)A failure where the target never answered asks a Cloudflare Worker for a second opinion; if Cloudflare gets the answer the monitor expects (a port for tcp/ping, the expected status for http), the failure is not counted, for 30 min at mostThe monitor stays as it was with a note; one "network path problem" line on the timeline per host per hour; after 30 min the failures count with a noteconfirm::tests::*, live on oregon
The Worker answers but the target is really broken (a Cloudflare 403, a redirect, an error page)Fails closed: anything other than the expected status means the failure countsA normal Downconfirm::tests::http_needs_the_expected_answer
Farsight's own internet failsThe canary (3 targets, 2 must fail) marks it blind; failures become UnknownA banner; afterwards one "Farsight was offline" noticecanary::canary_two_of_three, blind_failures_are_unknown
A server answers headers, then trickles its bodyThe whole check ends at the timeout, naming the phase"no answer in 10 s (reading the body)"probe::http::http_body_trickle_times_out
A huge responseReads 1 MB at mostRules judge the first 1 MBhttp_body_capped
Redirect loopStops after 5 hops"more than 5 redirects"http_redirect_loop_fails
Redirect to another siteFollows without credentialsNormal resultredirect_to_other_origin_drops_credentials
Expired or wrong certificateFails the check (unless TLS checks are off)"the certificate has expired"https_trusted_root_and_cert_expiry
Certificate near expiryOne warning per threshold (14, 7, 3, 1 days)A cert alertcert_thresholds_once_each
A monitor points at a private or metadata addressRefuses (private addresses can be allowed by the owner)"is a private address"probe::tests::blocklist, private_targets_blocked_by_default
Farsight restarts mid-incidentReloads state; no second Down alert; the incident closes onceNothing unusualengine::tests::restart_mid_incident
A heartbeat is lateDown after interval + grace"no ping in 1 h 1 m"scheduler::tests::heartbeat_late_goes_down_and_fail_endpoint
Brevo is downThe same email goes through ResendDelivery log shows "resend"alerts::email_brevo_then_resend_fallback
Every provider failsThe outbox retries for 24 h with backoffThe channel shows "failing"retry_backoff_schedule_and_four_xx_stops_after_three
An alert stormEmail capped at 80 a day, one notice"Email alerts paused until tomorrow"daily_email_cap_sends_notice_once
The disk failsResults stay in memory (10,000 at most), alerts keep workingDiagnostics: writer failingwriter (manual)
Password guessingPer-address blocks, a global breaker, TurnstileSecurity events and emailsauth::tests::*
Slow or idle clients try to fill the server (slowloris, idle keep-alives)Headers within 10 s, a TLS handshake within 10 s, 512 connections in all, 24 per address, 30 min per connection, 75 s per request, 1 MB bodies (4 KB for heartbeats), 64 live streamsThe dashboard and status page keep answeringserve::tests::*, live on oregon
A session cookie is stolenChanging the password or the authenticator signs out every other session at once (a reset signs out all); Settings also lists and revokes sessions"Password changed. 1 other session signed out."auth::tests::password_or_authenticator_change_signs_out_other_sessions
A leaked API keyScoped; cannot mint keys, open private targets or touch sign-inRevoke it in Settingse2e keys cannot mint keys
Someone edits a monitor to send its stored password elsewhereMasked secrets are kept only for the same origin"enter the password again"masked_secret_not_sent_to_a_new_host
The process diessystemd restarts it in 3 sA short gap in the charts
The VM diesNothing inside Farsight can tell youUse an outside watcher (for example a free heartbeat elsewhere)