How Farsight works

One process, farsightd (Rust, tokio, two worker threads), with SQLite on the blocking pool and the dashboard built into the binary.

 monitors (SQLite) --> scheduler: one small task per monitor
                          |  probes: http, tcp, ping, dns, tls (heartbeats arrive over HTTP)
                          |  canary: is Farsight's own internet up? (before counting a failure)
                          v
                       engine: state machine per monitor --> incidents, timeline, alerts, live feed
                          |                                        |
                          v                                        v
                       writer: results + rollups, one       alerts: 20 s grouping per channel,
                       SQLite transaction per 30 s          outbox, retries, Brevo then Resend,
                                                            Discord, signed webhooks
 API (axum): dashboard, agents, status page, heartbeats, Server-Sent Events, HTTPS with Let's Encrypt

Modules

ModuleJob
configFlags and FARSIGHT_* variables; nothing else reads the environment
store/*SQLite: schema migrations and every query, grouped by what they store
model/*Monitors (kinds, validation, the API shape, secret masking) and check results
rulesStatus code sets, JSON paths and operators
probe/*One check of one kind; shared resolve, connect and TLS steps with timings; the private-address guard
schedulerA task per monitor: spread start times, jitter, retry interval, 16 probes at most at once, "check now", heartbeat watchdogs
canaryThree TCP targets; two failing means Farsight is blind
confirmA second opinion from a Cloudflare Worker (worker/confirm) for failures where the target never answered: a path problem is not counted
engine/stateThe pure state machine (Up, Down, Slow, flapping, reminders, certificate warnings)
engineApplies results, opens and closes incidents, writes the timeline, emits alerts and live events
writer + rollupBatches results, hourly rollups with a latency histogram, daily rollups by local day
alerts/*Routing, grouping, rendering (email, Discord, webhook), the outbox and retries
auth/*Passkeys (WebAuthn, verified in-house), authenticator codes, sessions, step-up, API keys, Turnstile, brakes
api/*Routes, the caller check, JSON errors, idempotency, rate limits, security headers
status_pageThe public pages (main, per service, history), plain HTML in three layouts: settings, ranges and bars (pure), caches, renderers, the stylesheet and fonts
serveHTTP, HTTPS with automatic certificates (TLS-ALPN-01), the port-80 redirect
certsThe certificates HTTPS serves: the main one and one per custom domain of the status page, chosen by the name a browser asks for; custom ones start and stop while the server runs
domainsThe status page's custom domains: their DNS read at the zone's own nameservers, their state, and their certificates started or stopped (every 30 s until all are live, hourly after)
netA small HTTPS client for Farsight's own calls (providers, webhooks, Turnstile)

A check, start to finish

  1. The monitor's task wakes (its interval, spread by its id and jittered 5%).
  2. It takes one of 16 probe slots and runs the probe under the monitor's timeout. The HTTP probe resolves, connects, does TLS and sends the request, timing each step; follows redirects on new connections (never carrying credentials to another origin); reads at most 1 MB of the body when a rule needs it.
  3. A failure asks the canary for a reading at most 5 s old. If Farsight is blind, the result is recorded as Unknown and nothing changes state. A failure where the target never answered (timeout, no route, DNS, ping loss) then asks the second-opinion Worker on Cloudflare; if Cloudflare reaches the target, it is a network path problem, recorded as Unknown with a note.
  4. The engine applies the result: counts, transitions, an incident opened or closed, timeline entries, alerts, a live event. Transitions are saved at once so a crash never repeats an alert.
  5. The writer keeps the result and its rollups in memory and writes everything every 30 seconds in one transaction.
  6. The next run is the interval later, or the retry interval after a failure.

Data

TableKept
checks (every result with phases)7 days
rollup_hour (counts, latency sum, min, max, 16-bucket histogram)400 days
rollup_day (by local day)forever
incidentsforever
events (timeline)180 days
outbox, deliveries30 days

A snapshot of the database (VACUUM INTO) is written nightly at 03:00 local time to backups/; the last 7 are kept.

Budgets

On the free Google Cloud e2-micro (1 GB, slow disk): about 8 MB of memory and a few percent of a shared vCPU with 60 monitors; one SQLite transaction per 30 seconds; about 3 GB a month of outbound traffic (the free allowance is 200 GiB on the Standard network tier).