# How Farsight works One process, `farsightd` (Rust, tokio, two worker threads), with SQLite on the blocking pool and the dashboard built into the binary. ``` monitors (SQLite) --> scheduler: one small task per monitor | probes: http, tcp, ping, dns, tls (heartbeats arrive over HTTP) | canary: is Farsight's own internet up? (before counting a failure) v engine: state machine per monitor --> incidents, timeline, alerts, live feed | | v v writer: results + rollups, one alerts: 20 s grouping per channel, SQLite transaction per 30 s outbox, retries, Brevo then Resend, Discord, signed webhooks API (axum): dashboard, agents, status page, heartbeats, Server-Sent Events, HTTPS with Let's Encrypt ``` ## Modules | Module | Job | |---|---| | `config` | Flags and `FARSIGHT_*` variables; nothing else reads the environment | | `store/*` | SQLite: schema migrations and every query, grouped by what they store | | `model/*` | Monitors (kinds, validation, the API shape, secret masking) and check results | | `rules` | Status code sets, JSON paths and operators | | `probe/*` | One check of one kind; shared resolve, connect and TLS steps with timings; the private-address guard | | `scheduler` | A task per monitor: spread start times, jitter, retry interval, 16 probes at most at once, "check now", heartbeat watchdogs | | `canary` | Three TCP targets; two failing means Farsight is blind | | `confirm` | A second opinion from a Cloudflare Worker (`worker/confirm`) for failures where the target never answered: a path problem is not counted | | `engine/state` | The pure state machine (Up, Down, Slow, flapping, reminders, certificate warnings) | | `engine` | Applies results, opens and closes incidents, writes the timeline, emits alerts and live events | | `writer` + `rollup` | Batches results, hourly rollups with a latency histogram, daily rollups by local day | | `alerts/*` | Routing, grouping, rendering (email, Discord, webhook), the outbox and retries | | `auth/*` | Passkeys (WebAuthn, verified in-house), authenticator codes, sessions, step-up, API keys, Turnstile, brakes | | `api/*` | Routes, the caller check, JSON errors, idempotency, rate limits, security headers | | `status_page` | The public pages (main, per service, history), plain HTML in three layouts: settings, ranges and bars (pure), caches, renderers, the stylesheet and fonts | | `serve` | HTTP, HTTPS with automatic certificates (TLS-ALPN-01), the port-80 redirect | | `certs` | The certificates HTTPS serves: the main one and one per custom domain of the status page, chosen by the name a browser asks for; custom ones start and stop while the server runs | | `domains` | The status page's custom domains: their DNS read at the zone's own nameservers, their state, and their certificates started or stopped (every 30 s until all are live, hourly after) | | `net` | A small HTTPS client for Farsight's own calls (providers, webhooks, Turnstile) | ## A check, start to finish 1. The monitor's task wakes (its interval, spread by its id and jittered 5%). 2. It takes one of 16 probe slots and runs the probe under the monitor's timeout. The HTTP probe resolves, connects, does TLS and sends the request, timing each step; follows redirects on new connections (never carrying credentials to another origin); reads at most 1 MB of the body when a rule needs it. 3. A failure asks the canary for a reading at most 5 s old. If Farsight is blind, the result is recorded as Unknown and nothing changes state. A failure where the target never answered (timeout, no route, DNS, ping loss) then asks the second-opinion Worker on Cloudflare; if Cloudflare reaches the target, it is a network path problem, recorded as Unknown with a note. 4. The engine applies the result: counts, transitions, an incident opened or closed, timeline entries, alerts, a live event. Transitions are saved at once so a crash never repeats an alert. 5. The writer keeps the result and its rollups in memory and writes everything every 30 seconds in one transaction. 6. The next run is the interval later, or the retry interval after a failure. ## Data | Table | Kept | |---|---| | `checks` (every result with phases) | 7 days | | `rollup_hour` (counts, latency sum, min, max, 16-bucket histogram) | 400 days | | `rollup_day` (by local day) | forever | | `incidents` | forever | | `events` (timeline) | 180 days | | `outbox`, `deliveries` | 30 days | A snapshot of the database (`VACUUM INTO`) is written nightly at 03:00 local time to `backups/`; the last 7 are kept. ## Budgets On the free Google Cloud e2-micro (1 GB, slow disk): about 8 MB of memory and a few percent of a shared vCPU with 60 monitors; one SQLite transaction per 30 seconds; about 3 GB a month of outbound traffic (the free allowance is 200 GiB on the Standard network tier).