How Farsight works
One process, farsightd (Rust, tokio, two worker threads), with SQLite on the blocking pool and the dashboard built into the binary.
monitors (SQLite) --> scheduler: one small task per monitor
| probes: http, tcp, ping, dns, tls (heartbeats arrive over HTTP)
| canary: is Farsight's own internet up? (before counting a failure)
v
engine: state machine per monitor --> incidents, timeline, alerts, live feed
| |
v v
writer: results + rollups, one alerts: 20 s grouping per channel,
SQLite transaction per 30 s outbox, retries, Brevo then Resend,
Discord, signed webhooks
API (axum): dashboard, agents, status page, heartbeats, Server-Sent Events, HTTPS with Let's Encrypt
Modules
| Module | Job |
|---|---|
config | Flags and FARSIGHT_* variables; nothing else reads the environment |
store/* | SQLite: schema migrations and every query, grouped by what they store |
model/* | Monitors (kinds, validation, the API shape, secret masking) and check results |
rules | Status code sets, JSON paths and operators |
probe/* | One check of one kind; shared resolve, connect and TLS steps with timings; the private-address guard |
scheduler | A task per monitor: spread start times, jitter, retry interval, 16 probes at most at once, "check now", heartbeat watchdogs |
canary | Three TCP targets; two failing means Farsight is blind |
confirm | A second opinion from a Cloudflare Worker (worker/confirm) for failures where the target never answered: a path problem is not counted |
engine/state | The pure state machine (Up, Down, Slow, flapping, reminders, certificate warnings) |
engine | Applies results, opens and closes incidents, writes the timeline, emits alerts and live events |
writer + rollup | Batches results, hourly rollups with a latency histogram, daily rollups by local day |
alerts/* | Routing, grouping, rendering (email, Discord, webhook), the outbox and retries |
auth/* | Passkeys (WebAuthn, verified in-house), authenticator codes, sessions, step-up, API keys, Turnstile, brakes |
api/* | Routes, the caller check, JSON errors, idempotency, rate limits, security headers |
status_page | The public pages (main, per service, history), plain HTML in three layouts: settings, ranges and bars (pure), caches, renderers, the stylesheet and fonts |
serve | HTTP, HTTPS with automatic certificates (TLS-ALPN-01), the port-80 redirect |
certs | The certificates HTTPS serves: the main one and one per custom domain of the status page, chosen by the name a browser asks for; custom ones start and stop while the server runs |
domains | The status page's custom domains: their DNS read at the zone's own nameservers, their state, and their certificates started or stopped (every 30 s until all are live, hourly after) |
net | A small HTTPS client for Farsight's own calls (providers, webhooks, Turnstile) |
A check, start to finish
- The monitor's task wakes (its interval, spread by its id and jittered 5%).
- It takes one of 16 probe slots and runs the probe under the monitor's timeout. The HTTP probe resolves, connects, does TLS and sends the request, timing each step; follows redirects on new connections (never carrying credentials to another origin); reads at most 1 MB of the body when a rule needs it.
- A failure asks the canary for a reading at most 5 s old. If Farsight is blind, the result is recorded as Unknown and nothing changes state. A failure where the target never answered (timeout, no route, DNS, ping loss) then asks the second-opinion Worker on Cloudflare; if Cloudflare reaches the target, it is a network path problem, recorded as Unknown with a note.
- The engine applies the result: counts, transitions, an incident opened or closed, timeline entries, alerts, a live event. Transitions are saved at once so a crash never repeats an alert.
- The writer keeps the result and its rollups in memory and writes everything every 30 seconds in one transaction.
- The next run is the interval later, or the retry interval after a failure.
Data
| Table | Kept |
|---|---|
checks (every result with phases) | 7 days |
rollup_hour (counts, latency sum, min, max, 16-bucket histogram) | 400 days |
rollup_day (by local day) | forever |
incidents | forever |
events (timeline) | 180 days |
outbox, deliveries | 30 days |
A snapshot of the database (VACUUM INTO) is written nightly at 03:00 local time to backups/; the last 7 are kept.
Budgets
On the free Google Cloud e2-micro (1 GB, slow disk): about 8 MB of memory and a few percent of a shared vCPU with 60 monitors; one SQLite transaction per 30 seconds; about 3 GB a month of outbound traffic (the free allowance is 200 GiB on the Standard network tier).