# Monitors A monitor checks one thing on a schedule and decides whether it is **Up**, **Down**, **Slow**, **Paused** or **Unknown**. ## Kinds | Kind | Checks | Good for | |---|---|---| | HTTP(S) | A request and its answer: status code, text, JSON fields, a header, time | Websites, APIs, health endpoints | | TCP port | A connection, optionally the first line it sends | SSH, databases, mail, anything with a port | | Ping | ICMP echo replies and loss | Whole servers | | DNS | A record from a resolver, optionally containing a value | Catching DNS changes or deletions | | TLS certificate | The certificate on any TLS port | Expiry on non-web services | | Heartbeat | That something called Farsight on time | Backups, cron jobs, scripts | ### HTTP rules - **Expected status**: codes and ranges, like `200-299,401`. A gateway that answers 401 without a key is healthy, so `401` is a fine expectation for it. - **Body contains / does not contain**: plain text, searched in the first 1 MB. - **JSON rules**: a path, an operator and a value, for example `$.status == "ok"` or `$.queue.depth < 100`. Operators: `==` `!=` `<` `<=` `>` `>=` `contains` `exists` `not_exists`. - **Expected header**: for example `content-type` contains `json`. - **Maximum time**: slower than this fails the check. (**Slow above** only marks it slow.) - **Follow redirects** (on), **verify TLS** (on), **connect to IP** (`resolve_to`, to check an origin behind a CDN), **IP version**. - **Request**: method, headers (mark secret ones), body, Basic or Bearer auth. Credentials in the URL become Basic auth. Secrets are never shown again after saving. They are only sent to the monitor's own scheme, host and port, never on a redirect to somewhere else. ### Heartbeats Make a heartbeat monitor and Farsight gives it a secret URL. Have your job call it when it finishes: ``` pg_dump ... && curl -fsS https://farsight.example.com/hb/ ``` If no call arrives within the interval plus the grace time, the monitor goes down. A job can also report failure with its own words: `curl -fsS -d "disk full" https://farsight.example.com/hb//fail`. ## Timing | Setting | Default | Meaning | |---|---|---| | Interval | 60 s | Time between checks (10 s to a day) | | Timeout | 10 s | The whole check must finish in this time | | Confirm after | 2 | Failed checks in a row before Down | | Retry after | 10 s | After a failure, check again this soon | | Recover after | 1 | Good checks in a row before Up again | | Slow above / slow after | off / 3 | A good but slow check, so many times in a row, shows Slow | So with the defaults a dead site is confirmed down about a minute and ten seconds after it dies at the latest. ## What happens when - **A check fails**: it is counted and re-checked soon. Nothing is sent yet. - **Enough failures**: the monitor is Down, an incident opens, alerts go out. - **It comes back**: the incident closes with its duration, a recovery alert goes out. - **It flips up and down 4 times in 30 minutes**: it is **flapping**. One alert says so, then Farsight waits until it has stayed in one state for 15 minutes. - **A certificate nears expiry**: one warning at 14, 7, 3 and 1 days. - **Farsight's own internet fails**: failures are not counted (see [When things go wrong](failure-modes.md)). - **Only the path from Farsight is broken**: when a target never answers, Farsight asks a second checker on Cloudflare's network. If it can reach the target, the failure is a network path problem and is not counted. ## Reading what happened - **Recent checks** on a monitor's page come as runs: checks in a row with the same result fold into one line ("25 failed checks · 22:14 to 22:26 · got 502", "172 good checks · 321 to 402 ms"). Open a line to see its checks one by one. - **An incident** has its own page (open it from Incidents or from the monitor's page): when it started and ended, the cause, and **What happened** in plain sentences: the first failed check, when it was confirmed, the second opinion's verdict, every alert to every channel with whether it arrived, the acknowledgement and the end. Its failed checks are below as runs, and the note is kept there. ## Duplicate The monitor page's menu has **Duplicate**: a new-monitor form filled from that monitor, named " copy", not on the status page. Secrets (header values marked secret, passwords, tokens) are never copied; type them again. ## Pausing and muting **Pause** stops checks entirely (and closes an open incident). **Mute** keeps checking but holds alerts back for a while, during planned work for example. ## Groups, tags and order Groups organise the dashboard. Tags are free labels for filtering. The public status page uses its own **public name** and **public group** so internal names never leak.