Nanny is a monitoring tool that monitors the absence of activity.
Nanny runs an API server, which expects to be called every N seconds, and if no such call is made, Nanny notifies you.
Nanny can notify you via these channels (for now):
- print text to stderr
- sentry
- sms (twilio)
- slack (webhook)
- generic webhook (HTTP POST callback)
- xmpp (jabber)
Prometheus can already alert when a batch job stops succeeding, and Grafana Cloud accepts HTTP heartbeats too. If you already use them, you probably do not need Nanny.
Nanny is useful when you just want a small, self-hosted service that accepts one HTTP request from a cron job, backup, or anything else that should run regularly. Sending a request is often easier than exposing and maintaining metrics, and every signal can say when the next one is due.
Run API server:
$ LOGXI=* ./nanny
{"level":"INFO","msg":"Using config file","path":"nanny.toml"}
{"level":"INFO","msg":"Nanny listening","addr":"localhost:8080"}Call it via curl:
curl http://localhost:8080/api/v1/signal --data '{ "name": "my awesome program", "notifier": "stderr", "next_signal": "5s", "all_clear": false }'With this call, you tell nanny that if program named my awesome program does not call again within next_signal (5s), it should notify you using stderr notifier. Additionally, nanny appends the IP or X-Forwarded-For HTTP header to the program name. You can disable this behaviour by sending a X-Dont-Modify-Name along with the request. If you activate all_clear you will get an additional notification when the program sends a signal to nanny for the first time after an alert was sent.
After 5s pass, nanny prints to stderr:
2018-06-26T14:24:29+02:00: Nanny: I haven't heard from "my awesome program@127.0.0.1" in the last 5s! (Meta: map[])Download the Linux amd64 archive or Debian package from the releases.
Extract the archive, edit nanny.toml, and run Nanny directly:
tar -xzf nanny_0.5.0_linux_amd64.tar.gz
./nanny --config nanny.tomlInstall the downloaded package with APT:
sudo apt install ./nanny_0.5.0_linux_amd64.debThe package installs the configuration at /etc/nanny/nanny.toml, stores state
under /var/lib/nanny, and installs, enables, and starts the nanny.service
systemd unit. Edit the configuration and restart the service after making
changes.
git clone https://github.com/lunemec/nanny.git
cd nanny
make buildNanny requires Go >= 1.27 to build.
Starting with 0.5.0, Nanny is published to Docker Hub:
docker run -d -p 8080:8080 \
-e "NANNY_NAME=MyNanny" \
-v nanny-data:/var/lib/nanny \
docker.io/lunemec/nanny:latestNote:
- Use the
docker runenvironment variable parameter-ein combination withNANNY_<CONFIG_PROPERTY_HERE>to overridenanny.tomlfile configurations. - Optionally, mount your own configuration at
/etc/nanny/nanny.tomland persist state at/var/lib/nanny. - The
0.4container stored its binary, configuration, and database under/opt. Before upgrading a persistent0.4deployment, copy itsnanny.tomlto/etc/nanny/nanny.tomland its SQLite database to/var/lib/nanny/nanny.sqlite, keep the migrated files owned by UID/GID1000:1000, then update volume mounts.
Additionally, it's possible to run Nanny using the provided Docker Compose file (see docker-compose.yml):
docker-compose up -dSee nanny.toml for a configuration example. The fields are self-explanatory (I think). Please create an issue if anything does not make sense!
All enabled notifiers can be used via API, so enable only those you wish to allow.
ENV variables can be used to override the config file settings. They should be prefixed with NANNY_ and followed by same name as in nanny.toml.
Example:
NANNY_NAME="custom name" NANNY_ADDR="localhost:9090" LOGXI=* ./nanny
Print nanny version.
-
URL
/api/version
-
Method:
GET -
Success Response:
- Code: 200
Content:
Nanny vX.Y
- Code: 200
Content:
Signal Nanny to register notification with given parameters.
-
URL
/api/v1/signal
-
Method:
POST -
Headers:
X-Dont-Modify-Name: trueIf specified, Nanny won't modify thenamespecified in the JSON payload. Useful when your signals come from programs with dynamic IP addresses. -
Data Params
{ "name": "name of monitored program", "notifier": "stderr", # You can use only enabled notifiers, see config. "next_signal": "55s", # When to expect next call (or notify). "all_clear": false, # Optional all-clear notification when a call is received after an alert was sent "meta": { # Meta can contain any string:string values, "extra": "data" # they are passed to the notifiers and will eventually } # be passed to the user. }
-
Success Response:
- Code: 200
Content:
{"status_code":200, "status":"OK"}
- Code: 200
Content:
-
Error Response:
- Code: 400 Bad Request
Content:
{"status_code":400,"error":"unable to find notifier: "}
OR
- Code: 500 Internal Server Error
Content:
Message describing error, may be JSON or may be text.
- Code: 400 Bad Request
Content:
Return current signals as JSON.
-
URL
/api/v1/signals
-
Method:
GET -
Success Response:
- Code: 200
Content:
{ "nanny_name": "Nanny", "signals": [ { "name":"my awesome program", "notifier":"stderr", "next_signal":"2018-08-21T10:00:15+02:00", "all_clear":false, "meta": { "current-step": "loading" } }, { "name":"my awesome program without meta", "notifier":"email", "next_signal":"2018-08-21T09:45:00+02:00", "all_clear":false } ] }
- Code: 200
Content:
You can use one Nanny to monitor another Nanny or create a monitored Nanny-pair.
Run 1st nanny, on port 8080 that will use nanny at port 9090 as its monitor:
NANNY_ADDR="localhost:8080" LOGXI=* ./nanny --nanny "http://localhost:9090/api/v1/signal" --nanny-notifier "stderr"Run 2nd nanny, on port 9090 that will use 1st nanny on port 8080:
NANNY_ADDR="localhost:9090" LOGXI=* ./nanny --nanny "http://localhost:8080/api/v1/signal" --nanny-notifier "stderr"You may get some warnings until both Nannies are listening, but they will recover. If you stop one of them, the other will notify you.
Be sure to change nanny SQLite DB location! They would share the same DB and it could cause strange behavior.
This can be done in the config file or by setting NANNY_STORAGE_DSN ENV variable.
By default, nanny logs only errors. To enable more verbose logging, use LOGXI=* environment variable.
Global levels LOGXI=*=INF, LOGXI=*=WRN, and LOGXI=*=ERR are also supported.
Logs are written as JSON to standard output.
SMTP uses implicit TLS on port 465 and opportunistic STARTTLS on other ports. The
email.smtp_allow_insecure_auth option permits credentials to be sent without
TLS for trusted tunnels or legacy servers; enabling it can expose the SMTP
username and password on the network.
You can add extra meta-data to the API calls, which will be passed to all the notifiers. Metadata must conform to type map[string]string.
curl http://localhost:8080/api/v1/signal --data '{ "name": "my program", "notifier": "stderr", "next_signal": "5s", "all_clear": false, "meta":{"custom": "metadata"} }'These metadata will be displayed in the messages for stderr and email, and in tags for sentry.
Contributions welcome! Just be sure you run tests and lints.
$ make
Build
make build Build production binary.
make docker Build a Nanny container using Docker.
make snapshot Build release artifacts without publishing.
make release-check Validate the GoReleaser configuration.
make release-preflight Verify the release tag matches pushed master.
make release-verify Run every non-publishing release gate.
make release Verify, tag, and publish from clean master.
Dev
make run Run Nanny in dev mode, all logging and race detector ON.
make test Run tests.
make vet Run go vet.
make lint Run the pinned golangci-lint version.Releases are manual. From a clean master matching origin/master, log in to
Docker Hub once and export a GitHub token. Run one command, then enter an
unprefixed version such as 0.5.0 when prompted:
docker login docker.io
export GITHUB_TOKEN=...
make releaseFor a non-interactive run, use make release RELEASE_TAG=0.5.0. The
verification binary uses that version before the tag exists; GoReleaser reads
the published version from the tag. make release checks clean master, then
runs module consistency, preflight tests, build, vet, lint, race-enabled shuffled tests,
govulncheck, both Docker builds, GoReleaser validation, and a non-publishing
snapshot. Once every check passes, it creates and pushes the tag, checks it
again, and has GoReleaser create the GitHub release, Linux amd64 archive,
Debian package, SHA-256 checksums, and a Docker Hub image.
If publishing fails after the tag is pushed, rerun make release from the same
commit to reuse it.
Why write such a tool?
Sometimes you expect some job to run, say cron. But when someone messes up your crontab, or the machine is offline, you might not be notified.
Also often programs just log errors and fail silently, with nanny they fail loudly.
How do I secure my nanny?
To use HTTPS, or authentication you should use a reverse proxy like Apache or Nginx.
