/docs-archive/dev/misc/inventory/ · last edited 31 Dec 2020 · Marketing
We discussed today the need to improve the monitoring and alerting system for the Polygon server to ensure it remains stable and available for family and friends. We decided to implement a dashboard using Prometheus and Grafana for visualizing server health and resource usage, and to set up alerting via Twilio to notify me of any critical issues via text messages to my phone.
For the dashboard, we plan to include CPU and memory usage, disk space, network traffic, and any custom metrics we may add later. We also agreed to create alerts for disk usage exceeding 80%, sudden spikes in network traffic, and any critical errors in the logs.
Open questions we need to address include which specific Prometheus metrics to monitor, how often to refresh the dashboard, and whether we should include alerts for less critical issues that still warrant attention. Additionally, we need to confirm the setup process for Twilio and ensure that our phone numbers are correctly configured for receiving alerts.