You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
PolicyEngine needs functional monitoring across substantially more API behavior than the Better Stack free tier can cover. Existing host-level checks do not detect failures limited to particular route groups, and realistic calculation monitoring may require authenticated requests, response validation, and polling asynchronous operations.
Adding paid Better Stack coverage immediately would commit us to a service and recurring cost before we have compared it with an internally operated solution or other hosted products.
This issue supersedes #2155 and should preserve its required coverage while broadening the work to include option evaluation and platform design.
Objective
Evaluate and implement a cost-effective monitoring approach for PolicyEngine APIs. The selected approach may be:
a small custom service operated by PolicyEngine;
an existing self-hosted monitoring product; or
a paid hosted monitoring service, including Better Stack, if its cost and capabilities are the best fit.
Do not select an implementation until the alternatives have been compared against explicit requirements.
Investigation
Inventory the API endpoints and complete user operations that need monitoring, including metadata, household, policy, user-profile, household-calculation, economy, and budget-window behavior.
Define request frequency, acceptable response time, failure conditions, environments, and required geographic coverage.
Determine requirements for:
public and authenticated requests;
safe test data and credential storage;
multi-request and asynchronous operations;
response status, content, and schema validation;
Slack failure and recovery notifications, including duplicate suppression;
execution history, latency reporting, and diagnostic evidence;
configuration stored in version control;
maintenance ownership, security updates, reliability, and data retention.
Compare custom, self-hosted, and hosted approaches. Document expected recurring cost, implementation effort, operational burden, service limits, and vendor dependence.
Prototype the leading candidates where documentation alone does not establish whether they satisfy the requirements.
Implementation
After recording the decision, implement the selected approach and migrate or supplement the existing checks. Monitoring should exercise representative successful operations with safe inputs rather than only checking API root responses.
Completion criteria
The requirements and endpoint coverage inventory are documented.
The alternatives and their expected costs and maintenance requirements are compared in writing.
The selected approach and reasons for selecting it are documented.
Required public, authenticated, and asynchronous API operations can be checked where applicable.
Failures and recoveries produce actionable Slack notifications without excessive duplicate messages.
Monitor definitions, credentials guidance, operating instructions, and ownership are documented.
Existing checks are retained, migrated, or deliberately removed with the reason recorded.
Problem
PolicyEngine needs functional monitoring across substantially more API behavior than the Better Stack free tier can cover. Existing host-level checks do not detect failures limited to particular route groups, and realistic calculation monitoring may require authenticated requests, response validation, and polling asynchronous operations.
Adding paid Better Stack coverage immediately would commit us to a service and recurring cost before we have compared it with an internally operated solution or other hosted products.
This issue supersedes #2155 and should preserve its required coverage while broadening the work to include option evaluation and platform design.
Objective
Evaluate and implement a cost-effective monitoring approach for PolicyEngine APIs. The selected approach may be:
Do not select an implementation until the alternatives have been compared against explicit requirements.
Investigation
Implementation
After recording the decision, implement the selected approach and migrate or supplement the existing checks. Monitoring should exercise representative successful operations with safe inputs rather than only checking API root responses.
Completion criteria