Skip to main content

Responding to alerts

This page explains what to do when a Monitoring alert triggers: how to investigate it, which action to take for the affected resource, and how to confirm recovery.

Start with Investigate the alert, then go to the section for the resource named in the alert: Main Database, API Cluster, In-Memory Database, Installed apps: Compute, or Account: Billing. Confirm the observed value and user impact before changing capacity.

Every capacity change described here is made on a project page, saved, and applied with the Review & Deploy button that appears once changes are pending. While a deployment is in progress you cannot make further changes, and the project keeps running on the previous configuration until the deployment finishes. If a deployment fails, a warning appears at the top of every page in the TagoDeploy console; see Deployment history.

Investigate the alert​

  1. Read the notification. It names the rule, the metric, the value, the threshold, the severity, and the time. It carries no link, so sign in to TagoDeploy and open the project's Monitoring tab. With the recommended thresholds, Critical means users are affected or about to be, and Warning means you have time.
  2. Check the Overview, then the trend. The Alert activity list shows whether the alert is still active, the value that crossed, and when it started; the Notifications card shows whether each action was Sent or Failed, and why. Then open the Monitoring app's Logs page, search for the rule name, and read the last few Check Completed entries. A single crossing followed by an Alert Resolved entry was a spike. A run of rising values is a capacity problem.
  3. Look at the longer view. The project Overview and each service page have a section of charts, also labeled Monitoring, from 1 hour to 30 days. A value that has climbed for a week needs a different response than one that jumped an hour ago.
  4. Look for a cause. Open Deployments under Management for a recent deployment, a new app, or a feature change, and the project Logs page under Management, the runtime log rather than the Monitoring app's, for errors around the same time. Ask what changed in your application: new devices, a new dashboard, a new Analysis.
  5. Act. Use the section below for the resource. If repeat notifications get in the way while you work, pause the rule with its Enabled toggle; the active alert clears on the next check after you resume it and the value is back to normal. Resolve alert in the Overview row menu closes the current record, but it only holds the next notification until the rule's cooldown ends; while the value stays over the line the alert fires again. To stop notifications, pause the rule with Enabled.
  6. Confirm recovery. After your change is deployed, wait for the Alert Resolved entry in the Monitoring Logs. If the metric stays over the line, go one step further (a larger machine, a higher instance maximum) or contact TagoIO support.

A rising metric can indicate growing demand, inefficient processing, or a fault. Do not assume that every sustained alert requires more capacity.

warning

Larger machines, additional replicas, and higher instance limits can increase costs. Reducing capacity can affect availability or throughput. Confirm the service's restart and interruption behavior before changing production resources; do not assume a capacity change is interruption-free.

Manage repeat notifications​

Check the threshold's Repeat every setting if reminders are too frequent.

Do not use Enabled as a notification-only mute: disabling a rule stops its checks. If you disable a rule during an incident, arrange another way to watch the resource and record when the rule must be re-enabled.

Closing an alert record does not fix the underlying condition. Verify the metric and application behavior before treating an incident as resolved.

Main Database​

Settings are on the Main Database page: Machine controls the database server size, and Additional reader instances adds read-only instances for read traffic.

Before changing the machine size, review the deployment's availability impact and choose an appropriate maintenance window. Do not assume that adding a reader makes a machine change interruption-free.

CPU utilization is high​

Database CPU increases with both reads and writes. Heavy queries, such as aggregations over a full month of data, and updates across large device or entity datasets can cause spikes. Sustained growth can also follow an increase in connected devices, ingestion, or dashboard activity.

Compare CPU with Read IOPS and Write IOPS at the same timestamps in Check Completed details. Increased write activity can help identify ingestion or bulk updates; increased read activity can point toward dashboard and Analysis queries. These correlations guide the investigation but do not identify the responsible operation on their own.

Review Analysis scripts and application changes around the time the alert started. A scheduled query that repeatedly scans months of data may need a smaller date range or a less frequent schedule rather than more capacity.

For sustained write-heavy processing or large data submissions, consider a larger Machine. For read-heavy processing, consider an additional reader instance after confirming that the workload can use it.

Freeable memory is low​

Freeable memory is memory available to the database instance. Databases use memory for caching, so a low value alone does not establish that the instance is undersized. Compare it with the normal range and look for slower requests, increased I/O, or other signs of pressure.

Check whether a new dashboard or Analysis script retrieves substantially larger datasets than before. Reduce unnecessary retrieval or divide large operations into smaller batches where practical.

If expected workloads cause sustained memory pressure and degraded performance, consider a larger Machine.

Connections are high​

Connection count can rise when the API cluster scales out because API instances maintain database connections. It can also rise with concurrent work from other services.

Check whether API instance count increased around the time of the alert. Review Maximum instances on the API page alongside the database's supported connection limit. API capacity and database capacity need to support the same workload.

If the increase is expected, evaluate whether a larger database Machine provides the required connection capacity. If the count stays unexpectedly high, investigate application behavior rather than assuming that scaling is the only solution.

Do not reduce API capacity solely to lower database connections without considering the effect on requests. Contact TagoIO support if the applicable connection limit is unclear or the available machine sizes cannot support the workload.

Read IOPS, Write IOPS, or network throughput is high​

These metrics describe activity, not necessarily a fault. High values can be expected during ingestion, exports, imports, bulk updates, or scheduled Analysis runs.

Compare them with CPU, freeable memory, and request performance. If CPU is high or available memory is unusually low at the same time, investigate those conditions using the guidance above.

If a bulk job coincides with user-facing degradation, consider reducing its batch size or moving it to a quieter period. If performance remains normal, review whether the alert threshold reflects the resource's expected workload.

API Cluster​

Settings are on the API page: Machine, Minimum instances, Maximum instances, Scale on CPU utilization, Cooldown for scaling up, and Cooldown for scaling down.

The API Cluster serves requests to the API, Admin, and TagoRUN. Its autoscaling cooldowns control how quickly capacity changes; they are separate from an alert rule's cooldown, which controls alert behavior.

CPU utilization is high​

The cluster autoscales on CPU. Sustained utilization above the scaling target is a reason to check whether it has reached Maximum instances, whether scaling is still in progress, or whether individual requests require substantial processing.

Compare CPU with the Running Instances chart. If the instance count stays at the configured maximum while CPU remains high, the cluster has no room to scale out under its current settings.

Choose the response based on the workload:

  • High concurrency: consider raising Maximum instances. Check downstream database capacity before allowing more API instances.
  • Expensive individual requests: reduce request complexity or consider a larger Machine. Increasing instance count does not necessarily make an individual request less expensive.
  • Recurring bursts that outpace scaling: review Scale on CPU utilization, the scaling cooldowns, and Minimum instances. Earlier scaling or more capacity already running may help, but can increase cost.

Also investigate demand-side changes. Higher profile limits may allow a larger workload. Analysis scripts that make many API calls or dashboards that retrieve unnecessarily large datasets can increase load without a corresponding increase in users.

Reduce avoidable work before adding capacity where practical. For machine types beyond the available options, contact TagoIO support.

Memory utilization is high​

Requests that retrieve large datasets and many concurrent operations can increase memory usage. Compare the start of the increase with new dashboards, Analysis scripts, or changes in request size and frequency.

Reduce the amount of data retrieved and processed when the application does not need it. Check whether the increase is temporary or whether memory stays high under the expected workload.

If legitimate workloads still produce sustained memory pressure, consider a larger Machine. After the change, verify request performance and application errors, not just the memory metric.

In-Memory Database​

Settings are on the In-Memory Database page: Machine and Additional reader replicas. Replicas serve reads.

CPU utilization is high or freeable memory is low​

Heavy real-time activity, such as many dashboards retrieving large datasets at once, can increase processing demand and memory pressure. Compare CPU and freeable memory with ingestion volume, dashboard activity, and their normal ranges.

For sustained processing or memory pressure under expected workloads, consider a larger Machine. For read-heavy activity, consider reader replicas to distribute reads.

Match the change to the bottleneck. Reader replicas are not a general solution for write-heavy activity, and low freeable memory alone does not show that additional capacity is necessary. Check performance and workload changes before scaling.

Cache hits are dropping​

Cache hits count successful cache reads. A drop that follows a decrease in API traffic may be expected. A drop while traffic stays steady needs more investigation, but does not by itself prove that the cache is too small.

Check whether requests now access different data or whether a recent application change altered how the cache is used. Compare the drop with freeable memory and user-facing performance.

Do not choose a larger Machine based only on the hit count. Establish whether memory pressure or another resource constraint accompanies the change.

Network inbound or outbound is high​

High network traffic may reflect increased ingestion, larger payloads, or more real-time reads. Compare the timing with application activity and check whether users experience delays or errors.

Review the resource's throughput capabilities before selecting a larger Machine. A high traffic value without performance degradation may call for a threshold adjustment rather than a capacity change.

Installed apps: Compute​

Most installed apps have an Instances page with a Machine size and an instance range. Apps that autoscale also expose CPU scaling and cooldown settings. Available controls depend on the app type; see MQTT Instances or Middleware.

Review the app's deployment behavior before applying changes, especially when it maintains long-lived connections or processes messages in flight.

CPU utilization is high​

Compare CPU with message rate, request concurrency, and current instance count. A brief burst from an external network needs a different response from sustained growth during normal traffic.

If the app reaches Maximum Instances while CPU remains high under an expected concurrent workload, consider raising the maximum. If individual operations require substantial processing, review the workload and consider a larger Machine.

For middleware, compare the spike with external-network message timestamps. High CPU with unexpectedly low throughput can indicate repeated work or failures. Inspect runtime logs and check callback configuration, authentication, and retries before adding capacity.

For delayed downlinks, establish where the delay occurs. If app capacity is the bottleneck, review the instance maximum and scaling timing. Scaling will not correct a callback or external-network problem.

For MQTT, use the Connections page to compare live sessions with normal activity. A jump in connected devices can increase demand, but also check message rate and payload size before choosing a wider instance range.

Memory utilization is high​

Buffering, message aggregation, and large-payload parsing can increase memory usage. Compare spikes with incoming traffic and look for sustained usage near the limit on the app's charts.

A larger Machine may help when each instance needs more memory. Increasing instance count may help distribute concurrent work, but does not necessarily reduce the memory required by an individual operation.

Investigate application behavior, not just capacity. If memory keeps growing without a corresponding workload increase, adding resources may only delay the next alert.

Account: Billing​

A budget threshold is crossed​

Open Bills and review month-to-date spending, spending over time, and charges by service. Confirm that the billing view covers the same scope and period as the alert.

Look for changes around the start of the increase: a deployment, more connected devices, an ingestion burst, or additional resources. Use app-level filtering where available to isolate installed apps.

Then choose the response:

  • If growth is expected: review the approved budget before raising Budget (USD) on the rule.
  • If growth is unexpected: identify the services responsible before changing capacity. Review API instance ranges, Analysis runtime memory, feature instance counts, and installed-app machine sizes and instance ranges as applicable.
  • If resources appear unused: confirm whether they are still needed, including temporary or restored projects. Check availability requirements and normal peaks before reducing capacity or removing resources.
warning

A budget alert sends notifications; it does not cap spending. Raising Budget (USD) changes the alert threshold, not the amount being spent. Reducing capacity to control costs can affect availability and performance.

The bill summary is an estimate, so the final amount can change. Reducing usage affects future spending; it does not undo charges already accumulated. A budget alert may therefore remain active after corrective action.

Track the rate of new spending to evaluate the response. Review Repeat every if recurring reminders are useful while the budget condition remains active.

Confirm recovery​

After a workload change or successful deployment:

  1. Check new Check Completed entries for the affected rule.
  2. Confirm that the value no longer meets the threshold condition and look for Alert Resolved.
  3. Verify that the user-facing symptoms have recovered.
  4. Continue watching long enough to include the workload that caused the alert.
  5. Re-enable any disabled rules and restore any temporarily changed notification settings.

For budget alerts, confirm that the rate of new spending matches expectations rather than relying on immediate alert resolution.

If the metric remains beyond the threshold, reassess the cause before adding more capacity. A completed deployment is not proof that the incident is resolved.

When to contact TagoIO support​

Contact support when:

  • A capacity change fails to deploy.
  • The metric remains beyond its threshold after the expected workload and capacity changes.
  • You need a machine type beyond the available options.
  • Check Completed entries repeatedly report a resource-read error.
  • An enabled rule has an unexplained gap in checks, or Monitoring stops recording new entries.