Report
We have some metrics that scrape fine >99% of the time, but encounter an FTL error in the scraper code a few times per day for some reason. It's just a couple failed scrapes here and there, but it would be nice to know what condition causes this, and hopefully there is a way to improve it.
Example configuration that works >99% of the time
- name: azure_redis_respond_errortype_errors
description: "Errors."
resourceType: RedisCache
azureMetricConfiguration:
metricName: errors
aggregation:
type: Maximum
dimension:
name: ErrorType
resourceDiscoveryGroups:
- name: our-resource-group
Here's the error that gets logged when the scrape fails:
[00:25:01 FTL] Failed to scrape resource for metric 'azure_redis_respond_errortype_errors'
System.InvalidOperationException: Sequence contains no matching element
at System.Linq.ThrowHelper.ThrowNoMatchException()
at Promitor.Core.Metrics.MeasuredMetric.CreateForDimension(Nullable`1 value, String dimensionName, TimeSeriesElement timeseries) in /src/Promitor.Core/MeasuredMetric.cs:line 68
at Promitor.Integrations.AzureMonitor.AzureMonitorClient.QueryMetricAsync(String metricName, String metricDimension, AggregationType aggregationType, TimeSpan aggregationInterval, String resourceId, String metricFilter, Nullable`1 metricLimit) in /src/Promitor.Integrations.AzureMonitor/AzureMonitorClient.cs:line 117
at Promitor.Core.Scraping.AzureMonitorScraper`1.ScrapeResourceAsync(String subscriptionId, ScrapeDefinition`1 scrapeDefinition, TResourceDefinition resourceDefinition, AggregationType aggregationType, TimeSpan aggregationInterval) in /src/Promitor.Core.Scraping/AzureMonitorScraper.cs:line 72
at Promitor.Core.Scraping.Scraper`1.ScrapeAsync(ScrapeDefinition`1 scrapeDefinition) in /src/Promitor.Core.Scraping/Scraper.cs:line 103
Expected Behavior
Properly defined metrics should work consistently when Azure Monitor itself is in good condition.
Actual Behavior
Properly defined metrics experience intermittent fatal errors on scrapes, just a few per day always.
Steps to Reproduce the Problem
- I provided a metric definition in the report section.
- You could stand up an Azure Redis Cache to run that metric configuration against. Unfortunately, you would have to wait 24 hours to catch a few occurrences of the error.
Component
Scraper
Version
v2.9.1
Configuration
Configuration:
- name: azure_redis_respond_errortype_errors
description: "Errors."
resourceType: RedisCache
azureMetricConfiguration:
metricName: errors
aggregation:
type: Maximum
dimension:
name: ErrorType
resourceDiscoveryGroups:
- name: our-resource-group
Logs
[00:25:01 FTL] Failed to scrape resource for metric 'azure_redis_respond_errortype_errors'
System.InvalidOperationException: Sequence contains no matching element
at System.Linq.ThrowHelper.ThrowNoMatchException()
at Promitor.Core.Metrics.MeasuredMetric.CreateForDimension(Nullable`1 value, String dimensionName, TimeSeriesElement timeseries) in /src/Promitor.Core/MeasuredMetric.cs:line 68
at Promitor.Integrations.AzureMonitor.AzureMonitorClient.QueryMetricAsync(String metricName, String metricDimension, AggregationType aggregationType, TimeSpan aggregationInterval, String resourceId, String metricFilter, Nullable`1 metricLimit) in /src/Promitor.Integrations.AzureMonitor/AzureMonitorClient.cs:line 117
at Promitor.Core.Scraping.AzureMonitorScraper`1.ScrapeResourceAsync(String subscriptionId, ScrapeDefinition`1 scrapeDefinition, TResourceDefinition resourceDefinition, AggregationType aggregationType, TimeSpan aggregationInterval) in /src/Promitor.Core.Scraping/AzureMonitorScraper.cs:line 72
at Promitor.Core.Scraping.Scraper`1.ScrapeAsync(ScrapeDefinition`1 scrapeDefinition) in /src/Promitor.Core.Scraping/Scraper.cs:line 103
Platform
Microsoft Azure
Contact Details
No response
Report
We have some metrics that scrape fine >99% of the time, but encounter an FTL error in the scraper code a few times per day for some reason. It's just a couple failed scrapes here and there, but it would be nice to know what condition causes this, and hopefully there is a way to improve it.
Example configuration that works >99% of the time
Here's the error that gets logged when the scrape fails:
Expected Behavior
Properly defined metrics should work consistently when Azure Monitor itself is in good condition.
Actual Behavior
Properly defined metrics experience intermittent fatal errors on scrapes, just a few per day always.
Steps to Reproduce the Problem
Component
Scraper
Version
v2.9.1
Configuration
Configuration:
Logs
Platform
Microsoft Azure
Contact Details
No response