Overview
This page focuses on collecting and validating Lightning ASR metrics with Prometheus and exposing them through the Prometheus Adapter.Architecture
/* The original had a syntax error in Mermaid—edges must connect nodes, not labels. “Metrics” is now a node, and edge directions/names are consistent. */Components
Prometheus
Collects and stores metrics from Lightning ASR pods. Included in chart:values.yaml
ServiceMonitor
CRD that tells Prometheus which services to scrape. Enabled for Lightning ASR:values.yaml
Prometheus Adapter
Converts Prometheus metrics to Kubernetes custom metrics API. Configuration:values.yaml
Available Metrics
Lightning ASR exposes the following metrics:Verify Metrics Setup
Check Prometheus
Forward Prometheus port:- Status → Targets: Lightning ASR endpoints should be “UP”
- Graph: Query
asr_active_requestsorasr_batch_queue_depth- should return data - Status → Service Discovery: Should show ServiceMonitor
Check ServiceMonitor
Check Prometheus Adapter
Verify custom metrics are available:Custom Metric Configuration
Add New Custom Metrics
To expose additional metrics for your own autoscaling setup:values.yaml
Prometheus Configuration
Retention Policy
Configure how long metrics are stored:values.yaml
Storage
Persist Prometheus data:values.yaml
Scrape Interval
Adjust how frequently metrics are collected:values.yaml
Recording Rules
Pre-compute expensive queries:Alerting Rules
Create alerts for anomalies:Debugging Metrics
Check Metrics Endpoint
Directly query Lightning ASR metrics:Test Prometheus Query
Access Prometheus UI and test queries:Check Prometheus Targets
View Prometheus Logs
Troubleshooting
Metrics Not Appearing
Check ServiceMonitor is created:Custom Metrics Not Available
Check Prometheus Adapter logs:High Cardinality Issues
If Prometheus is using too much memory:- Reduce label cardinality
- Increase retention limits
- Use recording rules for complex queries
Best Practices
Use Recording Rules
Use Recording Rules
Pre-compute expensive queries:Then use this in your autoscaling logic instead of a raw query
Set Appropriate Scrape Intervals
Set Appropriate Scrape Intervals
Balance responsiveness vs storage:
- Fast autoscaling: 15s
- Normal: 30s
- Cost-optimized: 60s
Enable Persistence
Enable Persistence
Always persist Prometheus data:
Monitor Prometheus Itself
Monitor Prometheus Itself
Track Prometheus performance:
- Query duration
- Scrape duration
- Memory usage
- TSDB size
Use Grafana for Visualization
Use Grafana for Visualization
What’s Next?
Grafana Dashboards
Visualize metrics

