Skip to main content

Health and Limits

Health checks monitor device collector resources on a self-managed Director and are configured in vmetric.yml. The timing constants below are not configurable at all; they are documented for operational awareness.

Health Check Configuration

Health checks monitor device collector resources. A resource must exceed its threshold for the configured number of consecutive iterations before action is taken.

Check Toggles

FieldYAML PathTypeDefaultDescription
Check CPU Usagehealth.check_cpu_usageboolfalseMonitor CPU usage
Check Memory Usagehealth.check_memory_usageboolfalseMonitor memory usage
Check Log File Usagehealth.check_log_file_usagebooltrueMonitor log file sizes
Check TMP Usagehealth.check_tmp_usageboolfalseMonitor temp directory
Check Report Usagehealth.check_report_usagebooltrueMonitor report files
Check Simulator Usagehealth.check_simulator_usagebooltrueMonitor simulator
Check Hanged Processeshealth.check_hanged_processesbooltrueDetect hung processes
Check Stats Usagehealth.check_stats_usagebooltrueMonitor stats collection

Thresholds

FieldDefaultDescription
CPU Limit80%CPU usage threshold percentage
Memory Limit1024 MBMemory usage threshold
Max Log Size10 MBMaximum log file size
Stats Retention24 hoursStats retention period

Threshold Iterations

All threshold iteration defaults are 10. A resource must exceed its threshold for this many consecutive check iterations before the health system takes action.


Internal Timing Constants

These values are hardcoded and not user-configurable. They are documented here for operational awareness.

Manager Intervals

ConstantValuePurpose
Heartbeat Update Interval5 sDevice heartbeat publishing frequency
Leader Check Interval1 sLeadership election check frequency
KV Sync Interval5 sJetStream KV synchronization
Config Sync Interval1 sConfiguration change polling
Heartbeat Check Interval15 sDevice health check and node offline threshold

Cluster Update Intervals

ConstantValuePurpose
Cluster Update Check Interval10 sPoll for pending updates
Node Timeout30 minMaximum wait for a single node update
Health Check Wait30 sPost-update health check delay
Leader Stability30 sRequired leader stability before proceeding
Command Retry Interval30 sBetween retry attempts
Max Retries3Maximum update retries per node
Max Offline Retries5Maximum retries for offline nodes

Configuration Limits

ConstantValuePurpose
Config File Max Size1 MBMaximum configuration file size
Retention Interval~3 monthsDefault data retention period