PDF

Creating and Editing Alerting Rules in NetCrunch

This chapter provides a comprehensive guide on creating and editing alerting rules within NetCrunch. It covers the key properties that define an alert, such as Severity, Description, and Target State, and explores advanced configurations, including additional alerting conditions and automatic alert correlation. By understanding these concepts, network administrators can optimize monitoring strategies, reduce false positives, and ensure prompt and accurate incident responses.

Key Properties

When setting up alerting rules in NetCrunch, it's essential to comprehend the meaning of the fields that define an alert. You'll encounter three critical fields: Severity, Description, and Target State.

Severity

The Severity field signifies the importance level of the alert. It helps prioritize issues and determine the urgency of the response needed. The severity levels usually range from informational to critical. Here's a brief overview:

  • Critical - The severity represents serious issues that need immediate attention, as they can cause significant disruptions if not resolved quickly.
  • Warning - It indicates potential issues that may escalate if not addressed but are not critical.
  • Informational - Alerts that provide informative insights but do not require immediate action.
  • Minor - Minor severity alerts indicate issues that require attention but do not pose an immediate threat to operational functionality, serving as preventative signals to help maintain system health.

Choosing the correct severity level is crucial for proper incident management and ensuring that the right team members prioritize their responses effectively.

Description

The Description field should concisely explain the alert and its context. I.e. High Processor Utilization (> 90%)

Operational State

The Operational State field is essential for understanding the impact of alerts in NetCrunch. It distinguishes between a service merely experiencing issues and one that is entirely down, thus providing critical insight that supports effective incident management and swift resolution.

  • Operational - The object/service is functional. This includes scenarios where performance might be degraded, but the service remains available.

  • Non-Operational - The object/service is not functioning as expected, typically indicating that it is not responding or is entirely offline.

Event Condition

The Event Condition is the core element of an alerting rule in NetCrunch, defining the primary reason for triggering an alert. This condition specifies the exact circumstances under which an alert is generated. For example, it could be set to detect a state change—such as a service transitioning from responding to down—or to monitor a threshold on a particular metric, like CPU usage exceeding a predefined limit.

The alert will be triggered only when the Event Condition is met in conjunction with any configured Additional Alerting Conditions. This dual-layer approach ensures that alerts are both precise and contextually relevant. By accurately defining the Event Condition, you help ensure that the alert system responds only to significant events, thereby reducing false positives and alert fatigue while enabling prompt and effective incident response.

Additional Alerting Condition

NetCrunch allows you to define additional conditions for each alert, regardless of whether a node status change, an event log alert, or an SNMP trap trigger it. These additional conditions enable you to fine-tune the alerting process by triggering actions even when a primary event has not occurred. For example, you can specify conditions based on specific time intervals or the absence of an event, ensuring alerts are only activated under precise circumstances. The available additional conditions include:

  • On Event Condition: Trigger an alert when a specific event occurs.
  • Only if the time between: Execute the alert only if the event occurs within a designated time window.
  • Only if time not between: Trigger the alert only outside a defined time window.
  • The event did not happen in the time between: Ensure that the alert is only triggered if the expected event does not occur during a specified interval.
  • The event did not happen after a given time: Activate the alert if a particular event fails to occur after a set time.
  • Event pending for more than (a given time): Trigger an alert if an event remains unresolved for longer than a predefined duration.

Alert Close Correlation

NetCrunch's alert close correlation feature streamlines the management of alerts by automatically grouping related notifications. For internally triggered alerts—those generated by state changes or threshold breaches—NetCrunch is designed to close them automatically once the issue is resolved.

However, a closing correlation is required for external alerts originating from traps, syslog, or web messages. In these cases, external alerts are paired with corresponding resolution events to ensure they are appropriately dismissed. Alerts can be configured to close automatically after a set period, or they can be manually cleared by an operator once confirmed as resolved. This approach reduces alert fatigue and maintains clarity in the system, ensuring that the actual status of network issues is accurately represented.

Read more about conditions and correlation

Conclusion

Creating and editing alerting rules in NetCrunch requires a clear understanding of the Severity, Description, and Target State fields. Effectively using these fields, along with advanced configurations like additional alerting conditions and automatic alert close correlation, can enhance your monitoring strategy, streamline issue resolution, and ultimately ensure the reliability of your systems.