Skip to main content

Operations

This section describes how to operate Breeze Agent at scale: monitor health, tune performance, and troubleshoot issues across your estate.

Use these guides to validate host check-ins, understand fleet resource impact, plan larger rollouts, or investigate operational failures.

Monitor Agent Health

Use Health & Status to confirm that Breeze Agent is running and reporting data from installed hosts and Kubernetes nodes.

This guide covers:

  • Host-level health checks: Verifying Breeze Agent schedules, recent runs, and log activity on Linux, macOS, and Windows hosts.
  • Kubernetes health checks: Confirming DaemonSet readiness and reviewing Breeze pod logs for recurring errors.
  • Cloudaware status validation: Using CMDB list views and Breeze status fields to identify active, stale, or never-seen hosts.
  • Unhealthy host investigation: Checking scheduling, outbound connectivity, certificates, tags, and configuration before moving to troubleshooting procedures.

Evaluate Performance and Plan for Scale

Use Performance & Scale to understand Breeze Agent resource requirements and prepare for larger or performance-sensitive deployments.

This guide covers:

  • Runtime and resource usage: Typical cycle duration, CPU and memory footprint, scheduling cadence, and idle behavior between runs.
  • Network impact: OS facts payload size, bandwidth optimization, outbound connectivity requirements, and proxy or TunHub options for restricted networks.
  • Large-scale rollout planning: Pilot rings, phased deployment, device management tooling, maintenance windows, and capacity monitoring across large host estates.
  • Performance tuning: Cadence adjustments, plugin scheduling, and temporary logging changes for latency-sensitive systems or short-term troubleshooting.

Troubleshoot Breeze Agent

Use Troubleshooting to investigate agents that are not running, reporting, or connecting as expected.

This guide covers:

  • Check-in and stale-data issues: Verifying agent schedules, logs, and outbound connectivity when hosts stop reporting to Cloudaware.
  • Cloud metadata errors: Restoring access to instance metadata services and correcting proxy exclusions or role-related issues.
  • TLS, proxy, and server connectivity: Validating proxy settings, trusted CA bundles, Breeze Server endpoints, and outbound HTTPS access.
  • Plugin and data collection issues: Investigating missing facts or integration results, unsupported environments, and disabled or misconfigured plugins.
  • Kubernetes deployment issues: Diagnosing missing pods, DaemonSet scheduling problems, image, permission, and network failures.
  • Escalation preparation: Collecting logs, environment details, host IDs, and timestamps before contacting Cloudaware support.