Google Cloud Vertex AI monitoring integration
Monitor your Google Cloud Vertex AI resources to track online prediction volumes and errors, monitor training job resource utilization, and ensure your machine learning workloads run efficiently. Google Cloud Vertex AI monitoring lets you maintain visibility into model serving and training performance.
Benefits
Monitoring Vertex AI ensures that your service remains reliable, secure, and high-performing. Key benefits include:
- Identify prediction latency spikes or error surges before they impact model consumers.
- Track training job resource utilization to fine-tune compute allocation and reduce cost.
- Maintain SLA compliance and reduce downtime with automated alerts and IT automation.
- Consolidate Vertex AI metrics with other GCP services for a single pane of glass view.
- Ensure smooth monitoring across multiple models and training jobs with bulk configuration and agent upgrades.
Use cases
Here are some use cases where Google Cloud Vertex AI monitoring will be helpful:
- Monitor average prediction latency and error counts to detect model serving bottlenecks in real time.
- Capture training job resource utilization trends to quickly identify under- or over-provisioned jobs.
- Visualize prediction request volumes and peak usage periods to plan capacity and optimize scaling.
- Set thresholds for prediction error rates and training resource utilization to ensure contractual obligations are met.
- Track network throughput for training jobs to strengthen data pipeline reliability.
- Integrate Vertex AI monitoring with other GCP services (i.e., Cloud Storage, BigQuery, or Compute Engine) for holistic visibility.
Setup and configuration
- Adding Google Cloud Vertex AI while configuring a new Google Cloud monitor
If you have not configured a Google Cloud monitor yet, add one by following the steps below:
- Go to Cloud > GCP > Add GCP Monitor or Admin > Cloud Monitoring > Google Cloud Platform (GCP).
- Provide a unique display name for identification purposes.
- Upload the JSON file that contains the private key of the service account to authenticate and perform resource discovery.
- Select Vertex AI from the Select the Resources for Monitoring list.
- Select existing notification profiles, user alert groups, tags, and IT automation templates or add new ones. You can also integrate alarms with your preferred third-party service.
- Click Start GCP Monitoring.
- Adding Google Cloud Vertex AI to an existing Google Cloud monitor
If you already have a Google Cloud monitor configured for the service account, you can add Google Cloud Vertex AI by following the steps below:
- Go to Cloud > GCP and select your GCP monitor.
- Click the hamburger
icon next to Service View and select Edit, which brings you to the Edit GCP Monitor page. - On the Edit GCP Monitor page, select Vertex AI from the Select the Resources for Monitoring list and click Save.
- After successful configuration, go to Cloud > GCP > Vertex AI. Now you can view the discovered Google Cloud Vertex AI resources.
It will take approximately five minutes to discover new GCP resources.
Licensing
Each Google Cloud Vertex AI monitor consumes one basic monitor license.
Polling frequency
A Google Cloud Vertex AI monitor collects minute-wise metric data, and the statuses of your Google Vertex AI resources are reported every five minutes.
Supported metrics
| Metric name | Description | Statistic | Unit |
|---|---|---|---|
| Prediction Count | The number of online prediction requests processed by the deployed model version. | Total | Count |
| Prediction Error Count | The number of online prediction requests that resulted in an error. | Total | Count |
| Prediction Latencies | The latency distribution of online prediction requests processed by the deployed model version. | Average | Microseconds |
| Prediction Response Count | The number of responses returned for online prediction requests. | Total | Count |
| Training CPU Utilization | The percentage of allocated CPU resources currently used by the training job. | Average | Percentage (%) |
| Training Memory Utilization | The percentage of allocated memory currently used by the training job. | Average | Percentage (%) |
| Training Disk Utilization | The amount of allocated disk space currently used by the training job. | Average | Bytes |
| Training Network Bytes Received | The number of bytes received over the network by the training job. | Total | Bytes |
| Training Network Bytes Sent | The number of bytes sent over the network by the training job. | Total | Bytes |
Threshold configuration
- Global configuration
- In the web client, go to the Admin section on the left navigation pane.
- Select Configuration Profiles from the left pane and select Threshold and Availability from the drop-down menu.
- Click Add Threshold Profile in the top-right corner.
- For Monitor Type, select Vertex AI.
- Specify the threshold values for the required metrics, then click Save.
- Monitor-level configuration
- In the web client, go to Cloud > GCP > Vertex AI.
- Select a resource you would like to set a threshold for, then click the hamburger
icon. - Select Edit, which directs you to the Edit Vertex AI Monitor page.
- Specify the threshold values for the required metrics, then click Save.
IT automation
IT automation tools automatically resolve performance degradation issues. These tools react to events proactively rather than waiting for manual intervention. IT automation tools automate repetitive tasks and automatically remediate threshold breaches. The alarm engine continually evaluates system events for which thresholds are set and executes the mapped automation when there is a breach.
How to configure IT automation for a monitor
Configuration rules
Editing multiple monitors to associate different monitor groups or add a different tag can be a tedious process. With configuration rules, you can automate the configuration settings of your monitoring resources. Create custom rules to track configuration changes continuously and achieve the ideal configuration settings.
How to add configuration rules
Summary
The Summary tab will give you the performance data organized by time for the metrics listed above. To view the summary:
- Go to Cloud > GCP > Vertex AI.
- Select a resource.
- Click the Summary tab.
Configuration Details
The Configuration Details tab provides details on the configurations of application instances. To get the configuration details:
- Go to Cloud > GCP > Vertex AI.
- Select a resource.
- Click the Configuration Details tab.
Reports
Gain in-depth data about the various parameters of your monitored resources and accentuate your service performance using our insightful reports.
To view reports for a Google Vertex AI resource:
- Go to the Reports section on the left navigation pane.
- Select Vertex AI from the menu on the left.
- You can find the Availability Summary Report, Performance Report, and Inventory Report for one selected monitor. Or, you can get the Summary Report, Availability Summary Report, Health Trend Report, and Performance Report for all Google Vertex AI monitors.
You can also get reports from the Summary tab of the Google Vertex AI monitor:
- Click the Summary tab.
- Get the Availability Summary Report of the monitor by clicking Availability.
- You can also find the Performance Report of the monitor by clicking any chart title.
Related content
- Monitoring Google Cloud Platform
- List of Google Cloud services supported for monitoring
- Possible reasons why GCP resources are not added for monitoring
- What permissions should I have in my Google account to enable Google Cloud Platform (GCP) monitoring?
- How to create a service account JSON file to authenticate for the discovery of GCP resources
- How to create a service account in the GCP console
