
A. Role Purpose ( A High level Description of the Role)
The MLOps Engineer (GCP) is responsible for operationalizing, deploying, monitoring, and scaling production‑grade AI/ML solutions on Google Cloud Platform (GCP). The role focuses on building reliable, automated, and secure end‑to‑end ML platforms and pipelines, enabling seamless collaboration between Data Science, AI Engineering, Platform, and Operations teams. This role ensures that ML models developed using Vertex AI and other GCP services can be consistently trained, versioned, deployed, monitored, and governed throughout their lifecycle. Strong, ands‑on experience with GCP‑native MLOps services, CI/CD, and production operations is mandatory.
#
Key Accountabilities (WHAT)
Key Tasks
(HOW)
Key Deliverables(RESULT)
1
MLOps Platform & Architecture Design (GCP)
· Design end‑to‑end MLOps architectures using GCP‑native services.
· Define standardized patterns for training, deployment, monitoring, and retraining.
· Ensure high availability, scalability, security, and compliance.
· Align MLOps designs with enterprise cloud and security standards.
·
· MLOps reference architectures and design documents.
· Approved technical blueprints for ML platforms.
· Reusable MLOps architectural patterns.
2
Model Deployment & Serving Automation (Vertex AI)
· Deploy models to Vertex AI Endpoints (online and batch).
· Implement model versioning, rollout, rollback, and traffic splitting strategies.
· Automate inference pipelines and serving workflows.
· Optimize latency, scalability, and cost of inference workloads.
· Production‑ready mode deployments
· Automated serving pipelines.
· Versioned and auditable model endpoints.
3
CI/CD & ML Pipeline Automation
· Build CI/CD pipelines for ML training and deployment.
· Automate ML pipelines using Vertex AI Pipelines.
· Integrate source control, testing, and artifact registries.
· Enforce reproducibility across environments (dev, test, prod).
· Automated ML pipelines and CI/CD workflows.
· Reproducible builds and deployments.
· Pipeline execution and audit logs.
4
Monitoring, Observability & Reliability
· Implement monitoring for model performance, drift, and data quality.
· Set up logging, alerting, and SLOs for ML systems.
· Define retraining triggers based on drift or performance degradation.
· Support incident analysis and remediation for ML services.
· Model and pipeline monitoring dashboards.
· Drift and performance reports.
· Stable, highly available ML services.