What you’ll Do:
- Re-platform the ML Lifecycle: Migrate existing MLOps infrastructure, training pipelines, and CI/CD deployment workflows onto the new cloud environment's managed ML services, minimizing downtime and model drift during cutover.
- Drive Application Integration: Migrate APIs, microservices, and backend business logic connecting ML models to downstream applications, keeping prediction-serving SLAs intact through the migration.
- Migrate Data & Feature Pipelines: Re-platform data pipelines and feature stores onto the new cloud environment so training and real-time inference data flows continue uninterrupted.
- Own Production Health & Observability During Cutover: Establish monitoring, anomaly detection, and incident response for data quality, model drift, and A/B testing performance across both the current and new environments during the transition.
Optimize System Performance Post-Migration: Troubleshoot production bottlenecks and tune models for latency, throughput, and compute cost efficiency (GPU/CPU utilization, inference cost management). - Document & Hand Over: Produce migration runbooks and architecture documentation, and hand over ownership of migrated systems to the full-time team ahead of contract close.
What you’ll Need:
- 3+ years of professional experience in software engineering, machine learning engineering, or data science, with the ability to work independently from day one; this is a fixed-term engagement with limited onboarding runway.
- Working knowledge of core ML concepts and frameworks (PyTorch, TensorFlow, or Scikit-Learn), enough to validate that migrated models and pipelines produce correct output post-cutover.
- Strong proficiency with Python or Go in production environments, with exposure to distributed computing frameworks like Apache Spark considered a plus.
- Solid understanding of systems architecture and distributed systems, with the ability to maintain and re-platform complex ML pipelines and application logic against a defined timeline.
- Prior experience migrating infrastructure or workloads between environments (on-prem to cloud, or cloud to cloud) is a strong plus.
- Practical understanding of how MLOps pipelines interlock end to end, so migrating one component means tracing and moving everything upstream and downstream that depends on it.
- A pragmatic, high-velocity engineering approach, comfortable delivering against a fixed scope and timeline and balancing rapid delivery with a clean handover at contract end.
It'd be Great if you Have:
- Hands-on experience with major cloud platforms' ML and data services (for example AWS SageMaker, Glue, and EKS; GCP Vertex AI; or Azure ML).
- Experience with container orchestration tools like Kubernetes, or workflow orchestration tools like Apache Airflow, for managing scalable ML workloads.
