Qualifications
 / 
Full time

MLOps and AI Platform Engineer

About Us

Whizzbridge is hiring a Mid to Senior MLOps and AI Platform Engineer. WhizzBridge is a technology solutions provider and a talent enabler. On one hand, it offers clients access to best of breed engineering talent and delivers their mission critical projects using the industry's best practices. On the other hand, it attracts and trains engineering talent on cutting edge technologies, programming languages and project management practices that set them up for successful professional and financial growth.

What We Offer

  • Paid Leaves
  • Medical Insurance
  • Paid Udemy Courses and Certifications
  • Career Progression Program

Job Description

  1. Build and operate the infrastructure on which client AI and machine learning systems are deployed, monitored and maintained.
  2. Design and maintain data pipelines that move, schedule and transform data at volume, and own the data quality layer that AI systems depend on.
  3. Build reproducible training, fine tuning and deployment pipelines with versioned datasets, models and configurations.
  4. Deploy model serving infrastructure using managed endpoints or self hosted serving stacks, and manage scaling, cost and availability.
  5. Implement observability across AI systems, including tracing, latency and throughput monitoring, cost attribution per request and drift detection.
  6. Establish CI and CD for machine learning and LLM systems, including automated evaluation gates before deployment.
  7. Manage secrets, environment configuration and access control across client environments.
  8. Own incident response for production AI systems, including rollback procedures and root cause analysis.
  9. Containerise applications and manage orchestration across cloud platforms.
  10. Partner with AI Application and RAG engineers to move prototypes into supportable production systems.
  11. Document infrastructure, runbooks and handover material to a standard that allows a client team to operate the system independently.

Requirements

  1. Bachelor's degree in Computer Science, Software Engineering or a related field, or equivalent demonstrable experience.
  2. Three or more years of professional experience in DevOps, platform engineering, data engineering or MLOps, with direct exposure to machine learning or LLM workloads.
  3. Strong Python and proficiency in shell scripting.
  4. Production experience with Docker, and with container orchestration or managed container services.
  5. Hands on experience with at least one major cloud platform, including its machine learning services. AWS SageMaker or Bedrock, GCP Vertex AI, or Azure ML.
  6. Experience building and scheduling data pipelines using tools such as Airflow, Dagster, Prefect or dbt.
  7. Experience with experiment tracking and model registry tooling such as MLflow or Weights and Biases.
  8. Infrastructure as code experience, such as Terraform or CloudFormation.
  9. Solid understanding of SQL and data modelling.
  10. Strong written and verbal English.

Qualifications

  1. Experience with model serving frameworks such as vLLM, Triton or TorchServe.
  2. Experience with LLM observability platforms such as Langfuse, Phoenix or Helicone, and with OpenTelemetry.
  3. Experience with data warehouse and lakehouse platforms such as Snowflake, BigQuery or Databricks.
  4. Experience with feature stores.
  5. Experience with GPU provisioning, scheduling and cost optimisation.
  6. Exposure to security review processes for AI systems, including data residency and compliance requirements.
  7. Kubernetes experience.