Mimecast
Principal Machine Learning Engineer
Overview
As a Principal Machine Learning Engineer, you will set technical direction for Mimecast's ML capabilities, including our GCI products or advanced email threat-detection products. This is a senior individual-contributor role: you will own the hardest modeling and systems problems end to end, define the architecture other engineers build on, and be a technical authority for product and engineering leadership when the direction is not obvious.
About Mimecast
The work people build is worth protecting, and it's harder to protect than it used to be. AI agents now move at machine speed with human-level access, and a small slip-up can become a very public one. We disrupt cybercriminal activity before that happens.
Requirements & Eligibility
- Breadth across transformer architectures, RNNs, CNNs, generalized linear models, and gradient-boosted trees
- Deep Python proficiency and strong command of PyTorch, Hugging Face transformers, and NLP tooling
- Experience with dense and lexical retrieval, including embeddings, vector indexes and approximate nearest-neighbor search, BM25, TF-IDF, and hybrid approaches
- Experience working with datasets exceeding two million examples and highly imbalanced data
- A track record of owning production ML systems on AWS
- Working knowledge of model-serving frameworks such as TorchServe, FastAPI, and NVIDIA Triton Inference Server/KServe
- Hands-on experience running CUDA workloads in production
- Fluency with AI-native development tools and modern LLM application patterns
Key Responsibilities
- Develop and own ML systems end to end, from data sourcing, cleaning, and labeling strategy through feature engineering, model development, deployment, and monitoring.
- Set the ML architecture across model design, serving, and surrounding systems, optimizing accuracy, latency, and throughput for highly imbalanced threat-detection data.
- Set technical direction for production model serving using AWS SageMaker, NVIDIA Triton Inference Server, ensemble/KServe patterns, hardened container images, and integration with enrichment and gateway layers.
- Benchmark and prototype alternatives to de-risk major decisions, then give leadership defensible technical recommendations.
- Establish reproducible ML standards, including versioned datasets, region-partitioned data, and shared experimentation workflows.
- Make model observability and efficacy measurement first-class concerns through distributed tracing, threshold-independent metrics, raw-payload capture, and monitoring for real regressions.
- Own capacity planning and rollout strategy, including throughput per core or GPU, utilization headroom, peak-load provisioning, and phased regional canary or shadow deployments.
- Diagnose production incidents, close the structural gaps they expose, and act as a primary reviewer and mentor across the ML codebase.
Disclaimer: Trace Hiring is an independent job board. We are not directly affiliated with Mimecast. Please verify all details on the official company application portal.