← The wire

Principal ML Engineer

Mimecast · United States of America - Ohio - Columbus

Posted 22h ago · first seen by the radar 2h ago · last checked on the employer's board 3m ago

Staff · Onsite · Full time

Principal Machine Learning Engineer

About Mimecast

The work people build is worth protecting, and it's harder to protect than it used to be. AI agents now move at machine speed with human-level access, and a small slip-up can become a very public one. We disrupt cybercriminal activity before that happens. We think fast, go big, and always demand more of ourselves. We work hard, deliver, and repeat. We grow with real determination and put success well within reach. We push each other to be better and expect to be pushed back, in a community built on respect, where everyone is counted.
Work Protected.

Overview

As a Principal Machine Learning Engineer, you will set technical direction for Mimecast's ML capabilities, including our GCI products or advanced email threat-detection products. This is a senior individual-contributor role: you will own the hardest modeling and systems problems end to end, define the architecture other engineers build on, and be a technical authority for product and engineering leadership when the direction is not obvious.

Our threat ML runs on Mimecast's shared AI enrichment platform, the detection and extraction infrastructure serving models across the product suite. You will work at the intersection of applied ML, production model serving, and cross-team integration, making architectural decisions that hold up under production load and customer commitments.

Mimecast is an AI-first engineering organization. You will use AI development tools, including Claude Code, Cursor, and MCP integrations, in your daily work and establish the workflows, patterns, and quality bar for the team. You will also design product systems using LLM and agent-based patterns.

Employees are expected to work from the office at least two days per week. This fosters collaboration, communication, performance, and learning; drives innovation and creativity within and between teams; introduces employees to priorities beyond their immediate realm; and supports important interpersonal relationships and connections.

What You'll Do

  • Develop and own ML systems end to end, from data sourcing, cleaning, and labeling strategy through feature engineering, model development, deployment, and monitoring.

  • Set the ML architecture across model design, serving, and surrounding systems, optimizing accuracy, latency, and throughput for highly imbalanced threat-detection data.

  • Set technical direction for production model serving using AWS SageMaker, NVIDIA Triton Inference Server, ensemble/KServe patterns, hardened container images, and integration with enrichment and gateway layers.

  • Benchmark and prototype alternatives to de-risk major decisions, then give leadership defensible technical recommendations.

  • Establish reproducible ML standards, including versioned datasets, region-partitioned data, and shared experimentation workflows.

  • Make model observability and efficacy measurement first-class concerns through distributed tracing, threshold-independent metrics, raw-payload capture, and monitoring for real regressions.

  • Own capacity planning and rollout strategy, including throughput per core or GPU, utilization headroom, peak-load provisioning, and phased regional canary or shadow deployments.

  • Diagnose production incidents, close the structural gaps they expose, and act as a primary reviewer and mentor across the ML codebase.

  • Partner with Product, platform engineering, and adjacent teams to shape the roadmap and communicate technical complexity and business implications through engineering and product leadership.

What You'll Bring

  • Breadth across transformer architectures, RNNs, CNNs, generalized linear models, and gradient-boosted trees, with the judgment to select the right approach for the problem rather than defaulting to the largest model.

  • Deep Python proficiency and strong command of PyTorch, Hugging Face transformers, and NLP tooling, plus working knowledge of ONNX Runtime, quantization such as FP16, and inference acceleration.

  • Experience with dense and lexical retrieval, including embeddings, vector indexes and approximate nearest-neighbor search, BM25, TF-IDF, and hybrid approaches.

  • Experience working with datasets exceeding two million examples and highly imbalanced data, using rigorous evaluation methods for precision and recall trade-offs, threshold selection, and test-set leakage prevention.

  • A track record of owning production ML systems on AWS, including SageMaker, S3, Athena, Lambda, Glue, Kinesis, and Bedrock, with Terraform, IAM, containers, and Kubernetes-based deployment.

  • Working knowledge of model-serving frameworks such as TorchServe, FastAPI, and NVIDIA Triton Inference Server/KServe, and the trade-offs among throughput, GPU efficiency, flexibility, and speed of iteration.

  • Hands-on experience running CUDA workloads in production, including driver, toolkit, and runtime alignment; GPU passthrough in containers; debugging GPU failures; and improving GPU utilization.

  • Fluency with AI-native development tools and modern LLM application patterns, including OpenAI-style chat-completion and structured tool/function-calling APIs, MCP, and agent frameworks.

  • Demonstrated technical leadership as an individual contributor, including setting direction, mentoring engineers across seniority levels, and communicating technical decisions and their business implications to technical and executive audiences.

  • An understanding of handling sensitive data in accordance with Master Service Agreements and compliance requirements.

  • A Ph.D. or Master's degree in a quantitative discipline, such as computer science, statistics, or mathematics, with substantial experience applying advanced ML to production problems; or equivalent depth demonstrated through a Bachelor's degree and a longer track record. We value demonstrated technical authority over a specific year count.

The base salary range for this position is $172,000 - $258,000 USD plus benefits. This range represents the minimum and maximum new hire compensation for this role. The position may also be eligible for incentive plans and additional benefits, in accordance with company policy and local regulations. Our salary ranges are determined by role, level, and location with individual compensation also dependent on factors such as qualifications, experience, and skills. Final offers will reflect these considerations and may vary accordingly.

Belonging at Mimecast

Cybersecurity is a community effort. That’s why we’re committed to building an inclusive, diverse community that celebrates and welcomes everyone – unless they’re a cybercriminal, of course.

We’re proud to be an Equal Opportunity and Affirmative Action Employer, and we’d encourage you to join us whatever your background. We particularly welcome applicants from traditionally underrepresented groups.

We consider everyone equally: your race, age, religion, sexual orientation, gender identity, ability, marital status, nationality, or any other protected characteristic won’t affect your application.

If you require any adjustments or accommodations due to a disability, or any other reason that may help you in your interview process, please let us know by emailing [email protected].

Due to certain obligations to our customers, an offer of employment will be subject to your successful completion of applicable background checks, conducted in accordance with local law.

It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment.

Listing read directly from Mimecast's applicant tracking system. Check frequency varies by source. Listings are removed after successful checks confirm they are no longer present.