We are seeking an ML Engineer specializing in Python to package, serve, and operate AI capabilities across our client's shared platform. Their customers rely on fast, defensible answers across massive volumes of unstructured text, documents, audio, and video for critical missions in defence, intelligence, and law enforcement.
In this role, you will bridge the gap between research and production—taking models out of notebooks and packaging them into versioned, production-ready artifacts served behind standardized APIs. You will own inference performance, cost, and output provenance across cloud, on-premise, and fully air-gapped environments.
Key Responsibilities
- Production Serving: Package models as immutable versioned artifacts, deploy them onto shared runtime infrastructures (e.g., Triton, vLLM), and expose them behind standards-compliant contracts.
- Benchmarking & Model Selection: Evaluate and select models using a centralized testing framework with reproducible runs, clear task metrics, and complete documentation.
- Performance & Cost Optimization: Manage inference latency, throughput, dynamic batching, quantization, and GPU placement to keep performance high and cost accountable.
- Observability & Operations: Instrument dashboards, trace request pipelines, monitor for quality/drift, and participate in on-call rotations for served capabilities.
- Constrained & Offline Delivery: Prepare ML pipelines and OCI artifacts for unattended execution in air-gapped customer environments with strict audit and provenance constraints.
Requirements
- Depth in at least one key domain: Language & NLP, Speech & Audio, Vision & Document AI, or Agentic & LLM Systems.
- Advanced production skill in Python (3.12+, uv, PyTorch, Hugging Face Transformers) with a proven track record of reading, adapting, and evaluating model implementations.
- Hands-on experience with production model serving (vLLM, NVIDIA Triton), FastAPI, and streaming OpenAI-compatible API contracts.
- Experience deploying GPU workloads on Kubernetes (node selection, autoscaling, scheduling constraints) and managing observability (tracing, latency percentiles, GPU utilization).
- Strong evaluation hygiene with task-appropriate metrics (F1, mAP, WER, BLEU/COMET, DER), regression gating in CI, and experiment tracking.
- Firm grasp of ML provenance, reproducible builds, and immutable artifact packaging (OCI model supply chain).
Benefits
- Fully Remote Set-Up
- Company equipment and a home office budget
- Relocation Support
Education and experience
- Passionate about EU Security and Defence.