ML Inference / Serving Infrastructure Engineer
<div class="show-more-less-html__markup show-more-less-html__markup--clamp-after-5 relative overflow-hidden"> <p><strong>Location: </strong>San Francisco (US) or Berlin (GER)</p><p><strong>Type: </strong>Permanent (Full-time)</p><p><strong>Start Date: </strong>As soon as possible / by arrangement</p><p><br/></p><p><strong>ABOUT DAINA</strong></p><p>DAiNA is a precision-oncology company focused on enabling personalized cancer treatment for individual patients. We combine comprehensive molecular tumor data including genomics, transcriptomics (bulk, single cell and spatial), proteomics and epigenetics, with AI-driven analysis. Our platform connects multi-omic profiling with functional ex-vivo tumor models and personalized liquid-biopsy monitoring, creating a continuous workflow from biopsy and treatment selection through to therapy monitoring and adaptation. We also operate GMP manufacturing to produce individualized N=1 therapeutics. In short, we help physicians make more informed, personalized treatment decisions based on high-dimensional molecular tumor data.</p><p>For more information, visit: www.daina.com</p><p><br/></p><p><strong>THE ROLE</strong></p><p>As ML Inference / Serving Infrastructure Engineer, you will build and optimize the serving infrastructure that enables DAiNA to run AI and LLM workloads reliably, securely and efficiently.</p><p>You will focus on high-throughput inference, GPU orchestration, autoscaling, observability and deployment patterns for open-weight models. Your work will support DAiNA’s AI-driven reporting, retrieval, decision-support and computational oncology workflows.</p><p>A key part of the role is to ensure that model serving is performant, reproducible and suitable for secure biomedical environments with strong requirements around privacy, reliability and operational control.</p><p><br/></p><p><strong>WHAT YOU’LL DO</strong></p><ul><li>Build and optimize inference infrastructure for open-weight models using tools such as vLLM, TGI, TensorRT-LLM or comparable frameworks.</li><li>Design deployment patterns for secure cloud, isolated or sovereign environments.</li><li>Implement GPU orchestration, autoscaling, monitoring and observability for AI workloads.</li><li>Optimize model serving for latency, throughput, cost and reliability.</li><li>Define performance baselines, load-testing approaches and operational metrics.</li><li>Instrument systems for latency, throughput, error rates, utilization and drift.</li><li>Work with AI/ML, RAG, fine-tuning and platform teams to integrate serving infrastructure into DAiNA workflows.</li><li>Document deployment patterns and operational guidance clearly for production use.</li></ul><p><br/></p><p><br/></p><p><strong>WHAT YOU BRING</strong></p><ul><li>Strong infrastructure, MLOps or platform engineering background with experience serving ML or LLM systems in production.</li><li>Strong Python skills and solid systems engineering knowledge.</li><li>Hands-on experience with vLLM, TGI, TensorRT-LLM, Kubernetes-style orchestration or comparable tools.</li><li>Practical understanding of GPU workloads, autoscaling, model deployment and inference optimization.</li><li>Strong focus on observability, reliability, cost control and operational robustness.</li><li>Experience building clean, reproducible infrastructure components and documentation.</li><li>Ability to work closely with engineering, AI/ML and platform stakeholders.</li></ul><p><br/></p><p><strong>NICE TO HAVE</strong></p><ul><li>Experience with secure cloud, VPC-isolated, on-premise or sovereign deployment environments.</li><li>Experience tuning performance across different GPU hardware.</li><li>Familiarity with open-weight model families such as Llama, Mistral, Qwen or comparable models.</li><li>Experience with monitoring, logging, tracing, model versioning or deployment automation.</li><li>Background in healthcare, biotech, diagnostics, pharma or another sensitive-data environment.</li></ul><p><br/></p><p><strong>WHY DAINA</strong></p><ul><li>Employer contributions toward private health insurance or supplementary health coverage, as well as pension or retirement savings.</li><li>Hybrid model with ~60% of working time <strong>expected </strong>on-site and the rest <strong>remotely</strong>, depending on team and business needs.</li><li>Performance-based bonus opportunity, depending on company and individual performance.</li><li>A personal development budget for conferences, courses, certifications, and training.</li><li>The opportunity to work directly on real patient cases and help shape a first-in-class precision-oncology platform</li><li>A proactive, collaborative team with fast decision-making and strong ownership of your domain.</li></ul><p><br/></p><p><strong>HOW TO APPLY</strong></p><p>Send your CV and a short note on why this role fits you to <strong>recruiting@daina.com</strong>, quoting “ML Inference / Serving Infrastructure Engineer” in the subject line. We review applications on a rolling basis and aim to reply quickly.</p><p><br/></p><p><em>DAiNA is an equal-opportunity employer. We welcome applicants of every background and assess every candidate on merit, regardless of age, gender, ethnicity, religion, disability, sexual orientation or origin. If you need any adjustment to the process, let us know in your applicatio</em></p> </div>