This page was automatically translated and may contain errors. View in English.
함께 일하기

Machine Learning Engineer - Inference Optimization

Jobgether

Remote · 정규직

가장 먼저 지원하세요

경험
어느
샐러리
채용 공고
1
게시됨
3시간전
작업 모드
재택근무
재개하다
신청 시 필수 사항

직무 설명

Overview

This opportunity is offered through a partner organization looking for a Machine Learning Engineer focusing on inference optimization based in Ireland. The role centers on enhancing the performance of advanced ML systems deployed in production settings.

This position involves bridging research and engineering by transforming innovative models into swift, dependable, and cost-effective AI solutions. Your efforts will influence scalability, user experience, and the operational efficiency of AI-driven products.

You will engage deeply with optimization efforts from model structure, GPU execution, to managing large-scale inference frameworks. Collaborating with research, infrastructure, and product teams, you will enable advanced AI capabilities.

Key Responsibilities

  • Enhance inference systems to boost latency, throughput, scalability, and reduce operational expenses.
  • Analyze and profile GPU and CPU inference pipelines to detect bottlenecks in components like memory handling, kernel operations, data batching, and flow.
  • Apply sophisticated optimization techniques such as quantization, key-value caching improvements, speculative decoding, batching, streaming, and model simplification.
  • Work alongside research engineers to transition experimental model architectures into dependable production systems.
  • Develop and maintain inference-serving infrastructure with modern frameworks, custom runtimes, or specialized serving platforms.
  • Evaluate and benchmark performance across diverse hardware, including GPUs, CPUs, and cloud environments.
  • Enhance system dependability, monitoring, observability, and minimize costs during real production workloads.
  • Participate in refining engineering practices to enhance quality, scalability, and maintainability of ML infrastructure.

Qualifications and Experience

  • Substantial experience in machine learning inference optimization or building high-performance ML systems.
  • In-depth knowledge of ML principles including neural network designs, attention mechanisms, memory management, and compute graph theory.
  • Expertise in deep learning frameworks like PyTorch and deploying models in production contexts.
  • Strong background in GPU performance tuning with technologies such as CUDA, ROCm, Triton, or kernel-level optimization.
  • Proven track record of scaling inference setups beyond research prototypes for actual user scenarios.
  • Proficiency in programming spanning machine learning and systems engineering.
  • Ability to work autonomously in dynamic environments with shifting priorities.
  • Familiarity with inference platforms like TensorRT, ONNX Runtime, vLLM, or Triton is advantageous.
  • Knowledge of large language models, long-context inference, distributed architectures, low-latency services, or hardware optimization is a plus.
  • Contributions to open-source ML tools or inference-related projects are considered beneficial.

Benefits

  • Competitive salary complemented by significant equity shares.
  • Chance to work on performance-critical AI systems influencing real products.
  • High ownership level over infrastructure impacting scalability and operational efficiency.
  • Close collaboration with cutting-edge research, infrastructure, and product teams.
  • Engagement with advanced ML technologies and practical AI applications.
  • Culture centered on engineering excellence, experimentation, and quality assurance.
  • Flexible arrangements supporting remote work.
  • Opportunity to contribute to an innovative AI-driven organization’s expansion.

답변을 원하시면 남겨주세요. 다른 용도로는 사용하지 않습니다.

클릭하여 살펴보세요드래그 앤 드롭 또는 반죽 스크린샷

PNG, JPG, GIF, MP4, WebM, MOV · 파일당 최대 20MB · 최대 5개 파일

🤖
온라인 · 즉각적인 AI 도움말