This page was automatically translated and may contain errors. View in English.
Trabajo en conjunto

Machine Learning Engineer - Inference Optimization

Jobgether

Remote · Jornada completa

Sé el primero en postularte

Experiencia
Cualquier
Salario
Vacantes
1
Al corriente
Hace 7 horas
Modo de trabajo
Trabajar desde casa
Reanudar
Se requiere solicitud

Descripción del trabajo

Position Overview

This role is for a Machine Learning Engineer specializing in inference optimization based in Australia. The successful candidate will engage in enhancing the performance of sophisticated machine learning systems operating in real production environments. Serving at the intersection of research and software engineering, this position involves translating advanced model architectures into solutions that are fast, dependable, and cost-effective.

Your contributions will directly affect the scalability of models, consumer experience, and overall efficiency of AI-driven products. You will delve into optimizing performance aspects ranging from model design and GPU execution to large-scale inference infrastructure. Collaboration with cross-functional teams including research, infrastructure, and product groups will be integral to pushing AI system capabilities forward.

Key Responsibilities

  • Enhance ML inference pipelines focusing on reducing latency, boosting throughput, expanding scalability, and minimizing operational costs.
  • Perform comprehensive profiling to discover performance constraints across GPU and CPU inference workflows, examining memory consumption, kernel execution, batching strategies, and data flow.
  • Apply advanced optimization methods such as model quantization, KV-cache improvements, speculative decoding, batching, streaming, and model simplification.
  • Work alongside research engineers to bring new model architectures into production and convert experimental outcomes into stable, scalable systems.
  • Develop and maintain inference serving platforms leveraging current frameworks, custom runtime environments, or specialized serving technologies.
  • Conduct performance benchmarking across various hardware setups, including GPUs, CPUs, and cloud infrastructure.
  • Drive enhancements in system dependability, monitoring capabilities, observability, and manage costs effectively under real production loads.
  • Advance engineering practices to improve the robustness, scalability, and maintainability of machine learning infrastructure.

Candidate Requirements

  • Extensive professional experience in optimizing ML inference or building high-performance machine learning systems.
  • Strong expertise in machine learning principles, including neural network structures, attention mechanisms, efficient memory utilization, and computational graph optimization.
  • Practical experience with PyTorch or analogous deep learning frameworks for deployment of models in production settings.
  • Competence in GPU-centric performance tuning, with familiarity in CUDA, ROCm, Triton, or similar kernel-level optimizations.
  • Proven track record of scaling inference systems to serve actual user traffic beyond research or benchmark scenarios.
  • Advanced programming skills bridging machine learning and systems engineering.
  • Ability to thrive in fast-moving environments, with a self-driven mindset, ownership mentality, and ability to adapt to shifting priorities.
  • Experience using inference platforms like TensorRT, ONNX Runtime, vLLM, or Triton is advantageous.
  • Knowledge or practical work involving large language models, extended-context inference, distributed architectures, low-latency services, or hardware-level optimization is beneficial.
  • Contributions to open-source ML or inference tool projects are a plus.

Benefits

  • Competitive salary coupled with significant equity options.
  • Engagement in high-impact AI systems with tangible product influence.
  • Substantial autonomy over infrastructure that impacts scalability and operational efficiency.
  • Close-knit teamwork with research, infrastructure, and product development units.
  • Exposure to cutting-edge machine learning technologies and practical AI applications.
  • Workplace culture emphasizing technical excellence, experimentation, and quality outcomes.
  • Flexible arrangements supporting remote work.
  • Chance to contribute to a pioneering AI-driven company’s growth trajectory.

Additional Information

This recruitment is managed on behalf of a partner company who handle all application reviews and subsequent hiring steps. The company utilizes an AI-enhanced matching process to evaluate candidates objectively and efficiently, with final hiring choices made by their internal team.

Applicants should be aware that submitting their application entails consent to the processing of personal data for recruitment purposes under applicable laws (including GDPR). Artificial intelligence tools may be used in parts of the hiring process to assist recruiters, but decisions remain under human control.

Déjelo si desea una respuesta; no lo utilizaremos para ningún otro fin.

Haz clic para navegar, arrastrar y soltar, o pasta una captura de pantalla

PNG, JPG, GIF, MP4, WebM, MOV · Máximo 20 MB cada uno · Hasta 5 archivos

🤖
En línea · Ayuda instantánea con IA