This page was automatically translated and may contain errors. View in English.
f

Software Engineer, Distributed Systems

fal

Remote · 全职

抢先申请

经验
3年以上
薪水
USD 180,000 – USD 250,000 / year
职位空缺
1
发布
1 小时前
工作模式
在家办公
恢复
需要申请

职位描述

About fal

fal is at the forefront of the generative media ecosystem, enabling the development of next-generation AI products. We create the infrastructure, tools, and model access needed for teams to transition from concept to production efficiently and at scale, without compromises. By providing a unified platform that combines high-performance inference, orchestration, and monitoring, fal facilitates the creation of innovative AI-powered products for both developers and enterprises. As generative media transforms a multibillion-dollar market over the coming decade, fal stands as the essential ecosystem for ambitious teams.

Role Overview

We are seeking a seasoned software engineer proficient in constructing large-scale computing platforms. This role requires deep expertise in managing complex distributed systems that handle heavy traffic and data volumes. You will focus on achieving system reliability and scalability while minimizing operational complexity.

Key Responsibilities

  • Develop our core platform using Python and Rust, including components such as request routing, AI workload orchestration, scheduling, GPU autoscaling, large-scale file storage, and queueing systems.
  • Design forward-looking platform enhancements to accommodate a 100-fold increase in traffic and ensure low-latency service globally.
  • Utilize AI extensively to automate routine aspects of building complex, dependable systems.
  • Profile and optimize system CPU and memory performance at a low level.

Requirements

  • Minimum of three years building distributed compute and orchestration platforms with Python or Rust.
  • Solid grasp of distributed systems principles, including consensus algorithms, scheduling, fault tolerance, and capacity planning.
  • Deep knowledge of computational complexity and memory management.
  • Proven experience designing scalable systems that perform effectively under production workloads.
  • Expertise in employing observability tools to guide performance tuning and reliability improvements.
  • Strong communication skills with the capability to lead technical decisions across multiple teams.
  • A proactive self-starter who takes ownership, executes promptly, and continuously pursues improvements.

Preferred Qualifications

  • Experience with AI/ML inference or training infrastructure.
  • Background in high-performance systems programming, including asynchronous runtimes, zero-copy techniques, and memory-safe concurrency.
  • Familiarity with multi-tenant compute platform development.
  • Understanding of networking fundamentals and their impact on system performance.
  • Knowledge of GPU workload behavior and scheduling constraints.

Compensation

  • Annual salary range of $180,000 to $250,000, supplemented with equity and benefits. This range applies across Mid, Senior, and Staff levels.

Location & Remote Work

  • Position based in San Francisco, CA, with remote work options for Senior and Staff engineers.

What We Offer

  • Engaging and challenging projects.
  • Ample opportunities for personal and professional growth.
  • Relocation support for candidates moving to San Francisco.
  • Comprehensive health, dental, and vision insurance for US employees.
  • Regular team events and offsites to foster collaboration and culture.

他们正在寻找的工作方式

沟通 持续改进 Proactive Mindset

如果您希望收到回复,请留下您的信息——我们不会将您的信息用于其他用途。

点击浏览拖放,或 粘贴 截图

PNG、JPG、GIF、MP4、WebM、MOV 格式 · 每个文件最大 20MB · 最多 5 个文件

🤖
在线·即时人工智能帮助