Principal Software Engineer - ML Platform Engineer
Singapore · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 9 કલાક પેહલા
- Work mode
- In office
- Education
- Bachelor's degree in Computer Science or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Riot Games and the Role
Founded in 2006 by passionate gamers, Riot Games developed League of Legends, the world's most played PC game with over 100 million monthly players. Riot’s core is its player community, and the company strives to continually enhance the gaming experience. The AI Efficiency team focuses on providing the platforms and tools that enable safe and effective AI adoption to streamline workflows across Riot's creative and development teams.
As a Principal Platform Engineer within this team, you will be responsible for designing, developing, and evolving internal platforms, automation tools, and operational safeguards that improve the scalability, reliability, and usability of AI-related services and developer workflows. Your collaboration will span software engineers, infrastructure teams, and other stakeholders to enhance the developer experience and operational stability across Riot’s AI services and associated infrastructure.
Key Responsibilities
- Develop and advance internal platform features that simplify building, deploying, monitoring, securing, and operating AI Efficiency services.
- Create and uphold self-service workflows, reusable platform components, and best practice paths to boost developer efficiency while maintaining reliability, security, and compliance.
- Enhance platform dependability through improved monitoring, alerting, observability, safe deployment methods, regular release processes, and readiness for incidents.
- Establish and monitor health indicators, SLIs, SLOs, and reliability metrics to aid teams in balancing reliability, speed, and expense.
- Implement automation to reduce operational workload and accelerate incident detection, response, and resolution.
- Work alongside engineers at all stages of the software lifecycle to embed maintainability, operability, and production readiness into systems.
- Refine continuous integration and delivery pipelines and developer workflows to enable safer, faster, more predictable software releases.
- Identify and mitigate platform and reliability risks present in distributed systems, infrastructure, dependencies, and operational workflows.
- Diagnose and resolve AI model serving challenges across multiple frameworks, hardware setups, and GPU platforms, including model format conversions.
- Design and conduct resilience and failure-mode testing to ensure robustness before issues affect users.
- Explore, integrate, and manage AI-aided engineering tools that enhance code quality, security, performance, and developer productivity.
- Develop and maintain automation pipelines that merge traditional CI/CD with AI-driven workflows such as automated code review and bug remediation.
- Collaborate on the safe and accountable adoption of AI agents in tasks like pull request review, diagnostics, UI validation, accessibility checks, and production verification.
- Define safety guardrails, approval processes, observability, reporting tools, and escalation routes to keep AI-assisted automation reliable and trustworthy.
- Create evaluation frameworks and success metrics measuring effects on quality, false positives, latency, costs, risks, and engineering throughput.
- Participate in incident response and systemic improvements for critical internal platforms to assure long-term resilience.
- Promote operational excellence through detailed documentation, runbooks, standards, and tools to elevate engineering practices organization-wide.
Required Qualifications
- Bachelor’s degree in Computer Science or related discipline, or equivalent experience.
- At least five years’ background in platform, infrastructure, site reliability, DevOps, or related engineering roles supporting production environments.
- Proficient programming and automation skills in Python, Go, JavaScript, TypeScript, or similar languages.
- Experience with designing, creating, or managing internal platforms, developer tools, CI/CD systems, or multi-team infrastructure.
- Familiarity with cloud production systems such as AWS, GCP, or Azure.
- Solid grasp of observability principles including metrics, logs, tracing, dashboards, and alerting design.
- Proven experience improving reliability for distributed systems, microservices, APIs, or platform components.
- Background in incident handling, root cause analysis, and implementing durable operational enhancements.
- Understanding of container technologies and orchestration platforms like Kubernetes or ECS.
- Strong collaboration and communication skills across technical and non-technical stakeholders.
Preferred Qualifications
- Hands-on experience with AI/ML platforms, inference systems, model serving, data pipelines, or GPU-accelerated workloads.
- Usage of service-level objectives, error budgets, and reliability metrics for operational prioritization.
- Platform product mindset with experience in self-service designs, paved paths, or golden rails for users.
- Improvement of developer platforms or enterprise tools’ reliability and user experience.
- Familiarity with infrastructure-as-code and configuration management tools such as Terraform or Pulumi.
- Expertise in security practices, access control, secret management, and operational hardening.
- Balancing availability, latency, cost, efficiency, and usability for large-scale systems.
- Mentoring engineers and providing technical leadership to raise platform standards.
- Assessment and integration of AI-assisted engineering software for static analysis, code review, test generation, and incident investigation.
- Knowledge of agentic AI workflows including safe interactions with source control, CI/CD, browser automation, and platforms.
- Experience with browser automation frameworks like Playwright for UI validation, accessibility, regression, or workflow testing.
- Implementing governance, review cycles, and quality controls for AI-enhanced engineering systems.
- Comfort navigating the crossroads of platform reliability, developer productivity, and AI-native software delivery methods.
Additional Information
Success in this role requires technical mastery, collaboration, and decision-making focused on serving Riot engineers, who are the primary users of your work. Being a gaming enthusiast is not mandatory.
Employee Benefits
- Comprehensive relocation assistance.
- Health coverage including medical for employees and their families.
- Flexible paid vacation policies.
- Retirement plans with company matching contributions.
- Life insurance, maternity/paternity leave, and disability programs.
- Play Fund enabling deeper engagement with gamers and community.
- Matching programs for charitable donations of time and money.
Level
Lead
Minimum education
Bachelor's Degree