$97,600 - $142,783 USD yearly
Machine Learning is central to Quora's mission of growing the world's collective intelligence. We have 100+ Machine Learning models in production powering various product features. We use a variety of algorithms — everything from linear models to decision trees and deep neural networks. Our production models operate at a huge scale, serving hundreds of millions of people using Quora every month.
Our team owns Quora's ML platform and ranking infrastructure across four areas: serving reliability, ML engineer enablement and developer velocity, business impact, and cost efficiency. We want to empower all ML engineers at Quora to be as impactful as they can be in solving different ML problems at scale.
As a Software Engineer (New Grad) on this team, you'll work at the intersection of Machine Learning, Distributed Systems, and GPU Serving performance — and your work will have an enormous impact on Quora's long-term success.
No previous ML infrastructure experience is required for this role. You'll be joining a team of senior and staff engineers, learning this stack from the people who built it, with a dedicated mentor and strong technical guidance — and you'll be shipping to production in your first few weeks.
Stack: Python, Go, C++, PyTorch, Kubernetes/EKS, NVIDIA Triton, Ray, AWS
Responsibilities:
- Help build and maintain the core infrastructure that powers Quora's ML platform, ensuring high availability, scalability, and performance
- Build and improve the distributed systems that serve our ML models in production, from Large Recommendation Models (LRM) to Large Language Models (LLM)
- Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models
- Contribute to platform initiatives such as PyTorch-first standardization and ML ecosystem modernization
- Improve ML developer velocity by building tooling that helps ML engineers develop, test, and deploy models more efficiently
- Modernize our feature store so ML engineers can get new features into production faster
- Participate in the team's on-call rotation, helping resolve production issues as you grow your knowledge and ownership of the platform
Minimum Requirements:
- Availability for meetings and impromptu communication during Quora's (Mon-Fri: 9am-3pm Pacific Time)
- A 2025 or 2026 graduate with or pursuing a B.S., M.S., or Ph.D. in Computer Science, Engineering or a related technical field
- Genuine interest in large-scale distributed systems, infrastructure, and machine learning
- Knowledge of Python, Go or C++, or the ability to learn them quickly
- A passion for learning and always improving yourself and the team around you
Preferred Requirements:
- Previous software engineering experience via an internship, work experience, open-source contribution or coding competition
- Coursework or hands-on experience with ML frameworks such as PyTorch or TensorFlow
- Exposure to Kubernetes, Docker, or cloud technologies like AWS
- Experience with low-level performance work of any kind: profiling, benchmarking, optimization
- Passion for Quora's mission and goals