Engineering
$100,000 - $200,000 USD yearly
We're looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability of our platform, and help set the standard for how the rest of engineering builds, ships, and monitors software.
The range of problems is wide. You'll work on petabyte-scale data processing, SaaS problems like customer management and billing, and the query systems that power our APIs. Unlike most finance firms, we're open about our tech and how we build it.
Responsibilities
- Own uptime, SLAs, and SLOs across our API and platform services.
- Set reliability and operational best practices for other developers without slowing them down.
- Build and maintain observability across logging, metrics, and tracing.
- Design and run high-availability deployment and containerization strategies.
- Profile and optimize Python applications for throughput, latency, and cost.
- Debug production issues down to the OS level using tools like strace, perf, eBPF, ss, and gdb.
- Improve deployment and CI/CD workflows.
- Join the on-call rotation, lead incident response, and run post-incident reviews.
- Find what needs fixing on your own, then take projects from idea to completion.
Preferred background
- Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading firm, tech company, or high-growth startup.
- Hands-on experience with observability tooling for logging, metrics, and tracing (e.g. Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, Vector).
- Experience with containerization and high availability deployment (e.g. Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, k3s).
- Strong proficiency in Python, including application development and performance optimization.
- Comfortable with Linux debugging and profiling tools such as strace, perf, eBPF, ss, and gdb.
- A track record of measurable impact in a recent role, such as improving performance by X%, speeding something up Nx, or saving $Y per year.
- Experience with alerting and incident response best practices is a plus.
- Familiarity with configuration management or infrastructure-as-code tools (Ansible, Terraform) is also helpful.
- Bonus if you've done HTTP benchmarking, load testing, and capacity planning.
- Database schema design and query optimization skills are nice to have.
- Good communication skills and work ethic for a remote workplace.
- An interest in financial data or algorithmic trading.