Open Role
Braze

Remote opportunity at

Braze

Senior Site Reliability Engineer I

Braze is currently hiring a Senior Site Reliability Engineer I to join their remote team. This full-time position comes with a stated salary range of USD…

View Company

Role Snapshot

Hiring Now

Remote from

Remote

Salary

USD 153,815 - 277,000/yr

Department

General

Employment

Full-time

Experience

Not specified

Published1h ago
Listing Views1
Applications0
Apply BeforeNo deadline

Career Tools

About This Role

Braze is currently hiring a Senior Site Reliability Engineer I to join their remote team. This full-time position comes with a stated salary range of USD 153,815 to 277,000 per year. In this role, the successful candidate will work closely with product engineering teams to manage ingress fleets, API ingestion layers, and high-performance routing mechanisms. The position involves keeping internal platforms running smoothly, supporting the Ruby on Rails monolith, and…

Job Description

Braze is currently hiring a Senior Site Reliability Engineer I to join their remote team. This full-time position comes with a stated salary range of USD 153,815 to 277,000 per year.

In this role, the successful candidate will work closely with product engineering teams to manage ingress fleets, API ingestion layers, and high-performance routing mechanisms. The position involves keeping internal platforms running smoothly, supporting the Ruby on Rails monolith, and scaling a growing fleet of Go API services.

Braze operates a leading customer engagement platform powered by composable intelligence and AI tools. The company emphasizes a collaborative, transparent culture and provides comprehensive Total Rewards packages that include equity grants and flexible paid time off.

This opportunity suits engineering professionals who have extensive experience with NGINX, Kubernetes, and distributed systems, and who thrive in autonomous, globally distributed environments.

Responsibilities

  • Configure, tune, and operate high-performance NGINX routing, proxying, and ingress controller layers for real-time API traffic.
  • Manage and expand automated scaling routines for high-throughput API services using RED metrics and Horizontal Pod Autoscalers.
  • Collaborate with product engineering teams to translate feature requirements into scalable, highly available technology stacks.
  • Establish Service Level Indicators and Service Level Objectives while managing error budgets for API services.
  • Perform capacity planning, systems design, and bottleneck profiling to meet enterprise-grade SLAs.
  • Participate in a PagerDuty on-call rotation to resolve alerts, update runbooks, and prevent recurring incidents.
  • Lead root-cause analysis and blameless retrospectives for availability and performance incidents.

Requirements

  • Minimum of five years of professional experience as a DevOps or Site Reliability Engineer in high-scale production environments.
  • Extensive hands-on expertise with NGINX proxying, routing, and ingress controller layers under heavy traffic loads.
  • In-depth proficiency in Kubernetes administration, cluster networking, container orchestration, scheduling, and deployment.
  • Solid operating system-level understanding of Linux and Unix internals, including process management, memory allocation, TCP/IP networking, and disk I/O.
  • Strong programming or scripting skills with a preference for Ruby or Go, alongside experience in languages like Python or Java.
  • Practical experience managing infrastructure using Infrastructure as Code technologies such as Terraform or Ansible.

Qualifications

  • Familiarity with data-tier architectures including Redis, Kafka, Postgres, or MongoDB.
  • Experience using observability, monitoring, and profiling platforms like Prometheus, Grafana, or Datadog.
  • Practical deployment and management experience within major cloud environments such as AWS, GCP, or Azure.

Core Skills

Benefits

  • Competitive cash compensation that may include equity grants of restricted stock units.
  • Retirement and Employee Stock Purchase Plans.
  • Flexible paid time off.
  • Comprehensive medical, dental, vision, life, and disability benefit plans.
  • Family services featuring fertility benefits and equal paid parental leave.
  • Professional development support including formal career pathing, learning platforms, and a yearly learning stipend.
  • Curated in-office employee experiences and volunteer matching programs.

Frequently Asked Questions

What is the location and remote status for this job?

This position is fully remote.

What is the salary range for this role?

The salary ranges from USD 153,815 to 277,000 per year. For candidates based in Canada, the pay range is CA$153,815 to CA$277,000 per year with an expected On Target Earnings of CA$172,000 to CA$308,400 per year.

What level of experience is required?

Applicants need at least 5 years of professional experience as a DevOps or Site Reliability Engineer in a high-scale production setting, along with deep hands-on skills in NGINX and Kubernetes.

How can I apply for this position?

The provided source posting does not specify a direct application link or step-by-step application instructions beyond submitting information through their recruitment channels.

Sample Interview Questions

AI-generated questions tailored to this specific role — a preview of the full practice set.

Search similar jobs

Related Jobs

Advertisement
320 × 50

Posted by Braze

Source: Jobicy

Braze

Braze

2Open Jobs
—No reviews yet
View Company Profile